Codex 5.5 vs Claude Code: Who Won the Hyperliquid Trading Duel
Given the same prompt, a $100 budget, and one hour to trade perps, Codex 5.5 finished up nine points by actively cycling trades while Claude Code Opus 4.7 held one short and ended down about 3.93 percent.

OpenAI's Codex 5.5 and Anthropic's Claude Code Opus 4.7 were given the same account, the same $100 budget, and the same rules, and one of them finished nine percentage points ahead. The margin was not close. This is the second time Codex has beaten Opus 4.7 in a structured head-to-head challenge from the same creator, and the pattern that emerges across both results carries a practical implication for any business currently making a tooling decision.
I am Madhuranjan Kumar. What follows is my analysis of what actually happened in the Hyperliquid trading duel, why the margin reveals something meaningful about model behavior, what two wins in a row suggest about task-specific model matching, why the duel format is more useful than any benchmark for a real business decision, and what the concrete move is for anyone who needs to choose between two AI tools right now.
Codex 5.5 won a clean trading duel and the margin was not close
The setup was deliberately simple. Both models received an identical prompt and framework. Each got a $100 budget and one hour to trade perpetual contracts on Hyperliquid, restricted to non-crypto instruments: Brent oil, the S&P 500, individual stocks like Nvidia and Tesla. Crypto was deliberately excluded. The scoring metric was one number: total dollars left at the end of the hour.
Each model got a maximum of 15 minutes to research before trading. Research was grounded with a bash date command so each model knew the actual current time, and both had access to free web browsing during the research window. After research, each had to write a concrete trading plan before placing its first position. Both models also set up a live monitor running on an interval so they could cancel, adjust, or add trades across the full 60 minutes.
Because the test used a single account, the two runs happened one hour apart, which introduces some market movement between sessions. That is a legitimate limitation Madhuranjan Kumar acknowledged clearly. But the framework, identical prompt, identical rules, one scoreboard metric, is the important feature. It holds the variable to one thing: the model.
Codex finished with a nine-point gain. Claude ended down approximately 3.93 percent. The total spread between the two results was about 13 percentage points. On a fixed $100 budget that is the difference between walking away with $109 and walking away with $96. On a larger real-money basis, the behavioral difference that produced that gap would have significantly larger consequences.
The gap did not come from a single better call. It came from how each model behaved across the full hour. Codex used the live monitor actively, getting in and out of positions, taking profit when it appeared, adjusting exposure as conditions changed. Claude held a single short position on the S&P 500 for most of the session. That position started well, up around four percent, and then a specific short decision late in the session brought it down sharply. Commitment to a conviction while the market moved against it, rather than adapting, was the difference.
The results across both sessions also illustrate something worth noting about test design. Madhuranjan Kumar ran the two sessions one hour apart because only one account was available. A cleaner test would run both models simultaneously in separate accounts on the same market session. The one-hour gap means market conditions shifted slightly between runs. That is a real limitation. It does not invalidate the result, but it is exactly the kind of methodological detail that matters when you use a duel result to make a real decision. Document your own bake-off conditions with the same transparency, so you know how much confidence the result deserves.

Active position management beat conviction holding by nine percentage points
The behavioral difference between the two models is worth examining carefully, because it is not simply about trading. It is about how each model approaches an ongoing task when a live feedback signal is available and the optimal action changes over time.
Claude's approach was to form a plan, commit to it, and hold. In many contexts that is a reasonable strategy. In a fast-moving environment where information updates every few seconds and the optimal position changes across the hour, it is the wrong strategy. The monitor that both models set up existed precisely to allow mid-course adjustments. Claude used it minimally. Codex used it as its primary working mechanism.
Codex's approach resembles how a professional trader thinks about position management: a plan is a starting hypothesis, not a commitment. Positions get held when the evidence supports them and exited when they do not. The fact that you entered a position is not a reason to stay in it when conditions change. Codex appears to have implemented something closer to this adaptive posture, while Claude maintained a conviction-based approach and paid the price when conviction ran into a bad final call late in the session.
For businesses evaluating AI tools for any task with ongoing feedback, this behavioral difference is directly relevant. A tool that revises its output based on new information mid-task is different from a tool that commits to its first plan and holds. For tasks where the right answer is stable and does not change over time, commitment is fine. For tasks where the environment changes, like managing a meta-ads campaign over the course of a day, responding to new customer information as it comes in, or adjusting a proposed plan based on iterative feedback from a client, the adaptive model will outperform the conviction model over time.
The nine-point spread is the quantified version of that behavioral difference measured over one hour under controlled conditions. It is a clean result from a clean test, and the behavioral explanation for it is worth understanding before drawing any conclusions about which model to use for which type of work.

Codex has now won twice; the pattern suggests task-specific model matching
This was not the first time these two models faced each other in a structured challenge from the same creator. A Polymarket prediction challenge ran with the same format earlier, and Codex won that one too. Two wins across two different types of quantitative, decision-under-uncertainty tasks under the same controlled conditions begins to form a pattern worth taking seriously.
Madhuranjan Kumar's response to the second win is the instructive part. He switched to the Codex max plan for trading-related work. He did not switch for everything. He still rates Claude's Opus model well ahead on front-end development work. The conclusion he drew from two wins was not "Codex is better" in a global sense. It was that for this specific type of task, Codex is better, and he will use it for that task while keeping Opus for the tasks where it leads.
This is the correct response to any head-to-head evaluation, and it is the response that most tool evaluations fail to produce because most tool evaluations are not structured with a single task and a single metric. Benchmarks evaluate models across a range of standardized tasks and produce aggregate scores. Aggregate scores are useful for understanding a model's general capabilities. They are not useful for deciding which tool to use for the one specific task your business does ten times a day.
Two wins for Codex in task-specific quantitative challenges does not mean Codex is ahead overall. It means that for tasks requiring rapid iterative decisions based on a live quantitative signal, Codex appears to have an advantage over Opus 4.7 at this point in time. That is a specific, bounded, actionable finding. It is exactly the kind of finding that makes a tooling decision rather than just informing a debate.
For businesses using AI assistance with google-ads campaign optimization, or any analytical task with a clear success metric, the lesson is that the right model for quantitative optimization may not be the right model for creative copy generation or front-end code. Match the tool to the task category, and base that matching on task-specific evidence rather than aggregate scores.
The duel format is more useful than any benchmark for making a tooling decision
Benchmarks have a fundamental limitation: they measure performance across a standardized set of tasks designed to represent the space of possible tasks in the abstract. No standardized benchmark will perfectly represent the specific thing your business actually needs to do. A model that scores highest on a benchmark may underperform on your particular workflow, and a model that scores lower on the benchmark may outperform on the exact task you run every day.
The duel format bypasses this problem by substituting the benchmark with your actual task. Same prompt, same inputs, same success metric, run both models, observe the result. The only variable is the model. The result is not a general performance claim. It is a specific performance measurement on the exact task you care about.
The Hyperliquid duel ran both models in the same account, with the same tools available, on the same types of instruments, with the same scoring criterion. One hour of market movement between sessions is a legitimate limitation. But the framework is what matters and it is replicable on any task a business needs to decide on.
For a business choosing between two AI tools for handling customer support emails, the duel format looks like this: write one prompt, give both models twenty real past support emails, score each response on the metric your business actually cares about, whether that is resolution accuracy, tone consistency, or adherence to policy, and make the decision based on the scores. That result is worth more than any published benchmark, because it comes from your actual data and your actual metric.
For businesses that use AI assistance in a web-crm context, deciding which tool should draft follow-up emails or summarize call notes, a duel on twenty real past examples with a clear scoring rubric tells you more in two hours than a week of reading benchmark comparisons. The right metric might be: did the draft correctly identify the customer's expressed concern and address it specifically? Score both models on that question, across twenty examples, and pick the one that wins.
The concrete move: run your own bake-off on the task your business actually needs
The practical takeaway from the Hyperliquid duel is not to switch to Codex. It is to design a bake-off for the specific task your business needs solved and run it with the right structure.
A proper bake-off has four components. First, one prompt that both tools receive identically. Not similar prompts that you judge to be equivalent: the exact same text. Any difference in framing introduces a variable that could explain a performance gap and makes the result ambiguous. Second, one metric that both results are scored against. Not a vague judgment of which feels better: a specific, measurable criterion that does not shift after you see the outputs. Third, a sufficient number of examples. The trading duel ran for one hour across one market session. For most business tasks, twenty real examples is a reasonable minimum; fewer than ten is not enough to distinguish signal from noise. Fourth, a commitment to repeat it. One run generates a hypothesis. Two or three runs across different conditions generate confidence.
Madhuranjan Kumar's honest caveat is exactly right: one result is not proof. He plans more weekend tests to check whether the result holds across different market regimes. For a business, rerunning the bake-off on a different day with different inputs, and letting the pattern across multiple runs make the decision, is the responsible standard. A single duel result increases confidence in a hypothesis. Repeated results across varied conditions confirm it.
For any business currently paying for an AI subscription and using one tool for every task because no comparison has been run, the bake-off is a few hours of work that eliminates months of suboptimal tool allocation. Run it on your highest-frequency task first, where a performance difference compounds most rapidly and the payoff from picking the right tool is largest.
Madhuranjan Kumar works with businesses on designing and running these comparisons: defining the right metric for the task, structuring a fair test across both tools, interpreting the results in context, and building the workflow around the winning configuration for each task type. If you want to make a confident tooling decision based on evidence from your own work rather than someone else's benchmark, that conversation is worth having.
That is exactly what we do at AI DOERS. Book a private 30-minute call with Madhuranjan Kumar and we will map the fastest path to it for your specific business.
Book your call →
