Claude Code vs Codex: The Same Tests in Both

Claude Code and Codex are the two test setups here that run a model with its own tools around it. We sent the same prompting tips through both. No model was tested in both, so this page sets the two side by side and never ranks one above the other.

The same tips through both

TipIn Claude Code, helpedIn Codex, helped
Getting clean JSON back5 of 5 models5 of 5 models
Instructions before or after0 of 5 models0 of 5 models
Using tags and formatting0 of 5 models0 of 5 models
Showing examples: three shot format0 of 5 models0 of 5 models
Showing examples: zero vs one0 of 5 models1 of 5 models
Showing examples: zero vs three0 of 5 models1 of 5 models
Better step-by-step answers1 of 5 models0 of 5 models
Stop made-up answers5 of 5 models5 of 5 models
Instructions in long prompts0 of 5 models0 of 5 models
Offering the model money0 of 5 models0 of 5 models
Questions about recent events0 of 5 models1 of 2 models
Getting the length right5 of 5 models4 of 4 models

Where a result flips

Questions about recent eventsWith no helpWhat the tip changed
Via API85.6%−23.3 points (−41.6 to −5.0)
In Claude Code96.7%+2.5 points (−4.3 to +9.3)

On Claude Haiku 4.5, tested both ways, the tip that asks a model to name its knowledge date hurt through the API and made no clear change in Claude Code. Our guess is that Claude Code tells the model the date first. We have not shown that.

Blocked tools on questions about recent events

In Codex, some models tried to look things up when asked about recent events. The setup blocked the attempt. In Claude Code, tools were turned off for these tests, so there was nothing to block.

Model, in CodexBlocked attempts on this tipWhat happenedEffect on the result
GPT-5.4 mini1on a recent-events prompt the model reached for the runtime's own documentation and the sandbox refused it; no tool item completed and nothing was read, so the unit was invalidated rather than the instrument called changed38 of 60 answers measured on that tip, so it is withheld
GPT-5.6 Luna1616 invalid attempts on 10 cutoff-disclosure units were blocked before execution. C8 allowed one retry per unit; 6 units remained unmeasured after retry. C9 withholds that claim. No window-limit or model-availability refusal occurred.54 of 60 answers measured on that tip, so it is withheld
GPT-5.6 Terra11 invalid attempt on 1 cutoff-disclosure unit was policy-blocked. The C8 retry succeeded. All 280 planned units are valid after allowed retries; no claim is withheld under C9.every answer was measured after one retry
GPT-6 Sol3538 attempts were blocked by the sandbox policy before anything ran: 3 on 2 exact-length units and 35 on 20 cutoff-disclosure units. Each made its record invalid; the record was kept and never scored.45 of 60 answers measured on that tip, so it is withheld
GPT-6 Luna2323 attempts were blocked by the sandbox policy before anything ran, on 14 cutoff-disclosure units. Each made its record invalid; the record was kept and never scored.51 of 60 answers measured on that tip, so it is withheld

Which to use with which model

No model has been measured in both, so this site cannot say either one gets more out of the same model. What it can say: Claude Haiku 4.5, Claude Fable 5, Claude Fable 5.1, Claude Opus 5, Claude Opus 5.5 and Claude Sonnet 5 were measured in Claude Code, and GPT-5.4 mini, GPT-5.6 Luna, GPT-5.6 Terra, GPT-6 Luna and GPT-6 Sol in Codex. Use each with the models it was measured on.

How each setup was checked against the API

Each setup was checked by running the same tests on one model both ways and comparing the two results.

What was checkedResult
Single-turn claims, replayed on one model against their published API cellssecond test method, 1 of 2 tests
Multi-turn claims, replayed on one model against their published API cellscontinuous with panel v1, 3 of 3 tests
Deterministic skill classes, on claude-haiku-4-5, the one model the published API panel and this one share8 of 8 tests agree
Two single-turn claims, replayed on one model against a same-model API twin bought for the comparisoncontinuous with the API twin, 2 of 2 tests overlap

The session is isolated by flags rather than by `--bare`, because `--bare` forces API-key authentication and would defeat the purpose. Tools, slash commands and MCP servers are disabled, the working directory is outside any repository so no instruction file is discoverable above it, and the system prompt is replaced.

Models in Claude Code: Claude Haiku 4.5, Claude Fable 5, Claude Fable 5.1, Claude Opus 5, Claude Opus 5.5 and Claude Sonnet 5. Models in Codex: GPT-5.4 mini, GPT-5.6 Luna, GPT-5.6 Terra, GPT-6 Luna and GPT-6 Sol.