Claude Code vs Codex: The Same Tests in Both
Claude Code and Codex are the two test setups here that run a model with its own tools around it. We sent the same prompting tips through both. No model was tested in both, so this page sets the two side by side and never ranks one above the other.
The same tips through both
| Tip | In Claude Code, helped | In Codex, helped |
|---|---|---|
| Getting clean JSON back | 5 of 5 models | 5 of 5 models |
| Instructions before or after | 0 of 5 models | 0 of 5 models |
| Using tags and formatting | 0 of 5 models | 0 of 5 models |
| Showing examples: three shot format | 0 of 5 models | 0 of 5 models |
| Showing examples: zero vs one | 0 of 5 models | 1 of 5 models |
| Showing examples: zero vs three | 0 of 5 models | 1 of 5 models |
| Better step-by-step answers | 1 of 5 models | 0 of 5 models |
| Stop made-up answers | 5 of 5 models | 5 of 5 models |
| Instructions in long prompts | 0 of 5 models | 0 of 5 models |
| Offering the model money | 0 of 5 models | 0 of 5 models |
| Questions about recent events | 0 of 5 models | 1 of 2 models |
| Getting the length right | 5 of 5 models | 4 of 4 models |
Where a result flips
| Questions about recent events | With no help | What the tip changed |
|---|---|---|
| Via API | 85.6% | −23.3 points (−41.6 to −5.0) |
| In Claude Code | 96.7% | +2.5 points (−4.3 to +9.3) |
On Claude Haiku 4.5, tested both ways, the tip that asks a model to name its knowledge date hurt through the API and made no clear change in Claude Code. Our guess is that Claude Code tells the model the date first. We have not shown that.
Blocked tools on questions about recent events
In Codex, some models tried to look things up when asked about recent events. The setup blocked the attempt. In Claude Code, tools were turned off for these tests, so there was nothing to block.
| Model, in Codex | Blocked attempts on this tip | What happened | Effect on the result |
|---|---|---|---|
| GPT-5.4 mini | 1 | on a recent-events prompt the model reached for the runtime's own documentation and the sandbox refused it; no tool item completed and nothing was read, so the unit was invalidated rather than the instrument called changed | 38 of 60 answers measured on that tip, so it is withheld |
| GPT-5.6 Luna | 16 | 16 invalid attempts on 10 cutoff-disclosure units were blocked before execution. C8 allowed one retry per unit; 6 units remained unmeasured after retry. C9 withholds that claim. No window-limit or model-availability refusal occurred. | 54 of 60 answers measured on that tip, so it is withheld |
| GPT-5.6 Terra | 1 | 1 invalid attempt on 1 cutoff-disclosure unit was policy-blocked. The C8 retry succeeded. All 280 planned units are valid after allowed retries; no claim is withheld under C9. | every answer was measured after one retry |
| GPT-6 Sol | 35 | 38 attempts were blocked by the sandbox policy before anything ran: 3 on 2 exact-length units and 35 on 20 cutoff-disclosure units. Each made its record invalid; the record was kept and never scored. | 45 of 60 answers measured on that tip, so it is withheld |
| GPT-6 Luna | 23 | 23 attempts were blocked by the sandbox policy before anything ran, on 14 cutoff-disclosure units. Each made its record invalid; the record was kept and never scored. | 51 of 60 answers measured on that tip, so it is withheld |
Which to use with which model
No model has been measured in both, so this site cannot say either one gets more out of the same model. What it can say: Claude Haiku 4.5, Claude Fable 5, Claude Fable 5.1, Claude Opus 5, Claude Opus 5.5 and Claude Sonnet 5 were measured in Claude Code, and GPT-5.4 mini, GPT-5.6 Luna, GPT-5.6 Terra, GPT-6 Luna and GPT-6 Sol in Codex. Use each with the models it was measured on.
How each setup was checked against the API
Each setup was checked by running the same tests on one model both ways and comparing the two results.
| What was checked | Result |
|---|---|
| Single-turn claims, replayed on one model against their published API cells | second test method, 1 of 2 tests |
| Multi-turn claims, replayed on one model against their published API cells | continuous with panel v1, 3 of 3 tests |
| Deterministic skill classes, on claude-haiku-4-5, the one model the published API panel and this one share | 8 of 8 tests agree |
| Two single-turn claims, replayed on one model against a same-model API twin bought for the comparison | continuous with the API twin, 2 of 2 tests overlap |
The session is isolated by flags rather than by `--bare`, because `--bare` forces API-key authentication and would defeat the purpose. Tools, slash commands and MCP servers are disabled, the working directory is outside any repository so no instruction file is discoverable above it, and the system prompt is replaced.
Models in Claude Code: Claude Haiku 4.5, Claude Fable 5, Claude Fable 5.1, Claude Opus 5, Claude Opus 5.5 and Claude Sonnet 5. Models in Codex: GPT-5.4 mini, GPT-5.6 Luna, GPT-5.6 Terra, GPT-6 Luna and GPT-6 Sol.