Model under test
GPT-5.4 mini
Every result this site holds for GPT-5.4 mini. We tested this model two ways, through the API and inside Codex, and never average them.
Version strings recorded, as returned: gpt-5.4-mini, gpt-5.4-mini-2026-03-17. One answer names a dated build: gpt-5.4-mini-2026-03-17. That is the vendor resolving the name we asked for to its own build of it, so it is the model we asked for, written down. For the results tested in Codex, we cannot say which model answered. That test method returns no name for whatever produced the answer, so there is nothing to check against the name we asked for.
What was measured, and where
We ran 4 separate tests on this model, listed by test method below and never added together. A test run through the API and inside Codex are two different measurements, so a single figure across them would be one no run produced. How the test methods were compared.
Prompting tips, replayed through the API beside Codex
2 tests, tested via API, api twin for the codex run.
- Holds 2 of 2
Prompting tips, replayed in Codex
2 tests, tested in Codex.
- Holds 2 of 2
What these figures can be set beside
Every table on this page is one job, run one way, on one occasion. A model tested the other way is not on those tables and its figures are not set against these. We checked the two test methods against each other and they came out different, so they are reported side by side and never averaged.
- Measured on
- tested via API; tested in Codex.
- Not directly comparable
- Claude Fable 5; Claude Fable 5.1; Claude Opus 5; Claude Sonnet 5 were tested inside Claude Code, which this model did not run. No figure below is theirs and none of theirs is set against these.
This model is what connects through the API and inside Codex
GPT-5.4 mini was tested through the API and inside Codex, so its own readings are what connect them. They are listed job by job below and are never combined into one figure. It is one of two models that connect a pair of test methods this way. Each connects its own pair, and no chain is drawn through them.
- Questions about recent events GPT-5.4 mini scored 27.5% (no bracket published) tested via API, API twin for the Codex run, and 15.0% (no bracket published) tested in Codex, Codex run, replaying two earlier tests. They are stated side by side and are not differenced. How well the two paths agree here: continuous with the API twin, 2 of 2 tests overlap. What was compared.
- Getting the length right GPT-5.4 mini scored 15.7% (no bracket published) tested via API, API twin for the Codex run, and 18.6% (no bracket published) tested in Codex, Codex run, replaying two earlier tests. They are stated side by side and are not differenced. How well the two paths agree here: continuous with the API twin, 2 of 2 tests overlap. What was compared.
- The grader, separately
- a second test method, 92 of 120 picks, 10 of 12 tests. The grader itself, replayed on the committed judgements rather than on any subject. Both criteria we set in advance failed, and the graded classes stayed unrun on this path because of it.
What this model does on its own, one job at a time
One table per job and per test method, ranked by what the model did with no help. Nothing here is averaged across a job, a test method or a run. Every model’s tables sit together on the comparison page.
GPT-5.4 mini placed 1 of 1 on this work.
What one task costs with no help, and what it scores Questions about recent events
- tested via API
- tested in Claude Code
- tested in Codex
Across: what one task costs with no help, in US dollars at list price. Each gridline is ten times the one before it. Up: deterministic pass rate, 0 to 1
- Gemini 3.1 Flash Lite Main run, tip by tip tested via API $0.00015 per task scored 0.033, -0.032 to 0.099
- GPT-5 mini Main run, tip by tip tested via API $0.00027 per task scored 0.486, 0.318 to 0.654
- GPT-5.4 mini API twin for the Codex run tested via API $0.00048 per task scored 0.275, no range published
- Claude Haiku 4.5 Main run, tip by tip tested via API $0.00087 per task scored 0.856, 0.732 to 0.980
- Claude Haiku 4.5 tested in Claude Code $0.00220 per task scored 0.967, 0.901 to 1.032
- Claude Sonnet 5 tested in Claude Code $0.01136 per task scored 0.933, 0.843 to 1.024
- GPT-5.4 mini Codex run, replaying two earlier tests tested in Codex $0.01545 per task scored 0.150, no range published
- Claude Opus 5 tested in Claude Code $0.01566 per task scored 0.803, 0.674 to 0.931
- Claude Fable 5.1 tested in Claude Code $0.04170 per task scored 0.922, 0.830 to 1.015
- Claude Fable 5 tested in Claude Code $0.04206 per task scored 0.878, 0.775 to 0.980
This is one of 2 sets of work this model was measured on. The same chart for every one of them sits on the comparison page, and the tables below carry this model’s figures for all of them.
Questions about recent events
tested via API, API twin for the Codex run
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5.4 mini | 27.5% interval not measured 30 items | 86.7% C16-cutoff-disclosure | $0.00048 billed through the API | 52 in, 97 out | not measured |
0 of the 108 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Questions about recent events
tested in Codex, Codex run, replaying two earlier tests
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5.4 mini | 15.0% interval not measured 30 items | 69.2% C16-cutoff-disclosure | not measured | 19,960 in, 107 out | 0 |
Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Getting the length right
tested via API, API twin for the Codex run
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5.4 mini | 15.7% interval not measured 10 items | 96.9% C17-exact-length | $0.01442 billed through the API | 139 in, 3,182 out | not measured |
0 of the 70 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Getting the length right
tested in Codex, Codex run, replaying two earlier tests
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5.4 mini | 18.6% interval not measured 10 items | 96.7% C17-exact-length | not measured | 45,193 in, 2,211 out | 0 |
Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
This model was also run through Codex, on a path that bills nothing. How it was called, what it was carrying, and what it cannot tell us, is below.
Show the full breakdown, tested in Codex
How this model was called, tested in Codex
Through the Codex command line, on a subscription with no API key and nothing billed. No figure on this page prices a Codex result, not even as an estimate: this path returns no name for the model that answered, so there is no published rate we know applies to it. What that means, and how it was checked.
- What the model was carrying
- tested in Codex (replaced base instructions; runtime tools, permissions and system skills present) Its base instructions were replaced with ours, and the rest of the Codex setup was still there. A result from this method and a result tested via API differ by that as well as by the tip being tested.
- Which model answered
- served model not reported by this instrument We asked for GPT-5.4 mini and the request was accepted. We looked for a name in the reply twice, two different ways, and there is none, so we cannot check it. On the method that does return one, we found 71 answers that came from a different model than the one we asked for.
These results are tested in Codex and are never averaged with any figure tested via API or tested in Claude Code. Where this model was measured on more than one, each is reported as its own panel above.