Model under test

GPT-5.4 mini

Every result this site holds for GPT-5.4 mini. We tested this model two ways, through the API and inside Codex, and never average them.

Version strings recorded, as returned: gpt-5.4-mini, gpt-5.4-mini-2026-03-17. One answer names a dated build: gpt-5.4-mini-2026-03-17. That is the vendor resolving the name we asked for to its own build of it, so it is the model we asked for, written down. For the results tested in Codex, we cannot say which model answered. That test method returns no name for whatever produced the answer, so there is nothing to check against the name we asked for.

What was measured, and where

We ran 4 separate tests on this model, listed by test method below and never added together. A test run through the API and inside Codex are two different measurements, so a single figure across them would be one no run produced. How the test methods were compared.

Prompting tips, replayed through the API beside Codex

2 tests, tested via API, api twin for the codex run.

  • Holds 2 of 2

Prompting tips, replayed in Codex

2 tests, tested in Codex.

  • Holds 2 of 2

What these figures can be set beside

Every table on this page is one job, run one way, on one occasion. A model tested the other way is not on those tables and its figures are not set against these. We checked the two test methods against each other and they came out different, so they are reported side by side and never averaged.

Measured on
tested via API; tested in Codex.
Not directly comparable
Claude Fable 5; Claude Fable 5.1; Claude Opus 5; Claude Sonnet 5 were tested inside Claude Code, which this model did not run. No figure below is theirs and none of theirs is set against these.

This model is what connects through the API and inside Codex

GPT-5.4 mini was tested through the API and inside Codex, so its own readings are what connect them. They are listed job by job below and are never combined into one figure. It is one of two models that connect a pair of test methods this way. Each connects its own pair, and no chain is drawn through them.

  • Questions about recent events GPT-5.4 mini scored 27.5% (no bracket published) tested via API, API twin for the Codex run, and 15.0% (no bracket published) tested in Codex, Codex run, replaying two earlier tests. They are stated side by side and are not differenced. How well the two paths agree here: continuous with the API twin, 2 of 2 tests overlap. What was compared.
  • Getting the length right GPT-5.4 mini scored 15.7% (no bracket published) tested via API, API twin for the Codex run, and 18.6% (no bracket published) tested in Codex, Codex run, replaying two earlier tests. They are stated side by side and are not differenced. How well the two paths agree here: continuous with the API twin, 2 of 2 tests overlap. What was compared.
The grader, separately
a second test method, 92 of 120 picks, 10 of 12 tests. The grader itself, replayed on the committed judgements rather than on any subject. Both criteria we set in advance failed, and the graded classes stayed unrun on this path because of it.

What this model does on its own, one job at a time

One table per job and per test method, ranked by what the model did with no help. Nothing here is averaged across a job, a test method or a run. Every model’s tables sit together on the comparison page.

GPT-5.4 mini placed 1 of 1 on this work.

What one task costs with no help, and what it scores Questions about recent events

  • tested via API
  • tested in Claude Code
  • tested in Codex

Across: what one task costs with no help, in US dollars at list price. Each gridline is ten times the one before it. Up: deterministic pass rate, 0 to 1

  1. Gemini 3.1 Flash Lite Main run, tip by tip tested via API $0.00015 per task scored 0.033, -0.032 to 0.099
  2. GPT-5 mini Main run, tip by tip tested via API $0.00027 per task scored 0.486, 0.318 to 0.654
  3. GPT-5.4 mini API twin for the Codex run tested via API $0.00048 per task scored 0.275, no range published
  4. Claude Haiku 4.5 Main run, tip by tip tested via API $0.00087 per task scored 0.856, 0.732 to 0.980
  5. Claude Haiku 4.5 tested in Claude Code $0.00220 per task scored 0.967, 0.901 to 1.032
  6. Claude Sonnet 5 tested in Claude Code $0.01136 per task scored 0.933, 0.843 to 1.024
  7. GPT-5.4 mini Codex run, replaying two earlier tests tested in Codex $0.01545 per task scored 0.150, no range published
  8. Claude Opus 5 tested in Claude Code $0.01566 per task scored 0.803, 0.674 to 0.931
  9. Claude Fable 5.1 tested in Claude Code $0.04170 per task scored 0.922, 0.830 to 1.015
  10. Claude Fable 5 tested in Claude Code $0.04206 per task scored 0.878, 0.775 to 0.980
Each mark is one model on one job: what a task cost it with no help, against the score it got. The bar through a mark shows how much that score could move if we ran it again. Shape and colour say which test method the model was reached through. The two test methods are measured on their own and are never ranked against each other. Compare marks of one colour with each other; a mark of one colour sitting above a mark of another is not a result. The tables are the source, and every figure drawn here is printed above. A model measured more than once appears once per test, named by test. Two tests are two measurements and are never averaged into one mark. Names sit beside their marks where there is room for them, moved down where two would have overlapped, and on a narrow screen they are in the list above instead. The list pairs every name with its figures at every width. A mark with no bar is one whose test publishes no range; the reason is on that job’s own page. A bar wider than the scale is drawn to the edge; the tables give its two ends.

This is one of 2 sets of work this model was measured on. The same chart for every one of them sits on the comparison page, and the tables below carry this model’s figures for all of them.

Questions about recent events

tested via API, API twin for the Codex run

Questions about recent events, tested via API, API twin for the Codex run. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1GPT-5.4 mini 27.5% interval not measured 30 items 86.7% C16-cutoff-disclosure $0.00048 billed through the API 52 in, 97 out not measured

0 of the 108 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Questions about recent events

tested in Codex, Codex run, replaying two earlier tests

Questions about recent events, tested in Codex, Codex run, replaying two earlier tests. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1GPT-5.4 mini 15.0% interval not measured 30 items 69.2% C16-cutoff-disclosure not measured 19,960 in, 107 out 0

Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Getting the length right

tested via API, API twin for the Codex run

Getting the length right, tested via API, API twin for the Codex run. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1GPT-5.4 mini 15.7% interval not measured 10 items 96.9% C17-exact-length $0.01442 billed through the API 139 in, 3,182 out not measured

0 of the 70 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Getting the length right

tested in Codex, Codex run, replaying two earlier tests

Getting the length right, tested in Codex, Codex run, replaying two earlier tests. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1GPT-5.4 mini 18.6% interval not measured 10 items 96.7% C17-exact-length not measured 45,193 in, 2,211 out 0

Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

This model was also run through Codex, on a path that bills nothing. How it was called, what it was carrying, and what it cannot tell us, is below.

Show the full breakdown, tested in Codex

How this model was called, tested in Codex

Through the Codex command line, on a subscription with no API key and nothing billed. No figure on this page prices a Codex result, not even as an estimate: this path returns no name for the model that answered, so there is no published rate we know applies to it. What that means, and how it was checked.

What the model was carrying
tested in Codex (replaced base instructions; runtime tools, permissions and system skills present) Its base instructions were replaced with ours, and the rest of the Codex setup was still there. A result from this method and a result tested via API differ by that as well as by the tip being tested.
Which model answered
served model not reported by this instrument We asked for GPT-5.4 mini and the request was accepted. We looked for a name in the reply twice, two different ways, and there is none, so we cannot check it. On the method that does return one, we found 71 answers that came from a different model than the one we asked for.

These results are tested in Codex and are never averaged with any figure tested via API or tested in Claude Code. Where this model was measured on more than one, each is reported as its own panel above.

Every figure on this page is worked out when the site is built, from the committed answers. Nothing here averages a figure from one test method with a figure from another, and the page carries no spread across them, because a model measured two ways has two results and not one.