Model under test

GPT-5.6 Luna

Every result this site holds for GPT-5.6 Luna. We tested it inside Codex.

For the results tested in Codex, we cannot say which model answered. That test method returns no name for whatever produced the answer, so there is nothing to check against the name we asked for. No second reading through the API was bought for this model, so the figures below are a first measurement and there is nothing here they are checked against.

What was measured, and where

We ran 11 separate tests on this model, listed by test method below and never added together. How the test methods were compared.

Prompting tips, replayed in Codex

11 tests, tested in Codex, Codex run on GPT-5.6 Luna.

  • No measured effect 4 of 11
  • Could not measure 4 of 11
  • Holds 3 of 11

One tip was measured and not published. There is no result for it, and none is estimated from the answers that did arrive.

  • Questions about recent events: 6 units exhausted the single C8 retry after policy-blocked built-in tool attempts; no completed tool or MCP hit. Full one-pass coverage not reached.

What these figures can be set beside

Every table on this page is one job, run one way, on one occasion. A model tested the other way is not on those tables and its figures are not set against these. We checked the two test methods against each other and they came out different, so they are reported side by side and never averaged.

Measured on
tested in Codex.
Not directly comparable
Claude Haiku 4.5; Gemini 3.1 Flash Lite; GPT-5 mini were tested through the API, which this model did not run. No figure below is theirs and none of theirs is set against these.
Not directly comparable
Claude Fable 5; Claude Fable 5.1; Claude Haiku 4.5; Claude Opus 5; Claude Sonnet 5 were tested inside Claude Code, which this model did not run. No figure below is theirs and none of theirs is set against these.

The model measured on both, job by job

GPT-5.4 mini was tested through the API and inside Codex. Its readings are the only thing connecting them, and they are stated job by job, never combined into one figure. It is one of two models that connect a pair of test methods this way. Each connects its own pair, and no chain is drawn through them.

Show the readings, job by job (11 jobs)
  • Getting clean JSON back GPT-5.4 mini scored 0.0% (no bracket published) tested in Codex, Codex run, tip by tip. That is one test method only, so it bridges nothing on its own.
  • Instructions before or after GPT-5.4 mini scored 100.0% (no bracket published) tested in Codex, Codex run, tip by tip. That is one test method only, so it bridges nothing on its own.
  • Using tags and formatting GPT-5.4 mini scored 100.0% (no bracket published) tested in Codex, Codex run, tip by tip. That is one test method only, so it bridges nothing on its own.
  • Showing examples GPT-5.4 mini scored 100.0% (no bracket published) tested in Codex, Codex run, tip by tip. That is one test method only, so it bridges nothing on its own.
  • Showing examples GPT-5.4 mini scored 80.0% (no bracket published) tested in Codex, Codex run, tip by tip. That is one test method only, so it bridges nothing on its own.
  • Showing examples GPT-5.4 mini scored 78.6% (no bracket published) tested in Codex, Codex run, tip by tip. That is one test method only, so it bridges nothing on its own.
  • Better step-by-step answers GPT-5.4 mini scored 30.0% (no bracket published) tested in Codex, Codex run, tip by tip. That is one test method only, so it bridges nothing on its own.
  • Stop made-up answers GPT-5.4 mini scored 30.0% (no bracket published) tested in Codex, Codex run, tip by tip. That is one test method only, so it bridges nothing on its own.
  • Instructions in long prompts GPT-5.4 mini scored 100.0% (no bracket published) tested in Codex, Codex run, tip by tip. That is one test method only, so it bridges nothing on its own.
  • Offering the model money GPT-5.4 mini scored 193 words (no bracket published) tested in Codex, Codex run, tip by tip. That is one test method only, so it bridges nothing on its own.
  • Getting the length right GPT-5.4 mini scored 15.7% (no bracket published) tested via API, API twin for the Codex run, and 18.6% (no bracket published) tested in Codex, Codex run, replaying two earlier tests, and 20.6% (no bracket published) tested in Codex, Codex run, tip by tip. They are stated side by side and are not differenced. How well the two paths agree here: continuous with the API twin, 2 of 2 tests overlap. What was compared.
The grader, separately
a second test method, 92 of 120 picks, 10 of 12 tests. The grader itself, replayed on the committed judgements rather than on any subject. Both criteria we set in advance failed, and the graded classes stayed unrun on this path because of it.

What this model does on its own, one job at a time

We measured this model on 11 kinds of work, and the scores for every one of them are below.

Show the figures, one job at a time

One table per job and per test method, ranked by what the model did with no help. Nothing here is averaged across a job, a test method or a run. Every model’s tables sit together on the comparison page.

There is no cost-against-score chart for this model. This test did not report what a task cost on this model, so there is no place for it on the cost scale. It is left off rather than drawn at a guess.

Getting clean JSON back

tested in Codex, Codex run, tip by tip

Getting clean JSON back, tested in Codex, Codex run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1GPT-5.4 mini 0.0% interval not measured 10 items 100.0% C01-json-schema not measured 3,879 in, 49 out 0
1GPT-5.6 Luna 0.0% interval not measured 10 items 100.0% C01-json-schema not measured 4,620 in, 53 out 0
1GPT-5.6 Terra 0.0% interval not measured 10 items 100.0% C01-json-schema not measured 6,135 in, 44 out 0

Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Instructions before or after

tested in Codex, Codex run, tip by tip

Instructions before or after, tested in Codex, Codex run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1GPT-5.4 mini 100.0% interval not measured 10 items 100.0% C02-instruction-after not measured 4,010 in, 9 out 0
1GPT-5.6 Luna 100.0% interval not measured 10 items 100.0% C02-instruction-after not measured 4,754 in, 9 out 0
1GPT-5.6 Terra 100.0% interval not measured 10 items 100.0% C02-instruction-after not measured 6,248 in, 9 out 0

Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Using tags and formatting

tested in Codex, Codex run, tip by tip

Using tags and formatting, tested in Codex, Codex run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1GPT-5.4 mini 100.0% interval not measured 10 items 100.0% C03-xml-delimiters not measured 3,973 in, 37 out 0
1GPT-5.6 Luna 100.0% interval not measured 10 items 90.0% C03-xml-delimiters not measured 4,715 in, 35 out 0
1GPT-5.6 Terra 100.0% interval not measured 10 items 100.0% C03-xml-delimiters not measured 6,230 in, 33 out 0

Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Showing examples

tested in Codex, Codex run, tip by tip

Showing examples, tested in Codex, Codex run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1GPT-5.4 mini 100.0% interval not measured 10 items 100.0% C04-three-shot-format not measured 3,894 in, 22 out 0
1GPT-5.6 Luna 100.0% interval not measured 10 items 100.0% C04-three-shot-format not measured 4,637 in, 22 out 0
1GPT-5.6 Terra 100.0% interval not measured 10 items 100.0% C04-three-shot-format not measured 6,173 in, 22 out 0

Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Showing examples

tested in Codex, Codex run, tip by tip

Showing examples, tested in Codex, Codex run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1GPT-5.6 Luna 82.9% interval not measured 10 items 95.7% C04b-zero-vs-one not measured 4,720 in, 32 out 0
2GPT-5.6 Terra 81.4% interval not measured 10 items 100.0% C04b-zero-vs-one not measured 6,214 in, 31 out 0
3GPT-5.4 mini 80.0% interval not measured 10 items 100.0% C04b-zero-vs-one not measured 3,956 in, 31 out 0

Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Showing examples

tested in Codex, Codex run, tip by tip

Showing examples, tested in Codex, Codex run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1GPT-5.6 Luna 81.4% interval not measured 10 items 100.0% C04b-zero-vs-three not measured 4,719 in, 33 out 0
1GPT-5.6 Terra 81.4% interval not measured 10 items 100.0% C04b-zero-vs-three not measured 6,215 in, 31 out 0
2GPT-5.4 mini 78.6% interval not measured 10 items 100.0% C04b-zero-vs-three not measured 3,956 in, 31 out 0

Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Better step-by-step answers

tested in Codex, Codex run, tip by tip

Better step-by-step answers, tested in Codex, Codex run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1GPT-5.4 mini 30.0% interval not measured 10 items 40.0% C05-think-step-by-step not measured 3,879 in, 5 out 0
1GPT-5.6 Luna 30.0% interval not measured 10 items 40.0% C05-think-step-by-step not measured 4,599 in, 5 out 0
1GPT-5.6 Terra 30.0% interval not measured 10 items 40.0% C05-think-step-by-step not measured 6,135 in, 5 out 0

Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Stop made-up answers

tested in Codex, Codex run, tip by tip

Stop made-up answers, tested in Codex, Codex run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1GPT-5.4 mini 30.0% interval not measured 10 items 90.0% C06-permit-idk not measured 3,857 in, 25 out 0
1GPT-5.6 Luna 30.0% interval not measured 10 items 100.0% C06-permit-idk not measured 4,599 in, 17 out 0
2GPT-5.6 Terra 20.0% interval not measured 10 items 100.0% C06-permit-idk not measured 6,134 in, 15 out 0

Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Instructions in long prompts

tested in Codex, Codex run, tip by tip

Instructions in long prompts, tested in Codex, Codex run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1GPT-5.4 mini 100.0% interval not measured 10 items 100.0% C07-instruction-at-end not measured 4,086 in, 162 out 0
1GPT-5.6 Luna 100.0% interval not measured 10 items 100.0% C07-instruction-at-end not measured 4,829 in, 94 out 0
1GPT-5.6 Terra 100.0% interval not measured 10 items 100.0% C07-instruction-at-end not measured 6,343 in, 100 out 0

Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Offering the model money

tested in Codex, Codex run, tip by tip

Offering the model money, tested in Codex, Codex run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelAnswer length, on its ownBest assisted scoreCost per taskTokens per taskThinking per task
1GPT-5.4 mini 193 words interval not measured 10 items 296 words C08-tip-length not measured 3,850 in, 260 out 0
2GPT-5.6 Luna 109 words interval not measured 10 items 192 words C08-tip-length not measured 4,592 in, 148 out 0
3GPT-5.6 Terra 103 words interval not measured 10 items 185 words C08-tip-length not measured 6,106 in, 138 out 0

Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

Getting the length right

tested in Codex, Codex run, tip by tip

Getting the length right, tested in Codex, Codex run, tip by tip. Which model does this task best with no help, and what it costs.
RankModelUnaided scoreBest assisted scoreCost per taskTokens per taskThinking per task
1GPT-5.6 Terra 62.4% interval not measured 10 items 91.4% C17-exact-length not measured 6,108 in, 160 out 0
2GPT-5.6 Luna 57.7% interval not measured 10 items 94.6% C17-exact-length not measured 4,593 in, 154 out 0
3GPT-5.4 mini 20.6% interval not measured 10 items 95.4% C17-exact-length not measured 3,851 in, 326 out 0

Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.

Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.

This model was also run through Codex, on a path that bills nothing. How it was called, what it was carrying, and what it cannot tell us, is below.

Show the full breakdown, tested in Codex

How this model was called, tested in Codex

Through the Codex command line, on a subscription with no API key and nothing billed. No figure on this page prices a Codex result, not even as an estimate: this path returns no name for the model that answered, so there is no published rate we know applies to it. What that means, and how it was checked.

What the model was carrying
tested in Codex (replaced base instructions; runtime tools, permissions and system skills present) Its base instructions were replaced with ours, and the rest of the Codex setup was still there. A result from this method and a result tested via API differ by that as well as by the tip being tested.
Which model answered
served model not reported by this instrument We asked for GPT-5.6 Luna and the request was accepted. We looked for a name in the reply twice, two different ways, and there is none, so we cannot check it. On the method that does return one, we found 71 answers that came from a different model than the one we asked for.

These results are tested in Codex and are never averaged with any figure tested via API or tested in Claude Code. Where this model was measured on more than one, each is reported as its own panel above.

Every figure on this page is worked out when the site is built, from the committed answers. Nothing here averages a figure from one test method with a figure from another, and the page carries no spread across them, because a model measured two ways has two results and not one.