Model under test
GPT-5.6 Luna
Every result this site holds for GPT-5.6 Luna. We tested it inside Codex.
For the results tested in Codex, we cannot say which model answered. That test method returns no name for whatever produced the answer, so there is nothing to check against the name we asked for. No second reading through the API was bought for this model, so the figures below are a first measurement and there is nothing here they are checked against.
What was measured, and where
We ran 11 separate tests on this model, listed by test method below and never added together. How the test methods were compared.
Prompting tips, replayed in Codex
11 tests, tested in Codex, Codex run on GPT-5.6 Luna.
- No measured effect 4 of 11
- Could not measure 4 of 11
- Holds 3 of 11
One tip was measured and not published. There is no result for it, and none is estimated from the answers that did arrive.
- Questions about recent events: 6 units exhausted the single C8 retry after policy-blocked built-in tool attempts; no completed tool or MCP hit. Full one-pass coverage not reached.
What these figures can be set beside
Every table on this page is one job, run one way, on one occasion. A model tested the other way is not on those tables and its figures are not set against these. We checked the two test methods against each other and they came out different, so they are reported side by side and never averaged.
- Measured on
- tested in Codex.
- Not directly comparable
- Claude Haiku 4.5; Gemini 3.1 Flash Lite; GPT-5 mini were tested through the API, which this model did not run. No figure below is theirs and none of theirs is set against these.
- Not directly comparable
- Claude Fable 5; Claude Fable 5.1; Claude Haiku 4.5; Claude Opus 5; Claude Sonnet 5 were tested inside Claude Code, which this model did not run. No figure below is theirs and none of theirs is set against these.
The model measured on both, job by job
GPT-5.4 mini was tested through the API and inside Codex. Its readings are the only thing connecting them, and they are stated job by job, never combined into one figure. It is one of two models that connect a pair of test methods this way. Each connects its own pair, and no chain is drawn through them.
Show the readings, job by job (11 jobs)
- Getting clean JSON back GPT-5.4 mini scored 0.0% (no bracket published) tested in Codex, Codex run, tip by tip. That is one test method only, so it bridges nothing on its own.
- Instructions before or after GPT-5.4 mini scored 100.0% (no bracket published) tested in Codex, Codex run, tip by tip. That is one test method only, so it bridges nothing on its own.
- Using tags and formatting GPT-5.4 mini scored 100.0% (no bracket published) tested in Codex, Codex run, tip by tip. That is one test method only, so it bridges nothing on its own.
- Showing examples GPT-5.4 mini scored 100.0% (no bracket published) tested in Codex, Codex run, tip by tip. That is one test method only, so it bridges nothing on its own.
- Showing examples GPT-5.4 mini scored 80.0% (no bracket published) tested in Codex, Codex run, tip by tip. That is one test method only, so it bridges nothing on its own.
- Showing examples GPT-5.4 mini scored 78.6% (no bracket published) tested in Codex, Codex run, tip by tip. That is one test method only, so it bridges nothing on its own.
- Better step-by-step answers GPT-5.4 mini scored 30.0% (no bracket published) tested in Codex, Codex run, tip by tip. That is one test method only, so it bridges nothing on its own.
- Stop made-up answers GPT-5.4 mini scored 30.0% (no bracket published) tested in Codex, Codex run, tip by tip. That is one test method only, so it bridges nothing on its own.
- Instructions in long prompts GPT-5.4 mini scored 100.0% (no bracket published) tested in Codex, Codex run, tip by tip. That is one test method only, so it bridges nothing on its own.
- Offering the model money GPT-5.4 mini scored 193 words (no bracket published) tested in Codex, Codex run, tip by tip. That is one test method only, so it bridges nothing on its own.
- Getting the length right GPT-5.4 mini scored 15.7% (no bracket published) tested via API, API twin for the Codex run, and 18.6% (no bracket published) tested in Codex, Codex run, replaying two earlier tests, and 20.6% (no bracket published) tested in Codex, Codex run, tip by tip. They are stated side by side and are not differenced. How well the two paths agree here: continuous with the API twin, 2 of 2 tests overlap. What was compared.
- The grader, separately
- a second test method, 92 of 120 picks, 10 of 12 tests. The grader itself, replayed on the committed judgements rather than on any subject. Both criteria we set in advance failed, and the graded classes stayed unrun on this path because of it.
What this model does on its own, one job at a time
We measured this model on 11 kinds of work, and the scores for every one of them are below.
Show the figures, one job at a time
One table per job and per test method, ranked by what the model did with no help. Nothing here is averaged across a job, a test method or a run. Every model’s tables sit together on the comparison page.
There is no cost-against-score chart for this model. This test did not report what a task cost on this model, so there is no place for it on the cost scale. It is left off rather than drawn at a guess.
Getting clean JSON back
tested in Codex, Codex run, tip by tip
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5.4 mini | 0.0% interval not measured 10 items | 100.0% C01-json-schema | not measured | 3,879 in, 49 out | 0 |
| 1 | GPT-5.6 Luna | 0.0% interval not measured 10 items | 100.0% C01-json-schema | not measured | 4,620 in, 53 out | 0 |
| 1 | GPT-5.6 Terra | 0.0% interval not measured 10 items | 100.0% C01-json-schema | not measured | 6,135 in, 44 out | 0 |
Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Instructions before or after
tested in Codex, Codex run, tip by tip
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5.4 mini | 100.0% interval not measured 10 items | 100.0% C02-instruction-after | not measured | 4,010 in, 9 out | 0 |
| 1 | GPT-5.6 Luna | 100.0% interval not measured 10 items | 100.0% C02-instruction-after | not measured | 4,754 in, 9 out | 0 |
| 1 | GPT-5.6 Terra | 100.0% interval not measured 10 items | 100.0% C02-instruction-after | not measured | 6,248 in, 9 out | 0 |
Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Using tags and formatting
tested in Codex, Codex run, tip by tip
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5.4 mini | 100.0% interval not measured 10 items | 100.0% C03-xml-delimiters | not measured | 3,973 in, 37 out | 0 |
| 1 | GPT-5.6 Luna | 100.0% interval not measured 10 items | 90.0% C03-xml-delimiters | not measured | 4,715 in, 35 out | 0 |
| 1 | GPT-5.6 Terra | 100.0% interval not measured 10 items | 100.0% C03-xml-delimiters | not measured | 6,230 in, 33 out | 0 |
Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Showing examples
tested in Codex, Codex run, tip by tip
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5.4 mini | 100.0% interval not measured 10 items | 100.0% C04-three-shot-format | not measured | 3,894 in, 22 out | 0 |
| 1 | GPT-5.6 Luna | 100.0% interval not measured 10 items | 100.0% C04-three-shot-format | not measured | 4,637 in, 22 out | 0 |
| 1 | GPT-5.6 Terra | 100.0% interval not measured 10 items | 100.0% C04-three-shot-format | not measured | 6,173 in, 22 out | 0 |
Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Showing examples
tested in Codex, Codex run, tip by tip
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5.6 Luna | 82.9% interval not measured 10 items | 95.7% C04b-zero-vs-one | not measured | 4,720 in, 32 out | 0 |
| 2 | GPT-5.6 Terra | 81.4% interval not measured 10 items | 100.0% C04b-zero-vs-one | not measured | 6,214 in, 31 out | 0 |
| 3 | GPT-5.4 mini | 80.0% interval not measured 10 items | 100.0% C04b-zero-vs-one | not measured | 3,956 in, 31 out | 0 |
Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Showing examples
tested in Codex, Codex run, tip by tip
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5.6 Luna | 81.4% interval not measured 10 items | 100.0% C04b-zero-vs-three | not measured | 4,719 in, 33 out | 0 |
| 1 | GPT-5.6 Terra | 81.4% interval not measured 10 items | 100.0% C04b-zero-vs-three | not measured | 6,215 in, 31 out | 0 |
| 2 | GPT-5.4 mini | 78.6% interval not measured 10 items | 100.0% C04b-zero-vs-three | not measured | 3,956 in, 31 out | 0 |
Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Better step-by-step answers
tested in Codex, Codex run, tip by tip
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5.4 mini | 30.0% interval not measured 10 items | 40.0% C05-think-step-by-step | not measured | 3,879 in, 5 out | 0 |
| 1 | GPT-5.6 Luna | 30.0% interval not measured 10 items | 40.0% C05-think-step-by-step | not measured | 4,599 in, 5 out | 0 |
| 1 | GPT-5.6 Terra | 30.0% interval not measured 10 items | 40.0% C05-think-step-by-step | not measured | 6,135 in, 5 out | 0 |
Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Stop made-up answers
tested in Codex, Codex run, tip by tip
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5.4 mini | 30.0% interval not measured 10 items | 90.0% C06-permit-idk | not measured | 3,857 in, 25 out | 0 |
| 1 | GPT-5.6 Luna | 30.0% interval not measured 10 items | 100.0% C06-permit-idk | not measured | 4,599 in, 17 out | 0 |
| 2 | GPT-5.6 Terra | 20.0% interval not measured 10 items | 100.0% C06-permit-idk | not measured | 6,134 in, 15 out | 0 |
Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Instructions in long prompts
tested in Codex, Codex run, tip by tip
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5.4 mini | 100.0% interval not measured 10 items | 100.0% C07-instruction-at-end | not measured | 4,086 in, 162 out | 0 |
| 1 | GPT-5.6 Luna | 100.0% interval not measured 10 items | 100.0% C07-instruction-at-end | not measured | 4,829 in, 94 out | 0 |
| 1 | GPT-5.6 Terra | 100.0% interval not measured 10 items | 100.0% C07-instruction-at-end | not measured | 6,343 in, 100 out | 0 |
Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Offering the model money
tested in Codex, Codex run, tip by tip
| Rank | Model | Answer length, on its own | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5.4 mini | 193 words interval not measured 10 items | 296 words C08-tip-length | not measured | 3,850 in, 260 out | 0 |
| 2 | GPT-5.6 Luna | 109 words interval not measured 10 items | 192 words C08-tip-length | not measured | 4,592 in, 148 out | 0 |
| 3 | GPT-5.6 Terra | 103 words interval not measured 10 items | 185 words C08-tip-length | not measured | 6,106 in, 138 out | 0 |
Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Getting the length right
tested in Codex, Codex run, tip by tip
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5.6 Terra | 62.4% interval not measured 10 items | 91.4% C17-exact-length | not measured | 6,108 in, 160 out | 0 |
| 2 | GPT-5.6 Luna | 57.7% interval not measured 10 items | 94.6% C17-exact-length | not measured | 4,593 in, 154 out | 0 |
| 3 | GPT-5.4 mini | 20.6% interval not measured 10 items | 95.4% C17-exact-length | not measured | 3,851 in, 326 out | 0 |
Nothing was billed for this run, and no price is estimated for it either. This test method does not report which model answered, so there is no published rate that is known to apply to it.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
This model was also run through Codex, on a path that bills nothing. How it was called, what it was carrying, and what it cannot tell us, is below.
Show the full breakdown, tested in Codex
How this model was called, tested in Codex
Through the Codex command line, on a subscription with no API key and nothing billed. No figure on this page prices a Codex result, not even as an estimate: this path returns no name for the model that answered, so there is no published rate we know applies to it. What that means, and how it was checked.
- What the model was carrying
- tested in Codex (replaced base instructions; runtime tools, permissions and system skills present) Its base instructions were replaced with ours, and the rest of the Codex setup was still there. A result from this method and a result tested via API differ by that as well as by the tip being tested.
- Which model answered
- served model not reported by this instrument We asked for GPT-5.6 Luna and the request was accepted. We looked for a name in the reply twice, two different ways, and there is none, so we cannot check it. On the method that does return one, we found 71 answers that came from a different model than the one we asked for.
These results are tested in Codex and are never averaged with any figure tested via API or tested in Claude Code. Where this model was measured on more than one, each is reported as its own panel above.