Model under test
Gemini 3.1 Flash Lite
Every result this site holds for Gemini 3.1 Flash Lite. We tested it through the API.
Version string recorded, as returned: gemini-3.1-flash-lite. The answers carry the requested name back, with no dated snapshot behind it.
Of the 18 tips tested on this model via the API, 4 moved the score by more than the threshold, and the bracket around that move stays clear of zero. The other 14 are not a null result: they are the tips that could not be told apart from no effect, or could not be measured on their scale at all. The fact worth leading with here: fabricated 29 of 30 answers about recent events without a disclosure line.
What was measured, and where
We ran 55 separate tests on this model, listed by test method below and never added together. How the test methods were compared.
Prompting tips, tested via the API
18 tests, tested via API.
- No measured effect 10 of 18
- Holds 4 of 18
- Could not measure 4 of 18
Skills, tested via the API
37 tests, tested via API.
- No measured effect 23 of 37
- Holds 9 of 37
- Could not measure 4 of 37
- Scored worse 1 of 37
What these figures can be set beside
Every table on this page is one job, run one way, on one occasion. A model tested the other way is not on those tables and its figures are not set against these. We checked the two test methods against each other and they came out different, so they are reported side by side and never averaged.
- Measured on
- tested via API.
- Not directly comparable
- Claude Fable 5; Claude Fable 5.1; Claude Opus 5; Claude Sonnet 5 were tested inside Claude Code, which this model did not run. No figure below is theirs and none of theirs is set against these.
- Not directly comparable
- GPT-5.6 Luna; GPT-5.6 Terra were tested inside Codex, which this model did not run. No figure below is theirs and none of theirs is set against these.
The model measured on both, job by job
Claude Haiku 4.5 was tested through the API and inside Claude Code. Its readings are the only thing connecting them, and they are stated job by job, never combined into one figure. It is one of two models that connect a pair of test methods this way. Each connects its own pair, and no chain is drawn through them.
Show the readings, job by job (16 jobs)
- Accessibility audit Claude Haiku 4.5 scored 40.0% (26.3% to 53.7%) tested via API, First run, with and without the skill, and 45.5% (35.0% to 56.0%) tested in Claude Code, Second run, every task twice, and 46.1% (34.1% to 58.1%) tested in Claude Code, Wider run, three ways. They are stated side by side and are not differenced. How well the two paths agree here: 8 of 8 tests agree. What was compared.
- Code review Claude Haiku 4.5 scored 58.3% (47.8% to 68.7%) tested in Claude Code, Wider run, three ways. That is one test method only, so it bridges nothing on its own.
- On-page audit Claude Haiku 4.5 scored 23.9% (1.8% to 46.1%) tested via API, First run, with and without the skill, and 23.1% (6.6% to 39.6%) tested in Claude Code, Second run, every task twice, and 23.1% (3.6% to 42.6%) tested in Claude Code, Wider run, three ways. They are stated side by side and are not differenced. How well the two paths agree here: 8 of 8 tests agree. What was compared.
- Skill authoring Claude Haiku 4.5 scored 61.8% (52.3% to 71.3%) tested via API, First run, with and without the skill, and 70.0% (63.5% to 76.5%) tested in Claude Code, Second run, every task twice, and 67.3% (59.7% to 74.9%) tested in Claude Code, Wider run, three ways. They are stated side by side and are not differenced. How well the two paths agree here: 8 of 8 tests agree. What was compared.
- Spec writing Claude Haiku 4.5 scored 99.5% (98.7% to 100.0%†) tested via API, First run, with and without the skill, and 98.6% (97.7% to 99.6%) tested in Claude Code, Second run, every task twice, and 99.2% (98.7% to 99.7%) tested in Claude Code, Wider run, three ways. They are stated side by side and are not differenced. How well the two paths agree here: 8 of 8 tests agree. What was compared.
- Getting clean JSON back Claude Haiku 4.5 scored 0.0% (0.0% to 0.0%) tested via API, Main run, tip by tip, and 0.0% (0.0% to 0.0%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Instructions before or after Claude Haiku 4.5 scored 100.0% (100.0% to 100.0%) tested via API, Main run, tip by tip, and 100.0% (100.0% to 100.0%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Using tags and formatting Claude Haiku 4.5 scored 100.0% (100.0% to 100.0%) tested via API, Main run, tip by tip, and 100.0% (100.0% to 100.0%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Showing examples Claude Haiku 4.5 scored 81.4% (77.2% to 85.7%) tested via API, Main run, tip by tip, and 81.4% (77.2% to 85.7%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Showing examples Claude Haiku 4.5 scored 81.4% (77.2% to 85.7%) tested via API, Main run, tip by tip, and 81.4% (77.2% to 85.7%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Better step-by-step answers Claude Haiku 4.5 scored 30.0% (17.2% to 42.8%) tested via API, Main run, tip by tip, and 20.0% (8.8% to 31.2%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Stop made-up answers Claude Haiku 4.5 scored 0.0% (0.0% to 0.0%) tested via API, Main run, tip by tip, and 22.0% (10.4% to 33.6%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Instructions in long prompts Claude Haiku 4.5 scored 100.0% (100.0% to 100.0%) tested via API, Main run, tip by tip, and 100.0% (100.0% to 100.0%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Offering the model money Claude Haiku 4.5 scored 204 words (199 words to 209 words) tested via API, Main run, tip by tip, and 241 words (236 words to 247 words) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Questions about recent events Claude Haiku 4.5 scored 85.6% (73.2% to 98.0%) tested via API, Main run, tip by tip, and 96.7% (90.1% to 100.0%†) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
- Getting the length right Claude Haiku 4.5 scored 24.7% (1.2% to 48.2%) tested via API, Main run, tip by tip, and 27.5% (0.0% to 54.9%) tested in Claude Code. They are stated side by side and are not differenced. How well the two paths agree here: second test method, 1 of 2 tests. What was compared.
† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.
- The grader, separately
- a second test method, 92 of 120 picks, 10 of 12 tests. The grader itself, replayed on the committed judgements rather than on any subject. Both criteria we set in advance failed, and the graded classes stayed unrun on this path because of it.
What this model does on its own, one job at a time
We measured this model on 16 kinds of work, and the scores for every one of them are below.
Show the figures, one job at a time
One table per job and per test method, ranked by what the model did with no help. Nothing here is averaged across a job, a test method or a run. Every model’s tables sit together on the comparison page.
Of the 3 models measured on this work, Gemini 3.1 Flash Lite is the cheapest one whose score this test could not tell apart from the best score.
What one task costs with no help, and what it scores Showing examples
- tested via API
- tested in Claude Code
Across: what one task costs with no help, in US dollars at list price. Each gridline is ten times the one before it. Up: deterministic pass rate, 0 to 1.
- Gemini 3.1 Flash Lite Main run, tip by tip tested via API $0.00052 per task scored 80.0%, 75.4% to 84.6%
- GPT-5 mini Main run, tip by tip tested via API $0.00066 per task scored 78.3%, 73.0% to 83.6%
- Claude Haiku 4.5 Main run, tip by tip tested via API $0.00207 per task scored 81.4%, 77.2% to 85.7%
- Claude Haiku 4.5 tested in Claude Code $0.00311 per task scored 81.4%, 77.2% to 85.7%
- Claude Sonnet 5 tested in Claude Code $0.01713 per task scored 81.4%, 77.2% to 85.7%
- Claude Opus 5 tested in Claude Code $0.01882 per task scored 81.4%, 77.2% to 85.7%
- Claude Fable 5.1 tested in Claude Code $0.05013 per task scored 81.7%, 77.7% to 85.7%
- Claude Fable 5 tested in Claude Code $0.08170 per task scored 82.3%, 78.6% to 86.0%
Measured, and not on this chart
- GPT-5.4 mini Codex run, tip by tip tested in Codex This test did not report what a task cost on this model, so there is no place for it on the cost scale. It is left off rather than drawn at a guess. Why some figures are blank.
- GPT-5.6 Luna Codex run, tip by tip tested in Codex This test did not report what a task cost on this model, so there is no place for it on the cost scale. It is left off rather than drawn at a guess. Why some figures are blank.
- GPT-5.6 Terra Codex run, tip by tip tested in Codex This test did not report what a task cost on this model, so there is no place for it on the cost scale. It is left off rather than drawn at a guess. Why some figures are blank.
This is one of 16 sets of work this model was measured on. The same chart for every one of them sits on the comparison page, and the tables below carry this model’s figures for all of them.
Accessibility audit · seeded accessibility defects
tested via API, First run, with and without the skill
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5 mini | 69.9% 54.2% to 85.5% 20 records over 10 items, pooled from 2 entries | 68.9% rampstackco-accessibility-audit (of 2 entries) | $0.00035 billed through the API | 515 in, 111 out | not measured |
| 2 | Gemini 3.1 Flash Lite | 63.3% 46.0% to 80.7% 20 records over 10 items, pooled from 2 entries | 72.8% rampstackco-accessibility-audit (of 2 entries) | $0.00022 billed through the API | 528 in, 59 out | not measured |
| 3 | Claude Haiku 4.5 | 40.0% 26.3% to 53.7% 20 records over 10 items, pooled from 2 entries | 69.9% rampstackco-accessibility-audit (of 2 entries) | $0.00088 billed through the API | 600 in, 55 out | not measured |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
0 of the 20 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Accessibility audit · seeded accessibility defects
tested via API, Wider run, three ways
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Flash Lite | 63.3% 46.0% to 80.7% 40 records over 10 items, pooled from 4 entries | 76.3% alirezarezvani-a11y-audit (of 4 entries) | $0.00022 billed through the API | 528 in, 59 out | not measured |
| 2 | GPT-5 mini | 58.2% 47.9% to 68.4% 40 records over 10 items, pooled from 4 entries | 62.6% rampstackco-accessibility-audit (of 4 entries) | $0.00035 billed through the API | 515 in, 115 out | not measured |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
0 of the 40 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Code review · seeded code defects
tested via API, Wider run, three ways
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5 mini | 70.0% 59.2% to 80.7% 80 records over 10 items, pooled from 8 entries | 73.8% jeffallan-code-reviewer (of 8 entries) | not measured | 665 in, 70 out | not measured |
| 2 | Gemini 3.1 Flash Lite | 57.5% 40.1% to 74.9% 80 records over 10 items, pooled from 8 entries | 60.2% wshobson-code-review-excellence (of 8 entries) | not measured | 783 in, 34 out | not measured |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
0 of the 80 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Nothing was billed on this run for the version being priced here.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
On-page audit · seeded on-page issues
tested via API, First run, with and without the skill
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Flash Lite | 48.8% 27.1% to 70.5% 20 records over 10 items, pooled from 2 entries | 70.7% rampstackco-seo-onpage (of 2 entries) | $0.00023 billed through the API | 599 in, 53 out | not measured |
| 2 | GPT-5 mini | 44.9% 26.7% to 63.1% 20 records over 10 items, pooled from 2 entries | 49.2% rampstackco-seo-onpage (of 2 entries) | $0.00029 billed through the API | 566 in, 75 out | not measured |
| 3 | Claude Haiku 4.5 | 23.9% 1.8% to 46.1% 20 records over 10 items, pooled from 2 entries | 47.4% rampstackco-seo-onpage (of 2 entries) | $0.00085 billed through the API | 667 in, 37 out | not measured |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
0 of the 20 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
On-page audit · seeded on-page issues
tested via API, Wider run, three ways
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5 mini | 52.2% 37.4% to 67.0% 30 records over 10 items, pooled from 3 entries | 57.5% rampstackco-seo-onpage (of 3 entries) | $0.00029 billed through the API | 566 in, 76 out | not measured |
| 2 | Gemini 3.1 Flash Lite | 48.8% 27.1% to 70.5% 30 records over 10 items, pooled from 3 entries | 68.2% rampstackco-seo-onpage (of 3 entries) | $0.00023 billed through the API | 599 in, 53 out | not measured |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
0 of the 30 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
0 of the 33 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Skill authoring · specification requirements
tested via API, First run, with and without the skill
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Flash Lite | 87.3% 80.1% to 94.4% 20 records over 10 items, pooled from 2 entries | 100.0% tied: anthropics-skill-creator, rampstackco-skill-creation-walkthrough (of 2 entries) | $0.00088 billed through the API | 59 in, 580 out | not measured |
| 2 | GPT-5 mini | 73.2% 60.9% to 85.5% 20 records over 10 items, pooled from 2 entries | 100.0% tied: anthropics-skill-creator, rampstackco-skill-creation-walkthrough (of 2 entries) | $0.00297 billed through the API | 61 in, 1,478 out | not measured |
| 3 | Claude Haiku 4.5 | 61.8% 52.3% to 71.3% 20 records over 10 items, pooled from 2 entries | 100.0% tied: anthropics-skill-creator, rampstackco-skill-creation-walkthrough (of 2 entries) | $0.00618 billed through the API | 68 in, 1,222 out | not measured |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
0 of the 20 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Skill authoring · specification requirements
tested via API, Wider run, three ways
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Flash Lite | 87.3% 80.1% to 94.4% 40 records over 10 items, pooled from 4 entries | 100.0% tied: anthropics-skill-creator, openclaw-skill-creator, rampstackco-skill-creation-walkthrough, yusufkaraaslan-skill-builder (of 4 entries) | $0.00088 billed through the API | 59 in, 580 out | not measured |
| 2 | GPT-5 mini | 64.8% 56.1% to 73.4% 40 records over 10 items, pooled from 4 entries | 100.0% tied: anthropics-skill-creator, rampstackco-skill-creation-walkthrough (of 4 entries) | $0.00297 billed through the API | 64 in, 1,610 out | not measured |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
0 of the 40 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
0 of the 42 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Spec writing · required sections and acceptance criteria
tested via API, First run, with and without the skill
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 99.5% 98.7% to 100.0%† 20 records over 10 items, pooled from 2 entries | 99.1% rampstackco-pm-spec-writing (of 2 entries) | $0.00332 billed through the API | 118 in, 641 out | not measured |
| 2 | Gemini 3.1 Flash Lite | 99.1% 97.3% to 100.0%† 20 records over 10 items, pooled from 2 entries | 99.1% tied: affaan-m-product-capability, rampstackco-pm-spec-writing (of 2 entries) | $0.00080 billed through the API | 106 in, 513 out | not measured |
| 3 | GPT-5 mini | 94.5% 92.0% to 97.1% 20 records over 10 items, pooled from 2 entries | 90.0% affaan-m-product-capability (of 2 entries) | $0.00248 billed through the API | 107 in, 1,227 out | not measured |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.
0 of the 40 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Spec writing · required sections and acceptance criteria
tested via API, Wider run, three ways
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Flash Lite | 99.1% 97.3% to 100.0%† 60 records over 10 items, pooled from 6 entries | 100.0% tied: addyosmani-spec-driven-development, alirezarezvani-spec-driven-workflow, github-prd, phuryn-create-prd (of 6 entries) | $0.00080 billed through the API | 106 in, 513 out | not measured |
| 2 | GPT-5 mini | 96.4% 95.0% to 97.7% 60 records over 10 items, pooled from 6 entries | 97.3% affaan-m-product-capability (of 6 entries) | $0.00248 billed through the API | 107 in, 1,204 out | not measured |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.
0 of the 60 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Getting clean JSON back
tested via API, Main run, tip by tip
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 0.0% 0.0% to 0.0% 50 records | 100.0% C01-json-schema | not measured | 277 in, 415 out | not measured |
| 1 | Gemini 3.1 Flash Lite | 0.0% 0.0% to 0.0% 50 records | 100.0% C01-json-schema | not measured | 219 in, 386 out | not measured |
| 1 | GPT-5 mini | 0.0% 0.0% to 0.0% 50 records | 100.0% C01-json-schema | not measured | 235 in, 393 out | not measured |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
0 of the 50 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Nothing was billed on this run for the version being priced here.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Instructions before or after
tested via API, Main run, tip by tip
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C02-instruction-after | not measured | 981 in, 41 out | not measured |
| 1 | Gemini 3.1 Flash Lite | 100.0% 100.0% to 100.0% 50 records | 100.0% C02-instruction-after | not measured | 986 in, 41 out | not measured |
| 1 | GPT-5 mini | 100.0% 100.0% to 100.0% 50 records | 100.0% C02-instruction-after | not measured | 900 in, 71 out | not measured |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
0 of the 50 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Nothing was billed on this run for the version being priced here.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Using tags and formatting
tested via API, Main run, tip by tip
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C03-xml-delimiters | not measured | 749 in, 759 out | not measured |
| 1 | Gemini 3.1 Flash Lite | 100.0% 100.0% to 100.0% 50 records | 100.0% C03-xml-delimiters | not measured | 686 in, 368 out | not measured |
| 1 | GPT-5 mini | 100.0% 100.0% to 100.0% 50 records | 100.0% C03-xml-delimiters | not measured | 708 in, 771 out | not measured |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
0 of the 50 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Nothing was billed on this run for the version being priced here.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Showing examples
tested via API, Main run, tip by tip
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 81.4% 77.2% to 85.7% 10 items | 97.1% C04b-zero-vs-one | not measured | 971 in, 221 out | not measured |
| 2 | Gemini 3.1 Flash Lite | 80.0% 75.4% to 84.6% 10 items | 95.7% C04b-zero-vs-one | not measured | 871 in, 199 out | not measured |
| 3 | GPT-5 mini | 78.3% 73.0% to 83.6% 10 items | 96.9% C04b-zero-vs-one | not measured | 875 in, 218 out | not measured |
0 of the 60 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Nothing was billed on this run for the version being priced here.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Showing examples
tested via API, Main run, tip by tip
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 81.4% 77.2% to 85.7% 10 items | 100.0% C04b-zero-vs-three | not measured | 971 in, 221 out | not measured |
| 2 | Gemini 3.1 Flash Lite | 80.0% 75.4% to 84.6% 10 items | 98.6% C04b-zero-vs-three | not measured | 871 in, 199 out | not measured |
| 3 | GPT-5 mini | 79.4% 74.3% to 84.6% 10 items | 97.7% C04b-zero-vs-three | not measured | 875 in, 219 out | not measured |
0 of the 60 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Nothing was billed on this run for the version being priced here.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Better step-by-step answers
tested via API, Main run, tip by tip
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Flash Lite | 40.0% 26.3% to 53.7% 50 records | 50.0% C05-think-step-by-step | not measured | 220 in, 11 out | not measured |
| 2 | Claude Haiku 4.5 | 30.0% 17.2% to 42.8% 50 records | 50.0% C05-think-step-by-step | not measured | 273 in, 25 out | not measured |
| 3 | GPT-5 mini | 16.0% 5.7% to 26.3% 50 records | 50.0% C05-think-step-by-step | not measured | 237 in, 53 out | not measured |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
0 of the 50 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Nothing was billed on this run for the version being priced here.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Stop made-up answers
tested via API, Main run, tip by tip
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | GPT-5 mini | 20.0% 8.8% to 31.2% 50 records | 100.0% C06-permit-idk | not measured | 231 in, 165 out | not measured |
| 2 | Claude Haiku 4.5 | 0.0% 0.0% to 0.0% 50 records | 100.0% C06-permit-idk | not measured | 256 in, 271 out | not measured |
| 2 | Gemini 3.1 Flash Lite | 0.0% 0.0% to 0.0% 50 records | 100.0% C06-permit-idk | not measured | 222 in, 87 out | not measured |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
0 of the 50 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Nothing was billed on this run for the version being priced here.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Instructions in long prompts
tested via API, Main run, tip by tip
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 100.0% 100.0% to 100.0% 50 records | 100.0% C07-instruction-at-end | not measured | 1,386 in, 878 out | not measured |
| 1 | Gemini 3.1 Flash Lite | 100.0% 100.0% to 100.0% 50 records | 100.0% C07-instruction-at-end | not measured | 1,249 in, 398 out | not measured |
| 1 | GPT-5 mini | 100.0% 100.0% to 100.0% 50 records | 100.0% C07-instruction-at-end | not measured | 1,274 in, 845 out | not measured |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
0 of the 50 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Nothing was billed on this run for the version being priced here.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Offering the model money
tested via API, Main run, tip by tip
| Rank | Model | Answer length, on its own | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Flash Lite | 519 words 506 to 531 50 answers | 575 words C08-tip-length | not measured | 59 in, 3,513 out | not measured |
| 2 | GPT-5 mini | 355 words 334 to 375 50 answers | 534 words C08-tip-length | not measured | 90 in, 2,337 out | not measured |
| 3 | Claude Haiku 4.5 | 204 words 199 to 209 50 answers | 248 words C08-tip-length | not measured | 103 in, 1,619 out | not measured |
This model ran at its own default settings. We could not fix them: one vendor will not accept a setting of 0, and Claude Code has no setting to fix. Recorded as a difference, not a failure.
0 of the 50 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Nothing was billed on this run for the version being priced here.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Questions about recent events
tested via API, Main run, tip by tip
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Claude Haiku 4.5 | 85.6% 73.2% to 98.0% 30 items | 62.2% C16-cutoff-disclosure | not measured | 58 in, 162 out | not measured |
| 2 | GPT-5 mini | 48.6% 31.8% to 65.4% 30 items | 89.7% C16-cutoff-disclosure | not measured | 52 in, 131 out | not measured |
| 3 | Gemini 3.1 Flash Lite | 3.3% 0.0% to 9.9%† 30 items | 33.3% C16-cutoff-disclosure | not measured | 44 in, 90 out | not measured |
† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.
0 of the 108 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Nothing was billed on this run for the version being priced here.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
Getting the length right
tested via API, Main run, tip by tip
| Rank | Model | Unaided score | Best assisted score | Cost per task | Tokens per task | Thinking per task |
|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Flash Lite | 31.0% 5.8% to 56.1% 10 items | 99.4% C17-exact-length | not measured | 120 in, 3,398 out | not measured |
| 2 | Claude Haiku 4.5 | 24.7% 1.2% to 48.2% 10 items | 94.8% C17-exact-length | not measured | 164 in, 2,238 out | not measured |
| 3 | GPT-5 mini | 1.6% 0.0% to 4.6%† 10 items | 96.8% C17-exact-length | not measured | 139 in, 5,014 out | not measured |
† The bracket runs past the end of the scale. It is drawn to the end; the width beyond it is real and is what a small number of items buys.
0 of the 70 answers behind this figure came back with a thinking-token count. A run that reported no count did not report zero.
Nothing was billed on this run for the version being priced here.
Assisted scores name the entry that produced them; skills and tips are ranked on their own pages, not here.
We ran 18 tests on this model through the API. Every one is below, with the difference we measured and how sure we are of it.
Show the full breakdown, tested via API
What we found, tested via API
How this model was called, tested via API
- Vendor
- gemini
- Calls behind these figures
- 1,646 across 2 waves 1,040 in launch, on 2026-08-12; 606 in wave 2, 2026-08-15 to 2026-08-17.
- Sampling
- Sent and accepted Sent and accepted on every call, at temperature 0.
- Reasoning suppression sent
thinkingConfig.thinkingBudget=0- Thinking tokens billed
- Not reported This vendor reports no thinking token figure at all, so an accepted suppression setting is consistent with suppression but does not prove it.
- Launch: 1,040 calls, 2026-08-12, from this run’s committed answer file. The answers carry the requested name back, with no dated snapshot behind it.
- Wave 2: 606 calls, 2026-08-15 to 2026-08-17, from this run’s committed answer file. The answers carry the requested name back, with no dated snapshot behind it.
7 of 18 rows above come from the second wave, where the same task was run twice and the bracket is built from the two runs of each task rather than from every answer separately. Why the two differ.
Where this model is documented by the party that ships it: Google model documentation. Checked 2026-08-17. Get access: Google AI Studio. Checked 2026-08-17.