Best LLM for Writing Claude Skills, Tested
Writing a skill means turning a request into a skill file that meets a written list of needs. Each model was given the same ten tasks with no help, tested in the API and Claude Code.
Which model to use
Each pick is made inside one test method, and the run it comes from is named under it.
| Best on its own | Cheapest within range of the best | Most consistent | Fastest |
|---|---|---|---|
Via API
Gemini 3.1 Flash Lite 87.3%
Too close to call with GPT-5 mini
tier 1, First run, with and without the skill In Claude Code
Claude Fable 5.1 100.0%, Claude Opus 5 100.0%, Claude Sonnet 5 100.0%, Claude Fable 5 100.0%, tied
Too close to call with Claude Opus 5.5
tier 1, Wider run, three ways
Some of these models come from another run of the same tasks | Via API
Gemini 3.1 Flash Lite $0.00088 a task
The best score is also the cheapest in range
billed through the API In Claude Code
Claude Sonnet 5 $0.02027 a task
The best score is also the cheapest in range
run on subscription; shown at API list price for comparison | Via API
No run asked these tasks more than once this way, so there is no repeat to compare. In Claude Code
Claude Opus 5, same answer on 28 of 29
no sampling control here, Second run, every task twice | Via API
No run of this job timed its calls this way. In Claude Code
Only one model's calls were timed on this job this way. |
How each model scored with no help
Under each score is how often the model gave the same answer when asked the same task again.
Via API
| Model | Score | Same answer twice |
|---|---|---|
| Claude Haiku 4.5 | 61.8% tier 1 | one pass, not measured |
| GPT-5 mini | 73.2% tier 1 too close to call | one pass, not measured |
| Gemini 3.1 Flash Lite | 87.3% tier 1 best | one pass, not measured |
| GPT-5.4 mini | not run | |
In Claude Code
| Model | Score | Same answer twice |
|---|---|---|
| Claude Haiku 4.5 | 67.3% tier 1 | one pass, not measured |
| Claude Fable 5 | 100.0% tier 1 best | one pass, not measured |
| Claude Fable 5.1 | 100.0% tier 1 best | one pass, not measured |
| Claude Opus 5 | 100.0% tier 1 best | one pass, not measured |
| Claude Opus 5.5 | 99.8% tier 1 too close to call | one pass, not measured |
| Claude Sonnet 5 | 100.0% tier 1 best | one pass, not measured |
What measurably helps on this job
Each line helped on one model, by more than chance. The last column says whether it beat just telling the model what kind of task it was.
Via API
| Skill | Model | Better than no help by | Against one plain sentence |
|---|---|---|---|
| skills/skill-creation-walkthrough in the catalog | GPT-5 mini | +34.5 percentage points (+17.6 to +51.5) | beat the one-line instruction by +38.2 percentage points (+19.3 to +57.1) |
| skills/skill-creator | GPT-5 mini | +22.7 percentage points (+10.5 to +35.0) | beat the one-line instruction by +33.6 percentage points (+19.3 to +48.0) |
| skills/skill-creation-walkthrough in the catalog | Claude Haiku 4.5 | +38.2 percentage points (+28.7 to +47.7) | the one-line instruction was not asked on this run |
| skills/skill-creation-walkthrough in the catalog | GPT-5 mini | +28.2 percentage points (+12.8 to +43.5) | the one-line instruction was not asked on this run |
| distribution/claude-plugin/skills/skill-builder | GPT-5 mini | +44.5 percentage points (+26.8 to +62.3) | beat the one-line instruction by +27.3 percentage points (+14.8 to +39.7) |
| skills/skill-creator | Claude Haiku 4.5 | +38.2 percentage points (+28.7 to +47.7) | the one-line instruction was not asked on this run |
| skills/skill-creator | GPT-5 mini | +27.3 percentage points (+8.3 to +46.2) | the one-line instruction was not asked on this run |
| skills/skill-creator | GPT-5 mini | +35.6 percentage points (+19.8 to +51.4) | could not be told apart from the one-line instruction (+13.7 to +24.8) |
In Claude Code
| Skill | Model | Better than no help by | Against one plain sentence |
|---|---|---|---|
| skills/skill-creation-walkthrough in the catalog | Claude Haiku 4.5 | +30.9 percentage points (+22.0 to +39.8) | the one-line instruction was not asked on this run |
| skills/skill-creator | Claude Haiku 4.5 | +29.1 percentage points (+22.7 to +35.5) | the one-line instruction was not asked on this run |
| skills/skill-creation-walkthrough in the catalog | Claude Haiku 4.5 | 100.0% with it against 69.1% with no help | beat the one-line instruction by +45.5 percentage points (+45.5 to +45.5) |
| skills/skill-creator | Claude Haiku 4.5 | 100.0% with it against 69.1% with no help | beat the one-line instruction by +43.6 percentage points (+35.3 to +52.0) |
| skills/skill-creator | Claude Haiku 4.5 | 100.0% with it against 65.5% with no help | beat the one-line instruction by +34.5 percentage points (+23.7 to +45.4) |
| distribution/claude-plugin/skills/skill-builder | Claude Haiku 4.5 | 100.0% with it against 65.5% with no help | beat the one-line instruction by +40.0 percentage points (+29.3 to +50.7) |
| Model | Where | With no help |
|---|---|---|
| Claude Fable 5 | In Claude Code | 100.0%, tier 1 |
| Claude Fable 5.1 | In Claude Code | 100.0%, tier 1 |
| Claude Opus 5 | In Claude Code | 100.0%, tier 1 |
| Claude Opus 5.5 | In Claude Code | 99.8%, tier 1 |
| Claude Sonnet 5 | In Claude Code | 100.0%, tier 1 |
How hard the tasks were
These scores are from the first tasks we built for this job, tier 1, the set the models are read on.
The full tables behind this page
Via API
Read from First run, with and without the skill, 2026-08-30.
- Another run, Wider run, three ways, 2026-09-06: Gemini 3.1 Flash Lite 87.3% (tier 1), GPT-5 mini 64.8% (tier 1)
In Claude Code
Read from Wider run, three ways, 2026-09-06, with the same tasks also from Opus 5.5 series, wider run, three ways, 2026-09-26; Fable 5, wider run, three ways, 2026-09-19.
- Another run, Second run, every task twice, 2026-08-31: Claude Fable 5 100.0% (tier 1), Claude Opus 5 100.0% (tier 1), Claude Sonnet 5 100.0% (tier 1), Claude Haiku 4.5 70.0% (tier 1)
- Another run, First run, single attempt, 2026-08-31: Claude Fable 5 100.0% (tier 1), Claude Opus 5 100.0% (tier 1)
- Another run, Fable 5.1, single attempt, 2026-09-03: Claude Fable 5.1 100.0% (tier 1)
| Skill | Model | Test method | Run |
|---|---|---|---|
| skills/skill-creation-walkthrough in the catalog | GPT-5 mini | Via API | Wider run, three ways, tier 1 |
| skills/skill-creator | GPT-5 mini | Via API | Wider run, three ways, tier 1 |
| skills/skill-creation-walkthrough in the catalog | Claude Haiku 4.5 | Via API | First run, with and without the skill, tier 1 |
| skills/skill-creation-walkthrough in the catalog | GPT-5 mini | Via API | First run, with and without the skill, tier 1 |
| distribution/claude-plugin/skills/skill-builder | GPT-5 mini | Via API | Wider run, three ways, tier 1 |
| skills/skill-creator | Claude Haiku 4.5 | Via API | First run, with and without the skill, tier 1 |
| skills/skill-creator | GPT-5 mini | Via API | First run, with and without the skill, tier 1 |
| skills/skill-creator | GPT-5 mini | Via API | Wider run, three ways, tier 1 |
| skills/skill-creation-walkthrough in the catalog | Claude Haiku 4.5 | In Claude Code | Second run, every task twice, tier 1 |
| skills/skill-creator | Claude Haiku 4.5 | In Claude Code | Second run, every task twice, tier 1 |
| skills/skill-creation-walkthrough in the catalog | Claude Haiku 4.5 | In Claude Code | Wider run, three ways, tier 1 |
| skills/skill-creator | Claude Haiku 4.5 | In Claude Code | Wider run, three ways, tier 1 |
| skills/skill-creator | Claude Haiku 4.5 | In Claude Code | Wider run, three ways, tier 1 |
| distribution/claude-plugin/skills/skill-builder | Claude Haiku 4.5 | In Claude Code | Wider run, three ways, tier 1 |
Version pairs on this job
- On skill authoring, Claude Fable 5.1 finds 100% of specification requirements unaided against 100% for Claude Fable 5; could not be measured.
- On skill authoring, wider run, Claude Fable 5.1 finds 100% of specification requirements unaided against 100% for Claude Fable 5; could not be measured.
The job in full: Skill authoring. Every model on the same grid: the models page. Every job’s picks: the routing page. The same figures as data: /jobs/skill-authoring.json.