Best LLM for Writing Claude Skills, Tested

Writing a skill means turning a request into a skill file that meets a written list of needs. Each model was given the same ten tasks with no help, tested in the API and Claude Code.

Which model to use

Each pick is made inside one test method, and the run it comes from is named under it.

Best on its ownCheapest within range of the bestMost consistentFastest
Via API Gemini 3.1 Flash Lite 87.3% Too close to call with GPT-5 mini tier 1, First run, with and without the skill
In Claude Code Claude Fable 5.1 100.0%, Claude Opus 5 100.0%, Claude Sonnet 5 100.0%, Claude Fable 5 100.0%, tied Too close to call with Claude Opus 5.5 tier 1, Wider run, three ways Some of these models come from another run of the same tasks
Via API Gemini 3.1 Flash Lite $0.00088 a task The best score is also the cheapest in range billed through the API
In Claude Code Claude Sonnet 5 $0.02027 a task The best score is also the cheapest in range run on subscription; shown at API list price for comparison
Via API No run asked these tasks more than once this way, so there is no repeat to compare.
In Claude Code Claude Opus 5, same answer on 28 of 29 no sampling control here, Second run, every task twice
Via API No run of this job timed its calls this way.
In Claude Code Only one model's calls were timed on this job this way.

How each model scored with no help

Under each score is how often the model gave the same answer when asked the same task again.

Via API

ModelScoreSame answer twice
Claude Haiku 4.561.8% tier 1one pass, not measured
GPT-5 mini73.2% tier 1 too close to callone pass, not measured
Gemini 3.1 Flash Lite87.3% tier 1 bestone pass, not measured
GPT-5.4 mininot run

In Claude Code

ModelScoreSame answer twice
Claude Haiku 4.567.3% tier 1one pass, not measured
Claude Fable 5100.0% tier 1 bestone pass, not measured
Claude Fable 5.1100.0% tier 1 bestone pass, not measured
Claude Opus 5100.0% tier 1 bestone pass, not measured
Claude Opus 5.599.8% tier 1 too close to callone pass, not measured
Claude Sonnet 5100.0% tier 1 bestone pass, not measured

What measurably helps on this job

Each line helped on one model, by more than chance. The last column says whether it beat just telling the model what kind of task it was.

Via API

SkillModelBetter than no help byAgainst one plain sentence
skills/skill-creation-walkthrough in the catalogGPT-5 mini+34.5 percentage points (+17.6 to +51.5)beat the one-line instruction by +38.2 percentage points (+19.3 to +57.1)
skills/skill-creatorGPT-5 mini+22.7 percentage points (+10.5 to +35.0)beat the one-line instruction by +33.6 percentage points (+19.3 to +48.0)
skills/skill-creation-walkthrough in the catalogClaude Haiku 4.5+38.2 percentage points (+28.7 to +47.7)the one-line instruction was not asked on this run
skills/skill-creation-walkthrough in the catalogGPT-5 mini+28.2 percentage points (+12.8 to +43.5)the one-line instruction was not asked on this run
distribution/claude-plugin/skills/skill-builderGPT-5 mini+44.5 percentage points (+26.8 to +62.3)beat the one-line instruction by +27.3 percentage points (+14.8 to +39.7)
skills/skill-creatorClaude Haiku 4.5+38.2 percentage points (+28.7 to +47.7)the one-line instruction was not asked on this run
skills/skill-creatorGPT-5 mini+27.3 percentage points (+8.3 to +46.2)the one-line instruction was not asked on this run
skills/skill-creatorGPT-5 mini+35.6 percentage points (+19.8 to +51.4)could not be told apart from the one-line instruction (+13.7 to +24.8)

In Claude Code

SkillModelBetter than no help byAgainst one plain sentence
skills/skill-creation-walkthrough in the catalogClaude Haiku 4.5+30.9 percentage points (+22.0 to +39.8)the one-line instruction was not asked on this run
skills/skill-creatorClaude Haiku 4.5+29.1 percentage points (+22.7 to +35.5)the one-line instruction was not asked on this run
skills/skill-creation-walkthrough in the catalogClaude Haiku 4.5100.0% with it against 69.1% with no helpbeat the one-line instruction by +45.5 percentage points (+45.5 to +45.5)
skills/skill-creatorClaude Haiku 4.5100.0% with it against 69.1% with no helpbeat the one-line instruction by +43.6 percentage points (+35.3 to +52.0)
skills/skill-creatorClaude Haiku 4.5100.0% with it against 65.5% with no helpbeat the one-line instruction by +34.5 percentage points (+23.7 to +45.4)
distribution/claude-plugin/skills/skill-builderClaude Haiku 4.5100.0% with it against 65.5% with no helpbeat the one-line instruction by +40.0 percentage points (+29.3 to +50.7)
Nothing left to measure: these models already scored near the top with no help.
ModelWhereWith no help
Claude Fable 5In Claude Code100.0%, tier 1
Claude Fable 5.1In Claude Code100.0%, tier 1
Claude Opus 5In Claude Code100.0%, tier 1
Claude Opus 5.5In Claude Code99.8%, tier 1
Claude Sonnet 5In Claude Code100.0%, tier 1

How hard the tasks were

These scores are from the first tasks we built for this job, tier 1, the set the models are read on.

The full tables behind this page

Via API

Read from First run, with and without the skill, 2026-08-30.

  • Another run, Wider run, three ways, 2026-09-06: Gemini 3.1 Flash Lite 87.3% (tier 1), GPT-5 mini 64.8% (tier 1)

In Claude Code

Read from Wider run, three ways, 2026-09-06, with the same tasks also from Opus 5.5 series, wider run, three ways, 2026-09-26; Fable 5, wider run, three ways, 2026-09-19.

  • Another run, Second run, every task twice, 2026-08-31: Claude Fable 5 100.0% (tier 1), Claude Opus 5 100.0% (tier 1), Claude Sonnet 5 100.0% (tier 1), Claude Haiku 4.5 70.0% (tier 1)
  • Another run, First run, single attempt, 2026-08-31: Claude Fable 5 100.0% (tier 1), Claude Opus 5 100.0% (tier 1)
  • Another run, Fable 5.1, single attempt, 2026-09-03: Claude Fable 5.1 100.0% (tier 1)
Every result above, with the run it was read from.
SkillModelTest methodRun
skills/skill-creation-walkthrough in the catalogGPT-5 miniVia APIWider run, three ways, tier 1
skills/skill-creatorGPT-5 miniVia APIWider run, three ways, tier 1
skills/skill-creation-walkthrough in the catalogClaude Haiku 4.5Via APIFirst run, with and without the skill, tier 1
skills/skill-creation-walkthrough in the catalogGPT-5 miniVia APIFirst run, with and without the skill, tier 1
distribution/claude-plugin/skills/skill-builderGPT-5 miniVia APIWider run, three ways, tier 1
skills/skill-creatorClaude Haiku 4.5Via APIFirst run, with and without the skill, tier 1
skills/skill-creatorGPT-5 miniVia APIFirst run, with and without the skill, tier 1
skills/skill-creatorGPT-5 miniVia APIWider run, three ways, tier 1
skills/skill-creation-walkthrough in the catalogClaude Haiku 4.5In Claude CodeSecond run, every task twice, tier 1
skills/skill-creatorClaude Haiku 4.5In Claude CodeSecond run, every task twice, tier 1
skills/skill-creation-walkthrough in the catalogClaude Haiku 4.5In Claude CodeWider run, three ways, tier 1
skills/skill-creatorClaude Haiku 4.5In Claude CodeWider run, three ways, tier 1
skills/skill-creatorClaude Haiku 4.5In Claude CodeWider run, three ways, tier 1
distribution/claude-plugin/skills/skill-builderClaude Haiku 4.5In Claude CodeWider run, three ways, tier 1

Version pairs on this job

  • On skill authoring, Claude Fable 5.1 finds 100% of specification requirements unaided against 100% for Claude Fable 5; could not be measured.
  • On skill authoring, wider run, Claude Fable 5.1 finds 100% of specification requirements unaided against 100% for Claude Fable 5; could not be measured.

The job in full: Skill authoring. Every model on the same grid: the models page. Every job’s picks: the routing page. The same figures as data: /jobs/skill-authoring.json.