Claude skills for Code Review, tested

Skills for code review work, measured on code review tasks. Each was the same task, run twice, once with the skill loaded and once without it, and what is below is what changed. Ranked by how many models the rule found a clear effect on, never by the size of one.

Measured on 2 models (GPT & Gemini): GPT-5 mini, Gemini 3.1 Flash Lite.

10 fixture source files in 4 languages, each seeded with known bugs, security issues and correctness problems, 54 in total, drawn from a closed list of 18 defect types that both arms are given. The model is asked to name the symbol each defect sits on, never a line number; the score is the share of seeded defects it names. 50 are verified in the fixtures by mechanical check, 4 are judgement items listed as such, and 30 pieces of correct code that read as suspicious are recorded as traps and are deliberately not in the key.

On the models tested, no skill measurably helped with code review tasks; the unaided scores are the story.

What the models managed with no skill loaded
Task setUnaided score
Code review64%
Code review63%
Code review63%
Code review64%
Code review63%
Code review64%
Code review63%
Code review64%

The numbers

How to read these tables

How to read these tables

Without -> with
The score on the task set with no skill loaded, then with the skill loaded. For classes scored against an answer key, this is the share of seeded items found. For classes scored by blind comparison, it is how often a grader preferred the skill's output over the baseline's, with the baseline shown as the complement.
Effect (range)
With minus without, averaged over the models tested, with the lowest and highest per-model value in brackets. Per-model intervals are on each skill's page.
Orbit
  • Stable: delta at or above the pass threshold, and the interval excludes zero
  • Past the horizon: delta at or below the failure floor, and the interval excludes zero
  • In free drift: the interval spans zero, or the delta sits between the floor and the threshold
  • Unobservable: the cell cannot be classified: sub-case (a) an unbounded scale, or sub-case (b) both arms at the same bound

A fifth value, Decaying, is defined and appears once a skill has been re-tested: a cell previously Stable that has fallen below the pass threshold on a later run.

Cohort
The repository that publishes the skill. Rows marked with an asterisk come from the repository the operator of this site maintains, so those are the operator measuring its own work. The rest were chosen by a stated rule and not by preference.
Commit
The exact version tested. Later commits are not covered by this verdict.
Injected context tokens
How much text the skill adds to the model's context. This is cost, not quality; a bigger number is not a better one. Measured from the input tokens the API reported on the with arm, averaged over that skill's records.
Tested
The date. Verdicts age as models change.

Code review

With the skill and without it, per skill and model

  • Grey: the score with no skill loaded
  • Colour: the score with the skill loaded
  • Dashed tick: the score with the instruction only, where that was measured

Via API tested via API

skills/code-reviewer GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.693. With the skill: 0.738. Instruction only: 0.690. Change +0.045. It could sit between -0.032 and +0.122.

skills/code-review-web GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.710. With the skill: 0.702. Instruction only: 0.710. Change -0.008. It could sit between -0.100 and +0.083.

skills/code-review-and-quality GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.710. With the skill: 0.683. Instruction only: 0.727. Change -0.027. It could sit between -0.133 and +0.079.

skills/ce-code-review GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.710. With the skill: 0.677. Instruction only: 0.693. Change -0.033. It could sit between -0.077 and +0.010.

skills/gsd-code-review GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.677. With the skill: 0.660. Instruction only: 0.693. Change -0.017. It could sit between -0.124 and +0.091.

plugins/developer-essentials/skills/code-review-excellence GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.710. With the skill: 0.658. Instruction only: 0.643. Change -0.052. It could sit between -0.221 and +0.117.

skills/engineering/code-review GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.693. With the skill: 0.657. Instruction only: 0.792. Change -0.037. It could sit between -0.085 and +0.011.

engineering/skills/pr-review-expert GPT-5 mini Wider run, three ways, Deterministic

No measured effect

No skill: 0.693. With the skill: 0.628. Instruction only: 0.707. Change -0.065. It could sit between -0.171 and +0.041.

plugins/developer-essentials/skills/code-review-excellence Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.575. With the skill: 0.602. Instruction only: 0.592. Change +0.027. It could sit between -0.076 and +0.130.

engineering/skills/pr-review-expert Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.575. With the skill: 0.578. Instruction only: 0.592. Change +0.003. It could sit between -0.069 and +0.076.

skills/code-review-and-quality Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.575. With the skill: 0.542. Instruction only: 0.592. Change -0.033. It could sit between -0.077 and +0.010.

skills/code-reviewer Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.575. With the skill: 0.538. Instruction only: 0.592. Change -0.037. It could sit between -0.127 and +0.053.

skills/gsd-code-review Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.575. With the skill: 0.522. Instruction only: 0.592. Change -0.053. It could sit between -0.126 and +0.019.

skills/engineering/code-review Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.575. With the skill: 0.492. Instruction only: 0.592. Change -0.083. It could sit between -0.189 and +0.022.

skills/code-review-web Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

No measured effect

No skill: 0.575. With the skill: 0.453. Instruction only: 0.592. Change -0.122. It could sit between -0.239 and -0.005.

skills/ce-code-review Gemini 3.1 Flash Lite Wider run, three ways, Deterministic

Scored worse

No skill: 0.575. With the skill: 0.000. Instruction only: 0.592. Change -0.575. It could sit between -0.749 and -0.401.

In Claude Code tested in Claude Code

skills/code-review-and-quality Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.935. With the skill: 0.955. Change +0.020. It could sit between -0.019 and +0.059.

plugins/developer-essentials/skills/code-review-excellence Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.915. With the skill: 0.935. Change +0.020. It could sit between -0.019 and +0.059.

skills/ce-code-review Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.935. With the skill: 0.935. Change +0.000. It could sit between +0.000 and +0.000.

skills/code-review-web Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.935. With the skill: 0.935. Change +0.000. It could sit between +0.000 and +0.000.

skills/code-reviewer Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.935. With the skill: 0.935. Change +0.000. It could sit between +0.000 and +0.000.

skills/code-reviewer Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.915. With the skill: 0.935. Change +0.020. It could sit between -0.019 and +0.059.

skills/engineering/code-review Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.918. With the skill: 0.935. Change +0.017. It could sit between -0.016 and +0.049.

skills/engineering/code-review Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.915. With the skill: 0.935. Change +0.020. It could sit between -0.019 and +0.059.

skills/gsd-code-review Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.935. With the skill: 0.935. Change +0.000. It could sit between +0.000 and +0.000.

skills/gsd-code-review Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.935. With the skill: 0.935. Change +0.000. It could sit between +0.000 and +0.000.

skills/ce-code-review Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.918. With the skill: 0.927. Change +0.008. It could sit between -0.054 and +0.070.

engineering/skills/pr-review-expert Claude Opus 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.935. With the skill: 0.918. Change -0.017. It could sit between -0.049 and +0.016.

skills/code-review-and-quality Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.918. With the skill: 0.918. Change +0.000. It could sit between +0.000 and +0.000.

skills/code-review-web Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.902. With the skill: 0.902. Change +0.000. It could sit between +0.000 and +0.000.

engineering/skills/pr-review-expert Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.915. With the skill: 0.898. Change -0.017. It could sit between -0.049 and +0.016.

plugins/developer-essentials/skills/code-review-excellence Claude Fable 5.1 Wider run, three ways, Deterministic

No measured effect

No skill: 0.918. With the skill: 0.862. Change -0.057. It could sit between -0.113 and +0.000.

skills/code-review-web Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.793. With the skill: 0.833. Change +0.040. It could sit between -0.012 and +0.092.

skills/code-reviewer Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.777. With the skill: 0.773. Change -0.003. It could sit between -0.057 and +0.050.

skills/gsd-code-review Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.760. With the skill: 0.763. Change +0.003. It could sit between -0.050 and +0.057.

plugins/developer-essentials/skills/code-review-excellence Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.777. With the skill: 0.760. Change -0.017. It could sit between -0.084 and +0.050.

skills/engineering/code-review Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.760. With the skill: 0.760. Change +0.000. It could sit between -0.049 and +0.049.

skills/code-review-and-quality Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.770. With the skill: 0.757. Change -0.013. It could sit between -0.093 and +0.067.

engineering/skills/pr-review-expert Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.723. With the skill: 0.740. Change +0.017. It could sit between -0.066 and +0.099.

skills/ce-code-review Claude Sonnet 5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.777. With the skill: 0.710. Change -0.067. It could sit between -0.160 and +0.026.

skills/gsd-code-review Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.535. With the skill: 0.567. Change +0.032. It could sit between -0.104 and +0.167.

plugins/developer-essentials/skills/code-review-excellence Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.482. With the skill: 0.553. Change +0.072. It could sit between -0.133 and +0.276.

skills/code-review-web Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.498. With the skill: 0.517. Change +0.018. It could sit between -0.169 and +0.206.

skills/engineering/code-review Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.495. With the skill: 0.507. Change +0.012. It could sit between -0.077 and +0.100.

skills/code-reviewer Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.498. With the skill: 0.503. Change +0.005. It could sit between -0.096 and +0.106.

skills/code-review-and-quality Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.490. With the skill: 0.490. Change -0.000. It could sit between -0.103 and +0.103.

skills/ce-code-review Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.578. With the skill: 0.437. Change -0.142. It could sit between -0.319 and +0.036.

engineering/skills/pr-review-expert Claude Haiku 4.5 Wider run, three ways, Deterministic

No measured effect

No skill: 0.488. With the skill: 0.437. Change -0.052. It could sit between -0.317 and +0.214.

deterministic pass rate, 0 to 1 Each row is one skill on one model. The grey bar is the score with no skill loaded. The coloured bar is the score with it. The bracket shows how much the change could move if we ran it again. It starts at the grey bar, so it covers where the coloured bar could have ended. The two test methods are measured on their own and never ranked against each other. Every figure drawn here is printed beside its row. A dashed upright marks the score with the instruction only. Rows without one are rows where that was not measured, and no mark stands in for it. A bracket wider than the scale is drawn to the edge with its cap left off. The table below gives its two ends.
SkillCohortWithout -> withEffect (range)OrbitInjected context tokensTested
No skill loadedbaseline64%0 by definition02026-09-05
One-line instruction insteadbaseline65%0 by definition02026-09-05
plugins/developer-essentials/skills/code-review-excellence from wshobson/agents 38e19c20d2b1 Best result in this testwshobson/agents64% -> 63%+0.01 [+0.01, +0.02]In free drift3,9382026-09-05
skills/code-reviewer from Jeffallan/claude-skills 882ef55e377dJeffallan/claude-skills63% -> 64%-0.00 [-0.05, +0.05]In free drift8,2902026-09-05
engineering/skills/pr-review-expert from alirezarezvani/claude-skills 19392f7a0826alirezarezvani/claude-skills63% -> 60%-0.05 [-0.08, -0.01]In free drift4,2602026-09-05
skills/code-review-and-quality from addyosmani/agent-skills d2c37ef6225daddyosmani/agent-skills64% -> 61%-0.05 [-0.05, -0.04]In free drift5,2182026-09-05
skills/gsd-code-review from open-gsd/gsd-core 6beaa66b2587open-gsd/gsd-core63% -> 59%-0.05 [-0.07, -0.03]In free drift1,4202026-09-05
skills/code-review-web from rampstackco/claude-skills a67dd34c609frampstackco/claude-skills*64% -> 58%-0.07 [-0.14, -0.01]In free drift8,2062026-09-05
skills/engineering/code-review from mattpocock/skills 6654f6b60cd9mattpocock/skills63% -> 57%-0.12 [-0.14, -0.10]In free drift2,2782026-09-05
skills/ce-code-review from EveryInc/compound-engineering-plugin c9c10f8c7541EveryInc/compound-engineering-plugin64% -> 34%-0.30 [-0.59, -0.02]Mixed, see page73,7152026-09-05
Averaged over three models; per-model results on each skill's page. Ranked within this class only.

Ranked within this topic by how many models the rule found a clear effect on, which is a count of verdicts and never the size of one. Two classes are two answer keys on two scales, so no effect here is compared with another. Every class, with its full table.

* rampstackco/claude-skills is maintained by the operator of OpenAddict.com.