Does PR review expert help? Tested on code review tasks

PR review expert, from alirezarezvani/claude-skills. Its code review tasks were run with the skill loaded and without it. Each time it is the same task, run twice.

What we tested

Whether loading this skill helps a model review code.

What counts as helping

The skill has to find at least 10 percentage points more of the planted bugs than the model finds without it, on the same 10 source files.

How we scored it

We planted known bugs across the 10 source files and counted how many the model named. The score is the share it found, from 0 to 1.

Without to with, per model

Via API tested via API

GPT-5 mini

0.707
0.628

No measured effect

Gemini 3.1 Flash Lite

0.592
0.578

No measured effect

deterministic pass rate, 0 to 1

In Claude Code tested in Claude Code

Claude Opus 5 Wider run, three ways

0.935
0.918

No measured effect

Claude Fable 5.1 Wider run, three ways

0.915
0.898

No measured effect

Claude Sonnet 5 Wider run, three ways

0.723
0.740

No measured effect

Claude Haiku 4.5 Wider run, three ways

0.488
0.437

No measured effect

deterministic pass rate, 0 to 1

A picture of the per-model numbers, drawn from the same results. The tables are the source. Grey is the score without. Colour is the score with. The bracket shows how much the difference could move if we ran it again. Rows are ordered by the score with. A model measured at two versions keeps its versions next to each other. The two test methods are reported separately and never averaged.
Show per-task detail

Every item, without and with, per model

Via API tested via API

  • Gemini 3.1 Flash Lite
  • GPT-5 mini

deterministic pass rate, 0 to 1

Hover or focus an item to read its task and every model’s two values.

Task-paired: these item means are what this cell's interval is built from.

The item means behind this plot
Per-item arm means. Without is the unaided arm, with is the treated arm.
ItemGemini 3.1 Flash Lite GPT-5 mini
Task without with without with
c10 0.250 0.250 0.500 0.750
c07 0.400 0.600 0.400 0.200
c04 0.400 0.200 0.600 0.600
c05 0.500 0.333 0.500 0.333
c02 0.400 0.600 0.800 0.600
c01 0.800 0.800 0.600 0.800
c06 0.667 0.500 0.833 0.500
c08 0.667 0.667 0.833 0.667
c03 0.833 0.833 1.000 0.833
c09 1.000 1.000 1.000 1.000
A picture of the per-item means behind the per-model numbers. The tables are the source. Each model draws two lines over the same task set: a dashed line through its unaided scores and a solid line through its treated ones. Items are ordered by the mean unaided score across the models that measured them, lowest first. The two instruments are reported separately and never averaged.

Without to with, per model

In Claude Code tested in Claude Code

Claude Haiku 4.5 Expansion cohort (three arms)

No measured effect

Claude Sonnet 5 Expansion cohort (three arms)

No measured effect

Claude Fable 5.1 Expansion cohort (three arms)

No measured effect

Claude Opus 5 Expansion cohort (three arms)

No measured effect

deterministic pass rate, 0 to 1

A picture of the per-model numbers, drawn from the same results. The tables are the source. Each row runs from the score without to the score with. The thin bar beneath it is the bracket, and it shows how much the difference could move if we ran it again. The two test methods are reported separately and never averaged.

The Claude Code runs committed their cells as arm means and intervals rather than as per-item values, so this panel is drawn without to with rather than item by item.

What exactly was tested, and how it was scored
Repository
alirezarezvani/claude-skills
Path
engineering/skills/pr-review-expert
Commit
19392f7a08264ed00486a251f5b2098321771f94
Content hash
0b3ac983ef14cd4636638217ca86d6eb571bf6757eff5d34ae3aa009c53f3be3
Date tested
not yet tested

The pin is the whole of this skill's identity here. It resolves at https://github.com/alirezarezvani/claude-skills/tree/19392f7a08264ed00486a251f5b2098321771f94/engineering/skills/pr-review-expert, and the content hash is a sha256 over exactly the text the model was given with the skill loaded. Nothing else about the skill appears on this site.

S23-alirezarezvani-pr-review-expert

Loading engineering/skills/pr-review-expert from alirezarezvani/claude-skills improves outputs on code review tasks.

Pass criterion
With-arm mean coverage exceeds the without-arm by at least 0.10 on the same 10 fixture files at the same model and settings.
Scale
unit. Same scorer, same instrument and same scale as S11-code-review-web. Added 2026-08-31 by the source expansion under R23; adding a row to an instrument does not touch the instrument.
Task pairs planned per model
10
Notes
Shares the code-review task set with every other claim in this class, so all of its rows are measured on identical items against one answer key. The slot this claim scores was pinned by the 2026-08-31 source expansion under R23 and NO MODEL HAS BEEN RUN AGAINST IT. The claim exists so the class holds one claim per pinned slot, which is what makes the row comparable the day it is run.

Injected context tokens

Injected context tokens, per arm
ArmContext charactersInjected context tokens
without00
with12,5842,954

The character count is exact: it is the length of the text the with arm is given, and the content hash above is a sha256 over that same text. The token figure is an estimate at 4.26 characters per token, the ratio the phase 1 run measured over 3,120 calls, and it is labelled an estimate until a run reports its own token counts. The without arm is given the identical prompt and nothing else, so its zero is a measurement rather than a missing value.

Verdict per model

Effect per model
ModelReadingOrbitWithout -> withEffect95 percent intervalTask pairsModel version returned
Gemini 3.1 Flash LiteDeterministic, scored against a committed answer keyIn free drift unclear. The model scored 57% without it.0.5917 -> 0.5783-0.0133[-0.0995, 0.0728]10 pairs, 30 of 30 recordsgemini-3.1-flash-lite
GPT-5 miniDeterministic, scored against a committed answer keyIn free drift unclear. The model scored 69% without it.0.7067 -> 0.6283-0.0783[-0.1944, 0.0377]10 pairs, 30 of 30 recordsgpt-5-mini-2025-08-07
  • In free drift: no separation the design can resolve. Not evidence of no effect.

Cost

Defined in the metrics canon. Cost per task divides every dollar spent on an arm by the tasks attempted on it, including tasks whose call returned nothing, because a call that returned nothing was still billed.

Where the with arm did not win

Gemini 3.1 Flash Lite, Deterministic, scored against a committed answer key

Lost on 3 of 10 pairs: c04 (-0.2000), c05 (-0.1667), c06 (-0.1667).

Drew on 5 of 10 pairs: c01, c03, c08, c09, c10.

Findings not in the answer key: 7 with the skill, 7 without. These are counted and never netted off coverage.

GPT-5 mini, Deterministic, scored against a committed answer key

Lost on 6 of 10 pairs: c06 (-0.3333), c02 (-0.2000), c07 (-0.2000), c08 (-0.1667), c05 (-0.1667), c03 (-0.1667).

Drew on 2 of 10 pairs: c04, c09.

Findings not in the answer key: 32 with the skill, 49 without. These are counted and never netted off coverage.

In Claude Code

These cells are tested in Claude Code, on a subscription path with no API key. They are a second instrument: no figure here is averaged with one tested via API above. How the two were compared.

  • Claude Fable 5.1 · unclear in the three-arm cohort
  • Claude Haiku 4.5 · unclear in the three-arm cohort
  • Claude Opus 5 · unclear in the three-arm cohort
  • Claude Sonnet 5 · unclear in the three-arm cohort

Expansion cohort (three arms)

one pass over each committed item. Matrix hash 0078cfa8. Pre-registration · Run report.

One row per model. Every figure is read from this run’s committed cells.
ModelNo skillWith skillDeltaIntervalPairsOrbitInjected tokensFalse positives per page
Claude Fable 5.10.91500.8983-0.017-0.049 to 0.01610In free driftnot measured: this run records the dose per arm, in the doses list12 with, 9 without
Claude Haiku 4.50.48830.4367-0.052-0.317 to 0.21410In free driftnot measured: this run records the dose per arm, in the doses list15 with, 9 without
Claude Opus 50.93500.9183-0.017-0.049 to 0.01610In free driftnot measured: this run records the dose per arm, in the doses list21 with, 26 without
Claude Sonnet 50.72330.7400+0.017-0.066 to 0.09910In free driftnot measured: this run records the dose per arm, in the doses list10 with, 10 without

Every figure above is computed at build time from harness/results/runs-skills.jsonl and harness/results/ledger-skills.jsonl, both committed, by the frozen status_v1 rule and the metrics_v1 cost definitions. How a claim gets tested.

720 run records behind this page. Every verdict is computed at build time by the same frozen status_v1 rule that decides every other verdict on this site, and nothing here is written by hand.