skills/code-review-web from rampstackco/claude-skills

Code review tasks, run with the skill loaded and without it, on the same task pairs.

Identity pin

Repository
rampstackco/claude-skills
Path
skills/code-review-web
Commit
a67dd34c609f034c0cfd736a348659bbdf1605bf
Content hash
a27bba1351840ced78f98f02c94be30140cb159d940f866455ef0ca60c6183d5
Date tested
not yet tested

The pin is the whole of this skill's identity here. It resolves at https://github.com/rampstackco/claude-skills/tree/a67dd34c609f034c0cfd736a348659bbdf1605bf/skills/code-review-web, and the content hash is a sha256 over exactly the text the with arm was given. Nothing else about the skill appears on this site.

What was claimed

  • Loading skills/code-review-web from rampstackco/claude-skills improves outputs on code review tasks.
    Pass criterion: With-arm mean coverage exceeds the without-arm by at least 0.10 on the same 10 fixture files at the same model and settings. Scale: unit. Defects found over defects seeded, so a [0,1] fraction, which is the unit scale. The same scorer, the same scoring mode and the same scale as the accessibility and on-page audit classes.
    Ten fixture source files across TypeScript, JavaScript, Python, SQL, a Next.js Route Handler and a Supabase client module, each seeded with known defects drawn from a closed list of codes that is given to BOTH arms in the prompt, so the code list is not the manipulated variable. The ledger is at fixtures/code-review/defect-ledger.json and carries a predicate per entry, evaluated against the fixture bytes by the answer key test; entries no predicate decides are marked judgement and listed as unverified. The model is asked for the SYMBOL a defect sits on and never for a line number. False positives are counted on the record and never netted off coverage. The external slot for this class is unfilled and is recorded as unfilled in skill-cohort.json, so the class publishes one row rather than a pair until that selection is made.

Injected context tokens

Injected context tokens, per arm
ArmContext charactersInjected context tokens
without00
with31,1607,315

The character count is exact: it is the length of the text the with arm is given, and the content hash above is a sha256 over that same text. The token figure is an estimate at 4.26 characters per token, the ratio the phase 1 run measured over 3,120 calls, and it is labelled an estimate until a run reports its own token counts. The without arm is given the identical prompt and nothing else, so its zero is a measurement rather than a missing value.

Not yet measured

No run records exist for this skill. The instrument is built and committed, the task sets are frozen, and the projected cost of the run is published, but no model has been called. There is therefore no verdict, no effect, no interval and no cost per task on this page, and none is shown.

  • S11-code-review-web: 10 task pairs per model, as planned. Deterministic, scored against a committed answer key.

The committed instrument is at harness/claims/claims-skills.json, harness/claims/tasksets-skills.json and harness/claims/skill-cohort.json. The projection for the run that has not happened is at harness/results/dry-run-projection-skills.md. How a claim gets tested.

720 run records behind this page. Every verdict is computed at build time by the same frozen status_v1 rule that decides every other verdict on this site, and nothing here is written by hand.