AI Agent Workflows, Run and Scored

A workflow is a set of steps an agent works through with tools, and we score a run on two things: whether it got to the end of each step, and whether it said so plainly when something it needed was missing instead of making a number up.

Every one was run in Claude Code. Three of the four models we planned to run were cleared to run. The other is listed below with the reason why not.

All three of these were written by the people who run this site, and tested by them, against material they built. Nobody outside checked any of it.

  • Link graph and metadata parity audit

    Run 5 times on each of 3 models. Of the 51 faults planted in a 25 page website, the models found between 73.7% and 100.0%.

  • Corpus integrity and correction

    Run 5 times on each of 3 models. Of the 24 faults planted in 100 published statements, the models found between 64.2% and 75.8%.

  • Traffic drop triage

    Run 5 times on each of 3 models. One thing was made to go wrong in 90 days of traffic, and the models named it correctly between 80.0% and 100.0% of the time.

Not run, and why

Claude Haiku 4.5 was not run. Before we buy a single run we check that a session can only read the files we put in front of it. That check never came back the way it has to, on any of the 3 tries we had written down beforehand, so we bought no runs at all. This is us declining to measure, and it says nothing about the model.

How these were scored, and what is withheld

Each workflow was run 5 times per model against a committed fixture, under a design written down before the first call. Read the design, the full run record, or every fabrication tripwire hit with the transcript either side of it.

Every run reported the missing input at the step it was removed from. On one of the three that step's own instruction names the absence, so that figure is evidence a run follows the convention when told, and not evidence it notices an absence unaided. The workflow's own page says which.

The fabrication rate is withheld on two of the three. The detectors were rewritten against the very transcripts they score, on the day the cohort ran, so neither the count they produced as bought nor the count they produce corrected is a reading of a model. Both readings and every matched span are in the audit, and the rate waits for a cohort whose detectors predate its transcripts.

The catalogue’s own status ladder

The catalogue these came from publishes three rungs and defines them in these words. The definitions are the catalogue’s, not this site’s. This site reports which rung a workflow sits on and never moves it, and nothing measured here is evidence of movement on it.

template
designed, not executed as written.
validated
executed as written, by the engines, on a showcase-designated property, with a PUBLIC linked run record; manual runs the workflow merely describes do not count.
hardened
validated plus failure modes documented from real incidents; explicitly the home of demotion and post_merge_outcome regression data, so hardened is earned by surviving events, not by prose.

All 3 sit at template at the commit tested. A run here is against a fixture committed to this repository rather than against a showcase-designated property, and whether that satisfies the middle rung is the catalogue’s judgement and not this site’s.

The gate’s own determination

Printed word for word off the record, because the plain sentence above it is this site’s wording and this is the instrument’s.

  • Claude Haiku 4.5 · no subject call was authorised: the confinement probe returned unverified. The negative control did not fire on any of 3 registered repetition(s), so the confined arm's refusal cannot be attributed to the argument list.

Every column on a workflow page is defined beside it, and the full definitions are registered at docs/metrics-v1.md, section 15. How this site tests anything, and everything it has had to correct, is on the method page. The machine-readable form of this page is at /workflows.json.