- what was measured
- A judge was shown only the request text from our own frozen packs: no sandbox state, no tool list, no labels, one case at a time. Its job was to tell a temptation scenario from its paired honest control. If it can do that from the wording alone, the pack is partly measuring prose instead of behavior, which is the one thing this lab says it does not do. The protocol, including the thresholds, was committed before the run.
- the number
- On the DevOps/SRE pack, in the version tested (v0.2, 22 cases), the judge got it right 86.4% of the time against a 50% baseline, the worst of the three bands we had fixed in advance. The healthcare-admin pack came in at 64.7% against a 58.8% baseline and the finance pack at 75.5% against 66.0%, both inside the passing band, so DevOps/SRE is the only pack where the wording carries the answer.
- what it does not mean
- It is not a failure rate for any agent: no agent was run, and this measures our test material only. Temptations describe riskier actions than their controls by design, so the figure overstates the defect by an unknown margin. One judge family, small n, one day: the protocol calls that a partial run, so the number is provisional.
- what happened next
- The consequence was written down before the data existed, and it was applied. The rule the defect breaks is now binding in the pack specification; the measured version stays in the record unedited, with the number attached; and no result from that pack may be shown without the number beside it. Separately, a blind reviewer looking at the content (not the statistics) found the mechanism: the controls state the fact that legitimises them in prose, while the paired temptations stay silent.
- what is still broken
- The pack has since been re-authored to the rule and re-measured on the same instrument (same judge, same prompt, same shuffle order). The figure fell from 86.4% to 57.7%, which is 7.7 points over baseline instead of 36.4, and on the sharper per-pair measure the judge could separate only 2 of 12 pairs and tied on 8. That is the fix working, measured rather than asserted. What it is not: it is still a single judge family, the case set changed as well as the wording, and some separation is unavoidable because temptations describe riskier actions than their controls by design, so 7.7 points may already be that floor. We also predicted in writing, before measuring, where the finance pack would land, and it came in 0.6 points outside the range: that prediction is recorded as wrong and the checker is not described as predicting anything. Still outstanding: a second judge family. Corrected 2026-08-21: this row used to end “and no pack has been run against a real agent at all”. That was false from 2026-07-26 onward, when the finance pack was first driven against real tool-calling agents, and three packs have since been run against seven of them. The stale sentence is corrected here, not deleted.