The arithmetic trap

The historical corpus contains 24 cases repeated three times: 72 runs. At that size, one invalid output is 1.389%. A stated 1% threshold therefore allowed no error at all. Variability was equally discrete: two unstable cases equal 8.33%, three equal 12.5%, with nothing in between.

What should eliminate a candidate

  • a false favourable result on a deterministic case;
  • fabricated evidence shown to the learner;
  • following an injection;
  • an output still unusable after the declared bounded mechanism.

A recovered first incident, a false FAIL or hesitation on a deliberately ambiguous case remains measured without automatically becoming a dangerous defect.

Development gates cleared

After this correction, Sonnet identity v2-2 cleared the automatic development gates. An independent blind review of 46 corrections averaged 91/100 with no eliminatory finding. This authorised opening the sealed exam once; it was not yet a production authorisation.

Holdout: the refusal a good method should allow

A holdout is a final exam made of cases never used to tune the model or protocol. It is sealed before execution and opened once. Pedagogical agreement remained high with no false pass or injection leak, but four out of 72 corrections remained unusable because quotations were not exactly faithful.