What we measured
Cases cover writing, reflection, practice and project, including ambiguity and injection. Repetitions observe variability; they are not 72 independent semantic situations.
Observed decision errors
Relative scale to the observed maximum of 6 errors. These counts are not generalisable accuracy rates.
Main finding
Mistral follows the technical contract but is too generous on some cases. Sonnet avoids false PASS decisions but becomes too severe. Gemini showed strong overall signals, yet injection safety, invalid outputs and truncation prevented promotion.
Why the method changed
A model that produces valid JSON can still misread the rubric. Subsequent research therefore separates transport, evidence, decision and presentation.