What we measured

Cases cover writing, reflection, practice and project, including ambiguity and injection. Repetitions observe variability; they are not 72 independent semantic situations.

Main finding

Mistral follows the technical contract but is too generous on some cases. Sonnet avoids false PASS decisions but becomes too severe. Gemini showed strong overall signals, yet injection safety, invalid outputs and truncation prevented promotion.

Why the method changed

A model that produces valid JSON can still misread the rubric. Subsequent research therefore separates transport, evidence, decision and presentation.