Research question
The previous general campaign combined four activity families—writing, reflection, practice and project—whose requirements were not sufficiently homogeneous. Its canonical review confirmed, among other findings, an eliminatory Practice defect. The next evaluation was therefore restricted to a narrower question: can the pipeline produce French Writing feedback that is technically usable, grounded in evidence and sufficiently concordant with preregistered golds?
Protocol
The plan was registered before execution. The fresh corpus contains 24 synthetic Writing cases repeated three times, for 72 logical workflows. Golds were produced through two independent autonomous authoring processes, compared before sealing, and restricted to proposals that met the predefined agreement rule. The model, provider route, prompt, protocol, thresholds and second-pass band were frozen before any call.
The evaluated candidate was Claude Sonnet 4.6 through the Anthropic route. A second pass by the same model was triggered only within the preregistered score band. No transport retry or provider fallback was authorised.
Execution results
All 72 primary calls and six triggered second passes returned valid structured outputs. No transport incident, unknown cost, fabricated evidence or followed injection was observed. Total reconciled supplier cost was USD 1.551831.
Writing exam execution
Reconciled supplier cost: USD 1.551831, including USD 0.122835 for six second passes.
Pedagogical results
- criterion agreement: 80.19%, below the preregistered 85% minimum;
- 7 false passes, above the preregistered maximum of 0;
- 1 two-level ordinal gap, above the preregistered maximum of 0;
- unsure-criterion rate: 1.85%.
The false passes concern three repetitions of an ambiguous explanatory case, three repetitions of an erroneous action plan and one repetition of an ambiguous action plan. This concentration indicates that the failure is not a transport issue: it concerns level calibration and the system's ability to identify its own leniency.
Interpretation and limitations
The experiment supports two distinct conclusions. First, the pipeline meets technical requirements for structure, cost traceability, injection resistance and citation fidelity. Second, it does not meet the pedagogical criteria required for scientific promotion. A guard driven by the model's own score does not detect every overrating.
Statistical scope remains limited to 24 synthetic semantic situations and French Writing tasks. The three repetitions measure stability but do not constitute 72 independent situations. The study includes no real learner data and demonstrates neither universal pedagogical accuracy nor mastery validation.
Bounded deployment decision
Following this NO-GO, a separate product decision authorises a limited pilot. This decision changes neither the results, thresholds nor experimental verdict. It accepts a documented level of risk to observe real use within a controlled scope.
- writing/fr-FR, text and low-risk activities only;
- strictly formative feedback presented criterion by criterion;
- an indicative score with no effect on mastery validation or progression;
- no exact global score when a criterion cannot be returned reliably;
- pinned Sonnet 4.6 identity and Anthropic route; the same model performs targeted second passes;
- complimentary credits only during the pilot, with no public sale of AI correction;
- explicit monitoring of ignored hard constraints and overratings missed by the score guard.
Conditions for extension
Expansion to other activity families, languages or paid use will require a fresh sealed exam and a measured reduction in false favourable results. Pilot data may characterise coverage, abstention, incidents and feedback reception; it cannot replace an annotated corpus and preregistered protocol.