Research & methodology

Research programme on AI-assisted formative feedback

LearnX protocols, experimental results and methodological decisions, published by stage and preserved in their original context.

Last updated 24 August 2026. Campaigns use synthetic cases; no real learner data or human validation is claimed.

NO-GOscientific Writing verdict preserved
1published writing/fr-FR contract
Bounded pilotauthorised with complimentary credits

Sealed-protocol Writing evaluation: results and bounded deployment decision

Across 72 workflows, execution is technically complete but pedagogical gates fail: 80.19% criterion agreement, seven false passes and one two-level ordinal gap.

Read article →

Fixing the gates, then accepting the holdout refusal

Sample arithmetic had turned 1% into zero tolerance. Once corrected, Sonnet cleared development and blind review, then honestly failed the sealed exam.

Read article →

What worked and what still blocked us

The Gemini probe isolated a transport and reconciliation incident. The method progressed, but no pedagogical verdict was possible yet.

Read article →

From LLM judge to executable rubric engine

Models find or challenge exact evidence; LearnX then applies versioned rules and can abstain honestly.

Read article →

Why two models do not manufacture truth

The Mistral + Sonnet cascade reduces some technical risks but still creates false FAILs and abstentions. Disagreement becomes a signal, not an arbiter.

Read article →

Comparing models without confusing compliance and pedagogy

The 24 × 3 benchmark exposes Mistral, Sonnet and Gemini strengths, along with false verdicts and the limits of asking a model to choose a level directly.

Read article →

Read the full technical report

The dossier preserves tables, costs, gates, explored models, protocol changes and limitations at maximum detail.

Open dossier →