Research & methodology
Research programme on AI-assisted formative feedback
LearnX protocols, experimental results and methodological decisions, published by stage and preserved in their original context.
Last updated 24 August 2026. Campaigns use synthetic cases; no real learner data or human validation is claimed.
Sealed-protocol Writing evaluation: results and bounded deployment decision
Across 72 workflows, execution is technically complete but pedagogical gates fail: 80.19% criterion agreement, seven false passes and one two-level ordinal gap.
Read article →Fixing the gates, then accepting the holdout refusal
Sample arithmetic had turned 1% into zero tolerance. Once corrected, Sonnet cleared development and blind review, then honestly failed the sealed exam.
Read article →What worked and what still blocked us
The Gemini probe isolated a transport and reconciliation incident. The method progressed, but no pedagogical verdict was possible yet.
Read article →From LLM judge to executable rubric engine
Models find or challenge exact evidence; LearnX then applies versioned rules and can abstain honestly.
Read article →Why two models do not manufacture truth
The Mistral + Sonnet cascade reduces some technical risks but still creates false FAILs and abstentions. Disagreement becomes a signal, not an arbiter.
Read article →Comparing models without confusing compliance and pedagogy
The 24 × 3 benchmark exposes Mistral, Sonnet and Gemini strengths, along with false verdicts and the limits of asking a model to choose a level directly.
Read article →Read the full technical report
The dossier preserves tables, costs, gates, explored models, protocol changes and limitations at maximum detail.
Open dossier →