Submission
The full answer and versioned rubric are frozen.
LearnX · V4 research dossier
A documented account of model trials, observed limitations, and the move towards formative scoring verified by a composite pipeline.
Models can differ by one adjacent level while still producing useful feedback. LearnX must remain strict about security and evidence, without turning a grading nuance into a definitive academic verdict.
An indicative score, explained criterion by criterion, is more honest than an artificially certain “pass”.
Each stage answers a different question. Passing an initial test only proves that the model can communicate with the system; it is never enough to select it.
| Term | Plain-language definition | What it demonstrates |
|---|---|---|
| Smoke test | A very short test using one representative case. It checks that the provider route, model and response format work together. | The model may proceed to the next test. Its pedagogical quality is not yet demonstrated. |
| Smoke sequence | Three sequential cases: a clearly successful answer, an ambiguous answer and a prompt-injection attempt. Testing stops at the first material failure. | The model appears technically compatible and safe enough for a small panel. |
| Mini-panel | A small predefined set, usually six varied answers, reviewed blindly by a human evaluator. | It quickly exposes grading, tone or rubric errors before an expensive campaign. |
| Full panel | The complete campaign across the development corpus. Here, it covers all 24 calibrated answers. | It measures all planned task families, but not yet generalisation to unseen cases. |
| 24×3 | The same 24 answers are assessed three times, producing 72 logical runs. Repetition does not inflate the sample; it measures stability. | It reveals both average errors and changes in judgement on an identical answer. |
| Sealed holdout | A separate corpus never used to tune the prompt or select thresholds. It remains closed until the configuration is frozen. | It tests whether the solution generalises without having learned the development answers. |
Example. Two successful smokes do not equal a full panel. A successful full panel is still not production approval: blind human review, the sealed holdout and a limited pilot must follow.
Mistral is reliable and inexpensive but over-rated three answers. Sonnet is more conservative, with no observed false positive, but rejects more acceptable answers and experienced more format incidents.
Comparison limit. The cascade is a retrospective simulation combining campaigns that did not use exactly the same prompt and protocol. It informs architecture; it is not a promotion result or a final pricing basis.
Not every model reached the same evaluation depth. Full campaigns, panels, technical smokes and catalogue-only screening must not be treated as equivalent evidence.
| Model | Evidence level | Observed result | Lesson |
|---|---|---|---|
| Mistral Medium 3.5 | Several panels and 24×3 full runs | 72/72 usable; 3 false positives and 5 false negatives in campaign 1.6 | Reliable and economical; candidate primary, not sufficient alone for sensitive results. |
| Claude Sonnet 4.6 | Panels and protocol 3.0.1 full | 0 false positives, 6 false negatives, 1 finally unusable run | Conservative and strong on evidence; candidate verifier, not an absolute authority. |
| Gemini 3.6 Flash | Historical full then sequential retest | 90.4% criterion agreement, but 50% injection safety and 8.33% invalid outputs; later ambiguous case truncated | Promising general quality, insufficient safeguards under the tested protocol. |
| GPT-5.6 Terra | Technical smokes | Pinned structured OpenRouter route unavailable | No pedagogical verdict; a direct-API campaign would be a separate experiment. |
| GPT-5.6 Sol | Sequential smokes | Successful case valid; ambiguous case timed out or truncated | Capable, but the current profile is too slow or verbose for the bounded contract. |
| Claude Opus 4.8 | Non-comparable historical full + 2 current smokes | Two recent smokes valid; older data used a previous corpus and cost more | Possible qualitative reference, insufficient current evidence for promotion. |
| Kimi K2.5 | First smoke | Timeout before usable output | Stopped early under the sequential budget gate. |
| Command A | First smoke | Provider 400 before usable output | Technical compatibility not demonstrated. |
| Qwen 3.7 Max | Exploratory smokes | Invalid capped output, followed by timeout | Stopped before a panel due to cost/latency without enough pedagogical evidence. |
| DeepSeek V4 Flash | Catalogue only | Structured routes identified; no paid benchmark run | Reserve option only; no performance claim. |
Correct interpretation. A transport or format failure does not prove that a model is pedagogically poor. It only shows that it did not reliably execute the tested LearnX contract. Conversely, two successful smokes do not demonstrate production quality.
The second model is neither a vote nor a marketing claim. It is invoked only for sensitive first-pass results, while LearnX retains authority over calculation and presentation.
The full answer and versioned rubric are frozen.
Mistral proposes levels, evidence and feedback.
LearnX computes the indicative score and identifies sensitive cases.
Sonnet is invoked by a deterministic rule, not systematically.
Score, formative band or uncertain state, never a progression lock.
Initial experimental rule: verify primary scores from 60/100, sensitive criteria, security signals, and a random sample outside the trigger. A difference above 15 points or two levels on one criterion must not be silently averaged.
An invalid output is an unusable technical model response, not a mistake made by the learner. The new criteria separate a recovered first incident from a failure actually visible to the user.
The objective is no longer to find a perfect model. It is to demonstrate that the complete workflow is useful, honest, safe and economically predictable.
Define levels, formative bands, score presentation, disagreement language and the absence of any progression lock.
Pin models, providers, profiles, prompts, trigger rule, control sample, retry policy and gap resolution.
Compare the pipeline against baselines under the same protocol, then blindly review major gaps and a sample of agreements.
The French holdout must not become a tuning set or be replayed until the outcome looks favourable.
Reserve at a prudent workflow ceiling, charge actual aggregate cost, release the difference and present one understandable user action.