LearnX · V4 research dossier

Building useful assessment without manufacturing certainty.

A documented account of model trials, observed limitations, and the move towards formative scoring verified by a composite pipeline.

Status Exploration — no pipeline promotedData 24 synthetic French cases × 3 repetitionsUpdated 12 August 2026Version française
01 · Current decision

Move from a binary verdict to formative assessment.

Models can differ by one adjacent level while still producing useful feedback. LearnX must remain strict about security and evidence, without turning a grading nuance into a definitive academic verdict.

An indicative score, explained criterion by criterion, is more honest than an artificially certain “pass”.
02 · Understanding the method

Smoke, mini-panel, full and 24×3: what do they mean?

Each stage answers a different question. Passing an initial test only proves that the model can communicate with the system; it is never enough to select it.

Definitions of benchmark stages
TermPlain-language definitionWhat it demonstrates
Smoke testA very short test using one representative case. It checks that the provider route, model and response format work together.The model may proceed to the next test. Its pedagogical quality is not yet demonstrated.
Smoke sequenceThree sequential cases: a clearly successful answer, an ambiguous answer and a prompt-injection attempt. Testing stops at the first material failure.The model appears technically compatible and safe enough for a small panel.
Mini-panelA small predefined set, usually six varied answers, reviewed blindly by a human evaluator.It quickly exposes grading, tone or rubric errors before an expensive campaign.
Full panelThe complete campaign across the development corpus. Here, it covers all 24 calibrated answers.It measures all planned task families, but not yet generalisation to unseen cases.
24×3The same 24 answers are assessed three times, producing 72 logical runs. Repetition does not inflate the sample; it measures stability.It reveals both average errors and changes in judgement on an identical answer.
Sealed holdoutA separate corpus never used to tune the prompt or select thresholds. It remains closed until the configuration is frozen.It tests whether the solution generalises without having learned the development answers.

Example. Two successful smokes do not equal a full panel. A successful full panel is still not production approval: blind human review, the sealed holdout and a limited pilot must follow.

03 · Observed results

Two complementary profiles, no standalone winner.

Mistral is reliable and inexpensive but over-rated three answers. Sonnet is more conservative, with no observed false positive, but rejects more acceptable answers and experienced more format incidents.

72logical runs per full campaign
0observed injection leaks in retained campaigns
3observed Mistral-only false positives
0observed Sonnet-only false positives
Observed quality and stabilityPercentage — historical full campaigns
Mistral · criteria
87.5%
Sonnet · criteria
87.8%
Mistral · decision
88.9%
Sonnet · decision
91.5%
M. variability
12.5%
S. variability
12.5%
Average loaded cost + FX bufferUSD equivalent per assessment — historically recombined data
Mistral only
0.0077
Targeted cascade
0.0161
Sonnet only
0.0239

Comparison limit. The cascade is a retrospective simulation combining campaigns that did not use exactly the same prompt and protocol. It informs architecture; it is not a promotion result or a final pricing basis.

04 · Models explored

The search extended beyond the selected pair.

Not every model reached the same evaluation depth. Full campaigns, panels, technical smokes and catalogue-only screening must not be treated as equivalent evidence.

Models explored during LearnX research
ModelEvidence levelObserved resultLesson
Mistral Medium 3.5Several panels and 24×3 full runs72/72 usable; 3 false positives and 5 false negatives in campaign 1.6Reliable and economical; candidate primary, not sufficient alone for sensitive results.
Claude Sonnet 4.6Panels and protocol 3.0.1 full0 false positives, 6 false negatives, 1 finally unusable runConservative and strong on evidence; candidate verifier, not an absolute authority.
Gemini 3.6 FlashHistorical full then sequential retest90.4% criterion agreement, but 50% injection safety and 8.33% invalid outputs; later ambiguous case truncatedPromising general quality, insufficient safeguards under the tested protocol.
GPT-5.6 TerraTechnical smokesPinned structured OpenRouter route unavailableNo pedagogical verdict; a direct-API campaign would be a separate experiment.
GPT-5.6 SolSequential smokesSuccessful case valid; ambiguous case timed out or truncatedCapable, but the current profile is too slow or verbose for the bounded contract.
Claude Opus 4.8Non-comparable historical full + 2 current smokesTwo recent smokes valid; older data used a previous corpus and cost morePossible qualitative reference, insufficient current evidence for promotion.
Kimi K2.5First smokeTimeout before usable outputStopped early under the sequential budget gate.
Command AFirst smokeProvider 400 before usable outputTechnical compatibility not demonstrated.
Qwen 3.7 MaxExploratory smokesInvalid capped output, followed by timeoutStopped before a panel due to cost/latency without enough pedagogical evidence.
DeepSeek V4 FlashCatalogue onlyStructured routes identified; no paid benchmark runReserve option only; no performance claim.

Correct interpretation. A transport or format failure does not prove that a model is pedagogically poor. It only shows that it did not reliably execute the tested LearnX contract. Conversely, two successful smokes do not demonstrate production quality.

05 · Pipeline under consideration

Verify where uncertainty matters.

The second model is neither a vote nor a marketing claim. It is invoked only for sensitive first-pass results, while LearnX retains authority over calculation and presentation.

Step 1

Submission

The full answer and versioned rubric are frozen.

Step 2

Assessment

Mistral proposes levels, evidence and feedback.

Step 3

Calculation

LearnX computes the indicative score and identifies sensitive cases.

Step 4

Verification

Sonnet is invoked by a deterministic rule, not systematically.

Step 5

Delivery

Score, formative band or uncertain state, never a progression lock.

Initial experimental rule: verify primary scores from 60/100, sensitive criteria, security signals, and a random sample outside the trigger. A difference above 15 points or two levels on one criterion must not be silently averaged.

06 · Success criteria

Strict on risk, realistic about recovered incidents.

An invalid output is an unusable technical model response, not a mistake made by the learner. The new criteria separate a recovered first incident from a failure actually visible to the user.

Hard gates

  • No observed injection leakage.
  • No fabricated evidence shown to the learner.
  • No invalid output displayed or charged.
  • Final failure after retry ≤2% in beta.
  • Score and formative band computed server-side.
  • No effect on progression.

Monitored indicators

  • First invalid output: target ≤10%.
  • Exact criterion agreement: target ≥85%.
  • Score error and repetition stability.
  • Two-level gaps and cross-criterion contamination.
  • Feedback usefulness, tone and actionability.
  • Full-workflow P50/P90 cost and latency.
07 · Evidence still required

What remains to be demonstrated before beta.

The objective is no longer to find a perfect model. It is to demonstrate that the complete workflow is useful, honest, safe and economically predictable.

1. Freeze the formative scoring contract

Define levels, formative bands, score presentation, disagreement language and the absence of any progression lock.

2. Give the composite pipeline its own identity

Pin models, providers, profiles, prompts, trigger rule, control sample, retry policy and gap resolution.

3. Run a capped composite test

Compare the pipeline against baselines under the same protocol, then blindly review major gaps and a sample of agreements.

4. Keep the holdout sealed until the development gate passes

The French holdout must not become a tuning set or be replayed until the outcome looks favourable.

5. Calibrate price and experience

Reserve at a prudent workflow ceiling, charge actual aggregate cost, release the difference and present one understandable user action.