LearnX · V4 research dossier

Building useful assessment without manufacturing certainty.

A documented account of every model trial, the limitations encountered, and the move towards an engine where AI finds evidence while LearnX executes the rubric.

Status Interim research — no pipeline promotedData Synthetic French development corporaUpdated 20 August 2026Version française
01 · Current decision

Stop asking a model to decide the grade.

Models can be useful yet unstable at boundaries and can propagate one weakness across several criteria. LearnX therefore makes the rubric executable: the model only researches evidence, while versioned server rules determine the outcome.

AI finds evidence. LearnX certifies it, executes the rubric and chooses authored feedback.

Interim research. The French WRITING engine runs offline, but zero pipelines are promoted, zero contracts are published and zero AI assessment runtimes are active. Every paid experiment still requires a frozen identity, a budget decision and a separate owner GO.

02 · Work completed

From model-as-judge to model-as-evidence-researcher.

Nothing has been erased. Each campaign remains part of the audit trail and explains why responsibilities were progressively moved from the model to LearnX.

Chronology of LearnX AI assessment research
Date / phaseWhat was testedWhat changed
12 August · protocols 1–224 unique synthetic French cases across writing, reflection, practice and project; safety, format, evidence and blind-review gates.Provider failures, invalid outputs, fabricated citations, injection, variability and pedagogical disagreement became separate measurements.
12 August · protocol 3The model stopped producing a score or PASS/FAIL, but still produced levels, confidence and free feedback.The server became the scoring authority; later analysis showed that level and feedback judgement also needed to leave the model.
14 August · executable engineA French WRITING archetype with 3 criteria and 9 atomic elements.LearnX compiles ownership and level rules, checks evidence and selects authored remediation.
15 August · GeminiGate v2: 3/3 unique cases. Panel 10×2: 5/10 unique cases completed twice, producing 10/20 valid workflows, then a non-exact quote on call 11.The configuration is a genuine NO-GO; exact server-side quote resolution correctly rejected the fabricated substring.
15 August · Sonnet 5Screening: 3/3 unique cases, 27/27 statuses and 17 exact quotes. Panel 10×2: 5/10 unique cases completed twice, producing 10/20 valid and matching workflows, then no visible output on call 11.The request profile is a technical NO-GO: 2,500 reasoning tokens consumed the allowance and yielded 0 visible tokens. This is not a negative pedagogical verdict on Sonnet 5.
16 August · bounded profileA fresh identity reserves 1,024 tokens for reasoning and 1,800 for visible output, 2,824 total. Its first and only call yields 758 visible tokens but 1,082 reasoning tokens: 58 too many, or +5.66%. The gate stops before the second case: 1/4 call, 0/4 validated workflows.The explicit maximum was not honoured at runtime. The request profile is a technical NO-GO, with no pedagogical verdict on the visible raw output.
16 August · evidence-assistLearnX segments the answer and assigns opaque span identifiers. The model can only propose evidence for, evidence against or abstain.Candidate relations cannot produce levels, scores, progression or free feedback. Exact text and offsets stay under LearnX control.
20 August · four-case gateThe positive case matches 9/9 relations. On the negative case, Sonnet reports explicit evidence against a recommendation where the frozen gold expects only “not demonstrated”: 7/9, then mandatory stop.Two of four calls were made for an exact reconciled cost of USD 0.025622. Mutation and injection were not sent. The campaign is closed without replay.
20 August · current decisionThe latest divergence is semantically plausible and exposes an incomplete pseudo-oracle.Define absence, explicit refutation, contradiction and ambiguity offline before paying for another model run.
03 · Understanding the method

Smoke, gate, 10×2 panel, full and 24×3 are different evidence levels.

Each stage answers a different question. Passing an early test never selects or promotes a model.

Definitions of benchmark stages
TermPlain-language definitionWhat it demonstrates
Smoke testOne representative case checking that provider route, model and response format work together.Technical compatibility for the next test, not pedagogical quality.
GateA short preregistered sequence stopped at the first defect. The latest gate planned four cases: positive, negative, mutation and injection.The profile may advance only if every case passes. An early stop is neither a promotion nor a global verdict on the model.
Mini-panelA small predefined set, historically six varied answers, used to expose failures before a costly campaign.Early evidence only; size and repetition count must always be stated.
10×2 panel10 unique semantic cases, each run twice: 20 workflows maximum. The second ten runs measure repeatability; they are not ten new situations.Evidence extraction, negative controls, critical mutations and injection safety under repetition.
Full panelThe complete development corpus: 24 unique calibrated answers.Coverage of all planned development families, not unseen-case generalisation.
24×324 unique answers assessed three times: 72 logical workflows.Both average error and stability; it still contains only 24 semantic situations.
Sealed holdoutA separate corpus never used to tune prompts, profiles or gates.One-shot generalisation evidence after all development gates pass.

Reading rule. Results are always shown as n/N, with unique semantic cases separated from repetitions. Repetitions measure stability; they never inflate semantic diversity.

04 · Observed results

Historical judges failed; recent researchers have not yet completed a panel.

The earlier full campaigns remain useful baselines. The newer evidence-research campaigns test a narrower role and must not be compared as though they were the same experiment.

2/4calls made in the latest gate
16/18matching atomic relations
0.025622USD actually billed and reconciled
0promoted pipelines or published contracts
Historical quality and stabilityPercentage — 24 unique cases × 3 repetitions
Mistral · criteria
87.5%
Sonnet 4.6 · criteria
87.8%
Mistral · decision
88.9%
Sonnet 4.6 · decision
91.5%
M. variability
12.5%
S. variability
12.5%
Historical loaded cost + FX bufferUSD equivalent per assessment — retrospectively recombined data
Mistral only
0.0077
Targeted cascade
0.0161
Sonnet 4.6 only
0.0239
Recent evidence-researcher results
CampaignCoverageExact resultDecision
Gemini · gate v23/3 unique cases · 3/3 workflows27/27 elements; negative and injection controls matchedGate passed only; panel preparation allowed, no promotion
Gemini · 10×2 panel v25/10 unique cases · 10/20 valid workflows, then stop on call 11LearnX rejected a non-exact quote; 9/20 workflows were never called; cost USD 0.04345875Genuine NO-GO for this configuration; campaign closed without resume
Sonnet 5 · screening3/3 unique cases · 3/3 workflows27/27 statuses, 17 exact quotes and safe injection handlingScreening passed; not a promotion
Sonnet 5 · 10×2 panel5/10 unique cases · 10/20 valid and matching workflows, then stop on call 112,500 reasoning tokens and 0 visible tokens; 9/20 workflows were never called; cost USD 0.287208Technical NO-GO for the request profile, with no negative pedagogical verdict on Sonnet 5
Sonnet 5 · bounded four-case gate1/4 call · 0/4 validated workflows · 0 repetitionsReasoning maximum 1,024 + visible reserve 1,800 = total 2,824; observed 1,082 reasoning (+58 / +5.66%), 758 visible, total 1,840; Anthropic route and provider; cost USD 0.026104Technical NO-GO for the new profile; stopped before call two and no pedagogical verdict on Sonnet 5
Sonnet 5 · evidence-assist 3.02/4 calls · 1/4 fully matching workflow · stop before mutation and injectionPositive 9/9; negative 7/9; total 16/18. Two “against” relations diverged from a “not demonstrated” gold. Cost USD 0.025622, 100% reconciliation and no observed false support.Semantic NO-GO for this identity. The stop is valid, but it exposes an incomplete evidence ontology rather than an obvious pedagogical model error.

Status on 20 August 2026. Historical judge campaigns and recent evidence-research campaigns used different responsibilities and identities. The latest injection case was not sent, so this gate does not establish injection safety. Costs and outputs remain documented, but they cannot be merged into a promotion claim or product price.

05 · Models explored

The complete record remains visible.

Full campaigns, panels, technical smokes and catalogue-only screening are different evidence levels and are retained below without rewriting their historical conclusions.

Models explored during LearnX research
ModelEvidence levelObserved resultLesson
Mistral Medium 3.5Several panels and 24×3 full runs72/72 usable; 3 false positives and 5 false negatives in campaign 1.6Economical historical baseline; NO-GO as a judge. The simulated Mistral–Sonnet pipeline is not promoted.
Claude Sonnet 4.6Panels and protocol 3.0.1 full71/72 finally usable; 0 false positives and 6 false negativesConservative historical baseline; NO-GO as a judge.
Gemini 3.6 FlashHistorical full; researcher gate 3/3; interrupted 10×2 panelGate v2 passed 3/3, then 10/20 valid workflows before a non-exact quote on call 11; cost USD 0.04345875Genuine NO-GO for the panel v2 configuration. Ten simple results do not offset the evidence defect.
Claude Sonnet 5Screening 3/3; interrupted 10×2 panel; bounded gate 1/4; evidence-assist gate 2/4After two technical profile failures, evidence-assist 3.0 produced 9/9 on the positive case and 7/9 on the negative case before the “against” versus “not demonstrated” stop; latest cost USD 0.025622.The evidence-assist identity is a NO-GO, but the latest difference does not demonstrate an obvious pedagogical failure. Explicit refutation must be separated from absence of evidence.
GPT-5.6 TerraTechnical smokesPinned structured OpenRouter route unavailableNo pedagogical verdict; a direct-API campaign would be a separate experiment.
GPT-5.6 SolSequential smokesSuccessful case valid; ambiguous case timed out or truncatedCapable, but the tested profile was too slow or verbose for the bounded contract.
Claude Opus 4.8Non-comparable historical full + 2 current smokesTwo recent smokes valid; older data used a previous corpus and cost morePossible qualitative reference, insufficient current evidence for promotion.
Kimi K2.5First smokeTimeout before usable outputStopped early under the sequential budget gate.
Command AFirst smokeProvider 400 before usable outputTechnical compatibility not demonstrated.
Qwen 3.7 MaxExploratory smokesInvalid capped output, followed by timeoutStopped before a panel due to cost and latency.
DeepSeek V4 FlashCatalogue onlyStructured routes identified; no paid benchmark runReserve option only; no performance claim.

Correct interpretation. A transport, profile or format failure does not prove that a model is pedagogically poor. It shows that the tested identity did not execute the LearnX contract reliably. Conversely, a 3/3 screening cannot establish production quality.

06 · Architecture evaluated

Turn the rubric into an executable program without pretending that the researcher is validated.

LearnX segments the learner answer; one model proposes candidate relations on those spans; LearnX alone validates them, applies mechanical rules and controls restitution. The latest gate validated transport but exposed an incomplete negative-evidence boundary.

Step 1

Rubric

Atomic elements, owners, rules and authored remediation are compiled.

Step 2

Segmentation

LearnX creates opaque span identifiers, offsets and hashes before the call.

Step 3

Research

Sonnet proposes only “for”, “against” or “abstain” on allowed identifiers.

Step 4

Certification

LearnX validates candidate relations. They can never directly set a score, level or progression state.

Step 5

Restitution

LearnX selects authored feedback; the model writes no free feedback and cannot affect progression.

Strict boundary. The current problem is the evidence vocabulary: “nothing demonstrates X” and “the answer explicitly refutes X” must no longer be treated as the same semantic state.

07 · Success criteria

Block structural defects before measuring quality.

Evidence safety and the executable rubric remain deterministic. No threshold, corpus or pseudo-oracle may be opportunistically changed after a result.

Hard gates

  • No span outside the learner response and no canary leak.
  • No unknown element, wrong owner or forbidden field.
  • No false support on a mechanical control.
  • No model-produced level, score, PASS/FAIL or free feedback.
  • All dispatches and actual costs reconciled.
  • Zero effect on learner progression.

Gates currently closed

  • Evidence-assist gate: 2/4 calls, 1 matching workflow, then semantic stop.
  • Mutation and injection were not sent; no safety claim is made for this identity.
  • 10×2 panel: not authorised because the gate did not reach 4/4.
  • Sealed holdout: qualified but unopened.
  • V4-010: simulated flow only, hard-off in production.
  • A new experiment requires a new identity and approvals.
08 · Evidence still required

What remains before any beta activation.

The offline engine and evidence-assist transport work. The next task is to make the negative evidence ontology unambiguous before spending on another campaign.

1. Freeze evidence-assist 3.0

Retain the two calls, 16/18 result, cost and hashes as append-only evidence. Do not send mutation or injection under the closed identity.

2. Define four evidence states

Separate supported, explicitly refuted, not demonstrated and ambiguous. Internal contradiction remains a dedicated element rather than a synonym for refutation.

3. Test the ontology offline

Build minimal pairs for silence, abstention, explicit negation, contradiction and positive evidence. Verify that only the intended relation changes.

4. Complete telemetry and version everything

Persist input, cache, reasoning and visible-output tokens separately. Any mapping, oracle, runner or corpus change creates a new identity.

5. Restart with four cases

A new gate requires a new Finance decision and owner authorisation. The 10×2 panel and one-shot holdout remain conditional on a 4/4 result.

6. Keep V4-010 hard-off

No WRITING contract is published and no runtime is active. Product scaffolding remains testable only with simulated certificates.