Rubric
Atomic elements, owners, rules and authored remediation are compiled.
LearnX · V4 research dossier
A documented account of every model trial, the limitations encountered, and the move towards an engine where AI finds evidence while LearnX executes the rubric.
Models can be useful yet unstable at boundaries and can propagate one weakness across several criteria. LearnX therefore makes the rubric executable: the model only researches evidence, while versioned server rules determine the outcome.
AI finds evidence. LearnX certifies it, executes the rubric and chooses authored feedback.
Interim research. The French WRITING engine runs offline, but zero pipelines are promoted, zero contracts are published and zero AI assessment runtimes are active. Every paid experiment still requires a frozen identity, a budget decision and a separate owner GO.
Nothing has been erased. Each campaign remains part of the audit trail and explains why responsibilities were progressively moved from the model to LearnX.
| Date / phase | What was tested | What changed |
|---|---|---|
| 12 August · protocols 1–2 | 24 unique synthetic French cases across writing, reflection, practice and project; safety, format, evidence and blind-review gates. | Provider failures, invalid outputs, fabricated citations, injection, variability and pedagogical disagreement became separate measurements. |
| 12 August · protocol 3 | The model stopped producing a score or PASS/FAIL, but still produced levels, confidence and free feedback. | The server became the scoring authority; later analysis showed that level and feedback judgement also needed to leave the model. |
| 14 August · executable engine | A French WRITING archetype with 3 criteria and 9 atomic elements. | LearnX compiles ownership and level rules, checks evidence and selects authored remediation. |
| 15 August · Gemini | Gate v2: 3/3 unique cases. Panel 10×2: 5/10 unique cases completed twice, producing 10/20 valid workflows, then a non-exact quote on call 11. | The configuration is a genuine NO-GO; exact server-side quote resolution correctly rejected the fabricated substring. |
| 15 August · Sonnet 5 | Screening: 3/3 unique cases, 27/27 statuses and 17 exact quotes. Panel 10×2: 5/10 unique cases completed twice, producing 10/20 valid and matching workflows, then no visible output on call 11. | The request profile is a technical NO-GO: 2,500 reasoning tokens consumed the allowance and yielded 0 visible tokens. This is not a negative pedagogical verdict on Sonnet 5. |
| 16 August · bounded profile | A fresh identity reserves 1,024 tokens for reasoning and 1,800 for visible output, 2,824 total. Its first and only call yields 758 visible tokens but 1,082 reasoning tokens: 58 too many, or +5.66%. The gate stops before the second case: 1/4 call, 0/4 validated workflows. | The explicit maximum was not honoured at runtime. The request profile is a technical NO-GO, with no pedagogical verdict on the visible raw output. |
| 16 August · evidence-assist | LearnX segments the answer and assigns opaque span identifiers. The model can only propose evidence for, evidence against or abstain. | Candidate relations cannot produce levels, scores, progression or free feedback. Exact text and offsets stay under LearnX control. |
| 20 August · four-case gate | The positive case matches 9/9 relations. On the negative case, Sonnet reports explicit evidence against a recommendation where the frozen gold expects only “not demonstrated”: 7/9, then mandatory stop. | Two of four calls were made for an exact reconciled cost of USD 0.025622. Mutation and injection were not sent. The campaign is closed without replay. |
| 20 August · current decision | The latest divergence is semantically plausible and exposes an incomplete pseudo-oracle. | Define absence, explicit refutation, contradiction and ambiguity offline before paying for another model run. |
Each stage answers a different question. Passing an early test never selects or promotes a model.
| Term | Plain-language definition | What it demonstrates |
|---|---|---|
| Smoke test | One representative case checking that provider route, model and response format work together. | Technical compatibility for the next test, not pedagogical quality. |
| Gate | A short preregistered sequence stopped at the first defect. The latest gate planned four cases: positive, negative, mutation and injection. | The profile may advance only if every case passes. An early stop is neither a promotion nor a global verdict on the model. |
| Mini-panel | A small predefined set, historically six varied answers, used to expose failures before a costly campaign. | Early evidence only; size and repetition count must always be stated. |
| 10×2 panel | 10 unique semantic cases, each run twice: 20 workflows maximum. The second ten runs measure repeatability; they are not ten new situations. | Evidence extraction, negative controls, critical mutations and injection safety under repetition. |
| Full panel | The complete development corpus: 24 unique calibrated answers. | Coverage of all planned development families, not unseen-case generalisation. |
| 24×3 | 24 unique answers assessed three times: 72 logical workflows. | Both average error and stability; it still contains only 24 semantic situations. |
| Sealed holdout | A separate corpus never used to tune prompts, profiles or gates. | One-shot generalisation evidence after all development gates pass. |
Reading rule. Results are always shown as n/N, with unique semantic cases separated from repetitions. Repetitions measure stability; they never inflate semantic diversity.
The earlier full campaigns remain useful baselines. The newer evidence-research campaigns test a narrower role and must not be compared as though they were the same experiment.
| Campaign | Coverage | Exact result | Decision |
|---|---|---|---|
| Gemini · gate v2 | 3/3 unique cases · 3/3 workflows | 27/27 elements; negative and injection controls matched | Gate passed only; panel preparation allowed, no promotion |
| Gemini · 10×2 panel v2 | 5/10 unique cases · 10/20 valid workflows, then stop on call 11 | LearnX rejected a non-exact quote; 9/20 workflows were never called; cost USD 0.04345875 | Genuine NO-GO for this configuration; campaign closed without resume |
| Sonnet 5 · screening | 3/3 unique cases · 3/3 workflows | 27/27 statuses, 17 exact quotes and safe injection handling | Screening passed; not a promotion |
| Sonnet 5 · 10×2 panel | 5/10 unique cases · 10/20 valid and matching workflows, then stop on call 11 | 2,500 reasoning tokens and 0 visible tokens; 9/20 workflows were never called; cost USD 0.287208 | Technical NO-GO for the request profile, with no negative pedagogical verdict on Sonnet 5 |
| Sonnet 5 · bounded four-case gate | 1/4 call · 0/4 validated workflows · 0 repetitions | Reasoning maximum 1,024 + visible reserve 1,800 = total 2,824; observed 1,082 reasoning (+58 / +5.66%), 758 visible, total 1,840; Anthropic route and provider; cost USD 0.026104 | Technical NO-GO for the new profile; stopped before call two and no pedagogical verdict on Sonnet 5 |
| Sonnet 5 · evidence-assist 3.0 | 2/4 calls · 1/4 fully matching workflow · stop before mutation and injection | Positive 9/9; negative 7/9; total 16/18. Two “against” relations diverged from a “not demonstrated” gold. Cost USD 0.025622, 100% reconciliation and no observed false support. | Semantic NO-GO for this identity. The stop is valid, but it exposes an incomplete evidence ontology rather than an obvious pedagogical model error. |
Status on 20 August 2026. Historical judge campaigns and recent evidence-research campaigns used different responsibilities and identities. The latest injection case was not sent, so this gate does not establish injection safety. Costs and outputs remain documented, but they cannot be merged into a promotion claim or product price.
Full campaigns, panels, technical smokes and catalogue-only screening are different evidence levels and are retained below without rewriting their historical conclusions.
| Model | Evidence level | Observed result | Lesson |
|---|---|---|---|
| Mistral Medium 3.5 | Several panels and 24×3 full runs | 72/72 usable; 3 false positives and 5 false negatives in campaign 1.6 | Economical historical baseline; NO-GO as a judge. The simulated Mistral–Sonnet pipeline is not promoted. |
| Claude Sonnet 4.6 | Panels and protocol 3.0.1 full | 71/72 finally usable; 0 false positives and 6 false negatives | Conservative historical baseline; NO-GO as a judge. |
| Gemini 3.6 Flash | Historical full; researcher gate 3/3; interrupted 10×2 panel | Gate v2 passed 3/3, then 10/20 valid workflows before a non-exact quote on call 11; cost USD 0.04345875 | Genuine NO-GO for the panel v2 configuration. Ten simple results do not offset the evidence defect. |
| Claude Sonnet 5 | Screening 3/3; interrupted 10×2 panel; bounded gate 1/4; evidence-assist gate 2/4 | After two technical profile failures, evidence-assist 3.0 produced 9/9 on the positive case and 7/9 on the negative case before the “against” versus “not demonstrated” stop; latest cost USD 0.025622. | The evidence-assist identity is a NO-GO, but the latest difference does not demonstrate an obvious pedagogical failure. Explicit refutation must be separated from absence of evidence. |
| GPT-5.6 Terra | Technical smokes | Pinned structured OpenRouter route unavailable | No pedagogical verdict; a direct-API campaign would be a separate experiment. |
| GPT-5.6 Sol | Sequential smokes | Successful case valid; ambiguous case timed out or truncated | Capable, but the tested profile was too slow or verbose for the bounded contract. |
| Claude Opus 4.8 | Non-comparable historical full + 2 current smokes | Two recent smokes valid; older data used a previous corpus and cost more | Possible qualitative reference, insufficient current evidence for promotion. |
| Kimi K2.5 | First smoke | Timeout before usable output | Stopped early under the sequential budget gate. |
| Command A | First smoke | Provider 400 before usable output | Technical compatibility not demonstrated. |
| Qwen 3.7 Max | Exploratory smokes | Invalid capped output, followed by timeout | Stopped before a panel due to cost and latency. |
| DeepSeek V4 Flash | Catalogue only | Structured routes identified; no paid benchmark run | Reserve option only; no performance claim. |
Correct interpretation. A transport, profile or format failure does not prove that a model is pedagogically poor. It shows that the tested identity did not execute the LearnX contract reliably. Conversely, a 3/3 screening cannot establish production quality.
LearnX segments the learner answer; one model proposes candidate relations on those spans; LearnX alone validates them, applies mechanical rules and controls restitution. The latest gate validated transport but exposed an incomplete negative-evidence boundary.
Atomic elements, owners, rules and authored remediation are compiled.
LearnX creates opaque span identifiers, offsets and hashes before the call.
Sonnet proposes only “for”, “against” or “abstain” on allowed identifiers.
LearnX validates candidate relations. They can never directly set a score, level or progression state.
LearnX selects authored feedback; the model writes no free feedback and cannot affect progression.
Strict boundary. The current problem is the evidence vocabulary: “nothing demonstrates X” and “the answer explicitly refutes X” must no longer be treated as the same semantic state.
Evidence safety and the executable rubric remain deterministic. No threshold, corpus or pseudo-oracle may be opportunistically changed after a result.
The offline engine and evidence-assist transport work. The next task is to make the negative evidence ontology unambiguous before spending on another campaign.
Retain the two calls, 16/18 result, cost and hashes as append-only evidence. Do not send mutation or injection under the closed identity.
Separate supported, explicitly refuted, not demonstrated and ambiguous. Internal contradiction remains a dedicated element rather than a synonym for refutation.
Build minimal pairs for silence, abstention, explicit negation, contradiction and positive evidence. Verify that only the intended relation changes.
Persist input, cache, reasoning and visible-output tokens separately. Any mapping, oracle, runner or corpus change creates a new identity.
A new gate requires a new Finance decision and owner authorisation. The 10×2 panel and one-shot holdout remain conditional on a 4/4 result.
No WRITING contract is published and no runtime is active. Product scaffolding remains testable only with simulated certificates.