Pre-register
The question, the threshold, and the scoring were committed to version control before the outcome was known. No moving the goalpost from 14% to 15% after seeing a number.
Peel does not assume every problem needs a model. So we ran the experiment honestly: build increasingly difficult tasks, strengthen the deterministic opponent after every failure, freeze every result, and admit machine learning only when a pre-registered test says it adds measurable value. This is what happened.
Five deterministic wins, then one narrow opening — and even there, a real learner could not take it. The whole plot, in one line.
A pre-registered characterization program, not a leaderboard. Every claim below traces to a frozen result and a cryptographic digest.
The question, the threshold, and the scoring were committed to version control before the outcome was known. No moving the goalpost from 14% to 15% after seeing a number.
Every time a cheap deterministic solution existed, it became the new opponent. We even handed the model an oracle — a perfect version of itself — as an unbeatable upper bound.
Two scoreboard bugs of our own, a crash, and rejected hypotheses all remain in the record. When an attractive ML result appeared, we asked why it existed — twice it was our own measurement error.
A model may propose or strategize. What is admitted as true — verified state, identity, provenance — stays deterministic. A model in the truth path corrupts at its own error rate.
Neither system gets our loyalty. The experiment does.
Before any model, we characterized the deterministic floor across nine adversarial, cold-drawn biology runs — the foundation that makes it a strong opponent (it broke twice; we froze the wrecks and repaired generically).
Each row is a boundary where we thought a model might be needed. Each time, we asked one question: can deterministic machinery already do this? Each time it could — so ML was rejected.
| Boundary tested | Why ML might help | Could deterministic machinery do it? | ML |
|---|---|---|---|
| Operation selection | Maybe choosing which query to run is hard | Yes — every cheap operation was productive; run them all | REJECTED |
| Search depth / novelty | The search space explodes combinatorially | Yes — evidence was uniform; no selective signal to learn | REJECTED |
| Goal relevance | A goal makes some branches matter more | Yes — a cheap deterministic filter captured the relevant ones | REJECTED |
| Composition | Relevance needs multi-step, conjunctive reasoning | Yes — no reproducible ML advantage over the rule | REJECTED |
| Relational structure | Cheap category features finally fail here | Yes — a cheap graph-relation rule still routed better | REJECTED |
Five chances. Five rejections. We were not proving deterministic systems are universally better — we were repeatedly trying to falsify that, and failing.
Every rejection so far shared a trait: the deterministic model could represent everything that mattered. So we left that regime deliberately — a task with partially-observed state, hidden dependency structure, and costly mistakes, where correlations exist that an independence-assuming model cannot express.
Each cell: how much an ideal model would cut operational cost vs the deterministic planner, as hidden structure (ρ, down) and the price of a wrong answer (λ, across) rise. Green clears the pre-registered 15% bar.
A seat appeared in exactly one corner: where hidden structure is strong (ρ=0.8) and mistakes are expensive (λ≥50), the ideal model's advantage crosses the bar — 16.8%, then 19.8%. Everywhere else it isn't worth the inference. And pre-registration held: λ=25 lands at 14.3%, below the bar we fixed in advance.
ML did not become universally useful. A seat appeared only when relevant structure was latent — and being wrong became expensive.
That 16.8% belonged to a perfect model. So we trained an honest one from historical incidents — no oracle, no leakage — and asked whether it could actually take the seat. It could not.
We grew the learner's training data 16× in the one corner where the seat exists. It recovers more of the ideal model's edge — then plateaus.
Captured 1.8% → 25.7% → 42.8% → 47.8% of the ideal edge as data grew — but doubling from 12,000 to 24,000 incidents produced essentially no additional operational advantage, and it never cleared the bar.
The ideal model clears the bar, and the learner's accuracy climbs toward it. The structure is real.
The curve flattened while data still had headroom — the marginal value of more incidents collapsed near zero.
A generic learner recovers only half the structure and stops. The open question is model class — a learner whose shape matches the hidden structure. That is the next experiment.
Not a slogan — a division of labor that these experiments measured rather than assumed.
owns
must earn
decides
Use inference where inference is required. Don't spend inference on facts the computer can establish deterministically.
Peel does not ask whether AI can perform a task. It asks whether inference adds enough value that it should be allowed to perform the task at all.
Honest scope: results hold under these pre-registered tasks and protocols; the one theoretical seat is unearned by a realistic learner so far. For a lab this is the point — correctness comes from verified data and deterministic reasoning, on your own hardware, at zero model cost; a model is invited in only where an experiment proves it earns the room.