Predicts what sounds likely.
- Knowledge compressed into weights
- Fluency can outrun support
- Reasoning path reconstructed
Peel is a fail-first, weightless model. It builds an inspectable record from sources, tests what that record supports, and says unknown when it cannot prove the next step.
Drag to pan. Zoom in. Select a node to see what it does. Keyboard: Tab between nodes, Enter to inspect; arrow keys pan when the diagram is focused.
Our recorded Claude experiment used 160 kinase instances with 40% of the reference masked. These are in-house prototype results for one task, not a general hallucination leaderboard.
| Test condition | Correct identification | Measured task cost |
|---|---|---|
| Deterministic baseline, no fill | 100% | 19.66 |
| Claude guesses admitted as facts | 75.3% | 12.10 |
| Model only reorders questions | 100% | 18.86 |
| Retrieve verified records | 100% | 11.50 |
| GPT | No matched result established here | |
Task cost is the experiment’s scoring unit, not dollars or product credits. Not peer-reviewed. The model’s 85.8% accuracy at guessing masked facts is a different metric from end-to-end identification.
Read the benchmark conditions and results →Use the same source packet, date and tool access in each system. Record the exact model version and settings. These are visitor evaluation exercises; they do not reproduce the 160-instance study.
Source check: “Research P69905. List five claims with the exact database record and URL supporting each.” Open every link and count supported claims, unsupported claims and explicit unknowns.
Missing evidence: remove one supporting record from an identical source packet. Ask the same question again. Does the system identify the gap or supply an unsupported answer?
Hypothesis check: “What connects BRCA1 and TP53? Separate recorded relationships from hypotheses.” Verify each cited relationship and check whether a candidate is presented as a finding.
Handoff check: request a sourced report, reopen its citations and count how many findings another researcher can trace. Record completion time and cost separately from correctness.
Facts live as typed records with identifiers and sources—not as an answer recovered from hidden statistical memory.
Explore No weights between you and the evidenceA failed path is retained and can become a narrower rule. Learning cannot silently widen what the system may claim or do.
Explore Failure becomes a boundaryMissing evidence stays visible. That turns uncertainty into the next experiment instead of a polished guess.
Explore Unknown is a useful resultA language model may rank, plan or word an answer. It cannot admit a scientific fact; evidence standing belongs to the symbolic floor.
Explore Models propose. The floor decides.A short explanation of the fail-first learning model.
Recorded on a research prototype. The recording shows the scope demonstrated at that time.
Explore the research floor yourself →Start with a target, a dataset or a research question. Inspect what the evidence supports, what it rejects and what remains unknown.