PERSLIS SCIENCE / MEET THE MODEL

Meet the model that
refuses to make science up.

Peel is a fail-first, weightless model. It builds an inspectable record from sources, tests what that record supports, and says unknown when it cannot prove the next step.

◈ Evidence connected⌁ Unknowns stay visible
PEEL / FAIL-FIRST

A different answer to the model problem.

A generative model

Predicts what sounds likely.

  • Knowledge compressed into weights
  • Fluency can outrun support
  • Reasoning path reconstructed
VS
PEEL

Proves what the record supports.

  • Knowledge stays in typed records
  • Every fact keeps its source
  • Unknowns remain visible
160hard kinase benchmark instances100% → 75.3%correct identification when model guesses were allowed to become truthPrototype DGM-005C-HARD result; not peer-reviewed.
PUT YOUR HANDS ON THE SYSTEM

Follow what becomes knowledge.

Drag to pan. Zoom in. Select a node to see what it does. Keyboard: Tab between nodes, Enter to inspect; arrow keys pan when the diagram is focused.

100%Watch the walkthrough ↗
SCIENCENo weights between you and the evidenceFailure becomes a boundaryUnknown is a useful resultModels propose. The floor decides.
MEASURED RESULTS

Compare the evidence policy, then test it yourself.

Our recorded Claude experiment used 160 kinase instances with 40% of the reference masked. These are in-house prototype results for one task, not a general hallucination leaderboard.

Test conditionCorrect identificationMeasured task cost
Deterministic baseline, no fill100%19.66
Claude guesses admitted as facts75.3%12.10
Model only reorders questions100%18.86
Retrieve verified records100%11.50
GPTNo matched result established here

Task cost is the experiment’s scoring unit, not dollars or product credits. Not peer-reviewed. The model’s 85.8% accuracy at guessing masked facts is a different metric from end-to-end identification.

Read the benchmark conditions and results →
RUN YOUR OWN COMPARISON

Bring the same questions to Peel, GPT and Claude.

Use the same source packet, date and tool access in each system. Record the exact model version and settings. These are visitor evaluation exercises; they do not reproduce the 160-instance study.

  1. Source check: “Research P69905. List five claims with the exact database record and URL supporting each.” Open every link and count supported claims, unsupported claims and explicit unknowns.

  2. Missing evidence: remove one supporting record from an identical source packet. Ask the same question again. Does the system identify the gap or supply an unsupported answer?

  3. Hypothesis check: “What connects BRCA1 and TP53? Separate recorded relationships from hypotheses.” Verify each cited relationship and check whether a candidate is presented as a finding.

  4. Handoff check: request a sourced report, reopen its citations and count how many findings another researcher can trace. Record completion time and cost separately from correctness.

Try the evaluation on the research floor →
EXPLORE THE DETAILS

Every part of the work, connected.

02

Failure becomes a boundary

A failed path is retained and can become a narrower rule. Learning cannot silently widen what the system may claim or do.

Explore Failure becomes a boundary
SEE IT IN ACTION

See why failure makes the model safer.

A short explanation of the fail-first learning model.

Recorded on a research prototype. The recording shows the scope demonstrated at that time.

Explore the research floor yourself →
PERSLIS / SCIENCERECORDED DEMO
A short explanation of the fail-first learning model.
ONE CONTINUOUS THREAD

From a question to a record you can inspect.

01Ask
02Retrieve
03Pin
04Verify
05Retain
See how it works →
BRING LOIS A REAL QUESTION

Let your next discovery start here.

Start with a target, a dataset or a research question. Inspect what the evidence supports, what it rejects and what remains unknown.

Start a research pilot →

Everything Perslis