Perslis Defense

Drag to turn it. Scroll to walk through it.

PERSLIS DEFENSE · THE BRAIN

A brain that learns without a neural network.

Twelve regions, five systems. Each region is a part of the system you can open, read and delete. It learns from every failure, carries what it learned into a new machine, and cannot learn its way out of its orders.

+17%a brain trained in one game, moved into another
+66%the same brain after pruning, only with proof
1 / 37changes it proposed to itself that earned promotion

No weights, no gradients, no training run. The learned state is the explanation. PROTOTYPE: everything here was measured in games and simulators, negatives included.

Twelve regions, five systems

EYE · Visual cortex

The eye

It does not guess what is on screen. A claim arrives — the engine or a sensor says a target is at this distance and bearing — and the eye checks that claim against the pixels, or abstains and says why.

Measured Gate passed on 7,863 frames, four-fold cross-validated: recall 0.984, false confirmations 0.48%. The margin is thin, and we say so.

code: vcard-eye

EYE · Thalamus

Facts

Pixels and telemetry become a short vocabulary of named facts — the only part of the brain that is specific to a domain. It audits itself too: candidate facts outside the vocabulary are tested against the moments it got hurt, to find the lessons it is not getting.

Measured Between Atari games the fact extractor is literally one codebase: Freeway imports the Invaders failure memory unchanged.

code: situation · discovery

PEEL · Amygdala

Failure memory

It remembers what hurt it, and how often, against the base rate. One unlucky death is not a lesson: a 95% Wilson lower bound must clear the base rate before a situation becomes a rule.

Measured An absolute threshold produced 0 rules from 268 real failures; the relative bar is a measured choice. A rule reads: died 13 of 19 times here, 3.6× the base rate.

code: fail-first memory

PEEL · Hippocampus

Memory and map

Every decision is an experience it can cite. It draws its own map from what its eye has seen — walls, secret doors and the places it died — and walks around the killing grounds.

Measured Doom: over 1.3 million decisions logged so far. The memory the brain carries is a table of situation → goal → outcome counts you can open and read. There are no weights to inspect.

code: memory · automap

PEEL · Synaptic pruning

Pruning

It removes what a brain never uses in a game, and what hurts it there — only with proof, and never the protected core.

Measured BattleZone: 7 behaviours removed (3 never fired in 69,469 decisions, 3 could never trigger, 1 proven harmful). The pruned brain scored 48,588 against the native rules’ 29,275 over 80 fresh seeds.

code: vdsg_brain.prune

FAIL-FIRST · Association cortex

The evolver

Every death is explained, one change is proposed, and it is tested in paired attempts on the same level and seed. Promoted only at z ≥ 2, rejected at z ≤ −1. Everything else stays “not proven”.

Measured Doom: 37 changes proposed, 1 promoted (z 2.26 over 11 pairs; route progress 0.28 → 0.41), 9 rejected, 26 not proven, 1 inert. Our own Peel-trace floor was tested as a change — and rejected.

code: evolve · mutate

FAIL-FIRST · Cingulate cortex

The doctor

Error detection. When it stops making progress it asks why, relaxes one assumption, tries a remedy and writes the result into its own rule book, with the evidence. What it cannot solve becomes a question.

Measured Live on a Doom level it wrote learned walls → replan from 4 successes in 5, and raised a line to use as an open question after 16 tries gained no ground.

code: doctor · containment

FLOOR · Prefrontal cortex

The floor

Before anything is chosen, it computes what is permitted. Orders can only narrow it. Learning can never widen it: the learner may rewrite what it believes works, never what it is authorised to do.

Measured CARLA: 64,952 of 64,952 hostile commands overridden over 40 km, 0 collisions. With the floor removed, the same commands caused 2 collisions within seconds.

code: floor · orders

FLOOR · Basal ganglia

Selection

It chooses only among the actions the floor admitted — the way the basal ganglia release one action by holding back all the others.

Measured Doom, rules with no learning at all: 17.9 against random play’s 3.2.

code: pilot · objective

TRACE · White matter

The trace

Every veto, promotion and rejection is written down with the events that caused it. One decision walks back to every experience that justifies it.

Measured 30 rules expand to 203 chained evidence tiles. Delete a row and the behaviour changes — there is nothing else in there.

code: peel trace · receipts

BODY · Cerebellum

Aim and motion

Fine control, trained per weapon: shots, hits, aim error and dodges are counted from the engine, not guessed.

Measured Tank eye aim error: 95th percentile 0.55°.

code: combat stats · armoury

BODY · Motor cortex

Same brain, new body

The brain is one sealed file — checksummed, signed and verified before every run. Only the adapter to the body changes.

Measured A tank brain trained in BZFlag, moved into Atari BattleZone: 33,575 against the native rules’ 28,600 over 80 seeds (+17%); on the held-out half +4,775, z 3.00. A second export was even on fresh seeds — transfer is not automatic.

code: vdsg-brain

Same brain. Harder body. Harder world.

Doom and the humanoid are not two projects. The timeline removes the artificial advantages from the same architecture, one at a time:

  • Simulation gives you perfect resets. Remove that.
  • Engine truth gives you authoritative state. Remove that.
  • Controlled environments give you predictable conditions. Remove that.
  • Cheap hardware limits the financial consequences. Eventually, remove that.
  • Known opponents let you prepare. Remove that.

01RUNNING

Doom / simulation

Doom, Wolfenstein, Fallout, GoldenEye, tanks, a drone flight stack in software, photoreal driving.

Advantage removed Nothing yet: perfect resets, engine truth, known opponents.

02NEXT

Hardware-in-the-loop

The same brain on a real flight controller and motor driver; the world is still simulated.

Advantage removed Perfect timing: real latency, a real bus, a real clock.

03GATED

Low-cost physical platforms

Small ground and air robots, where a crash costs a part, not a program.

Advantage removed Engine truth: the eye has only its own camera.

04GATED

Humanoid testbed

A commercially available humanoid in our own lab. It never spars against people.

Advantage removed Cheap consequences: balance, falls and wear are real.

05GATED

Robot-vs-robot circuit

Sanctioned robot-vs-robot bouts, the remote pilot replaced by the brain.

Advantage removed Known opponents: the opponent controls half the experiment.

The experiment, in full ↓

06GATED

Independent validation

Someone else runs the protocol, on their robot, against their opponents.

Advantage removed Us: the people who wrote the experiment.

WHERE THIS BRAIN CAME FROM

It started as a toy’s offline brain.

The first question was whether a child’s toy could have its own small brain, offline. Every step since has been the same question at a larger scale: how do you build intelligence you actually own?

  1. The brain factory
  2. TinkyBrains
  3. SML — the smallest language model
  4. Structured knowledge
  5. Trace — Peel
  6. Fail-First
  7. Peel — a self-evolving symbolic brain
  8. One brain, many worlds
  9. The embodied brain

The Brain Factory: how a toy brain became this one →

Where a neural model fits

  1. Neural modelperception · language · ambiguity
  2. Peelexplicit state · relationships · hypotheses · learned rules
  3. Floorauthority · invariants · execution
  4. Worlda game · a legacy system · a machine

A neural model can have billions of weights. Peel does not need to change them to learn something new: adaptive computation without weight updates. Learning without weight updates has prior art — symbolic learning, inductive logic programming, program synthesis, shielding; what we claim is this architecture and the evidence behind each change it makes. The full comparison →

HOW IT LEARNS

Every incident becomes an experiment.

  1. Impact / failure
  2. State trace
  3. Hypothesis
  4. Controlled change
  5. Retest
  6. Reject · not proven · promote

On Doom so far: 37 hypotheses, 1 promoted, 9 rejected, 26 not proven. The bar is high on purpose: a result that only ever improves is less trustworthy, not more.

PROOF · GAME BY GAME

The evolution of learning, in every game we run.

14 lanes, extracted from each one’s own evidence files on 2026-09-28 19:48. 7 show the brain getting better; the rest are flat, mixed or not learners at all — and they are here too. A section that only ever shows improvement would be a highlight reel, not proof.

A frame from Doom

Doommixed

On E1M5 route progress more than doubled with the pilot unchanged (z +6.35) — the memory and the map it draws did it. On E1M3 it is flat at both skills.

route progress (mean per block) by blocks of control attempts: E1M5 · normal skill 0.2 → 0.43; E1M3 · normal skill 0.42 → 0.33; E1M3 · hard skill 0.08 → 0.0600.210.43E1M5 · normal skill 0.43E1M3 · normal skill 0.33E1M3 · hard skill 0.06blocks of control attemptsroute progress (mean per block)
  • climbed E1M1 → E1M2 → E1M4 quickly; E1M3 took 335 attempts
  • no exit on E1M5 yet
  • orders, weapons and the hazard floor run live in the console

evidence: doom-floor/evidence/evolve/doom-s1/attempts.jsonl

A frame from Fallout (1997)

Fallout (1997)learns, then plateaus

XP 300 → 2,150 and a level gained from its own play. Learning on vs off, same save and seed, 12 minutes each: 15 vs 80 deaths per hour.

experience points held by hours of play (logged): XP 300 → 2,15001,0752,150XP 2,150hours of play (logged)experience points helddeaths per hour (A/B, 12 min each): learning on · live memory 15; learning off · live memory 80; learning on · empty memory 30; learning off · empty memory 80learning on · live memory15learning off · live memory80learning on · empty memory30learning off · empty memory80deaths per hour (A/B, 12 min each)
  • 11 campaign steps completed by the pilot
  • now stalled on one quest for many hours — shown, not hidden

evidence: fallout-floor/maps/journal.json

Drone course (MuJoCo)learns

Best clean lap 39.4 s → 27.8 s over 87 rounds. From held-out starts: learned 10/10 clean laps with 0 contacts; fresh 0/10 with 18.

best clean lap (s) by round: lap 39 → 2802039lap 28roundbest clean lap (s)
  • at chase speeds of 2–3 m/s the learned route still gets stuck or crashes

evidence: drone-floor/evidence/trial_rounds.jsonl

BZFlaglearns, still loses

Net kills per minute against BZFlag's own AI, pilot by pilot: v0 -2.05, v2 -1.17, v3 -0.43, v4 -0.26. Each promotion passed a paired test; it still loses.

net kills per minute vs BZFlag's AI: v0 (1 match) -2.05; v2 (8 matches) -1.17; v3 (1 match) -0.43; v4 (3 matches) -0.26v0 (1 match)-2.05v2 (8 matches)-1.17v3 (1 match)-0.43v4 (3 matches)-0.26net kills per minute vs BZFlag's AI
  • our pilot drives through the client's autopilot hook — the on-screen label is BZFlag's
  • after v4, 24 more trials were never proven and the loop stopped itself

evidence: bzflag-floor/evidence/pilots/evolution.jsonl

The eye (BZFlag frames)learns, not in a straight line

Six rounds of the eye learning to witness targets on the same 7,863 frames. Recall fell while false confirmations fell; only the last round passed the gate.

recall (%) by learning round: recall 96 → 9804998gate 98recall 98learning roundrecall (%)
roundrecallfalse confirmationsgate
first witness96.1%1.14%fail
+ glow clues79.5%0.53%fail
+ shots masked78.3%0.15%fail
+ engine flags78.3%0.15%fail
+ decision trees92.7%0.23%fail
+ turret layout98.4%0.46%PASS
  • gate: recall ≥ 98% and false confirmations ≤ 0.5%

evidence: bzflag-floor/evidence/eye-gate/evolution-teacher4.json

BattleZone (Atari)improves by transfer and pruning

A brain trained in BZFlag, moved into BattleZone: 33,575 vs the native rules' 28,600 over 80 seeds. Pruned with proof: 48,588 vs 29,275 on 80 fresh seeds.

mean score, 80 fresh seeds: native rules 29,275; transferred brain 28,975; transferred + pruned 48,588native rules29,275transferred brain28,975transferred + pruned48,588mean score, 80 fresh seeds
  • the second export (v4) was even with the native rules on fresh seeds — transfer is not automatic
  • the evolver's own parameter search promoted nothing here

evidence: vdsg-brain/evidence/battlezone/transfer.jsonl

Space Invaders (Atari, raw pixels)up, then down

Mean score 162 with no rules → 276 at its best → 247 at 100 episodes. More rules is not always better.

mean score, 12 held-out seeds by episodes of experience: score 162 → 2470138276score 247episodes of experiencemean score, 12 held-out seeds
  • Freeway runs the same mechanism and loses 12% — published, not tuned away

evidence: invaders-floor/web/curve.json

Doom · the evolver's own ledgerstrict

37 changes proposed by the system to itself; 1 earned promotion. The rest were rejected or stayed not proven — including our own Peel-trace floor.

changes: promoted 1; rejected 9; not proven 26; inert 1promoted1rejected9not proven26inert1changes
  • promotion needs z ≥ 2 in paired attempts on the same level and seed
  • rejection at z ≤ −1

evidence: doom-floor/evidence/evolve/doom-s1/rounds.jsonl

GoldenEye 007 (N64)flat

Progress on Dam stays near 0.06; 0 of 72 attempts survived. It did invent and promote 2 tactics over the rules' default.

route progress by blocks of 10 control attempts: Dam 0.06 → 0.0600.030.07Dam 0.06blocks of 10 control attemptsroute progress
  • promoted: “keep aim on the target” (z 2.03)
  • promoted: “strafe away from the target, keep aim on the target, fire when on target” (z 2.13)

evidence: goldeneye-floor (lane home)/evidence/evolve/goldeneye-dam-s3-doom/attempts.jsonl

A frame from Robot Tank (Atari)

Robot Tank (Atari)rules, not learning

Hand-written rules against random play: 5.67 vs 2.33 tanks destroyed (6 held-out games each). A comparison between arms, not a learning curve.

enemy tanks destroyed: rules 5.67; random 2.33; never fire 0rules5.67random2.33never fire0enemy tanks destroyed

evidence: robotank-floor/evidence/eval.jsonl

Quake III (OpenArena)not shown

Replaying maps with the memory from the first pass: t = 2.02, p = 0.079. Learning is not shown, and we say so.

  • the floor under the bots caused lava deaths in v1 before it was fixed

evidence: quake-floor/RESULTS.md

A frame from Gorillas (QBasic, 1991)

Gorillas (QBasic, 1991)no learning trend

The physics solver hits first time; the learner that brackets its throws hit 2 of 10 and shows no trend.

  • the real GORILLA.BAS, driven through a tap we added to it

evidence: gorillas-floor/tests/fixtures/throws.json

Wolfenstein 3Dno learning record

Rules and navigation only: it plays through levels, but this lane keeps no evidence memory yet.

evidence: wolf-floor/wolf_floor/console.py

Photoreal driving (CARLA)a shield, not a learner

64,952 of 64,952 hostile commands overridden over 40 km, 0 collisions. Remove the floor and the same commands crash within seconds. Nothing here learns — it is the boundary.

evidence: drive-floor/web/runs/manifest.json

Game names and imagery belong to their owners and are shown for research commentary. Frames are from our own runs.

MILESTONE 05 · EMBODIED ADVERSARIAL VALIDATION

Can Peel survive contact with a physical world we do not control?

Perslis plans to acquire a commercially available humanoid fighting robot and enter sanctioned robot-vs-robot competitions. We do not develop the robot hardware or any system that targets people. The experiment does one thing: it replaces continuous remote piloting with the Perslis autonomy architecture.

  1. Robot sensors
  2. Eye
  3. Peel
  4. Floor
  5. Robot actuators

A human keeps emergency and safety authority, but does not continuously pilot the robot during an autonomous benchmark run. The vendor’s own safety systems stay on.

Why not a lab-only benchmark?

LabWe control the environment.
CircuitThe opponent controls half the experiment.

Laboratory tests are controlled by us. A competition brings independently developed opponents, real contact physics, balance disruption, occlusion, sensor degradation, hardware wear, unpredictable behaviour, and situations the Perslis team did not script.

The benchmark is not “does our robot win?”

Failures are kept in public. A lost match is experimental evidence, not automatically a failed benchmark: every impact goes down the chain above — state trace, hypothesis, controlled change, retest, reject / not proven / promote.

Gated, not a first seed purchase

The expensive humanoid is bought only after the earlier milestones pass their gates: hardware-in-the-loop, low-cost physical platforms, then the humanoid testbed in our own lab. The gates are written before the run, like every Perslis experiment card.

Where the field is today

Humanoid bouts today are mostly piloted by people — VR rigs and gamepads — with short autonomous sequences; a robot in the Unitree G1 class starts at about US$13,500, research configurations cost more. The vendor has also shown its own learned model sparring without a pilot. Our question is different: not whether a robot can fight on its own, but whether a brain that learns without a neural network, under a floor, with a trace for every decision, survives contact.

Scope

Sanctioned robot-vs-robot competition and non-weaponised embodied-autonomy research only. No bouts against people, no weapons, no targeting of people, no hardware development.

The robot is not the product. It is the instrument. Peel is the experiment.

We know what this architecture does when we control the experiment. Seed funding lets us progressively stop controlling it.

WHAT THIS IS NOT

A prototype. Nothing is qualified.

Talk to us ↗