What this is not.
A safety claim invites the strongest scrutiny available, and it should. Here is everything we would rather a reviewer hear from us than find on their own.
Status
PROTOTYPE. Research-grade. Nothing here has been validated for a safety-critical path. It carries no functional-safety qualification. It has not flown, has not driven a road vehicle, and has not been exercised on a range. The demonstrations are a simulator, four games and a document corpus. Anyone citing the +32% without the −12% is misreading the work.
The open problem
Our own frozen result shows the mechanism fails when risk and objective are coupled — when the action that accomplishes the task is the action that exposes you.
That is Freeway: up is both the only scoring action and the dangerous one, three rules formed, every one blocking up, and the agent correctly concluded that moving is lethal and correctly stopped playing. It is also the ordinary condition of a contested environment. A floor that refuses what killed it before will refuse the mission.
Adding a catastrophic-outcome floor on top of a utility ranking did not fix it — it reconstructed the original veto and made both games worse. There is no catastrophic tail to exclude when the fatal action is the only useful one.
The work is to price recoverability rather than fatality: refuse what cannot be undone, and permit everything else right up to that edge. A mechanism that understood recoverable failure should beat the unassisted baseline on Freeway outright, not merely reach parity — that is the benchmark, and it is not met yet. Our irrecoverability research is the formal side of the question.
Other limits worth naming
- Composition still needs bounding. Retirement takes dead ends to zero on the measured set; it has not been shown at squad scale, where the between-element constraints dominate.
- The learned floor needs exposure to learn. A rule is earned by observed failures. In a domain where the first failure is unacceptable, the floor must be specified rather than learned — which is the classical shielding case, and a different product posture.
- Bucketing is a modelling choice. The abstraction that decides when two situations are “the same” is declared in one place and is arguable. Get it wrong and rules generalise to situations they should not.
- We have no third-party evaluation. Every figure on this site is ours. That is a limitation, and independent replication is something we want rather than something we would resist.
Where to start anyway
None of the above blocks an evaluation, because an evaluation does not require the system to be trusted. Shadow mode against a recorded mission is inert by construction: read the counterfactual log and decide for yourself what it would have refused and whether it was right.