So we don't make it. We report what an agent DEMONSTRATED. On a date. Against a fixed set of checks you can read. In conditions you can inspect.
The difference matters in the room that counts. A “safe” badge is a claim you can't defend the first time the agent surprises you. A demonstrated result — this agent passed these checks, across these six domains, on this date, with this evidence — is what an auditor recognizes: a signal, scoped and dated, defensible precisely because it doesn't pretend to cover the future.
How the result is produced, so it can't be waved away: (1) An administered exam, not a self-assessment — 126 checks across six domains (ownership, identity, task, safety, governance, runtime); no domain buys back a critical failure in another. (2) Deterministic — the same evidence produces the same score, every time; the scoring core is rule-based. Model-consensus review of judgment checks is built and tested, and switches on when behavioral evaluation goes live. (3) Tested, and marked as such — a designation says whether the behavioral suite was actually administered, not merely that documents were reviewed. We show the honest number, including a low one.
And because “demonstrated on a date” is honest, it carries an honest obligation: check again. An agent that behaved in March can drift by June. So the score isn't a sticker — we re-administer on a cadence: baseline today, first drift reading at day 30. Never “live monitoring,” and never “stable” when we don't yet have the signal.
None of this is anti-AI. But autonomy should expand only as fast as you can evidence control — and evidence is a dated, inspectable record of what an agent did when tested. Not a promise it's safe. A durable answer to the only defensible question: what did it demonstrate, and when?