
Open your AI risk register and find two rows.
Row one: the prompt-injection guardrail on the support agent. Documented, deployed, never tested. Likelihood 3, impact 4. Score 12. Amber.
Row two: the same guardrail, tested in March. Bypassed on the fourth attempt. Ticket open. Likelihood 3, impact 4. Score 12. Amber.
One row is a question. The other is a route someone can walk today. The register cannot tell them apart, and the remediation clock hanging off amber ticks at the same speed for both.
This is not a failure of judgment. Most risk scoring is likelihood times impact adjusted by a control-effectiveness rating — a model assembled when controls were static and testing them was rare. Likelihood was a forecast, because for most rows evidence was never going to arrive. The register's job was to rank guesses.
Two things changed. AI controls became behavioral rather than configurational — whether a guardrail holds depends on what the model was asked, what it retrieved, and which tool it reached for, none of which you can read off a config file. And testing got cheap enough to repeat, turning "we tried it and it failed" into an ordinary event.
The register never grew a field for that. It has one for how bad, one for how likely, and none for how we know.

Impact is identical across those rows — same agent, same refund path, same blast radius — so the difference must land in likelihood or control effectiveness, and both punish you for testing.
Likelihood is scored as a prediction, and a careful scorer hedges a prediction. But the tester who bypassed the guardrail did not predict anything. They produced the event, and the scale has no way to say so. The untested control is scored on design intent: documented, so it probably works, so likelihood sits mid-range.
Control effectiveness compounds it. A failed test lands on "partially effective," because "ineffective" reads as an accusation; "not tested" lands near the middle for the same reason. They reach amber from opposite directions, and nobody is pushed to separate them: an untested control has no owner of bad news, while the proven failure, the row where the attacker's clock is already running, has a name attached.
Illustrative composite, not a client. A mid-market SaaS company runs a support agent with refund authority — the scenario from our post on proving an autonomous action is real. Forty AI rows, thirty-one amber. Two of them: *human approval required for refunds over $500*, never exercised, and *agent cannot be induced to split a refund into sub-threshold amounts*, attempted in the last assessment and defeated. Same score, same color, same SLA.

Tested-and-failed is not a bad prediction. It is a demonstrated capability under stated conditions, carrying the four things that separate evidence from a claim: the attempt, the result, the artifact, the conditions. Its uncertainty is bounded — you know a path exists; you do not know whether anyone is walking it.
Security agreed with this elsewhere and never carried it into AI governance. CISA's Known Exploited Vulnerabilities catalog exists because observed exploitation outranks theoretical severity. The SSVC decision model makes exploitation its own axis, defining active as "shared, observable, reliable evidence that the exploit is being used in the wild." NIST SP 800-53A does it quietly: each determination is recorded as satisfied or other than satisfied, and "not assessed" is not an option.
The honest objection is a good one. An unknown can hide something worse than anything you have proven: the control you never tested might fail on the first attempt, in a way that makes the March bypass look minor. Unknown is not a synonym for fine.
NIST's AI Risk Management Framework says so plainly on risk measurement: "The inability to appropriately measure AI risks does not imply that an AI system necessarily poses either a high or low risk." Scoring unknowns down produces a worse register, not a better one.
So the claim is narrower than "failing beats unknown." Among rows of comparable impact, a proven failure should outrank an unknown, because one is a bounded demonstrated path and the other is a range. But the ordering matters less than this: no register should give the two the same value, because they need different work. An unknown needs a test, on a clock. A proven failure needs a fix, on a clock. Collapse them and both land in the remediation queue, where the unknown always loses.

The fix is not a better formula. It is a second field: a state, defined strictly by the evidence that set it, carrying that evidence's date.
Unknown. Nobody has looked. Evidence: none.
Asserted. Someone says it works — a policy, a vendor questionnaire, an architecture diagram. Evidence: a statement. Set by governance review.
Validated. The running system was inspected and matches intent: tenant settings, identity scopes, tool permissions. Evidence: the configuration, captured. Behavior under attack still unproven.
Proven-holding. Someone attempted what the control is supposed to stop, under agreed scope, and it held. Evidence: the attempt, the result, the artifact, the conditions. Set by adversarial testing.
Proven-exposed. The same attempt, the opposite result.
Four rules make the field do work. Proven-exposed is the loudest signal in the register, and cannot be quieted by a compensating control that is itself only Asserted — offset proof with proof. State decays: a change to the model version, system prompt, retrieval corpus, or tool scope returns a Proven-holding row to Unknown, and drift is the ordinary case for AI. Unknown gets an assessment SLA, not a remediation SLA. And do not average state into the score — sort by impact, then by state, because blending re-hides what the field reveals, which is how check-the-box work survives.
Scoring does not prioritize anything. People do, and a state field only makes a good argument harder to lose. It adds work: every Proven-holding row needs an expiry, and expiries create queue pressure. It will also make the register look worse before better, since the first honest pass turns confident green into Unknown and the data-leakage controls usually turn out never to have been tested.
Framework alignment will not rescue you either. ISO/IEC 42001 is certifiable, and certification means a management system met a standard — not that a control held under attack. CSA's AI Controls Matrix and the OWASP Top 10 for Agentic Applications give you a control set and a threat vocabulary, but neither sets a state. Only a test does.
None of this eliminates risk. It lowers the odds that the loudest row in your register is the one nobody checked.
For a self-check, take five minutes with your register. Pick any amber row: without opening a ticket, can you say whether it is Unknown, Asserted, Validated, Proven-holding, or Proven-exposed? For every Proven-holding row, what is the evidence date, and what has changed since? And is there a row where a control observed to fail scores the same as one nobody has looked at — would you fix them in the register's order?
If those are hard to answer, the problem is not your scoring formula. The register has no field for how you know.
We are a CREST-accredited, ISO/IEC 27001 certified offensive security firm, and our deliverables set a state rather than fill a cell. Our AI and ML penetration testing attempts the control and hands back the attempt, the result, the artifact, and the conditions. Where the system changes faster than an annual cycle, we run it continuously so the states carry current dates. Independent adversarial testing produces that evidence, whoever does it.
If you want a second opinion on how your AI risks are currently scored, get in touch.