A prospect sends you a 47-question AI security assessment. Question 12 asks whether you have an AI acceptable use policy. You do, so you tick yes. Question 19 asks whether access to AI systems is role-based. It is, so you tick yes. Question 31 asks whether AI activity is logged. It is, so you tick yes.
You send it back. Three weeks later the deal clears security review.
Now here is the question nobody asked. Of those 47 controls, how many were ever attempted by anyone trying to get past them? That gap is the whole problem with how an AI security assessment usually gets scoped.
This is not a complaint about governance. Questionnaires exist because they work reasonably well for a specific kind of control: one that is static, configuration-based, and either present or absent. Is disk encryption on? Is MFA enforced? Those questions have real answers, and the answer stays true tomorrow.
AI systems broke that assumption in a way worth naming plainly. The controls that matter around them are behavioral rather than configurational. Whether an AI access control holds depends on what the model was asked, what it retrieved, and which tool it reached for. Whether a data boundary holds depends on paths nobody drew on the diagram. The control does not have a fixed state you can read off a config file.
So a yes-or-no answer about existence has never been weaker evidence than it is right now. "We have a policy" and "the policy holds" have quietly become different claims, and the questionnaire only ever captured the first one.

Here is a composite of what we see, illustrative rather than any one client.
A company scores well on an AI readiness assessment. Every control is marked in place. Then someone tests four of them.
The policy exists and is genuinely well written, but nobody has checked whether the paths it names are the paths the data takes, so it governs a set of tools that is not the set in use. Access is role-based, and the roles were assigned correctly, but the AI agent runs under a service account that inherits a role granted eighteen months ago for a different purpose. Activity is logged, and the logs faithfully record every tool call, but not the retrieved content that triggered them, so nothing in them supports a reconstruction. Reviews happen quarterly, which is true, and the vendor ships changes weekly, which is also true, which is the argument for testing on the system's schedule rather than the calendar's.
Every answer on that questionnaire was honest. Nobody lied. The assessment simply asked a question that existence can satisfy, and existence is not the property anyone actually cares about.

This is where the vocabulary matters, because "AI security assessment" gets used for three things that produce three different deliverables.
Does the control exist? That is governance. You answer it by reading policies, interviewing owners, and mapping to a framework. It is legitimate, necessary work, and it produces a documented position. What it cannot produce is any statement about whether the control functions, because nothing was exercised.
Is the control configured the way it is supposed to be? That is control validation. You answer it by inspecting the running system rather than the document — tenant settings, identity scopes, retention flags, what is actually switched on. This is a real step up, and it catches the gap between the policy and the deployment. It is roughly the difference between a vulnerability assessment and a penetration test, applied to AI. What it still does not tell you is what happens when someone works against the control on purpose.
Does the control hold when someone attacks it? That is adversarial testing. You answer it by attempting the thing the control is supposed to stop, under an agreed scope, and recording what happened. This is the only one of the three that produces evidence in the ordinary sense of the word.
Most programs need all three, and it is fine to buy them separately. The failure is buying the first and believing you received the third, which is how check-the-box testing survives.
If you take one practical thing from this, take this. When you read an assessment deliverable, a finding is evidence only if it has four things attached. Anything short of that is a claim wearing a finding's clothes.
The attempt. What exactly was tried, in enough detail that someone else could do it. "Tested access controls" is not an attempt. "Authenticated as a support-tier service account and requested a record outside its assigned scope" is.
The observed result. What the system did, not what the tester concluded it would do. There is a real difference between "the boundary would likely fail" and "the request returned the record."
The artifact. The log line, the response, the screenshot, the captured traffic. Something a third party can look at without taking anyone's word for it. This is the part most often missing, and its absence is usually the tell.
The conditions. When it was run, against which version, under what scope. A control that held in March against a build that no longer exists is a historical note, not a current assurance.
Run that test on the last AI assessment report you received. Count how many findings have all four. In our experience the number is lower than people expect, and it is lowest in exactly the reports that scored best.

Use them, and be precise about what they do. NIST's AI RMF gives you a governance structure. CSA's AI Controls Matrix gives you a control set to organize around. MITRE ATLAS gives you shared language for adversary techniques, which is genuinely useful when you write a test plan. ISO/IEC 42001 is certifiable, and certification means a management system met a standard — not that any particular control held under attack.
None of them is a test. Mapping your controls to a framework produces a map, and a map is not a report from the territory. That distinction is the whole argument of this post.

Adversarial testing does not eliminate risk, and no assessment of any kind makes a system safe. What it does is convert assumptions into observations, which is a smaller claim and a much more useful one. It is also more expensive and slower than a questionnaire, which is why the honest recommendation is not to test everything. It is to know which of your answers are claims, and to buy evidence for the ones where being wrong would actually hurt.
And to be clear about our own position: independent adversarial testing is what produces this kind of evidence, whoever performs it. Plenty of good firms do it, and the methodology for testing an AI system is not a secret. We think the distinction between asserting and proving is worth insisting on regardless of who you hire.
If you want a self-check, pull your most recent AI security assessment and pick the three controls you would least like to be wrong about. For each one, find the attempt, the observed result, the artifact, and the conditions. If you can find all four, you have evidence. If you cannot, you have a well-organized set of claims — which is a normal place to start, and a good place to stop being satisfied with.
We are a CREST-accredited, ISO/IEC 27001 certified offensive security firm, and every engagement we run is built to produce the four things above. Our AI and ML penetration testing attempts the control rather than reading about it, and hands you the attempt, the result, the artifact, and the conditions. Where the system changes faster than an annual cycle, we run it continuously instead. We are services-first, so we assure the stack you already chose.
If you want a second opinion on an assessment report you already have, get in touch.