
Three proposals are open on your desk, all titled some version of *AI Security Assessment*, all naming the same system. Two use the phrase *red team* in the methodology. The cheapest is a fifth of the price of the most expensive, and nothing in the documents explains what the extra money buys.
That spread is not vendors being greedy. It is one word covering three different amounts of work.
We separated the three exercises in Prove, Don't Assert: governance asks whether a control exists, validation whether it is configured correctly, adversarial testing whether it holds under attack. Three deliverables, three prices — and the difference between red teaming and AI penetration testing sits inside the third.
NIST drew that line long before AI. SP 800-115 separates *examination* — checking, inspecting, reviewing — from *testing*, exercising an object under specified conditions and comparing actual behavior against expected. Reading and attempting have never cost the same. The buyer's question is not which exercise is best. It is which one each line item actually is.

Five variables move the number, and model size is not one of them.
What the system can do. Scope scales with actions, not parameters. An assistant that retrieves documents and returns text has a narrow surface. An agent with five tools, two of which write to other systems, has five scope items and the chains between them.
How much access the tester gets. Black-box testing against a public endpoint is not more realistic than testing with the system prompt, tool manifests, retrieval configuration, and logs in hand. It is slower: days go into rediscovering architecture you could have handed over.
Whether a safe place to attack exists. Adversarial testing needs somewhere the tester can push, usually a staging replica with synthetic data. Where none exists, building it is often the largest line item — and testing in production under tight rules of engagement trades money for operational risk.
Repetition. These systems are not deterministic. An injection that works once in twenty attempts is still a finding, but establishing "one in twenty" costs twenty attempts, and the report has to state the rate. A screenshot with no rate attached is a demo.
Who is on the bench. Validation can be run by a competent engineer working from a control list; attacking an agent chain is specialist work, and specialist days are the cost. That is usually what a cheap penetration test really is — the bench changed, not the efficiency.
Published prices for the same three words vary widely, which is also true of conventional pentest pricing. So ask for a ratio rather than a price: on comparable engagements, tester-days spent per finding that came back with an artifact attached.
Four conditions. If none holds for a system, validation is genuinely enough this year.
One: the system can take an action you cannot take back. Money moves, data leaves, a record changes, a message goes out under your brand.
Two: the control you rely on is behavioral, not configurational. If a setting settles it, read the setting — that is validation, and it is cheaper. If it holds or fails depending on what the model was asked and what it retrieved, no export answers the question.
Three: the blast radius crosses a boundary you do not own — a third-party tool, an MCP server, another vendor's agent.
Four: someone outside your company will rely on the answer — a customer's security review, a regulator, an insurer, an acquirer in diligence.
Now the commercially inconvenient part. We sell adversarial testing, and it would suit us to say every AI system needs it. It does not. Most generative AI in a mid-market company is a read-only assistant over documents the same staff could already open, and the questions that matter there — who can query it, what ended up in the index, what the vendor retains — are answered by validation plus a data-leakage test. Adversarial days buy reassurance there, not information.

An illustrative composite, not a client. A mid-market SaaS company runs two AI systems: an internal assistant over policy documents with no write access, and a customer-facing agent that issues account credits below a threshold. There is budget for one engagement, and splitting it evenly is the wrong answer.
The assistant fails the first two conditions. It takes no irreversible action, and the controls that matter on it are readable — identity scope, index contents, retention flag, log destination — so validation settles them, with an export apiece as evidence.
The credit agent hits three of the four. It moves money, its authorization control is behavioral, and its tool chain reaches a billing system nobody in security owns. That is where the adversarial days go, on the questions from proving an autonomous action is real: was the instruction genuine, was the action authorized, was it untampered, can it be reconstructed.
Governance produces a control register mapped to a framework — the CSA AI Controls Matrix, 243 control objectives across 18 domains, or the structure in NIST's AI RMF — with a gap list, an owner, and a remediation plan. No artifacts, correctly, because nothing was exercised. A proposal implying otherwise is check-the-box testing in better formatting.
Control validation produces an inventory of what is actually configured, with per-setting evidence and a delta against intended state. It also produces the list buyers overlook: the controls that could not be verified from configuration alone. That list is the scope of your adversarial engagement, which is why validation first is cheaper overall — the logic behind vulnerability assessment before a pentest.
Adversarial testing produces the test cases attempted, including the ones that did not work — a report where everything succeeded is a report with a narrow scope. Each finding carries the attempt, the observed result, the artifact, and the conditions; chains appear as chains; versions and dates are stated; there is a retest. Techniques map to shared vocabulary: MITRE ATLAS for adversary techniques against AI systems, the OWASP GenAI Red Teaming Guide for the four areas a red team should span — model evaluation, implementation testing, infrastructure assessment, runtime behavior. That mapping matters less for what it claims than for what it reveals: the parts nobody covered.

Ask these in writing, and read the answers as part of the proposal.
1. Which of the three exercises is each line item? If the methodology will not say plainly, assume governance priced as testing.
2. What is the unit of scope? "One AI system" is not a scope. An endpoint, an application, an agent and its declared tools — counted, in the document.
3. What access do we grant, and what drops out of scope if we do not?
4. Can we see a redacted sample finding first? This one does most of the work of choosing an AI testing vendor. If the sample arrives with no artifact and no conditions, price is the least of your problems.
5. What does the result expire against? A version, a date, and the changes that reopen it — a new tool, a model upgrade, a prompt edit. If the answer is "annually," you are being priced on the calendar rather than the system's rate of change.
None of this eliminates risk. Validation converts a document into an observation about configuration; adversarial testing converts an assumption into an observation about behavior. Both are smaller claims than "secure," and both are more useful than one.
For a self-check before the next budget cycle, build three columns for every AI system you run: which of the three exercises you last bought, whether the system can take an action you cannot take back, and whether the control you rely on is configurational or behavioral. Rows reading *governance / yes / behavioral* are where the adversarial money belongs; *governance / no / configurational* needs validation first. If you cannot fill in the middle column, that is not a testing problem — it is an inventory problem, and a cheaper one to fix.
We are a CREST-accredited, ISO/IEC 27001 certified offensive security firm, and we scope AI engagements the way this post describes: validation on every engagement, adversarial days where an irreversible action or a behavioral control justifies them. Our AI and ML penetration testing attempts the control rather than reading about it, and runs continuously where the system changes faster than an annual cycle. We are services-first, so we assure the stack you already chose.
If you have proposals in hand and want a second opinion on what they are actually offering, get in touch.