Flat vector illustration of an audit period timeline with an orange flag marking an untracked change, beside a binder and a clipboard
TABLE Of CONTENTS

What AI Evidence Does a SOC 2 Auditor Want?

Omair
2026-10-23
14
min read

Your company shipped an AI feature in March. Your SOC 2 Type 2 window runs January to December. In October your auditor asks how changes to that feature were authorized, tested and approved across the period — and someone has to explain that the prompt was edited in a config file eleven times, the model version changed when the provider deprecated the old one, and none of it went through the change process because nobody thought of a prompt as code.

‍

That conversation is the whole of AI evidence in SOC 2. There is no AI criterion to fail. There is an existing criterion that your AI system quietly stopped satisfying.

‍

How an untracked AI configuration change becomes a SOC 2 change management finding

‍

Is there a SOC 2 for AI?

No. There is no AI-specific SOC 2, no AI trust services category, and no separate AI examination type. AI systems are assessed against the same Trust Services Criteria as everything else in scope.

‍

The document in force is the 2017 Trust Services Criteria for Security, Availability, Processing Integrity, Confidentiality, and Privacy (With Revised Points of Focus — 2022). Note the structure, because it is frequently misdescribed: the criteria are still the 2017 criteria, and what was revised in 2022 was the points of focus beneath them.

‍

We looked for AI-specific attestation criteria from the AICPA and did not find any. The AI material the AICPA publishes concerns practitioners using AI in their own work, not AI systems as the subject matter of an attestation.

‍

That does not make every "SOC 2 for AI" offering a fiction, but it does mean you should ask what you are buying. Most are an ordinary SOC 2 examination with AI systems in scope. Some are a SOC 2 examination extended with additional subject matter evaluated against additional criteria — a form the AICPA does provide for. In that second case the additional criteria are somebody else's, typically ISO/IEC 42001 or a control matrix, not the AICPA's. Ask which one is on the engagement letter, and whose criteria are being used.

‍

This is good news and bad news. The good news is that you do not need a new framework. The bad news is that the criteria were written for systems that change through a release process, and a great deal of AI behavior changes through configuration — which is precisely where evidence goes missing. It is the same divergence we keep describing between what a control asserts and what it does.

‍

Which SOC 2 criteria does an AI system actually land under?

Four of the common criteria carry most of the weight. The series titles are worth getting right, because invented criterion numbers circulate freely.

‍

CC8 Change Management is the one that bites hardest, and it contains a single criterion. CC8.1 requires that the entity "authorizes, designs, develops or acquires, configures, documents, tests, approves, and implements changes to infrastructure, data, software, and procedures to meet its objectives" — emphasis ours. Data is enumerated alongside software in the criterion itself. (The AICPA gates the criteria document behind a login; NIST republishes the full criterion text for free in its AICPA crosswalk workbook, which is where these quotations come from.) A system prompt, a fine-tuning dataset, a retrieval corpus and a model version are all within reach of that sentence. If they change outside your change process, you have a CC8.1 gap and not a novel AI problem.

‍

CC9 Risk Mitigation is where a third-party model provider lands. CC9.2 requires that the entity "assesses and manages risks associated with vendors and business partners." Calling an external model API is a vendor relationship, and the evidence your auditor wants is the assessment you performed and the risk decisions you recorded — not the provider's marketing page. What that assessment should actually probe is a question of its own.

‍

CC7 System Operations is where testing evidence lives. CC7.1 covers using detection and monitoring procedures to identify "(1) changes to configurations that result in the introduction of new vulnerabilities, and (2) susceptibilities to newly discovered vulnerabilities." An AI system is a configuration-heavy system, which makes the first limb unusually live.

‍

CC6 Logical and Physical Access Controls covers the boundary. CC6.6 concerns logical access measures protecting against threats from outside system boundaries — and for an agentic system the interesting access is not the user's, it is the identity the agent acts under and what that identity can reach.

‍

For products where the AI output is the service, the Processing Integrity category becomes relevant on its own terms, covering completeness and accuracy of inputs and delivery of output in accordance with specifications. That is a harder commitment to make about a probabilistic system than about a billing engine, and it should be scoped deliberately rather than by default.

‍

Four common criteria that govern AI systems in a SOC 2 examination

‍

What does "evidence over a period" mean for a system that changes weekly?

This is the structural difficulty, and it comes straight from the difference between report types. A Type 1 addresses the description and the suitability of design of controls as of a specified date. A Type 2 addresses design and operating effectiveness throughout the specified period, and the report includes a description of the tests of controls and the results thereof.

‍

"Throughout the specified period" is the phrase to sit with. If your model version changed in April and your prompt changed in July, the auditor is not asking whether your control works today. They are asking whether it worked in April and in July, and the only acceptable answer is a record made at the time.

‍

In practice that means three things. Version the artifacts — prompts, model identifiers, tool permissions and retrieval sources — in the same system of record as code, so a change produces a ticket and an approval rather than a commit nobody reviewed. Define a change taxonomy that says which AI changes are material enough to require testing before release, and apply it consistently, because an inconsistently applied policy is worse evidence than a narrow one. And tie your testing evidence to versions, so that a report can be matched to the configuration it examined. Model drift is not only a quality problem; it is an evidence problem, because the thing you tested in February may not be the thing running in September.

‍

How is a third-party model provider handled?

This is the crux, and it is where most AI SOC 2 conversations go wrong.

‍

When a service organization relies on a subservice organization, its description either carves that provider out of scope or includes it. The carve-out method is the ordinary choice, and it is what almost every company calling a commercial model API will do. Under carve-out, the provider's controls are excluded from the description and from the scope of the service auditor's engagement — the provider sits outside the engagement, and the report says so.

‍

Read that consequence plainly. If you carve out your model provider, your SOC 2 report contains no assurance about that provider, by design. What it does contain is your own CC9.2 vendor risk evidence and your disclosure of the complementary controls you are assuming the provider implements. Those two things are the entire story your report tells about the model — so they are worth doing properly rather than filling in at the end.

‍

Two follow-on points. Your customers reading your report inherit that boundary, so if they are relying on your report for assurance about the model, they are relying on something that is not there. And the controls you own at the integration — what you send, what you log, what you allow the output to trigger — are inside your scope whether or not you tested them. That surface is where data leakage paths actually open up.

‍

Carve-out method: what your SOC 2 report covers and what it does not

‍

A composite example, drawn from patterns rather than any single organization: a Series B SaaS company adds an AI assistant, carves out its model provider, and maps the feature to its existing access and change criteria. During fieldwork the auditor asks for approvals covering the period. The prompt lives in an environment variable, changed by whoever was on call, with no record. The control was well designed. It was never operating.

‍

Does SOC 2 require penetration testing of AI systems?

SOC 2 does not prescribe penetration testing as a named, mandatory procedure. The criteria state outcomes and leave the entity to select the controls that achieve them.

‍

What it does is create places where test evidence is the most natural thing to put. CC7.1 asks you to identify susceptibilities to newly discovered vulnerabilities, and for AI systems the relevant vulnerability classes are not the ones a network scanner reports. CC4.1 contemplates ongoing and/or separate evaluations, and an independent assessment is a separate evaluation. CC9.2 asks you to assess vendor risk, and a scoped test against your own tenant is a stronger assessment than a returned questionnaire.

‍

So the useful framing for a security lead is not "does SOC 2 require this." It is: when the auditor asks how you identified AI-specific weaknesses in your system during the period, what do you hand them? A dated report with reproduction steps answers the question. A policy saying you take AI security seriously does not. This is the same argument as penetration testing for SOC compliance generally, with a scope that most testing programs have not caught up with yet.

‍

One caution on framework mapping. NIST publishes a register of AI Risk Management Framework crosswalks — to ISO/IEC 23894, ISO/IEC 42005, ISO/IEC 42001 and others — and none of them maps to the Trust Services Criteria. The AICPA publishes its own TSC mappings to various frameworks behind a member login, so check that list yourself rather than taking a vendor's word for what exists. Mappings in general circulation are vendors' own constructions. They can be useful working documents; they are not authority, and an auditor is entitled to ignore one.

‍

Frequently asked questions

Does using AI automatically bring it into SOC 2 scope? It depends on whether the AI system is part of the system described in your report. If it processes in-scope data or supports an in-scope commitment, it is in. If it is an internal productivity tool outside the described system, it may not be — but you need to know which, and shadow AI is exactly the case where nobody does.

‍

Do we need ISO/IEC 42001 as well? They answer different questions. ISO/IEC 42001:2023 sets requirements for an AI management system and organizations are certified against it — evidence that governance exists and is managed. SOC 2 reports on controls relevant to the trust services criteria over a period. Neither demonstrates that a given control held under adversarial conditions, which is a third question again.

‍

Is a prompt change a change under CC8.1? CC8.1 names data alongside infrastructure, software and procedures, and a prompt materially determines system behavior. Treating prompt changes as out of scope for change management is a position you would have to defend to your auditor, and we would not want to defend it.

‍

Our model provider has its own SOC 2. Can we rely on it? For its own controls, subject to its scope, period and complementary user entity controls — read all three. It says nothing about your integration, and if you carved the provider out, your report says nothing about the provider.

‍

When should AI testing happen relative to the audit window? Early enough that findings can be remediated and retested inside the period, rather than producing a dated report that documents an unfixed weakness. Testing a week before fieldwork mostly generates evidence of a problem.

‍

ioSENTRIX Can Help

ioSENTRIX is a CREST-accredited, ISO/IEC 27001 certified offensive security firm, and we already produce audit-ready penetration testing deliverables for SOC 2 examinations. Extending that to AI systems means testing the parts a conventional application test does not reach — the prompt and orchestration layer, the retrieval boundary, tool permissions and the identities behind them — and writing it up so it maps cleanly onto the criteria your auditor is working through. Our AI and ML penetration testing produces dated, reproducible findings, which is the form evidence has to take to be worth anything in a Type 2.

‍

If your next window includes an AI feature for the first time, the cheapest hour you will spend is the one that establishes what is in scope and what your change record currently proves. Talk to us.

‍

Keep reading

#
AI Compliance
#
Penetration Testing
Contact us

Similar Blogs

View All