Flat vector illustration of NIST AI RMF evidence: an assessment report carrying pass and fail marks and an orange signed seal, beside four tiles representing the govern, map, measure and manage functions
TABLE Of CONTENTS

What Evidence Maps to Each NIST AI RMF Function?

Omair
2026-11-06
24
min read

Your auditor asks whether you are aligned with the NIST AI Risk Management Framework. You say yes — you have the crosswalk spreadsheet, the policy set references it, and someone has mapped your controls to the subcategories.

‍

Then they ask the second question. Show me the evidence for MEASURE 2.7.

‍

If the honest answer is a policy document, a diagram and a vendor questionnaire, you are not aligned with MEASURE 2.7. You are aligned with the idea of it. MEASURE 2.7 says "AI system security and resilience — as identified in the MAP function — are evaluated and documented." Evaluated is a verb with an output. A policy stating that evaluation happens is not that output.

‍

This is the most common gap we see in AI governance programs, and it is not a failure of effort. It is a structural feature of the framework, and NIST says so itself.

‍

What is the NIST AI RMF, and what does it actually ask for?

The NIST AI Risk Management Framework is a voluntary catalog of risk-management outcomes for AI systems, organized into four functions — GOVERN, MAP, MEASURE and MANAGE — that break down into 72 subcategories describing what a mature program achieves, without prescribing how to achieve it. It is NIST AI 100-1, version 1.0, published January 26, 2023.

‍

The framework is unusually candid about what it is not. Its own Appendix D lists, among its design attributes, "Be outcome-focused and non-prescriptive. The Framework should provide a catalog of outcomes and approaches rather than prescribe one-size-fits-all requirements." Section 5 adds: "Actions do not constitute a checklist, nor are they necessarily an ordered set of steps." The companion Playbook repeats it: "The Playbook is neither a checklist nor set of steps to be followed in its entirety."

‍

So the framework tells you what good looks like. It does not tell you what to put in the evidence folder. That translation is the reader's job, and this post is an attempt at it.

‍

One currency note: AI RMF 1.0 remains the current released version, but NIST states that "The AI RMF 1.0 is being revised as part of the White House AI Action Plan." No draft of a revised version has been published as of this writing, so 1.0 is what you map against today.

‍

The four NIST AI RMF functions as a pipeline: GOVERN with 19 subcategories, MAP with 18, MEASURE with 22 and MANAGE with 13, and what each hands the next

‍

Why doesn't framework alignment produce evidence?

Because the framework is deliberately not a control catalog, and the gap between the two has been measured.

‍

The Cloud Security Alliance's AI Controls Matrix v1.1, released June 2026, contains 247 control objectives across 18 domains. In a July 2026 post announcing that release, CSA published a gap analysis of those 247 controls against each major target framework. Against the NIST AI RMF and the Generative AI Profile together, 18 controls (7%) had no gap, 119 (48%) had a partial gap, and 110 (45%) had a full gap. For comparison, against ISO/IEC 42001 the full-gap figure was 2%.

‍

CSA's own reading of that result is the important part, and it is not a criticism: "The 45% Full Gap with NIST frameworks is expected and by design. NIST AI RMF and AI 600-1 are risk frameworks, not control catalogs. They focus on AI lifecycle governance and intentionally delegate infrastructure security to companion frameworks like NIST SP 800-53."

‍

Read that alongside the framework's own non-prescriptive design attribute and the conclusion is clear. Nearly half of what an implementable control catalog specifies is simply not in the AI RMF, because the AI RMF was never trying to specify it. An organization that maps its controls to subcategories and stops has done real and useful work on the governance half, and has produced no assurance evidence at all.

‍

Which is the same argument ioSENTRIX makes about questionnaires: governance establishes that a control exists, and assurance establishes that it works. A crosswalk is a map. A map is not a test.

‍

What evidence maps to GOVERN?

GOVERN is where documents are legitimately the evidence — it is the function about policies, roles and accountability, and a policy genuinely is the artifact. Nineteen subcategories sit here, and for most of them a signed, versioned, dated document with an accountable owner is the right answer.

‍

Three exceptions matter to a tester.

‍

GOVERN 1.6 — "Mechanisms are in place to inventory AI systems and are resourced according to organizational risk priorities." The evidence is not the inventory. The evidence is a discovery exercise that found something the inventory did not contain. An inventory nobody has tried to break is an assumption. The Generative AI Profile is more specific than the framework here: its suggested action GV-1.6-003 lists what an entry should carry, including "Data provenance information (e.g., source, signatures, versioning, watermarks)" and "Underlying foundation models, versions of underlying models, and access modes." Our experience is that the AI hiding inside already-approved SaaS is what a first discovery pass turns up, and it is rarely on the list.

‍

GOVERN 4.3 — "Organizational practices are in place to enable AI testing, identification of incidents, and information sharing." The Playbook's suggested actions include "Establish policies and procedures to facilitate and equip AI system testing." Equip is doing real work in that sentence. The checkable evidence is a scoping document, a test environment that exists, and rules of engagement somebody has actually signed — not a paragraph asserting that testing is supported.

‍

GOVERN 6.1 and 6.2 — third-party risk, and contingency processes for failures or incidents in third-party data or AI systems "deemed to be high-risk." The evidence for 6.2 is a tested contingency, which means a documented exercise where a third-party model or data feed was withdrawn or degraded and the system's behavior was observed. A runbook is the plan; the exercise log is the evidence.

‍

What evidence maps to MAP?

MAP is context and categorization — eighteen subcategories. Most of its output is documentation, but two produce artifacts a tester consumes directly, and one produces an artifact a tester makes.

‍

MAP 2.3 — "Scientific integrity and TEVV considerations are identified and documented, including those related to experimental design, data collection and selection… system trustworthiness, and construct validation." Among its suggested actions: "Identify testing modules that can be incorporated throughout the AI lifecycle, and verify that processes enable corroboration by independent evaluators." That phrase — corroboration by independent evaluators — is the framework inviting exactly the kind of external testing this post is about.

‍

MAP 4.1 and 4.2 sit under the category "Risks and benefits are mapped for all components of the AI system including third-party software and data," and 4.2 reads: "Internal risk controls for components of the AI system, including third-party AI technologies, are identified and documented." This is where a component inventory becomes a security artifact rather than a procurement one, and it is the natural home for an AI bill of materials with integrity evidence attached — provided the fields in it are the checkable kind. The same category is why securing the AI supply chain belongs in MAP rather than in procurement.

‍

The artifact a tester makes here is the threat model. MAP asks what could go wrong across the system's components and context; AI system threat modeling is how that question gets answered in a form that later drives test cases. MITRE ATLAS is the right external reference for adversary behavior against AI systems — note that it now ships monthly content updates, so date-stamp the version you worked from; the release current at the time of writing is content version 2026.09, dated 2026-09-15.

‍

The link between MAP and MEASURE is load-bearing and easy to miss: eight MEASURE subcategories contain the clause "as identified in the MAP function" — 2.4, 2.6, 2.7, 2.8, 2.9, 2.10, 2.11 and 2.12. If your MAP output does not enumerate the risks, the corresponding MEASURE evidence has nothing to be evidence of.

‍

What evidence maps to MEASURE?

This is where the assurance work lives. MEASURE has 22 subcategories — the largest function — and the framework's summary of what completing it means is the clearest statement of intent in the document: "After completing the MEASURE function, objective, repeatable, or scalable test, evaluation, verification, and validation (TEVV) processes including metrics, methods, and methodologies are in place, followed, and documented."

‍

Four subcategories carry most of the weight.

‍

MEASURE 2.7 — "AI system security and resilience — as identified in the MAP function — are evaluated and documented." This is the closest thing in the AI RMF to a penetration-test deliverable, and the Playbook's suggested actions read like a test plan. Verbatim: "Establish and track AI system security tests and metrics (e.g., red-teaming activities, frequency and rate of anomalous events, system down-time, incident response times, time-to-bypass, etc.)"; "Use red-team exercises to actively test the system under adversarial or stress conditions, measure system response, assess failure modes or determine if system can return to normal function after an unexpected adverse event"; and the one worth framing, "Use red-teaming exercises to evaluate potential mismatches between claimed and actual system performance."

‍

The Generative AI Profile (NIST AI 600-1, July 2024) goes further and names the attack classes. Its action MS-2.7-007: "Perform AI red-teaming to assess resilience against: Abuse to facilitate attacks on other systems (e.g., malicious code generation, enhanced phishing content), GAI attacks (e.g., prompt injection), ML attacks (e.g., adversarial examples/prompts, data poisoning, membership inference, model extraction, sponge examples)." Two adjacent actions are the ones programs forget: MS-2.7-008, "Verify fine-tuning does not compromise safety and security controls," and MS-2.7-009, "Regularly assess and verify that security measures remain effective and have not been compromised."

‍

The evidence for MEASURE 2.7 is therefore a dated report with reproducible findings, the scope and conditions under which testing ran, the attack classes attempted and their outcomes, and the residual risk that remains — including what could not be tested. The current OWASP Top 10 for LLM Applications, reranked in August 2026, is a reasonable coverage checklist for the application layer, with Prompt Injection at LLM01, Sensitive Information Disclosure at LLM02 and Excessive Agency now at LLM03. NIST AI 100-2e2025, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations (March 2025), is the reference for the model layer.

‍

MEASURE 1.3 — "Internal experts who did not serve as front-line developers for the system and/or independent assessors are involved in regular assessments and updates." This subcategory is the framework asking for independence explicitly, and its suggested actions sharpen it: "Utilize separate testing teams… to enable independent decisions and course-correction for AI systems" and "Assess independence and stature of TEVV and oversight AI actors, to ensure they have the required levels of independence and resources to perform assurance, compliance, and feedback tasks effectively." The MEASURE narrative gives the reason: "Processes for independent review can improve the effectiveness of testing and can mitigate internal biases and potential conflicts of interest."

‍

The evidence is an engagement record from a party that does not report to the team that built the system. Note that this can be an internal team, provided the independence is real and documented — the framework does not require an external firm, it requires separation.

‍

MEASURE 2.1 — "Test sets, metrics, and details about the tools used during TEVV are documented." This is a reproducibility requirement, and it is the one most often failed by otherwise good testing. An assessment that reports findings without the prompts, payloads, datasets, tool versions and configuration that produced them cannot be re-run, which means the fix cannot be verified either.

‍

MEASURE 2.13 — "Effectiveness of the employed TEVV metrics and processes in the MEASURE function are evaluated and documented." NIST is asking you to test your testing. The practical evidence is a comparison: findings from a new method or a new tester against what the previous approach found, with the delta recorded. Rotating who tests is one way this evidence gets produced as a side effect.

‍

Also here: MEASURE 2.5 (validity and reliability demonstrated, with generalizability limits documented), MEASURE 2.6 (safety risks evaluated regularly, with the system "demonstrated to be safe" against a stated risk tolerance, and able to fail safely), MEASURE 2.10 (privacy risk examined — where AI data leakage testing produces the artifact), and MEASURE 2.4 (functionality and behavior monitored in production, which is a continuous testing argument rather than an annual one).

‍

What evidence maps to MANAGE?

MANAGE is the thirteen subcategories about what you do with what you found. The evidence is decisions and their consequences, which makes it the function most improved by having real findings to manage.

‍

MANAGE 1.1 — "A determination is made as to whether the AI system achieves its intended purposes and stated objectives and whether its development or deployment should proceed." The evidence is a dated go/no-go decision with a named owner and the assessment it was based on.

‍

MANAGE 1.4 — "Negative residual risks (defined as the sum of all unmitigated risks) to both downstream acquirers of AI systems and end users are documented." This is the subcategory that makes an honest test report more valuable than a clean one. A report with no residual risk section cannot satisfy MANAGE 1.4. It is also where the distinction between a tested and failing control and an untested one matters, which is the argument for scoring unknowns and failures differently in a risk register.

‍

MANAGE 3.2 — "Pre-trained models which are used for development are monitored as part of AI system regular monitoring and maintenance." The evidence is a monitoring record showing that a supplier-side model change was detected. Not a policy saying it would be.

‍

MANAGE 4.1 and 4.3 — post-deployment monitoring plans implemented, and incidents and errors tracked, responded to and documented. The evidence is an incident record, including a retrospective on something that actually happened, however small.

‍

MANAGE 2.4 — mechanisms in place "to supersede, disengage, or deactivate AI systems that demonstrate performance or outcomes inconsistent with intended use." The evidence is a kill-switch that somebody has pulled in a test window and timed.

‍

Six NIST AI RMF subcategories mapped to the evidence artifact each one requires: GOVERN 1.6, GOVERN 4.3, MAP 2.3, MEASURE 2.7, MEASURE 1.3 and MANAGE 1.4

‍

Where the mapping breaks down in practice

A synthetic composite, for illustration — not a real company or client.

‍

A healthcare SaaS provider builds a genuinely careful AI RMF program for a clinical-documentation assistant. Policies map to all 19 GOVERN subcategories. A risk register enumerates 40 risks against MAP. A model-evaluation suite runs accuracy, hallucination and bias metrics on every release and files results against MEASURE. An incident process exists and has been rehearsed.

‍

A customer's security team asks for the MEASURE 2.7 evidence. What comes back is the model-evaluation suite output.

‍

It is real testing, carefully done, and it measures the wrong axis. Accuracy, hallucination rate and bias are quality properties. MEASURE 2.7 asks about security and resilience: whether the assistant can be induced to disclose another patient's record through the retrieval layer, whether the agent's document-export tool can be driven by injected content, whether the guardrails hold under adversarial input rather than under representative input, whether the system fails safely when the model provider degrades. None of those appear in a quality suite, because a quality suite tests the system doing its job — and an attacker is not interested in the system doing its job.

‍

The program was not lazy. It conflated evaluated with evaluated for security, which the subcategory's own wording distinguishes. This is the same failure mode as a model that is compliant without being correct: the measurement ran, and it measured something other than the thing at risk.

‍

Does NIST AI RMF certification exist?

No. NIST publishes the AI RMF for voluntary use, and publishes no conformity-assessment, certification or attestation scheme against it — none appears among its AI RMF resources, and the framework's own text disclaims any such role. NIST's FAQ puts the voluntariness as a direct question and answer: "Will private or public sector organizations be required to use the Framework? No. NIST has produced the AI RMF as a voluntary Framework." The framework text says it is "intended to be voluntary, rights-preserving, non-sector-specific, and use-case agnostic." There is no NIST-recognized basis for a claim of being "NIST AI RMF certified," and a vendor making that claim is describing something that does not exist.

‍

ISO/IEC 42001 is the deliberate contrast, and the contrast is instructive rather than a point in either standard's favor. ISO/IEC 42001:2023 is certifiable, through a full conformity-assessment stack: a requirements standard, ISO/IEC 42006:2025 governing the bodies that audit against it, national accreditation bodies above those, and certificates issued by independent third parties. ISO itself issues none of them — in ISO's words, "ISO does not perform certification or issue certificates."

‍

But be precise about what a 42001 certificate attests. It certifies an AI management system — policies, roles, risk processes — not that any particular model is safe, accurate, or resistant to attack. So the three states are genuinely distinct, and conflating them is how programs end up surprised:

‍

Aligned with the AI RMF means you have mapped outcomes you intend to achieve. Certified to ISO/IEC 42001 means an accredited body confirmed your management system meets the requirements. Tested means somebody adversarial tried to break a specific control on a specific system and wrote down what happened. You can hold the first two and have never done the third.

‍

Three distinct states compared: aligned to the NIST AI RMF, certified to ISO/IEC 42001, and adversarially tested

‍

Frequently asked questions

Is NIST AI RMF mandatory? No. NIST states plainly that it produced the AI RMF as a voluntary framework, and the framework describes itself as voluntary and non-prescriptive. Some organizations adopt it because customers or insurers ask for it, and some sector regulators reference it, but NIST imposes no obligation and runs no compliance regime.

‍

What is the difference between NIST AI RMF and ISO/IEC 42001? The AI RMF is a voluntary catalog of risk-management outcomes with no certification scheme. ISO/IEC 42001:2023 is a management-system standard that accredited third parties certify against, governed for auditors by ISO/IEC 42006:2025. They answer different questions: the AI RMF asks what risks you manage and how you know; 42001 asks whether you run a conforming management system. Neither one certifies that a given AI system withstood an attack.

‍

Which NIST AI RMF subcategory does a penetration test satisfy? Primarily MEASURE 2.7, "AI system security and resilience… are evaluated and documented." A well-scoped engagement also produces evidence for MEASURE 1.3 (independent assessors involved), MEASURE 2.1 (test sets, metrics and tools documented), MANAGE 1.4 (residual risk documented) and, where testing surfaces something new, MANAGE 2.3 and MANAGE 4.3. One engagement does not satisfy the framework, and no single engagement covers all 72 subcategories.

‍

What does TEVV mean in the NIST AI RMF? Test, evaluation, verification and validation. The framework treats it as a set of tasks performed across the whole AI lifecycle rather than as one of the four functions, and it appears explicitly in MAP 1.1, MAP 2.3, MEASURE 2.1 and MEASURE 2.13. Notably, the framework says "Ideally, AI actors carrying out verification and validation tasks are distinct from those who perform test and evaluation actions."

‍

Do we need the Generative AI Profile as well? If you operate generative or foundation-model systems, it is the more actionable document. NIST AI 600-1 (July 2024) defines 12 GAI-specific risks and maps suggested actions to AI RMF subcategories with their own identifiers, and its MEASURE 2.7 actions name the attack classes to test for. One caveat worth knowing: it was issued pursuant to an executive order that has since been revoked; NIST has not withdrawn or superseded the publication, and it remains listed as a current resource.

‍

ioSENTRIX can help

ioSENTRIX is a CREST-accredited, ISO/IEC 27001 certified offensive security firm. Our AI and ML penetration testing engagements are built to produce the evidence the MEASURE function asks for and the MANAGE function consumes: a dated report scoped to the risks your MAP output identified, the attack classes attempted and their outcomes, the test sets, payloads and tool versions needed to re-run any finding, and an explicit residual-risk section covering what remains and what could not be tested. Because we do not report to the team that built the system, the engagement also stands as MEASURE 1.3 evidence. We assure the stack you have chosen rather than selling a competing platform, and independent adversarial testing produces this evidence whoever performs it.

‍

If you want to check your own position before talking to anyone, take the four subcategories this post treats as load-bearing — GOVERN 1.6, MAP 2.3, MEASURE 2.7 and MANAGE 1.4 — and for each one name the file you would hand an auditor. If any of the four is a policy rather than a result, that is where the gap is.

‍

Keep reading

#
AI Risk Assessment
#
AI Compliance
Contact us

Similar Blogs

View All