
A national competent authority sends your company a reasoned request. It wants you to demonstrate that your high-risk AI system meets the cybersecurity requirement in Article 15. You have an ISO/IEC 42001 certificate, a model provider's security whitepaper, and a risk register. Does any of that answer the question that was actually asked?
Probably not. Those documents describe how you govern AI risk. The request is about whether a specific system resisted specific attacks. Those are different claims, and only one of them is settled by a test.

Article 15 requires high-risk AI systems to achieve an appropriate level of accuracy, robustness and cybersecurity and to perform consistently in those respects throughout their lifecycle. It states an outcome, not a method — it names no test, tool or technique. The one procedural duty it adds is in Article 15(3): accuracy levels and the relevant accuracy metrics must be declared in the instructions for use.
The operative cybersecurity wording is Article 15(5), and most summaries soften it:
"High-risk AI systems shall be resilient against attempts by unauthorised third parties to alter their use, outputs or performance by exploiting system vulnerabilities."
The same paragraph names the AI-specific vulnerabilities that technical solutions must address, where appropriate: attacks trying to manipulate the training data set (data poisoning), or pre-trained components used in training (model poisoning), inputs designed to cause the model to make a mistake (adversarial examples or model evasion), confidentiality attacks, and model flaws. Five named vulnerability classes — and note that the last is not an attack at all, but a weakness you are expected to go looking for. The qualifier matters too: the obligation is proportionate to the circumstances and the risks, not a fixed checklist to run in every case.
Two scoping points get lost constantly. Article 15 applies to high-risk AI systems as defined in Article 6, not to everything with a model behind it, which is why knowing what you have actually deployed comes first. And the obligation falls on the provider, not the deployer: Article 16(a) puts compliance with Section 2 on providers, while Article 26 gives deployers a narrower set of duties around using the system as instructed.
No. The phrase "penetration testing" does not appear anywhere in the AI Act — not in the articles, not in the annexes, not in the recitals. Anyone telling you Article 15 mandates a pen test is selling something.
What the Act requires is a demonstrable outcome plus documentation. Article 16(k) obliges providers, on a reasoned request from a national competent authority, to "demonstrate the conformity of the high-risk AI system with the requirements set out in Section 2." Article 11 requires technical documentation drawn up before the system is placed on the market or put into service, kept up to date, and containing at minimum the elements in Annex IV. Small and mid-sized providers get a concession on form, not on substance: SMEs, start-ups and small mid-caps may supply those elements in a simplified manner on a Commission-issued template that notified bodies must accept. The content still has to be there.
That is where testing evidence lands, and Annex IV is far more specific than Article 15 itself. Point 2(g) requires the validation and testing procedures used, the metrics used to measure accuracy and robustness, and — the phrase to underline — "test logs and all test reports dated and signed by the responsible persons, including with regard to pre-determined changes as referred to under point (f)." That closing cross-reference is what gives the phrase teeth: point (f) requires pre-determined changes to be described in advance, not reconstructed afterwards. Point 2(h) separately requires the cybersecurity measures put in place.
So the honest framing: the law demands signed, dated, retained test reports evidencing resilience, and leaves the method to you. Adversarial testing is one way to produce that evidence. It is not a named legal requirement, and the difference between red teaming and a structured assessment matters when you decide which one to commission.
One adjacent provision is often misquoted into this discussion. Article 55(1)(a) does require adversarial testing by name — but it applies to general-purpose AI models with systemic risk, not to high-risk AI systems under Article 15.

The high-risk timeline moved in 2026, and a great deal of published commentary is still running the old dates.
Regulation (EU) 2026/1744, adopted 8 July 2026 and in force from 27 July 2026, amended the application timetable for the high-risk chapter. On the European Commission's current implementation timeline, obligations for stand-alone high-risk AI systems listed in Annex III apply from 2 December 2027, and for high-risk systems embedded in regulated products under Annex I from 2 August 2028. The earlier high-risk dates — 2 August 2026 and 2 August 2027 — are superseded. Read that narrowly: the Act's other 2 August 2026 obligations still stand, and the prohibitions and general-purpose AI model rules took effect earlier still. The Commission tied the change to giving companies the implementation support — chiefly standards — that did not yet exist.
That is not spare time. Annex IV documentation has to exist before a system is placed on the market, so the window is when the evidence gets generated.
Not yet — and in fact not for anything. Article 40(1) grants a presumption of conformity to high-risk AI systems or general-purpose AI models conforming to harmonised standards, but only where the references have been published in the Official Journal. A standard existing is not enough; citation is the trigger.
EN 18286:2026 is the first European standard supporting the AI Act, and it addresses Article 17 — quality management systems. It does not cover Article 15 accuracy, robustness or cybersecurity, and CEN-CENELEC has described a broader suite still to come. It is also not yet cited in the Official Journal: as things stand, no AI Act harmonised standard is, so no presumption of conformity is available to anyone for any requirement.
The practical consequence is that there is no safe-harbour checklist to point at for cybersecurity. Conformity has to be argued from your own evidence. This is the pattern we keep running into: alignment with a framework is not the same thing as a control holding under attack. ISO/IEC 42001:2023 is genuinely valuable, and what it certifies is an AI management system. It does not evidence that a specific model resisted a specific attack, which is the gap a test report fills.
An Article 15 evidence report is a dated, signed technical report that states what was tested, which of the five named vulnerability classes each test exercised, what the system did, and what weakness remains. Most security summaries we are shown contain the first of those and none of the rest. Five things make one hold up.
Scope tied to the declared system. The boundary in the report has to be the boundary in your Annex IV description — model version, prompt and orchestration layer, retrieval sources, the tools an agent may call, and the identities it uses. A report on "the chatbot" cannot evidence a system the documentation defines more broadly, and non-human identities are the part most often left outside the line.
Coverage mapped to the five named classes. Data poisoning, model poisoning, adversarial examples and model evasion, confidentiality attacks, and model flaws. Where a class is out of scope — you did not train the model, so training-data poisoning sits largely with the model provider — say so and say why. A documented, reasoned exclusion is evidence. Silence is a gap.
Reproducible method, not a score. Each finding needs the input, the observed output, and the conditions under which it reproduced. A severity rating with no reproduction steps is an assertion, and reproduction is what makes retesting meaningful when the model version changes.
Results against the specific claim. "The guardrail blocked 47 of 52 attempts, and the 5 that succeeded used this technique" is evidence. "The system is secure" is not, and no responsible report should contain it.
Signature, date and change linkage. Annex IV asks for dated and signed reports "including with regard to pre-determined changes." A system whose prompt, model version or tool permissions change monthly needs its evidence tied to a version and a defined re-test trigger — the same discipline model drift already demands.

A composite example, drawn from patterns we see rather than from any single organization: a mid-market lender deploys a credit-decisioning assistant. The team holds an ISO/IEC 42001 certificate and a model card from its model provider. Under test, the retrieval layer returns documents from a shared index that includes other tenants' uploaded files — a confidentiality attack path neither artifact would have surfaced, because neither artifact was produced by attacking the system. The certificate was accurate. It was answering a different question.
They give the report its spine, not its authority. The NIST AI Risk Management Framework 1.0 organizes work under GOVERN, MAP, MEASURE and MANAGE; filing findings under MEASURE makes a technical report legible to a governance audience. MITRE ATLAS supplies adversary tactics and techniques for AI systems, turning "we tested prompt injection" into a traceable technique reference. The OWASP Top 10 for LLM Applications gives a shared vocabulary for application-layer findings. None is cited in the Official Journal, so none creates a presumption of conformity — and a structured AI security assessment is what turns that structure into findings.
Does Article 15 apply to a company that only uses a high-risk AI system? The Article 15 requirements fall on the provider under Article 16(a). Deployers have separate obligations under Article 26 — mainly to use the system as instructed, ensure human oversight, and retain logs. A deployer can become a provider in some circumstances, such as putting its own name on a system or substantially modifying it, so check that against Article 25 rather than assuming.
Does an ISO/IEC 42001 certificate satisfy Article 15? No. ISO/IEC 42001:2023 certifies an AI management system — how an organization governs AI. Article 15 concerns the technical resilience of a specific system. The certificate is good evidence of governance and no evidence of resilience under attack.
Is a model provider's SOC 2 report enough for a model we call via an API? Not on its own. It describes controls at that provider, while your Annex IV documentation covers the system you place on the market — including how you integrate the model and what the system may do with its answer.
How often does testing need to be repeated? The Act sets no interval. Annex IV's reference to "pre-determined changes" is the anchor: define which changes — model version, prompt, tool permissions, retrieval sources — trigger a retest, and record that rule alongside the report.
What if we cannot test one of the named vulnerability classes? Document the reasoning and who holds the risk. For a model you did not train, training-data poisoning sits largely with the provider, and the evidence is contractual rather than a test you run. An unexamined gap is a finding; a reasoned exclusion is part of the record.
ioSENTRIX is a CREST-accredited, ISO/IEC 27001 certified offensive security firm. Our AI and ML penetration testing is built to produce the artifact Annex IV asks for: a dated, signed report with reproducible findings, scoped to the system as your technical documentation defines it, and mapped to the five vulnerability classes Article 15(5) names. We assure the stack you chose rather than selling one that competes with it, and we say plainly what held and what did not. Independent adversarial testing is the position that produces this evidence, whoever performs it; what matters is that somebody does it before a competent authority asks.
If you are working toward the December 2027 date, the useful first step is smaller than a full engagement: a scoping conversation about which of your AI systems are likely to be high-risk, and what evidence you could produce today. Talk to us about where your documentation currently stops.