Compliance

DORA Compliance for AI Systems: Fintech Security Guide 2026

BT

BeyondScale Team

AI Security Team

17 min read

With DORA enforcement underway since January 2025 and the EU AI Act high-risk deadline arriving August 2, 2026, EU financial institutions face a dual regulatory crunch that most compliance programs are not prepared for. DORA compliance requirements apply directly to AI and LLM systems, yet the vast majority of financial entities have yet to classify their AI models within their ICT risk management frameworks, negotiate DORA-compliant addenda with LLM vendors, or scope AI into their TLPT programs.

This guide covers what DORA compliance means specifically for AI and LLM deployments: how to integrate AI into ICT risk management, what threat-led penetration testing looks like for LLMs, which AI failures trigger major incident reporting, and how to build a compliant third-party AI vendor register before the August deadline.

Key Takeaways

    • DORA treats AI and LLM systems as full ICT systems, subject to all five pillars: risk management, resilience testing, incident reporting, third-party risk, and information sharing.
    • The August 2, 2026 EU AI Act high-risk enforcement date coincides with active DORA enforcement, creating compounded obligations for credit scoring, insurance pricing, and other regulated AI use cases.
    • Standard LLM vendor terms (OpenAI, Anthropic, Google Gemini) do not meet DORA Article 30(3) requirements. Renegotiation or documented compensating controls are required for every AI vendor supporting a critical function.
    • Only 6.5% of approximately 1,000 firms passed all 116 ESA data quality checks in the 2024 Register of Information dry-run, indicating widespread gaps in AI vendor documentation.
    • TLPT scope must include any LLM supporting a critical or important function. External threat intelligence providers must produce AI-specific target threat intelligence, which requires specialized competency most traditional TI providers lack.
    • A supply chain attack on an LLM gateway (like the March 2026 LiteLLM compromise, which exposed 40,000+ AI pipelines in approximately 40 minutes) triggers DORA major incident reporting independently through the data loss criterion.

What DORA Means for AI and LLM Systems

DORA (Regulation (EU) 2022/2554) entered full application on January 17, 2025, covering approximately 22,000 EU financial entities including banks, investment firms, insurance undertakings, payment institutions, and trading venues. AI systems are not exempt and are not addressed in a separate annex. They fall under the existing ICT governance framework as ICT systems, assets, and third-party services.

In January 2026, Germany's BaFin issued non-binding guidance explicitly confirming that generative AI and LLMs must be integrated into DORA-compliant ICT risk management. The BaFin guidance addressed a gap that many compliance teams had hoped to use: the assumption that AI was not yet formally in scope. It is.

ICT Risk Management (Articles 5-16): Financial entities must identify, protect against, detect, respond to, and recover from all ICT risks. For AI systems, this includes model availability failures, hallucinations affecting decision integrity, model drift causing behavioral change over time, adversarial manipulation of model inputs or outputs, and unauthorized model access or exfiltration. Risk classification decisions made before AI deployment are mandatory. An LLM integrated into a loan processing workflow has a fundamentally different ICT risk profile than one used as an internal knowledge base, and DORA requires documented evidence that the distinction has been assessed and governed.

Register of Information (Article 28(3)): Governed by ITS (EU) 2024/2956, which defines 15 inter-linked templates in xBRL-CSV format. Every LLM vendor with a contractual arrangement must appear in the register. The 2024 ESA dry-run exercise found that only 6.5% of nearly 1,000 firms passed all 116 data quality checks on first submission. Shadow AI deployments, where business units adopt LLM tools outside procurement, constitute an Article 28 violation before any contract deficiency is even assessed.

TLPT (Articles 26-27): Threat-led penetration testing is mandatory for systemically important institutions and NCA-designated entities. AI systems supporting critical or important functions are in scope. Section 3 covers TLPT in detail.

Incident Reporting (Articles 17-23): LLM failures that meet classification criteria trigger binding 4-hour initial notification requirements. Section 4 covers classification criteria and how specific AI failure modes map to them.

Third-Party Risk (Articles 28-44): All LLM vendors are third-party ICT providers under DORA, regardless of whether they appear on the CTPP designation list. Section 5 covers the vendor risk obligations.

The August 2026 Compliance Crunch

The EU AI Act's high-risk enforcement deadline of August 2, 2026 creates a compressed window for financial services firms that deploy AI systems classified as high-risk under Annex III. As of this writing, no formal delay has been enacted.

Under EU AI Act Annex III, the following are high-risk AI systems in financial services:

  • AI systems that evaluate the creditworthiness of natural persons or establish credit scores
  • AI systems used for risk assessment and pricing in life and health insurance
  • AI systems evaluating eligibility for essential services
The EBA's November 2025 mapping exercise found that the majority of AI use cases at EBA-supervised institutions fall into the high-risk category. This means most production AI deployments at EU banks are now subject to both DORA (from January 2025) and the EU AI Act (from August 2026) simultaneously.

EU AI Act Article 9(10) explicitly permits integrating AI risk management into DORA's existing ICT risk management procedures. This is the recommended approach: a single unified risk register that addresses both DORA ICT risk and EU AI Act Article 9 requirements for each AI system, rather than running two parallel compliance programs.

By August 2, 2026, high-risk AI systems must have:

  • Documented risk management frameworks with AI-specific controls
  • Data governance protocols including training data provenance documentation
  • Complete audit trails for all AI decisions, stored outside the application layer per AI Act Article 19
  • Registration in the EU AI database
  • Ongoing monitoring for accuracy, robustness, and cybersecurity (the AI Act uses "robustness" in its own technical sense, referring to resistance to adversarial inputs, not operational reliability)
Penalties for EU AI Act non-compliance reach €15M or 3% of global annual turnover for high-risk system violations, and up to €30M or 7% for prohibited AI systems.

The practical implication for any institution reading this in early July 2026: less than four weeks remain before the AI Act enforcement date. Entities that have deferred AI governance work must prioritize the combined DORA and AI Act documentation program immediately.

TLPT for LLMs: Adapting TIBER-EU Methodology

Threat-Led Penetration Testing is DORA's most demanding requirement. TLPT applies to globally systemically important institutions automatically, and to other entities at NCA discretion based on systemic character and ICT risk profile. The methodology aligns to TIBER-EU, which was updated in February 2025 specifically to align with DORA's RTS requirements.

Core TLPT requirements:

  • Frequency: Minimum every three years
  • Environment: Live production systems only, not test environments
  • Threat intelligence: An external, independent threat intelligence provider is required for every exercise (internal TI is not permitted)
  • Blue Team awareness: The internal detection and response team must be unaware until the post-exercise debrief
  • Purple teaming: Compulsory, integrated into the regulation
  • Team rotation: Internal red teams permitted for two of three cycles; the third cycle requires an external red team
  • Cost and duration: EUR 200,000-500,000 per exercise; 3-6 months duration
For AI-enabled financial institutions, TLPT scope determination comes down to one question: which AI systems support a critical or important function? The February 2025 TIBER-EU update broadened scope from isolated technical systems to "people, processes, technology across the organization." Under this framing, any LLM that an employee or customer interacts with as part of a critical process is in TLPT scope.

The external threat intelligence provider must produce AI-specific Target Threat Intelligence. This requires the TI provider to have competency in AI attack vectors, not just traditional financial sector threat actors. This is a meaningful constraint: the pool of threat intelligence providers with documented AI/LLM TTI capability is small.

Mapping OWASP LLM Top 10 to TIBER-EU scenarios:

The OWASP Top 10 for LLM Applications 2025 provides the foundational taxonomy of AI attack vectors, and each maps to specific financial services threat scenarios for TIBER-EU planning:

LLM01: Prompt Injection requires adversarial input scenarios in TLPT scope. A 2025 study by NTU, IBM Research, and UIUC found direct prompt injection attacks against AI trading agents succeeded over 79% of the time, with indirect injection succeeding between 41.7% and 68.2% across 3,168 simulated attacks. The researchers documented "stealthy parasitism," where an agent completes legitimate tasks while executing hidden attacker objectives, including unauthorized payment authorization and order routing manipulation. In a public red-team competition, 60,000+ out of 1.8 million injection attempts succeeded. Any LLM processing external inputs (emails, documents, web content, customer queries) in connection with a critical function requires a prompt injection scenario in TLPT.

LLM03: Supply Chain maps to supply chain compromise scenarios. The March 2026 LiteLLM supply chain attack, where a compromised PyPI maintainer account exposed 40,000+ AI pipelines in approximately 40 minutes, provides the concrete template. CISA issued a KEV advisory on the incident. For TIBER-EU scoping, the red team should assess whether the financial entity's AI pipeline has transitive dependency exposure of this type.

LLM06: Excessive Agency maps to privilege escalation and lateral movement scenarios. An AI compliance assistant with write access to production systems, prompted to execute bulk account actions without human approval, represents the core DORA risk: AI systems executing irreversible financial operations without adequate authorization gates. We have seen agentic workflows in production financial environments with tool access to core banking APIs and no human-in-the-loop gate for destructive operations.

LLM08: Vector and Embedding Weaknesses maps to data integrity attacks via RAG poisoning. Attackers upload fabricated "internal policy" documents to a RAG vector database; the AI assistant then recommends unsuitable high-fee products as "compliance required," creating mis-selling liability. The integrity breach criterion triggers DORA incident reporting independently of the financial impact.

LLM04: Data and Model Poisoning maps to insider threat scenarios. Systematic bias introduced into a credit scoring model during fine-tuning can generate discriminatory or systemically incorrect risk assessments. Unlike traditional software bugs, poisoning introduces changes in decision-making behavior without any code modification, making it invisible to standard application monitoring.

For a structured assessment of your AI penetration testing readiness and TLPT preparation, see BeyondScale's AI penetration testing service.

AI Incident Reporting Under DORA: Thresholds and Timelines

DORA's major incident classification criteria, defined in RTS 2024/1772, apply fully to AI-related incidents. An incident is classified as major if the data loss criterion is met alone, or if at least two of the remaining criteria are met simultaneously.

Major incident classification criteria:

| Criterion | Threshold | |---|---| | Clients Affected | More than 10% of clients using the service, or more than 100,000 clients | | Economic Impact | Direct costs exceeding €100,000; total impact exceeding €500,000 | | Service Duration | More than 24 hours total; more than 2 hours downtime for payment-critical functions | | Geographical Spread | More than one EU Member State affected | | Data Loss | Any confidentiality, integrity, or authenticity breach | | Critical Function Impact | Services supporting a Critical or Important Function disrupted | | Reputational Impact | Media coverage, regulatory action, or significant client departure |

Reporting timeline:

  • Initial notification: within 4 hours of classifying the incident as major
  • Intermediate report: within 72 hours, including preliminary root cause, business impact, containment measures, and indicators of compromise
  • Final report: within 1 month, covering full analysis, total impact, corrective actions, and lessons learned
How common AI failure modes map to classification:

An LLM hallucination in a credit decision affecting more than 10% of customers assessed during the period meets the clients affected criterion. If decisions cannot be reconstructed because span-level audit data was not retained, the integrity breach criterion triggers independently. Both criteria are met, major incident classification is required, and the 4-hour notification clock starts.

A prompt injection in a payment processing AI causing unauthorized transactions exceeding €100,000 meets the economic impact criterion and the data loss criterion simultaneously. Major incident.

An LLM gateway supply chain compromise (LiteLLM pattern) affecting the financial entity's AI traffic meets the data loss criterion alone, which is sufficient for major incident classification without a second criterion. Major incident with 4-hour notification.

An LLM-dependent settlement screening service unavailable for more than 2 hours meets the critical function downtime criterion. If the outage spans operations in multiple member states, the geographic spread criterion is also met. Major incident.

The most significant compliance gap for most institutions is not the reporting timeline itself but the prerequisite: the institution must have the telemetry to know an AI incident occurred. An LLM that logs only at the application layer, without span-level audit trails retained outside the application, cannot satisfy the intermediate report requirement for root cause analysis or the final report requirement for reconstructing which decisions were affected. EU AI Act Article 19 creates a parallel obligation for high-risk systems: audit trails must be maintained independently of the application layer and must cover the full decision chain.

In practice, this means that the neobank running a credit decisioning AI with 30-day span retention against a 7-year record retention requirement is in violation of both frameworks simultaneously, and would be unable to produce the required DORA intermediate report if that system experienced an integrity incident.

For a full review of your DORA compliance posture for AI systems, including incident detection capabilities, see our AI security and compliance services.

Third-Party AI Vendor Risk and the Article 28(3) Register

The November 2025 CTPP designation list includes 19 providers: hyperscalers (AWS, Azure, Google Cloud), enterprise software (SAP, Salesforce, Temenos), financial infrastructure (SWIFT, Worldline, Euroclear/Clearstream), and trading platforms (Murex, SIX Group). OpenAI, Anthropic, Mistral, Cohere, and the Google Gemini API are not on the list.

This does not reduce the compliance burden. EU financial entities carry Article 28 obligations for all ICT third-party providers regardless of CTPP status.

Register of Information obligations:

Every LLM API contract must appear in the Article 28(3) register using ITS (EU) 2024/2956 templates. Shadow AI deployments, where business units adopt LLM tools through personal accounts or outside procurement, violate register completeness before any contract review occurs. Business analysts using an enterprise ChatGPT subscription that was not routed through the ICT vendor management process are operating a DORA-non-compliant ICT arrangement.

Critical or Important Function assessment:

The entity must determine whether each LLM vendor supports a critical or important function. This is a documented self-assessment. An LLM making consumer credit decisions supports a critical function. An LLM used for internal scheduling does not. The assessment must be on file and auditable.

DORA-compliant contract terms:

Standard LLM vendor terms do not meet Article 30(3) requirements. For AI providers supporting critical or important functions, mandatory contract clauses include:

  • Quantitative performance targets with availability and latency commitments
  • Unrestricted audit rights, including for supervisory authorities
  • Incident notification timelines aligned to DORA's 4-hour initial notification window
  • Mandatory exit strategies with documented data portability procedures
  • Prior notification of subcontracting changes, including model version changes and inference infrastructure migrations
Commission Delegated Regulation EU 2025/532 permits documented deviations from mandatory Article 30 clauses when standard terms are non-negotiable (as is the case for most hyperscaler and foundation model provider terms), provided compensating controls are implemented and documented. This is not a permanent exemption; it is a risk acceptance framework that must be approved at the appropriate governance level and reviewed annually.

Concentration risk:

The institution's board-approved ICT strategy must explicitly address AI vendor concentration risk. Substituting between foundation model providers for a production credit decisioning workflow requires model comparison testing, prompt template revalidation, and downstream processing adjustment. This is a material operational dependency. The DORA concentration risk analysis must reflect it.

Sub-processor chain mapping:

LLM providers use sub-processors for inference compute, model training, and trust and safety filtering. For a critical function AI deployment, each material sub-processor must be identified and documented in the register. This requires engaging the LLM vendor's sub-processor disclosure process, which most major providers support through their DPA or enterprise agreement documentation.

The EBA Guidelines on ICT and Security Risk Management (EBA/GL/2025/02) provide detailed implementation guidance for the third-party risk framework, including documentation standards for AI system contracts.

Building a DORA-Ready AI Security Program: Practical Checklist

With the August 2026 deadline weeks away, the following checklist represents the minimum viable compliance program for EU financial entities with AI deployments. It is structured by urgency.

Governance (Complete before August 2)

  • Assign a named accountable owner for AI ICT risk within the ICT risk management function
  • Produce an AI system inventory classifying each deployment by: whether it supports a critical or important function, risk tier, vendor, and model version
  • Integrate the AI inventory into the existing ICT asset register
  • Identify shadow AI deployments and initiate procurement procedures
Register of Information (Complete before August 2)
  • Audit all LLM vendor contracts against ITS 2024/2956 template requirements
  • Complete sub-processor chain mapping for all AI vendors supporting critical or important functions
  • Populate all 116 required data fields (recall the 6.5% first-submission pass rate in the ESA dry-run)
  • Document critical or important function assessments for each AI vendor
EU AI Act High-Risk Systems (Complete before August 2)
  • Classify each AI system against Annex III high-risk criteria
  • For high-risk systems: complete data governance documentation, establish audit trails outside the application layer, prepare conformity assessment documentation
  • Register high-risk systems in the EU AI database per Article 49
ICT Risk Management for AI (Q3 2026)
  • Define AI-specific risk scenarios: prompt injection, model poisoning, agent authorization escalation, supply chain compromise
  • Implement behavioral drift monitoring for production AI systems. This requires statistical divergence tracking from established decision baselines, not just availability monitoring
  • Establish span-level audit logging with retention periods meeting both DORA intermediate report requirements and EU AI Act Article 19 traceability requirements
  • Update incident classification runbooks to include AI failure modes, with classification decision points for hallucination incidents, supply chain compromise, and model availability events
TLPT Planning (Q3-Q4 2026)
  • Determine whether your institution is in TLPT scope
  • Identify AI systems supporting critical or important functions for TLPT scoping documents
  • Begin external threat intelligence provider procurement (provider must have documented AI/LLM TTI competency)
  • Target TLPT completion timeline accounting for the 3-6 month exercise duration
Vendor Contract Remediation (Q3 2026)
  • Identify all AI vendor contracts requiring DORA-compliant addenda
  • For critical function AI vendors: initiate contract renegotiation or document compensating controls under Delegated Regulation EU 2025/532
  • Board-approve the concentration risk analysis for foundation model dependencies

Conclusion: DORA Compliance for AI Systems Starts Now

DORA compliance for AI systems is the application of an existing regulatory framework to a class of technology that most financial entities deployed faster than their governance processes could follow. The concurrent August 2026 EU AI Act deadline closes the window for iterative catch-up.

The compliance gaps are technical as much as procedural. Standard LLM vendor terms do not satisfy Article 30(3). Shadow AI deployments violate Article 28 before any contract review. AI incident detection requires span-level telemetry that most current deployments do not produce. And TLPT for LLMs requires threat intelligence providers with AI-specific expertise that most traditional providers have not yet built.

In practice, every institution's compliance program will look different based on which AI systems it has deployed and whether those systems support critical or important functions. The starting point is the same for all: a complete AI system inventory, an honest critical function assessment for each system, and a gap analysis against the five DORA pillars.

The TIBER-EU 2025 framework and the OWASP LLM Top 10 together provide the technical foundation for scoping AI into your TLPT program. The EBA ICT Risk Management Guidelines provide the governance framework. What most institutions need is the specialist AI security competency to execute against them.

Book an AI security assessment with BeyondScale to assess your institution's DORA readiness for AI and LLM systems, covering ICT risk management integration, TLPT scoping, incident detection capabilities, and third-party AI vendor risk. Or run a Securetom scan to identify unregistered AI endpoints and API connections across your environment before the August deadline.

Check your AI endpoint against these findings

SecureTom runs a free quick scan on any AI endpoint in about a minute. No signup needed.

Run a free scan
BT

BeyondScale Team

AI Security Team

The SecureTom research team at BeyondScale Technologies, an ISO 27001 certified company. We build the scanner and publish what we learn testing production AI systems.