AI Security

OWASP AISVS Enterprise Implementation Guide

BT

BeyondScale Team

AI Security Team

11 min read

If your security program already uses OWASP ASVS, the OWASP AISVS enterprise implementation question is the obvious next one: how do you verify the AI parts of your stack with the same discipline? AISVS 1.0, the OWASP Artificial Intelligence Security Verification Standard, released in June 2026, is the first community standard that turns AI security into a checklist an auditor can actually test. This guide explains the structure, what the three levels mean in practice, where assessments tend to find gaps, and how to run a first AISVS-mapped assessment in 30 days.

Key Takeaways

    • AISVS 1.0 has 12 requirement chapters and three verification levels, mirroring the ASVS model.
    • It is a verification standard, not a risk list. Every requirement is meant to be testable as pass or fail.
    • Level 2 is the sensible default for production systems that touch sensitive data or make consequential decisions.
    • Chapters 8 (vector databases and memory), 9 (agentic orchestration), and 10 (MCP) cover the newest territory and are where most enterprises have the least existing evidence.
    • AISVS complements OWASP ASVS, NIST AI RMF, ISO/IEC 42001, and the OWASP LLM Top 10. It does not replace them.
    • Scope by system and risk tier, not by organization. A first assessment should cover one or two high-risk AI systems, not everything.
    • AISVS is not a certification. Treat it as a structured evidence framework.

What AISVS Is and How It Complements ASVS

OWASP ASVS has long given application security teams a shared vocabulary: a numbered requirement, a level, and an expected test. AISVS applies the same idea to systems where a model sits in the data path. The AISVS project page describes the standard as a structured set of requirements for AI systems, and the project's own guidance says it is designed to be used together with ASVS rather than instead of it.

That division of labor matters. Your AI gateway still needs TLS, session management, and input handling that ASVS covers. What ASVS does not ask is whether the training data has provenance, whether a retrieval index can be poisoned, or whether an agent's tool calls are authorized outside the model. AISVS fills that layer.

Requirement identifiers are stable across versions, for example v1.0-C9.4.3, which makes them practical to cite in audit workpapers and ticket trackers.

A note on counts: published descriptions of AISVS 1.0 vary. The research wiki is described as covering 191 requirements across 60 pages, while other write-ups cite larger totals that appear to include additional appendices. Before you build a tracking spreadsheet, pull the requirement list from the repository tag you are assessing against and treat that as your source of truth.

The 12 Chapters Mapped to a Typical Enterprise AI Stack

The 12 chapters in version 1.0 are:

  • Training Data Integrity and Traceability
  • Input Validation
  • Model Lifecycle Management and Change Control
  • Infrastructure, Configuration and Deployment Security
  • Access Control and Identity for AI Components and Users
  • Supply Chain Security for Models
  • Model Behavior, Output Control and Safety Assurance
  • Memory, Embeddings and Vector Database Security
  • Orchestration and Agentic Security
  • Model Context Protocol (MCP) Security
  • Adversarial resilience of models and pipelines
  • Monitoring, Logging and Anomaly Detection
  • Not every chapter applies to every system. A practical way to scope is to map chapters to the components you actually run:

    | If you run... | Prioritize chapters | | --- | --- | | A hosted model API behind a chat UI | C2, C5, C7, C12 | | A RAG application over internal documents | C2, C5, C7, C8, C12 | | Fine-tuned or self-hosted models | C1, C3, C4, C6, C11 | | Autonomous or tool-using agents | C5, C9, C10, C12 | | MCP servers exposed to agents | C5, C10, C12 |

    If you are also working from the OWASP threat lists, our OWASP LLM Top 10 2026 enterprise guide and OWASP Agentic AI implementation guide cover the risk side. AISVS is where you record the evidence that a control exists.

    Verification Levels in Practice

    The project defines the levels this way: Level 1 is the essential baseline every AI system should implement; Level 2 covers systems handling sensitive data or making consequential decisions; Level 3 is for high-assurance environments that must defend against sophisticated attacks.

    Here is how we translate that for scoping an assessment. This is our practitioner interpretation, not official AISVS wording:

    • Level 1, baseline hygiene. Mostly confirmable by reviewing configuration, architecture documents, and a small number of spot tests. Examples: input length limits exist, API keys are not embedded in prompts, model artifacts come from a known registry.
    • Level 2, tested controls. The control must be shown to work, not just documented. Expect prompt injection test cases against your real endpoints, authorization tests on retrieval, and evidence that logs capture tool calls.
    • Level 3, adversarial assurance. Controls are exercised by a red team that assumes a capable attacker, including multi-step attacks across components. Few systems need this across the board. Reserve it for the highest-impact pipelines.
    The common mistake is picking one level for the whole company. A customer-facing agent with write access to billing is not in the same tier as an internal summarization tool. Tier the systems first, then assign levels.

    Deep Dive: C9 Orchestration and Agentic Security

    C9 is where agent risk gets translated into requirements. The control themes, as summarized by independent reviewers of the standard, include execution budgets, loop control, circuit breakers, human approval for high-impact actions, tool isolation, agent identity, runtime authorization, and kill switches.

    One requirement is worth quoting in spirit because it exposes a lot of designs: access control decisions must be enforced by application logic or a policy engine, never by the AI model itself (referenced as C9.5.3 in public write-ups). In practice we see the opposite pattern often. A system prompt says "only call the refund tool for orders under 100 dollars," and that sentence is the entire control. A prompt injection in a support ticket makes the model ignore it, and nothing downstream checks.

    What to verify for C9 at Level 2:

    • Every tool call passes through a policy layer that evaluates the agent identity, the tool, and the arguments, independent of model output.
    • High-impact actions (payments, deletions, external email, permission changes) require out-of-band human approval.
    • There is a documented limit on steps, spend, and wall-clock time per task, and a tested way to stop a runaway agent.
    • Agents run with their own scoped identities, not a shared service account with broad rights.
    For design patterns, see our guide on containing the blast radius of agentic AI.

    Deep Dive: C10 MCP Security

    C10 is the newest area, and the one most teams have not yet inventoried. The Model Context Protocol lets agents discover and call tools on external servers, which means tool descriptions and tool responses become an attack surface. Summaries of the chapter list requirements around trusted MCP components, allow-listed servers, per-request access token validation, OAuth 2.1 claim validation, tool-level authorization, secure transport, schema validation, and screening of tool responses.

    A test plan for C10 usually starts with inventory, because shadow MCP servers are common:

  • List every MCP server your agents and developer tools can reach, including those installed locally by engineers.
  • For each, confirm it is on an allow list and pinned to a reviewed version.
  • Confirm tokens are validated on every request, with audience and scope checks, rather than trusted once at connection time.
  • Review tool descriptions for hidden instructions and test whether tool responses can inject instructions into the agent context.
  • Check that tool-level permissions match the least privilege the task needs.
  • We cover the attack side in MCP tool poisoning defense and the control side in the MCP security enterprise guide. Both map directly onto C10 evidence.

    Deep Dive: C8 Memory, Embeddings and Vector Database Security

    C8 addresses the retrieval layer: how embeddings are generated, stored, access controlled, and retrieved. The failure modes are familiar from RAG incidents. Documents are indexed without carrying their original permissions, so a user can retrieve content they could never open directly. Poisoned documents enter the index and later steer answers. Agent memory persists attacker-written content across sessions.

    Questions an assessor should be able to answer from evidence:

    • Does retrieval enforce the requesting user's entitlements at query time, not just at ingestion?
    • Is there provenance on indexed content, so a poisoned source can be identified and removed?
    • Are tenants separated at the index or namespace level, and is that tested?
    • Can persistent agent memory be reviewed, expired, and purged?
    Our vector database security guide goes through the hardening steps in detail.

    Cross-Referencing AISVS to NIST AI RMF, EU AI Act, and the LLM Top 10

    AISVS works best as the technical evidence layer under governance frameworks. A rough mapping:

    • NIST AI RMF. The AI RMF defines Govern, Map, Measure, and Manage functions. AISVS results feed Measure and Manage by providing concrete test results per system.
    • ISO/IEC 42001. The management system standard asks you to operate controls and keep records. AISVS requirements give you a defensible control catalogue to point at.
    • OWASP LLM Top 10. Each Top 10 risk category should trace to one or more AISVS chapters. Prompt injection lands mainly in C2 and C9, supply chain in C6, vector weaknesses in C8.
    • EU AI Act. For high-risk systems, AISVS evidence can support the technical documentation, logging (C12), and accuracy and resilience obligations. It does not by itself demonstrate conformity, and legal interpretation should come from counsel.
    Treat any mapping as a starting point. Requirements rarely align one to one with regulatory text, and you should record where your mapping is a judgment call.

    What Fails in Production: Patterns We See

    Across AI security assessments, a few gaps show up repeatedly when mapped to AISVS:

    • Authorization delegated to the model (C9, C5). Policies exist only as prompt text.
    • No tool-call logging (C12). Teams log prompts and responses but not which tools ran with which arguments, which makes incident reconstruction slow.
    • Unscoped retrieval (C8). The index holds more than any single user should see.
    • Unmanaged MCP and plugin sprawl (C10, C6). Nobody owns the inventory.
    • No change control on prompts and models (C3). A system prompt edit ships without review or regression tests.
    These are tendencies, not universal findings. Your results will depend on how mature your platform engineering is.

    Run Your First AISVS-Mapped Gap Assessment in 30 Days

    Days 1 to 5: scope. Inventory AI systems. Assign each a risk tier based on data sensitivity and the actions it can take. Pick one or two Tier 1 systems and set a target level, normally Level 2.

    Days 6 to 10: select chapters. Use the mapping table above to choose applicable chapters. Export the requirements from the AISVS repository tag into a tracker with columns for status, evidence, owner, and test method.

    Days 11 to 22: collect evidence and test. For each requirement, mark it as verified by document review, configuration inspection, or hands-on test. Run prompt injection, authorization, and retrieval-entitlement tests against real endpoints. Record failures with reproduction steps.

    Days 23 to 27: triage. Group failures by chapter and by root cause. Many will collapse into a handful of fixes, such as adding a policy layer in front of tools.

    Days 28 to 30: report. Produce a per-system scorecard (met, partially met, not met by chapter), a prioritized remediation plan, and a retest date. Be explicit about what was not assessed.

    Expect limitations. A first pass will mostly show you where evidence is missing, not that controls are broken. That is still valuable, because you cannot fix what you cannot see.

    Conclusion

    The OWASP AISVS enterprise implementation path is straightforward once you tier your systems: pick a level, pick the chapters that apply, and collect real evidence. Start with the chapters where you have the least visibility, usually C8, C9, and C10, and treat the standard as a living checklist inside engineering workflows rather than an annual audit artifact.

    To see what an external attacker can already reach on your AI endpoints, run a free Securetom scan. You can also compare approaches on the compare page or browse more guides on the blog.

    Check your AI endpoint against these findings

    SecureTom runs a free quick scan on any AI endpoint in about a minute. No signup needed.

    Run a free scan
    BT

    BeyondScale Team

    AI Security Team

    The SecureTom research team at BeyondScale Technologies, an ISO 27001 certified company. We build the scanner and publish what we learn testing production AI systems.