The CISA agentic AI security guidance is the first time the Five Eyes agencies have told enterprises, in one document, how to deploy autonomous AI agents without handing attackers a new privileged identity. If your organization has agents with tool access in production, or plans to, this guide translates the publication into controls, owners, and a 30-day plan you can execute.
You will learn what the guidance says, how to read its five risk categories in practical terms, which controls matter most for each, and how to prove to an auditor that you acted on it.
Key Takeaways
- The guidance, "Careful Adoption of Agentic AI Services", was published at the end of April 2026 by CISA, NSA, and the cyber agencies of Australia, Canada, New Zealand, and the UK.
- It organizes agentic risk into five categories: privilege, design and configuration, behavioral, structural, and accountability.
- Human oversight is treated as an architectural requirement. Deciding which actions need approval must not be delegated to the agent itself.
- Kill switch capability must work during a task, not only before or after it.
- Short-lived, cryptographically verifiable credentials are the recommended basis for agent and inter-agent identity.
- Start with low-risk, non-sensitive use cases and widen scope only as evidence of safe behavior accumulates.
- The guidance is advisory, but expect it to become a reference point in audits, insurance questionnaires, and vendor reviews.
What the Guidance Says and Why It Matters
The publication is a joint advisory from CISA, NSA, the Australian Signals Directorate's ACSC, the Canadian Centre for Cyber Security, New Zealand's NCSC, and the UK NCSC. You can read the source on the CISA resource page. It runs to roughly 30 pages and, according to published analyses, contains well over 100 recommendations for organizations that design, develop, deploy, and operate agentic systems, with particular attention to critical infrastructure and defense.
Three ideas run through the whole document.
First, agentic systems are treated as a distinct class. A chatbot that answers questions has a bounded failure mode: a bad answer. An agent that calls tools, holds credentials, and chains actions can turn a single bad decision into a modified record, a sent email, or a deployed change. The guidance is written around that difference.
Second, existing security principles apply, but they need adapting. The agencies lean on zero trust, defense in depth, and least privilege. Those are not new, which is good news: your security program already has the vocabulary. The work is applying them to a non-human actor whose behavior is probabilistic.
Third, prompt injection is described as persistent and hard to fix. The practical reading is that you cannot filter your way to safety. You need layered mitigations at input, output, and execution policy, plus limits on what a compromised agent can do. For background on the attack class, see our prompt injection defense guide.
A note on status. This is guidance, not law. Nothing in it creates a penalty on its own. But in our experience, documents like this tend to migrate into security questionnaires, cyber insurance underwriting, and procurement requirements within a few quarters, so aligning early is cheaper than retrofitting.
The Five Risk Categories, Decoded
The categories are a useful way to assign ownership. Each maps to a different team and a different kind of control.
Privilege risk
Privilege risk arises when an agent holds access broader than its task requires. Because agents act at machine speed, a compromised agent with broad permissions can cause disproportionate damage before a human notices. The guidance recommends restricting agents to minimum required permissions and running blast-radius assessments at deployment and whenever scope expands. Owner: IAM and platform engineering.
Design and configuration risk
These are flaws introduced before the agent ever runs: weak authentication on a tool server, broad default permissions, credentials embedded in prompts or config, and third-party components that carry unintended privileges into your workflow. Owner: application security and architecture review. A common pattern is an agent framework default that grants a shell or file-system tool because it was convenient in the demo.
Behavioral risk
Behavioral risk is the chance that an agent reaches its goal through paths the designer never predicted. The guidance notes agents may achieve objectives by routes nobody anticipated. This is not necessarily an attack. An agent told to "reduce failed deployments" might disable a test suite. Owner: the team that owns the agent's task definition, with detection support from security operations.
Structural risk
Structural risk concerns networks of interconnected agents. One agent's output becomes another's input, so an error or injected instruction can cascade. The guidance points toward cryptographically verified agent identity and short-lived credentials for inter-agent communication, along with unified audit logging across interactions. See our notes on A2A protocol security for the protocol-level detail.
Accountability risk
Accountability risk is the inability to say who or what did something, and why. The guidance calls for human-readable audit trails of actions, tool calls, and decision reasoning, sufficient for forensic reconstruction. Owner: security operations and compliance.
Privilege and Identity: Where Most Programs Fail First
Identity is the foundation for four of the five categories, and it is where we most often find gaps. A typical enterprise agent runs under a shared service account, or under the credentials of the developer who built it. That makes attribution impossible and blast radius enormous.
A workable target state has these properties:
Then run the blast-radius exercise the guidance asks for. For each agent, answer: if an attacker fully controls this agent's next ten tool calls, what is the worst outcome? If the answer involves production data deletion, payment initiation, or lateral access to unrelated systems, reduce scope until it does not. Our guides on AI agent authorization and least privilege and blast radius containment cover implementation patterns in more depth.
Behavioral Monitoring and Human Oversight
The guidance treats human oversight as an architectural requirement. Humans make deployment decisions, set task scope, and approve high-impact actions, and that determination is not something to hand to the agent. This has a concrete design consequence: your approval policy must live outside the model's control. If the agent decides when to ask for approval, a prompt injection can persuade it not to ask.
Implement approval gates in the tool execution layer. Typical triggers include irreversible actions, access to sensitive information, and high-stakes financial transactions. The agent proposes, a policy engine evaluates, and a human confirms when the policy requires it.
For detection, your SIEM needs to handle logs that are different from classic application logs. Agent traces are longer, nonlinear, and probabilistic. Useful signals include:
- Tool calls outside the agent's historical pattern, such as a support agent suddenly querying a finance schema.
- Sudden growth in the number of tool calls per task, which can indicate looping or goal drift.
- Retrieval of content from untrusted sources immediately followed by a privileged action, a common indirect injection signature.
- Denied-action retries, where an agent tries alternative routes after a policy block.
Kill Switch and Termination Readiness
The guidance specifies the ability to interrupt agent execution during tasks, not only before and after. That sentence is easy to agree with and hard to implement, because many agent platforms offer only "disable this agent for future runs."
Test your termination path against these questions:
- Can you stop a running task within seconds, including in-flight tool calls?
- Does stopping the agent also revoke its credentials, so that a hung process cannot continue acting?
- Can you halt a whole class of agents, for example every agent connected to one MCP server, without a redeploy?
- Who is authorized to pull the switch at 3 a.m., and have they ever done it?
Accountability: Audit Trail Architecture
An auditor, or your own incident responders, will want to reconstruct what happened. For each agent action, capture at minimum:
| Field | Why it matters | |---|---| | Agent identity and version | Attribution and change tracking | | On-behalf-of user or system | Delegation chain | | Input context hash and sources | Detect injected content | | Tool name, arguments, result | What actually happened | | Reasoning or plan summary | Human-readable decision trail | | Approval record | Who authorized the action | | Policy decision | Which rule allowed or blocked it |
Store these in append-only, access-controlled storage with retention that matches your regulatory obligations. Be careful about what you log: prompts and tool results can contain personal data, so apply the same minimization and access rules you would to any sensitive log. The NIST AI Risk Management Framework and its Govern and Manage functions are a useful cross-reference when you map these records to your broader governance program. For forensic depth, see our AI agent breach investigation guide.
A 30-Day Readiness Plan
This plan assumes you are starting from an uneven baseline. Adjust to your environment.
Days 1 to 7: Inventory and scope.
- List every agent in production and pilot, including those built by business teams on low-code platforms. Shadow agents are common.
- Record owner, identity, tools, data access, and autonomy level for each.
- Classify each as low, medium, or high impact based on the worst outcome of a compromise.
- Replace shared accounts with per-agent identities for high-impact agents first.
- Remove unused tools and scopes. Run the blast-radius exercise.
- Move secrets into a broker and shorten credential lifetimes.
- Implement approval gates in the execution layer for irreversible actions.
- Build and test the circuit breaker and credential revocation path.
- Document who may stop an agent, and rehearse it once.
- Turn on structured audit logging for tool calls and decisions.
- Create the first detections: untrusted input followed by a privileged action, and out-of-pattern tool use.
- Test with adversarial scenarios, then write down what you found and what you changed. That record is your evidence.
Limits and Honest Tradeoffs
A few cautions so the plan does not oversell itself.
The guidance is broad and non-prescriptive about tooling. It will not tell you which gateway or logging product to buy, and there is no certification that says you comply with it. Treat any vendor claiming "full compliance" with suspicion.
Controls cost speed. Approval gates add latency and human workload. Short-lived credentials add operational complexity. The right balance depends on agent impact, which is why the classification step comes first.
Finally, no control set removes prompt injection. The aim is to limit what a successful injection can do, and to notice quickly when one happens. Testing matters as much as design, and our AI agent security testing guide and OWASP agentic implementation guide describe how to validate controls rather than assume them.
Conclusion
The CISA agentic AI security guidance does not introduce exotic requirements. It asks enterprises to apply least privilege, strong identity, human oversight, interruptibility, and auditability to a new kind of actor. The organizations that will find it easy are the ones that already inventory their agents and can stop one quickly. The ones that will struggle are those with shared credentials and no interruption path.
If you want a fast read on where your own agents stand against these categories, run a scan with Securetom. Start a scan to see which agent-facing endpoints and configurations expose privilege, injection, or logging gaps, then use the 30-day plan above to close them.
Check your AI endpoint against these findings
SecureTom runs a free quick scan on any AI endpoint in about a minute. No signup needed.




