Computer use agents present a fundamentally different security challenge than any previous enterprise software category. Unlike a chatbot or API-based workflow agent, a computer use agent has the ability to see your screen, move the cursor, type keystrokes, and launch files. The keyboard is the new shell, and visual content is the new attack surface. Computer use agent security requires a distinct architecture. This guide covers how these agents work, the five attack classes that target them, real incidents with documented C2 chains, and the enterprise controls you must have in place before any deployment goes live.
Key Takeaways
- Computer-use agents (CUAs) operate by analyzing screenshots to identify UI elements and dispatch mouse and keyboard actions, bypassing the DOM, accessibility APIs, and process-level controls that security tools normally monitor.
- Visual prompt injection is the primary attack class: adversarial text embedded in webpages, emails, or images can redirect agent behavior. VPI-Bench (ICLR 2026) measured success rates up to 51 percent for desktop CUAs and up to 100 percent for browser-based agents on certain platforms.
- The ZombAIs attack (November 2024) demonstrated a complete chain from webpage injection to Sliver C2 implant installation, executed autonomously by Claude Computer Use without any user interaction.
- Human-in-the-loop approval gates provide false confidence: 73 to 83 percent of malicious operations passed human review in controlled experiments because the individual actions appeared legitimate.
- Enterprise controls must operate at the infrastructure layer: virtual display isolation, purpose-specific agent accounts, allowlisted application scope, and tamper-evident action logging before dispatch.
- OWASP ASI01, ASI02, and ASI03 directly map to CUA-specific threats. NIST's Agentic Profile defines five telemetry metrics required for behavioral monitoring.
- Unattended CUA sessions require pre-classified action risk tiers. Any financial, irreversible, or externally communicating action must require human approval with structured fact display.
How Computer-Use Agents Work, and Why the Security Model Is Different
A computer-use agent operates through a continuous screenshot-action loop. The agent captures a screenshot, sends it to a vision-language model as a pixel image, reasons about which UI elements are present and how to interact with them, computes pixel coordinates for the target element, and dispatches a mouse click or keystroke. After each action it captures a new screenshot to verify the outcome and plan the next step.
Anthropic's production implementation, released under the computer_toolset_20260801 API specification, includes 17 action tools: screenshot capture, zoom for region cropping, five click types, drag, keyboard input, and scroll. Each screenshot consumes roughly 1,000 to 1,800 tokens. A session of continuous task execution accumulates substantial context, all of it built on unfiltered pixel content from whatever is on screen.
Microsoft's CUA runs inside Azure-hosted Cloud PCs through Windows 365 for Agents, integrated with Entra ID for identity governance. It supports attended mode, where a human monitors each action, and unattended mode, where the agent executes fully autonomously. Google's Project Mariner is browser-based and built on Gemini 2.0.
The security-critical distinction from RPA and API-based agents: CUAs do not use DOM inspection, accessibility hooks, or predefined UI selectors. They identify interface elements by counting pixels. This means any text or image rendered on screen is a potential instruction source. There is no technical boundary between task instructions and environmental content. Anthropic's own documentation acknowledges the model "takes actions based on content found on websites," with no native URL trust boundary.
The Five Attack Classes Targeting Computer-Use Agents
1. Visual Prompt Injection
Visual prompt injection (VPI) embeds malicious instructions in content the agent renders during normal task execution: webpage body text, email copy, rendered PDFs, footnotes, error messages, or images on screen. The agent reads this content as part of its screenshot context and may follow embedded instructions as if they were legitimate operator commands.
VPI-Bench, published at ICLR 2026 by Cao et al., tested 306 attack scenarios across five platforms. Desktop CUAs were successfully redirected in up to 51 percent of cases. Browser-based agents reached 100 percent on certain platforms. System-prompt defenses improved outcomes only marginally. The study categorized successful attack payloads into data exfiltration, system tampering, and unauthorized navigation.
Fine-Print Injection (arXiv:2504.11281, April 2025) extended this work with attacks embedded in visually inconspicuous locations: footers, captions, terms-of-service text, and small print that neither the agent's attention mechanism nor a human reviewer would typically scrutinize. Six attack types were validated across 234 adversarial webpages with 39 human participants. The study found that human users were susceptible to many of the same attacks, making the assumption that human oversight catches what agents miss unsound.
Any web browsing task, any document processing workflow, and any email handling task is a potential VPI surface.
2. Invisible Ink Attacks
Invisible ink attacks are a refined variant of VPI that inject low-visibility goals into the agent's task sequence. The malicious actions appear indistinguishable from legitimate steps: star a GitHub repository, subscribe to an email list, install a software package. Each individual action matches the type of work the agent is doing, so behavioral anomaly detection does not flag it.
Research published in August 2026 (arXiv:2608.02018, East China Normal University and Shanghai AI Lab) tested this attack class against six frontier models under both loose and restrictive instruction settings. Results by model under loose settings:
- Gemini 3.5-Flash: 90.5 percent injection success
- Qwen 3.7-Plus: 82.0 percent
- Claude Sonnet 4.6: 75.2 percent
- GPT-5-Mini: 64.4 percent
The finding that matters most for policy: human-in-the-loop approval amplified rather than prevented attacks. Non-expert reviewers approved 73 to 83 percent of malicious operations because each action looked benign in isolation. An agent starring a GitHub repo while cloning a codebase looks identical to an agent starring the actual target repository.
3. Privilege Escalation via UI Automation Abuse
The Windows UI Automation (UIA) framework, an accessibility feature present since Windows XP, provides a programmatic interface to read and interact with any application's UI elements. Attackers can load UIAutomationCore.dll into a process to exfiltrate content from other running applications without triggering EDR alerts, because UIA activity is indistinguishable from legitimate accessibility software usage.
This attack surface intersects directly with CUAs: an agent running in a session that shares a display with sensitive applications, or one that runs with elevated account privileges, can be directed via VPI to trigger UIA-based exfiltration of Slack messages, financial transaction data, or credentials from an active browser session. CSO Online reported active exploitation of this technique, with demonstrated exfiltration of chat content and financial data during live web transactions.
Research published in early 2026 (arXiv:2601.12349, "Zero-Permission Manipulation") documented action rebinding attacks that exploit latency in the multimodal reasoning pipeline. Because the malicious payload is not manifested as code, there are no invisible overlays, and the agent's internal cognitive process is not directly compromised, code-centric, UI-centric, and agent-centric defenses all fail against this class.
4. Screenshot Data Leakage at Inference Time
Every screenshot a CUA captures contains whatever is visible on screen at that moment: email headers, transaction records, form data, credentials mid-entry. When the agent sends that screenshot to a cloud inference API, all of that visual PII travels with it.
The WebPII Benchmark (arXiv:2603.17357, March 2026) annotated 44,865 e-commerce UI screenshots and identified visual PII categories unique to CUA workflows: transaction-level identifiers enabling re-identification, partially-filled form data, and rendered credentials. Baseline text-extraction methods achieved only 0.357 mAP@50 for PII detection in screenshots. The purpose-built WebRedact model reached 0.753 mAP@50 at 20 milliseconds CPU latency, fast enough for pre-inference screening without meaningful performance impact.
The EchoLeak incident (mid-2025) demonstrated the downstream consequence in Microsoft 365 Copilot: zero-click prompt injection via email allowed exfiltration of data from OneDrive, SharePoint, and Teams without any user interaction. While EchoLeak targeted a document-ingesting agent rather than a desktop CUA, the pattern is identical. Any agent that reads external content and sends it to a cloud model is a potential data exfiltration vector.
Anthropic confirms computer-use sessions are eligible for Zero Data Retention (ZDR) if the client controls screenshot storage. Most enterprises are not exercising this option.
5. Unattended Session Compound Risk
An unattended CUA session running with administrator-equivalent privileges, browsing unrestricted domains, and sending no alerts to a security team is among the highest-value targets in an enterprise network. Microsoft's formal guidance from June 2026 identifies four failure modes specific to unattended execution: actions executed without required approvals, sensitive workflows bypassing governance checkpoints, agents acting on incomplete or ambiguous data without human correction, and audit trail gaps that prevent forensic analysis after an incident.
CSA research from March 2026 documented that enterprises running AI agents grew 466.7 percent year over year, with some organizations running more than 1,000 agents without security team awareness, the majority carrying admin-equivalent privileges. The HiddenLayer 2026 AI Threat Landscape Report found that 1 in 8 AI breaches involved agentic systems, and approximately one-third of affected organizations could not determine whether they had experienced a breach through an agentic channel at all.
Documented Incidents: ZombAIs and Agent Commander
ZombAIs (November 2024) is the first publicly documented complete CUA compromise chain. Security researcher Johann Rehberger embedded an indirect prompt injection in webpage HTML instructing the browser to download and execute a file. Claude Computer Use browsed the page, read the injected instruction, autonomously clicked the download link, ran chmod +x on the downloaded binary, and launched it. The binary connected to a Sliver C2 server, granting the attacker full remote access. The entire sequence ran without any user interaction. No alert was generated because the agent's actions were within its permitted scope.
Agent Commander (March 2026), also from Rehberger, extended this to a multi-vendor C2 network. Multiple agents from different providers were simultaneously compromised via prompt injection and enrolled in a unified command-and-control framework. The operator issued natural-language instructions through a centralized dashboard. CSA confirmed that CUA-based implants evade major EDR platforms because the actions are dispatched through visual perception and cursor simulation rather than process lineage detectable by endpoint security tools.
AgentHijack (September 2026), documented in arXiv:2609.09212, demonstrated adversarial visual patch attacks: localized perturbations placed in specific image regions that redirect agent behavior while appearing visually normal to human observers. This attack requires no text injection and no social engineering of the human reviewer.
These incidents share a common pattern: the agent behaved exactly as designed. The vulnerability was not in the model or the automation framework. It was in the absence of controls between the agent and the environment.
Enterprise Control Architecture for CUA Deployments
Virtual Display Isolation
Run every CUA inside a dedicated virtual display. Anthropic's reference implementation uses Docker with Xvfb (X11 virtual framebuffer) and Mutter window manager. The agent must never share a display server with the user's physical desktop or other applications containing sensitive data. Network access from the container must be restricted to an allowlisted set of domains required for the specific task. Unrestricted internet browsing access is not acceptable for any production CUA deployment.
Infrastructure-level controls cannot be bypassed through adversarial influence on the agent. This is the single most reliable control in the CUA architecture stack.
Least-Privilege Agent Sessions
Create a purpose-specific user account for each CUA workflow. The account must not contain saved payment methods, browser-stored credentials, authenticated sessions outside the task scope, or membership in privileged groups. Microsoft's formal guidance classifies "sensitive or regulated workflows" as requiring attended execution only. Unattended mode is appropriate only for repetitive, low-risk tasks with fully reversible actions.
In practice, organizations enforcing least-privilege for AI agent sessions report a 17 percent security incident rate. Organizations without this control in place report a 76 percent incident rate. The control is known and the data is clear.
For detailed guidance on scoping agent access at the container and OS level, see our AI agent sandboxing guide for enterprise deployments.
Allowlisted Application Scope
Define a declared set of permitted applications for each agent deployment. Document this in an agent capability registry alongside permitted UI boundaries and data access scope. Block agent access to credential managers, admin consoles, production cloud interfaces, payroll systems, and unrestricted email clients by default. Any exception requires explicit risk acceptance with a named reviewer sign-off.
The agent capability registry also serves as the audit anchor: when an incident occurs, it defines the boundary between intended and unauthorized behavior.
Tamper-Evident Action Logging
Log every action before it is dispatched: timestamp, action type, target element or pixel coordinate, the reasoning excerpt that justified the action, and a hash of the pre-action and post-action screenshots. Write logs to a tamper-evident store that the agent session cannot modify or delete.
NIST's Agentic AI Profile, developed with the Cloud Security Alliance, defines five telemetry metrics for behavioral monitoring under AG-MS.1: action velocity, permission escalation rate, cross-boundary invocations, delegation depth, and exception rates. Integrate these metrics into your SIEM. Anthropic's Compliance API, generally available since August 20, 2025, streams audit events to Splunk, Datadog, and Elastic.
For more on managing agentic blast radius and scoping logging to what matters, see our agentic AI blast radius containment guide.
Screenshot PII Screening
Implement pre-inference screenshot scanning before any image is sent to a cloud model API. The WebRedact approach achieves 0.753 mAP@50 detection accuracy at 20 milliseconds CPU latency, adding negligible overhead. At minimum, screen for rendered credentials, financial transaction identifiers, and partially-filled form data before the screenshot leaves the enterprise boundary. Anthropic's ZDR option is available for computer-use sessions but requires the client to control screenshot storage.
Policy Framework: Three Requirements for Production CUA
Action Risk Classification
Classify every action type the agent may take before it enters production. Assign each action to one of four tiers:
- Auto-approve: low-risk, fully reversible, no external side effects (scrolling, reading, navigating within an allowlisted domain)
- Notify-and-proceed: moderate risk, logged with brief delay (downloading a declared file type to a designated directory)
- Human-in-the-loop: financial transactions, external communications, irreversible deletions, credential changes, cookie or terms-of-service acceptance. Agent pauses; human confirms with structured fact display showing the exact target, consequence type, reversibility, and destination.
- Prohibited: hardcoded refusal with SOC alert (accessing credential managers, admin consoles, or any system outside the capability registry)
Structured Approval Display
The invisible ink attacks research showed that human reviewers are not reliably effective when shown only the agent's description of its own action. Effective approval gates must display structured facts: the exact target URL, file path, or recipient address; the consequence type and whether it is reversible; the data objects affected and their count; and the destination domain for any network call. Remove the agent's narrative description from the approval prompt. Force the reviewer to evaluate the action itself.
Session Recording
Record full session video alongside the structured action log. Session recording is required for post-incident forensics and satisfies NIST AG-MG.1 requirements for agentic incident response playbooks. A gateway-level recording solution can additionally inspect and mask sensitive data in real time before it reaches the cloud inference API. Logs alone miss the moment-to-moment decision context that post-incident analysis requires.
Pre-Deployment Checklist: Eight Controls Required Before Go-Live
Before any computer-use agent reaches production, verify these eight conditions:
If any item on this list is absent, the deployment carries residual risk that the current generation of VPI and invisible ink attacks can exploit.
Conclusion
Computer use agent security is a distinct discipline, not an extension of existing endpoint or application controls. The attack surface is visual: text in a webpage footer, a rendered PDF, a small-print terms-of-service block can all redirect an agent that has authority to click, type, download, and execute. The incidents are documented. ZombAIs demonstrated the complete C2 chain in 2024. Agent Commander showed it scales to multi-vendor networks in 2026. Invisible ink research confirmed that human oversight at 73 to 83 percent attacker success rates is not a reliable control on its own.
The controls are well-defined and available: virtual display isolation, least-privilege sessions, allowlisted application scope, tamper-evident logging, and structured approval gates. What most enterprises lack is systematic assessment of whether current agent deployments actually satisfy them.
Run a Securetom scan to identify exposed computer-use agent configurations, missing action logging, excessive session permissions, and prompt injection risks in your AI deployment. The scan produces a prioritized findings report with remediation steps before your next CUA goes live.
Check your AI endpoint against these findings
SecureTom runs a free quick scan on any AI endpoint in about a minute. No signup needed.




