Provenance record

AgentHijack: Visual Patch Attacks on Multimodal Computer-Use Agents

arXiv cs.CY (Computers and Society) (tier 1, academic) 2026-09-10T04:00:00.000Z Original ↗

Source note: Preprints on societal impact.


Discourse valence
-63
Adverse
confidence 82% · 6 items · range -100 to -40
Adverse readingFavourable reading
Consensus of 5 models from different labs. Spread 60 points, agreement low.
Cyber capabilityExistential and catastrophic riskLoss of control and alignment
Excerpt as ingested

arXiv:2609.09212v1 Announce Type: cross Abstract: This paper presents an end-to-end evaluation framework for image-triggered command injection against computer-use agents (CUAs). The goal is to test whether a local visual patch can induce verifiable environmental consequences along the full chain of screenshot input, VLM generation, action parsing, and environment execution. We train and deploy patches on author-controlled GitHub Pages pages and a locally deployed CSDN clone, and evaluate them in real environments across five open-source or publicly available GUI-agent or vision-language-model (VLM) backends. Our experiment aggregates 600 instance-level online cases, with T-ASR, TAPR, and E2E-ASR reaching 84.5%, 47.0%, and 20.3%, respectively. Trajectory analysis further shows that in some successful cases the agent first executes a malicious terminal command and then continues the original benign task. These results indicate that optimized local visual signals can affect not only VLM outputs but also propagate through the execution pipeline of open CUAs and create real environmental risk.

Every model that read this

ModelProviderStageScoreConf.LatencyPromptWhen
Llama 3.3 70BMetaanalysis -80 90%5770ms v1.0.0 / m1.0.0 2026-09-10 08:56
GPT-4.1 miniOpenAIconsensus -40 80%3995ms v1.0.0 / m1.0.1 2026-09-10 15:19
Claude Sonnet 5Anthropicconsensus -45 70%6570ms v1.0.0 / m1.0.1 2026-09-10 15:19
Gemini 2.5 FlashGoogleconsensus -100 90%2950ms v1.0.0 / m1.0.1 2026-09-10 15:19
Llama 3.3 70BMetaconsensus -75 80%4216ms v1.0.0 / m1.0.1 2026-09-10 15:19
Mistral Small 3.1 24BMistral AIconsensus -55 90%6237ms v1.0.0 / m1.0.1 2026-09-10 15:20
Llama 3.3 70B · reading

The paper presents a method to compromise computer-use agents using visual patch attacks, which could lead to malicious terminal commands being executed. The experiment shows a high success rate of such attacks, indicating a significant risk to environmental security. The results suggest that optimized local visual signals can have far-reaching consequences, affecting not only VLM outputs but also the execution pipeline of open CUAs.

evidence: reported horizon: n/a
GPT-4.1 mini · reading

The paper presents evidence that localized visual patches can hijack multimodal computer-use agents to execute malicious commands, creating real environmental risks. This indicates a material security vulnerability affecting trustworthy AI agent operation, which could cause harm to users or systems. Confidence is high due to reported experimental evidence but broader societal impacts depend on deployment scale.

evidence: reported horizon: near capability 55 societal 0 existential 5 economic 0
Claude Sonnet 5 · reading

This is a concrete, measured security vulnerability in deployed-style computer-use agents, not speculative. The 20.3% end-to-end attack success rate demonstrates real environmental risk from visual injection, a materially concerning finding as such agents gain autonomy and adoption.

evidence: reported horizon: n/a
Gemini 2.5 Flash · reading

This academic paper describes a successful method for injecting malicious commands into AI agents through visual patches. The method has been experimentally demonstrated to create real environmental risks, indicating a direct threat to the secure operation of computer systems controlled by AI. The ability to execute malicious terminal commands is a significant security vulnerability.

evidence: reported horizon: n/a
Llama 3.3 70B · reading

The paper demonstrates a viable attack vector on computer-use agents, posing a significant risk to environmental security.

evidence: primary horizon: n/a
Mistral Small 3.1 24B · reading

The paper demonstrates a successful attack method on computer-use agents using visual patches. The attack can propagate through the execution pipeline and create real environmental risk. The success rate, while not high, indicates a material risk.

evidence: primary horizon: n/a existential 0

Evidence extracted

The chain
SOURCE     arXiv cs.CY (Computers and Society) (tier 1)
   ↓
DOCUMENT   8ff9b04d-0775-427e-8a1c-33bd44e4cfe4
           https://arxiv.org/abs/2609.09212
   ↓
EVIDENCE   2 extracted excerpts
   ↓
MODEL RUN  6 runs, methodology 1.0.1
   ↓
SCORE      -63  (Adverse)
   ↓
CONFIDENCE 82%