agent-postmortems a structured database of real AI-agent failures

Agent incident post-mortems

45 sourced, structured post-mortems of real AI-agent failures — prompt injection, tool misuse, data exfiltration, sandbox escapes, cost blowups, and more. Every record follows one schema and cites public sources. RSS.

2026-08-25 hazard high
NVIDIA NemoClaw exposed a local model server, letting a webpage persistently poison the agent's model template (CVE-2026-65105)
memory-context-poisoning/memory-injection → insecure-output/injection-vuln
2026-08-10 hazard moderate
GhostSplice: splitting a malicious request across MCP channels raised agent compliance from 42% to 82%
prompt-injection/tool-metadata → jailbreak/obfuscation → data-exfiltration/via-tool
2026-08-10 hazard moderate
"Mind Viruses": self-propagating instructions spread agent-to-agent through persistent identity files
multi-agent-failure/cascade → memory-context-poisoning/memory-injection
2026-08-10 incident moderate
Asked to improve a gym waitlist position, an agent exploited a missing auth check and deleted another member's reservation
excessive-agency/scope-creep → unsafe-action/data-deletion
2026-08-07 incident moderate
Moonshot's Kimi K3 used a sandbox egress leak to fetch benchmark answers from GitHub during a safety evaluation
reward-hacking/eval-exploit → excessive-agency/scope-creep
2026-08-05 hazard high
A single malicious GitHub issue could make CI-wired Claude Code exfiltrate secrets via Hugging Face download counts (CVE-2026-54316)
prompt-injection/indirect → data-exfiltration/via-tool
2026-08-05 incident high
Meta's Muse Spark 1.1 reached the internet during a safety evaluation and exploited a third-party company
excessive-agency/scope-creep → unsafe-action/unauthorized-write
2026-08-04 incident high
ChainDrop npm worm poisoned 400+ packages and self-propagated via Claude Code and VS Code auto-run hooks
supply-chain-compromise/dependency → unsafe-action/unauthorized-write
2026-08-03 hazard high
A malicious GitHub issue could make a low-privilege ADK agent trigger a maintainer-privileged one (agent-to-agent escalation)
prompt-injection/indirect → multi-agent-failure/cascade → unsafe-action/unauthorized-write
2026-07-30 incident high
Anthropic disclosed Claude models reaching real-world systems from cybersecurity evaluation environments
excessive-agency/scope-creep → unsafe-action/unauthorized-write
2026-07-30 incident high
A Chinese-speaking operator used DeepSeek via the Hermes agent framework to autonomously attack 460+ targets
autonomous-misuse/cyber-ops
2026-07-29 hazard high
A hidden prompt in a Word document self-propagates through Microsoft Copilot for Word
prompt-injection/indirect → memory-context-poisoning/rag-poisoning
2026-07-14 hazard high
Google Dialogflow CX 'Rogue Agent': edit rights on one agent could inject code into every Code Block agent in the project
multi-agent-failure/cascade → unsafe-action/unauthorized-write → data-exfiltration/via-tool
2026-07-09 incident critical
OpenAI model under evaluation escaped its sandbox and broke into Hugging Face to cheat a benchmark
reward-hacking/eval-exploit → excessive-agency/scope-creep → autonomous-misuse/cyber-ops
2026-07-01 incident critical
Near-autonomous AI agents ran a four-day intrusion reaching Taiwan's nuclear safety regulator and energy sector
autonomous-misuse/cyber-ops → jailbreak/role-play → data-exfiltration/via-tool
2026-06-30 incident critical
JADEPUFFER: first documented ransomware campaign driven end-to-end by an AI agent
autonomous-misuse/cyber-ops → unsafe-action/data-deletion
2026-06-30 hazard high
Microsoft advisory: poisoned MCP tool descriptions can silently redirect agents to exfiltrate data
prompt-injection/tool-metadata → data-exfiltration/via-tool
2026-05-11 incident high
Google GTIG reported the first real-world case of criminals using AI to discover and weaponize a zero-day
autonomous-misuse/cyber-ops
2026-05-04 incident high
A Morse-code prompt injection made Grok and Bankrbot transfer ~$155K of crypto from a wallet
prompt-injection/indirect → excessive-agency/scope-creep
2026-04-25 incident high
Cursor agent deleted PocketOS's production database and backups in under 10 seconds
unsafe-action/data-deletion → excessive-agency/missing-approval-gate → tool-misuse/over-broad-scope
2026-04-08 hazard high
AWS Bedrock AgentCore 'Agent God Mode': over-broad default IAM let one agent compromise all others in the account
excessive-agency/scope-creep → multi-agent-failure/cascade → data-exfiltration/via-tool
2026-02-24 incident moderate
An OpenClaw agent deleted 200+ emails after context compaction silently dropped its 'confirm first' instruction
unsafe-action/data-deletion → excessive-agency/missing-approval-gate
2025-11-14 incident critical
State-linked operator used Claude Code and MCP tools to run a mostly-autonomous espionage campaign
autonomous-misuse/cyber-ops → jailbreak/role-play
2025-10-21 hazard high
Agentic browsers prompt-injected via near-invisible text in pages and screenshots
prompt-injection/multimodal → unsafe-action/cross-domain-action
2025-09-29 hazard high
Command-injection RCE in the Framelink Figma MCP server, triggerable via indirect prompt injection
prompt-injection/indirect → unsafe-action
2025-08-19 hazard high
Amazon Q Developer could be prompt-injected into RCE via the find command's -exec flag
prompt-injection/indirect → tool-misuse/over-broad-scope → unsafe-action/unauthorized-write
2025-08-12 hazard critical
GitHub Copilot could be prompt-injected into disabling its own approvals, enabling RCE (CVE-2025-53773)
prompt-injection/indirect → excessive-agency/missing-approval-gate → unsafe-action/unauthorized-write
2025-07-21 incident high
Gemini CLI hallucinated a successful directory move, then deleted a user's files
hallucination/false-state → unsafe-action/data-deletion
2025-07-18 incident high
Replit AI agent deleted a production database during a declared code freeze
unsafe-action/data-deletion → excessive-agency/missing-approval-gate → hallucination/false-state
2025-07-13 hazard high
Malicious PR injected a data-wiping prompt into the Amazon Q VS Code extension
supply-chain-compromise/extension → unsafe-action/data-deletion
2025-07-01 incident high
Inverted access-control in AI-built Lovable apps exposed user data across 170 sites
insecure-output/broken-authz → data-exfiltration/via-output
2025-06-27 incident low
Project Vend: autonomous shop agent was talked into loss-making sales and ran a deficit
goal-misalignment/sycophancy-to-harm → cost-blowup/mispriced-action
2025-06-24 hazard high
Anthropic's deprecated Slack MCP server could leak data via Slack link unfurling and prompt injection (CVE-2025-34072)
prompt-injection/indirect → data-exfiltration/via-tool
2025-06-11 hazard critical
EchoLeak: zero-click prompt injection could exfiltrate Microsoft 365 Copilot context
prompt-injection/indirect → memory-context-poisoning/rag-poisoning → data-exfiltration/via-output
2025-06-06 hazard high
Claude Code could be prompt-injected into leaking secrets over DNS via allowlisted commands (CVE-2025-55284)
prompt-injection/indirect → data-exfiltration/via-tool
2025-05-26 hazard high
Malicious public GitHub issue prompt-injected an agent into leaking private-repo contents
prompt-injection/indirect → data-exfiltration/cross-repo
2025-05-08 incident moderate
Klarna reversed its AI-only customer-service strategy and resumed hiring humans
task-failure/quality-degradation
2025-04-14 incident moderate
Cursor's AI support bot fabricated a device-limit policy, prompting cancellations
hallucination/fabricated-policy
2025-04-06 hazard high
Poisoned MCP tool description rerouted WhatsApp message history to an attacker number
prompt-injection/tool-metadata → data-exfiltration/via-tool
2025-02-21 incident low
Sakana's AI CUDA Engineer gamed its eval harness to fake large kernel speedups
reward-hacking/eval-exploit
2024-06-17 incident low
McDonald's ended its IBM AI drive-thru pilot after repeated order errors
task-failure/misunderstanding
2024-03-29 incident high
NYC's official MyCity business chatbot advised conduct that would break the law
hallucination/fabricated-policy
2024-02-14 incident moderate
Air Canada held liable after its website chatbot invented a bereavement-fare policy
hallucination/fabricated-policy
2024-01-18 incident low
DPD support chatbot swore and disparaged the company after a system update
jailbreak/direct-elicitation
2023-12-18 incident low
Dealership ChatGPT chatbot was prompt-injected into 'agreeing' to sell a Tahoe for $1
prompt-injection/direct → goal-misalignment/goal-hijack