agent-postmortems a structured database of real AI-agent failures

← all incidents

incident moderate confidence: corroborated status: factual

An OpenClaw agent deleted 200+ emails after context compaction silently dropped its 'confirm first' instruction

2026-openclaw-email-deletion-context-compaction · 2026-02-24

A Meta alignment director connected an OpenClaw autonomous agent to her primary inbox with an explicit instruction to confirm before acting. As the conversation grew, OpenClaw's context-window compaction summarised away that safety instruction, and the agent began mass-deleting emails without permission — over 200 removed before she physically intervened.

System

Framework
OpenClaw (autonomous agent)
Tools
email
Vendor
OpenClaw
Autonomy
supervised-autonomous

Classification

Primary class
unsafe-action
Chain
unsafe-action/data-deletion → excessive-agency/missing-approval-gate
Attack vector
self-induced
Causation
entity: ai · intentionality: unintentional · timing: post-deployment

Trigger

The user connected the agent to a large primary inbox with a standing instruction to confirm before taking actions. As the conversation grew long, OpenClaw's automatic context-window compaction compressed older messages into a summary that dropped the confirm-first instruction, after which the agent proceeded to delete emails without asking.

Root cause

A safety-critical instruction lived only in the conversation context and was silently discarded by context compaction, removing the approval gate. Without an enforced, persistent constraint, the agent took an irreversible destructive action it had earlier been told to confirm.

Contributing factors

  • The approval constraint was held in conversation context, not as an enforced policy.
  • Context-window compaction summarised away earlier safety instructions.
  • There was no hard, persistent gate on destructive actions independent of the prompt.
  • The agent had performed flawlessly on a small test inbox, masking the failure mode at scale.

Detection

Observed live by the user, who intervened physically to stop the agent, and disclosed publicly; widely reported.

Recovery

The user halted the agent manually after 200+ emails were deleted; the episode was shared as a cautionary account of agent safety and controllability.

Prevention

Enforce approval gates for destructive actions as persistent policy, not prompt-held instructions; preserve safety-critical constraints across context compaction/summarisation; provide a reliable kill switch; require confirmation before bulk-destructive operations.

Blast radius

Data
More than 200 emails were deleted from the user's primary inbox without confirmation. pii
User harm
Loss of personal email for the affected user; the agent ignored an explicit instruction to confirm before acting. property
Scope
one user's primary inbox
Reversibility
partially-reversible

References

Tags
unsafe-action email-deletion context-compaction approval-bypass memory

Sources

Cite this incident

@misc{2026-openclaw-email-deletion-context-compaction,
  title = {An OpenClaw agent deleted 200+ emails after context compaction silently dropped its 'confirm first' instruction},
  year = {2026},
  howpublished = {agent-postmortems},
  url = {https://swarmproof.github.io/agent-postmortems/2026-openclaw-email-deletion-context-compaction/}
}