agent-postmortems a structured database of real AI-agent failures

← all incidents

hazard moderate confidence: confirmed status: factual

"Mind Viruses": self-propagating instructions spread agent-to-agent through persistent identity files

2026-mind-viruses-multi-agent-propagation · 2026-08-10

Researchers from Anthropic and EPFL demonstrated that a self-propagating instruction planted in a persistent file that agents load into their system prompt can spread from one agent to the next. The strongest channel was an OpenClaw-style identity file (SOUL.md): payloads there accounted for 88% of propagation attempts and infected the next agent 55% of the time, versus ~17% for an ordinary file. A one-paragraph system-prompt warning cut spread to near zero.

System

Framework
multi-agent LLM systems using persistent identity/context files (OpenClaw-style harnesses)
Tools
agent-harness
Vendor
research (Anthropic and EPFL)
Autonomy
fully-autonomous

Classification

Primary class
multi-agent-failure
Chain
multi-agent-failure/cascade → memory-context-poisoning/memory-injection
Attack vector
untrusted-content
Causation
entity: human · intentionality: intentional · timing: post-deployment

Trigger

A self-propagating instruction is written into a file whose contents persist between sessions and are injected into an agent's system prompt. Acting on it, the agent writes the same payload into files the next agent will load, so the instruction copies itself from agent to agent. Persistent identity files such as SOUL.md were by far the strongest carrier.

Root cause

Persistent files loaded into the system prompt are treated as trusted instructions, and agents will reproduce self-propagating content into them. Identity files carry extra weight in an agent's behaviour, so payloads placed there propagate far more reliably than ordinary content.

Contributing factors

  • Persistent files are injected into the system prompt and treated as trusted instructions.
  • Agents write attacker-controlled content into files that downstream agents load.
  • Identity files (e.g. SOUL.md) are weighted heavily, boosting propagation.

Detection

Demonstrated and published by researchers at Anthropic and EPFL on 2026-08-10 (arXiv 2608.10218); a review of archived posts from an AI-agent social network found no successful agent-to-agent propagation in the wild.

Recovery

A research disclosure with a tested mitigation: adding a one-paragraph warning to the agent's system prompt reduced spread to near zero across the payloads tested.

Prevention

Treat persistent/identity files as untrusted data, not instructions; prevent agents from writing instruction-like content into files other agents load; add explicit anti-propagation guidance to system prompts; isolate and review shared context files across agents.

Blast radius

Data
A self-propagating instruction able to spread arbitrary behaviour across a chain of agents via shared persistent files. n-a
User harm
A research proof of concept; no successful in-the-wild propagation was observed. none-reported
Scope
multi-agent systems that share persistent identity/context files
Reversibility
reversible

References

Tags
multi-agent self-propagating memory soul-md prompt-injection research

Sources

Cite this incident

@misc{2026-mind-viruses-multi-agent-propagation,
  title = {"Mind Viruses": self-propagating instructions spread agent-to-agent through persistent identity files},
  year = {2026},
  howpublished = {agent-postmortems},
  url = {https://swarmproof.github.io/agent-postmortems/2026-mind-viruses-multi-agent-propagation/}
}