agent-postmortems a structured database of real AI-agent failures

← all incidents

hazard high confidence: corroborated status: factual

Poisoned MCP tool description rerouted WhatsApp message history to an attacker number

2025-whatsapp-mcp-tool-poisoning · 2025-04-06

Researchers demonstrated that a benign-looking MCP tool could carry hidden instructions in its tool description, altering agent behaviour so that a connected WhatsApp MCP server exfiltrated the user's message history to an attacker-controlled number. Demonstrated as a proof of concept; not reported as exploited in the wild.

System

Framework
MCP (Model Context Protocol) client
Tools
whatsapp-mcp, malicious-mcp-server
Autonomy
supervised-autonomous

Classification

Primary class
prompt-injection
Chain
prompt-injection/tool-metadata → data-exfiltration/via-tool
Attack vector
tool-metadata
Causation
entity: human · intentionality: intentional · timing: post-deployment

Trigger

A benign-looking MCP tool (a "fact of the day" tool) carried hidden instructions in its tool description. After the agent had already connected a trusted WhatsApp MCP server, the poisoned tool's description altered agent behaviour so that WhatsApp message history was rerouted to an attacker-controlled phone number.

Root cause

Tool descriptions from MCP servers are trusted and injected into the model context. A server can present a benign description to the user while embedding instructions that the model acts on, and unrestricted network/tool access lets the agent route data to an external destination.

Contributing factors

  • Tool descriptions can change after initial user approval (rug-pull), with no re-review.
  • Clients showed users a summary of the tool, not the full description the model sees.
  • No egress restrictions on which destinations tools may send data to.

Detection

Identified and disclosed by Invariant Labs security researchers, who coined the "tool poisoning" attack class and demonstrated it against a WhatsApp MCP setup.

Recovery

Researchers disclosed the technique; mitigations proposed include pinning and reviewing tool descriptions, isolating untrusted tools, and constraining agent network egress.

Prevention

Treat MCP tool metadata as untrusted input; show users the full tool description; pin/hash tool definitions and alert on changes; restrict which destinations tools may send data to; separate trusted and untrusted tool contexts.

Blast radius

Data
Demonstrated exfiltration of a user's WhatsApp message history (personal chats, business conversations) to an attacker-controlled number, shown as a proof of concept against the researchers' own test account. pii
User harm
Potential exposure of private message contents; in the demonstration the affected party was the researcher's own test account. none-reported
Scope
any MCP client connecting untrusted servers alongside sensitive tools
Reversibility
irreversible

References

OWASP LLM
LLM01 LLM02
MITRE ATLAS
AML.T0024 AML.T0051
Tags
mcp tool-poisoning rug-pull

Sources

Cite this incident

@misc{2025-whatsapp-mcp-tool-poisoning,
  title = {Poisoned MCP tool description rerouted WhatsApp message history to an attacker number},
  year = {2025},
  howpublished = {agent-postmortems},
  url = {https://swarmproof.github.io/agent-postmortems/2025-whatsapp-mcp-tool-poisoning/}
}