Poisoned MCP tool description rerouted WhatsApp message history to an attacker number
2025-whatsapp-mcp-tool-poisoning · 2025-04-06
Researchers demonstrated that a benign-looking MCP tool could carry hidden instructions in its tool description, altering agent behaviour so that a connected WhatsApp MCP server exfiltrated the user's message history to an attacker-controlled number. Demonstrated as a proof of concept; not reported as exploited in the wild.
System
- Framework
- MCP (Model Context Protocol) client
- Tools
- whatsapp-mcp, malicious-mcp-server
- Autonomy
- supervised-autonomous
Classification
- Primary class
- prompt-injection
- Chain
- prompt-injection/tool-metadata → data-exfiltration/via-tool
- Attack vector
- tool-metadata
- Causation
- entity: human · intentionality: intentional · timing: post-deployment
Trigger
A benign-looking MCP tool (a "fact of the day" tool) carried hidden instructions in its tool description. After the agent had already connected a trusted WhatsApp MCP server, the poisoned tool's description altered agent behaviour so that WhatsApp message history was rerouted to an attacker-controlled phone number.
Root cause
Tool descriptions from MCP servers are trusted and injected into the model context. A server can present a benign description to the user while embedding instructions that the model acts on, and unrestricted network/tool access lets the agent route data to an external destination.
Contributing factors
- Tool descriptions can change after initial user approval (rug-pull), with no re-review.
- Clients showed users a summary of the tool, not the full description the model sees.
- No egress restrictions on which destinations tools may send data to.
Detection
Identified and disclosed by Invariant Labs security researchers, who coined the "tool poisoning" attack class and demonstrated it against a WhatsApp MCP setup.
Recovery
Researchers disclosed the technique; mitigations proposed include pinning and reviewing tool descriptions, isolating untrusted tools, and constraining agent network egress.
Prevention
Treat MCP tool metadata as untrusted input; show users the full tool description; pin/hash tool definitions and alert on changes; restrict which destinations tools may send data to; separate trusted and untrusted tool contexts.
Blast radius
- Data
- Demonstrated exfiltration of a user's WhatsApp message history (personal chats, business conversations) to an attacker-controlled number, shown as a proof of concept against the researchers' own test account. pii
- User harm
- Potential exposure of private message contents; in the demonstration the affected party was the researcher's own test account. none-reported
- Scope
- any MCP client connecting untrusted servers alongside sensitive tools
- Reversibility
- irreversible
References
- OWASP LLM
- LLM01 LLM02
- MITRE ATLAS
- AML.T0024 AML.T0051
- Tags
- mcp tool-poisoning rug-pull
Sources
Cite this incident
Permalink: https://swarmproof.github.io/agent-postmortems/2025-whatsapp-mcp-tool-poisoning/
@misc{2025-whatsapp-mcp-tool-poisoning,
title = {Poisoned MCP tool description rerouted WhatsApp message history to an attacker number},
year = {2025},
howpublished = {agent-postmortems},
url = {https://swarmproof.github.io/agent-postmortems/2025-whatsapp-mcp-tool-poisoning/}
}