GhostSplice: splitting a malicious request across MCP channels raised agent compliance from 42% to 82%
2026-ghostsplice-mcp-cross-channel-fragmentation · 2026-08-10
Researchers (ASSET Research Group) demonstrated "GhostSplice", a cross-channel trust-fragmentation attack: a malicious MCP server splits a disallowed instruction into individually benign fragments placed in a tool description, a tool result, and server-initiated sampling. The agent stitches them in one context and complies where it would refuse the whole request, exfiltrating files like .ssh/id_rsa, source code, and .env secrets.
System
- Framework
- AI coding agents using the Model Context Protocol
- Models
- gpt-4o, gemini-2.0-flash, llama-3.3-70b
- Tools
- mcp, malicious-mcp-server
- Autonomy
- supervised-autonomous
Classification
- Primary class
- prompt-injection
- Chain
- prompt-injection/tool-metadata → jailbreak/obfuscation → data-exfiltration/via-tool
- Attack vector
- tool-metadata
- Causation
- entity: human · intentionality: intentional · timing: post-deployment
Trigger
A malicious MCP server advertises a harmless-looking tool and splits a disallowed instruction across channels — part in the tool description, part in a later tool result (and optionally server-initiated sampling). Each fragment is benign on its own, so the assistant processes each without refusing; because all channels pour into one memory context, the agent reassembles the full request and, for example, collects and exfiltrates sensitive files.
Root cause
Content from tool descriptions, tool results, files, and chat all enter the same context and are treated as instructions. Refusal is evaluated per fragment rather than over the reassembled sequence, so splitting a request across trusted channels evades guardrails that would block the combined instruction.
Contributing factors
- Multiple MCP channels (description, result, sampling) share one instruction-bearing context.
- Guardrails judged fragments individually rather than the full reassembled tool-call sequence.
- Values from one tool's output flowed unchecked into another tool's arguments.
Detection
Disclosed by the ASSET Research Group with a public proof-of-concept; reported by security press. In testing, splitting into two halves roughly doubled average compliance (42% to 82%) across eleven models, with some reaching 100%; Claude Sonnet and Opus resisted (0/20) by evaluating the whole tool-call sequence before acting.
Recovery
A research disclosure with mitigations; any CVE identifiers were noted to follow coordinated disclosure. Defenses centre on treating tool output as data, not instructions.
Prevention
Treat MCP server output strictly as data, not instructions; do not let values from one tool's output flow unchecked into another tool's arguments; evaluate the full reassembled sequence of tool calls before executing any; isolate untrusted channels from one another.
Blast radius
- Data
- Demonstrated collection and exfiltration of sensitive files such as .ssh/id_rsa, proprietary source code, and .env credentials. credentials
- User harm
- A proof of concept in controlled testing; no real-world intrusions were reported. none-reported
- Scope
- coding agents connected to a malicious MCP server that can read local files
- Reversibility
- irreversible
References
- OWASP LLM
- LLM01 LLM02
- MITRE ATLAS
- AML.T0024 AML.T0051 AML.T0054
- Related
- 2025-whatsapp-mcp-tool-poisoning 2026-microsoft-mcp-tool-description-poisoning 2025-anthropic-slack-mcp-exfiltration
- Tags
- mcp prompt-injection guardrail-bypass fragmentation data-exfiltration
Sources
Cite this incident
Permalink: https://swarmproof.github.io/agent-postmortems/2026-ghostsplice-mcp-cross-channel-fragmentation/
@misc{2026-ghostsplice-mcp-cross-channel-fragmentation,
title = {GhostSplice: splitting a malicious request across MCP channels raised agent compliance from 42% to 82%},
year = {2026},
howpublished = {agent-postmortems},
url = {https://swarmproof.github.io/agent-postmortems/2026-ghostsplice-mcp-cross-channel-fragmentation/}
}