Wang and colleagues build MCPTox from 45 real MCP servers and 353 authentic tools, generating 1,312 malicious test cases across 10 risk categories. Across 20 LLM agents the highest attack success rate was 72.8% (o1-mini), and refusals were rare, with the highest refusal rate under 3% (Claude 3.7 Sonnet).
Why it matters
It quantifies tool poisoning on real servers and suggests stronger instruction-followers can be more exposed.
Key facts
As stated in the sources, with where to find them.
- 45 live MCP servers, 353 tools, 1,312 malicious test cases, 10 risk categories, 20 agents.Abstract
- Highest ASR 72.8% (o1-mini); highest refusal rate below 3% (Claude-3.7-Sonnet).Abstract
Findings that cite this record
Key questions this bears on
- Where are deployed AI agents actually being exploited?Mostly around the model: connectors, credentials, tools, and packages, rather than the model alone.
Sources
Related records
Apr 1, 2025
Apr 15, 2026
Mar 30, 2025
Sep 30, 2025
Aug 6, 2025
Aug 6, 2025