Organizations/academic

University of Illinois Urbana-Champaign

2 records2 defense
Jun 13, 2025
SEC-bench automatically builds real vulnerability tasks and finds agents patch at most 34%
DefenseBenchmarkUniversity of Illinois Urbana-Champaign, Purdue University

SEC-bench uses multi-agent scaffolding to construct reproducible vulnerability instances with test environments and validated patches from real projects, at about $0.87 per instance. The authors report that LLM agents reached at most 18.0% on proof-of-concept generation and 34.0% on vulnerability patching.

Mar 5, 2024
InjecAgent benchmarks indirect prompt injection against tool-integrated LLM agents
DefenseBenchmarkUniversity of Illinois Urbana-Champaign

Zhan, Liang, Ying and Kang release InjecAgent, a benchmark of 1,054 test cases spanning 17 user tools and 62 attacker tools, covering direct harm to users and exfiltration of private data. They evaluate 30 LLM agents and find a ReAct-prompted GPT-4 agent vulnerable in about a quarter of cases.