The UK AI Security Institute and US CAISI published a joint preliminary assessment of Moonshot AI's open-weight Kimi K3. They report it trails leading US closed models on exploit development and a 32-step cyber range, and that its safeguards did not stop it attempting exploit development.
ExploitBench
Benchmark that separates crashes from real exploitation progress on browser-engine targets.
Records citing ExploitBench
Jul 23, 2026
UK AISI and US CAISI jointly assess Kimi K3 cyber capability as trailing US frontier models
May 13, 2026
ExploitBench grades AI exploit development as a 16-step capability ladder on V8 bugs
Carnegie Mellon researchers released ExploitBench, which scores exploitation progress on 41 V8 JavaScript-engine vulnerabilities across 16 flags from reaching the bug through arbitrary read/write, control-flow hijack and code execution. The paper reports that public models routinely reach and crash vulnerable code but rarely achieve arbitrary code execution, while one private frontier model succeeded on roughly half of cases.