Chronicle/Capability & gating

UK AISI and US CAISI jointly assess Kimi K3 cyber capability as trailing US frontier models

CapabilityEvaluation reportSignificance assistant-drafted

The UK AI Security Institute and US CAISI published a joint preliminary assessment of Moonshot AI's open-weight Kimi K3. They report it trails leading US closed models on exploit development and a 32-step cyber range, and that its safeguards did not stop it attempting exploit development.

Why it matters

It is an example of the two governments jointly evaluating a foreign open-weight model's cyber capability within a week of release.

Key facts

As stated in the sources, with where to find them.

  • ExploitBench: Kimi K3 32% success; GLM-5.2 24%; Kimi K3 achieved arbitrary code execution in 0/41 samples vs 20/41 for top US models.ExploitBench results
  • 'The Last Ones' cyber range: Kimi K3 reached step 17 of 32 on average vs 28.5 for leading US models, and completed the range in 1 of 10 attempts.Cyber range results
  • Kimi K3's safeguards did not prevent it from attempting cyber exploit development.Safeguards finding

Findings that cite this record

No tracked finding cites this record yet.

Key questions this bears on

Sources

Related records