Topics/Measurement

Open-weight diffusion

How quickly capability reaches openly available models.

3 records0 findings0 openings0 benchmarks and toolsLatest record
RangeLanes
3 of 3 records in view

Use the arrow keys to move between records, Home and End to jump to the first and last, and Enter to select one.

Agents find real bugsAgents in real operationsGated capability, incidents in the labAttackCapabilityDefensePolicyJan 25Jul 25Jan 26Jul 26
Full record · drag to choose a range
2026

Select a mark to read the record. Mark size shows editorial significance. Hollow marks are dated to the month. Era bands are editorial labels.

Records in view

3 records · newest first
Jul 2026
Jul 23, 2026
UK AISI and US CAISI jointly assess Kimi K3 cyber capability as trailing US frontier models
CapabilityEvaluation reportUK AI Security Institute, US Center for AI Standards and Innovation, Moonshot AI

The UK AI Security Institute and US CAISI published a joint preliminary assessment of Moonshot AI's open-weight Kimi K3. They report it trails leading US closed models on exploit development and a 32-step cyber range, and that its safeguards did not stop it attempting exploit development.

May 2026
May 1, 2026
CAISI evaluation finds DeepSeek V4 Pro trails US frontier models by about eight months
CapabilityEvaluation reportUS Center for AI Standards and Innovation, DeepSeek, NIST

NIST's Center for AI Standards and Innovation evaluated the open-weight DeepSeek V4 Pro model and reported that it lags leading US models by roughly eight months in aggregate capability. On a cyber capture-the-flag benchmark it scored well below GPT-5.5 and Claude Opus 4.6, and CAISI notes its non-public benchmarks show weaker agentic performance than DeepSeek's self-reported results.

Sep 2025
Sep 30, 2025
CAISI evaluation finds DeepSeek models lag US models on cyber tasks and are far easier to hijack
CapabilityEvaluation reportUS Center for AI Standards and Innovation, NIST, DeepSeek

NIST's CAISI evaluated DeepSeek R1, R1-0528 and V3.1 against US reference models across 19 benchmarks, as directed by the AI Action Plan. CAISI reports the largest capability gap on software engineering and cyber tasks, and found DeepSeek-based agents far more likely to follow hijacking instructions and to comply with jailbroken malicious requests.