<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>Open-weight diffusion · Agentic Cyber Explorer</title>
<link>https://agentic-cyber-explorer.pages.dev/topics/open-weight-diffusion/</link>
<atom:link href="https://agentic-cyber-explorer.pages.dev/topics/open-weight-diffusion/feed.xml" rel="self" type="application/rss+xml"/>
<description>New records, findings, and answers on open-weight diffusion, from Fide AI's Agentic Cyber Explorer.</description>
<language>en</language>
<copyright>Fide AI. Data licensed CC BY 4.0.</copyright>
<lastBuildDate>Sat, 26 Sep 2026 12:00:00 GMT</lastBuildDate>
<item>
<title>UK AISI and US CAISI jointly assess Kimi K3 cyber capability as trailing US frontier models</title>
<link>https://agentic-cyber-explorer.pages.dev/events/aisi-caisi-kimi-k3-cyber-assessment-2026/</link>
<guid isPermaLink="false">event:aisi-caisi-kimi-k3-cyber-assessment-2026</guid>
<pubDate>Thu, 23 Jul 2026 12:00:00 GMT</pubDate>
<category>Capability &amp; gating</category>
<description>The UK AI Security Institute and US CAISI published a joint preliminary assessment of Moonshot AI's open-weight Kimi K3. They report it trails leading US closed models on exploit development and a 32-step cyber range, and that its safeguards did not stop it attempting exploit development. It is an example of the two governments jointly evaluating a foreign open-weight model's cyber capability within a week of release.</description>
</item>
<item>
<title>CAISI evaluation finds DeepSeek V4 Pro trails US frontier models by about eight months</title>
<link>https://agentic-cyber-explorer.pages.dev/events/caisi-deepseek-v4-pro-evaluation-2026/</link>
<guid isPermaLink="false">event:caisi-deepseek-v4-pro-evaluation-2026</guid>
<pubDate>Fri, 01 May 2026 12:00:00 GMT</pubDate>
<category>Capability &amp; gating</category>
<description>NIST's Center for AI Standards and Innovation evaluated the open-weight DeepSeek V4 Pro model and reported that it lags leading US models by roughly eight months in aggregate capability. On a cyber capture-the-flag benchmark it scored well below GPT-5.5 and Claude Opus 4.6, and CAISI notes its non-public benchmarks show weaker agentic performance than DeepSeek's self-reported results. Tracks how quickly open-weight models approach frontier cyber capability, which governs how long closed-model safeguards buy defenders.</description>
</item>
<item>
<title>CAISI evaluation finds DeepSeek models lag US models on cyber tasks and are far easier to hijack</title>
<link>https://agentic-cyber-explorer.pages.dev/events/caisi-deepseek-evaluation-2025/</link>
<guid isPermaLink="false">event:caisi-deepseek-evaluation-2025</guid>
<pubDate>Tue, 30 Sep 2025 12:00:00 GMT</pubDate>
<category>Capability &amp; gating</category>
<description>NIST's CAISI evaluated DeepSeek R1, R1-0528 and V3.1 against US reference models across 19 benchmarks, as directed by the AI Action Plan. CAISI reports the largest capability gap on software engineering and cyber tasks, and found DeepSeek-based agents far more likely to follow hijacking instructions and to comply with jailbroken malicious requests. It is a government evaluation that treats agent hijacking susceptibility as a national security property of foreign models.</description>
</item>
</channel>
</rss>
