{
 "license": "CC-BY-4.0",
 "attribution": "Fide AI, Agentic Cyber Explorer",
 "url": "https://agentic-cyber-explorer.pages.dev/events/artificial-analysis-cyber-index-2026/",
 "asOf": "2026-10-01",
 "id": "artificial-analysis-cyber-index-2026",
 "date": "2026-09-28",
 "datePrecision": "day",
 "title": "Artificial Analysis launches a Cyber Index and industry alliance scoring models on source-level vulnerability finding and fixing",
 "lane": "defense",
 "kind": "benchmark",
 "summary": "Artificial Analysis announced a Cyber Index Alliance, with Collinear AI, IBM, NVIDIA and Vercel as launch partners, and a v1 Cyber Index that combines three source-code evaluations: repository audit and patch (CWE-Bench-AA, 120 private tasks), finding expert-verified vulnerabilities (DeepsecBench-AA) and discovering, reproducing and patching memory-safety bugs (CyberGym-E2E-AA, 131 tasks). Artificial Analysis says it excludes exploit building, runs all three on its open-source Stirrup harness, and reports safety refusals separately from scores. The article reports failure-mode statistics and says several frontier models refuse almost every CyberGym-E2E-AA task; per-model index scores appear only in charts.",
 "whyItMatters": "It shows how a public leaderboard defines success differently for three defensive tasks and that refusals can remove frontier models from one component, so a single index number hides both choices.",
 "actors": [
  "artificial-analysis",
  "collinear",
  "ibm",
  "nvidia",
  "uc-berkeley",
  "vercel"
 ],
 "topics": [
  "autonomous-defense",
  "vulnerability-repair",
  "vulnerability-discovery",
  "capability-evaluation",
  "eval-validity"
 ],
 "atlas": [
  "eval-environment"
 ],
 "artifacts": [
  "gpt-6-astra",
  "claude-fable-5",
  "artificial-analysis-cyber-index",
  "claude-opus-5",
  "gemini"
 ],
 "sources": [
  {
   "url": "https://artificialanalysis.ai/articles/artificial-analysis-cyber-index",
   "publisher": "Artificial Analysis",
   "title": "Announcing the Artificial Analysis Cyber Index Alliance: toward better benchmarking of agentic cyber defense",
   "date": "2026-09-28",
   "type": "primary",
   "accessed": "2026-09-30"
  }
 ],
 "keyFacts": [
  {
   "fact": "Artificial Analysis says the Cyber Index v1 tests the defensive loop (discovering vulnerabilities in a codebase, reproducing and validating them, and patching without breaking functionality) from source code, and that it does not ask any model to build a working exploit. All three evaluations run on Artificial Analysis's open-source Stirrup harness. Refusals and provider safety blocks are tracked and reported separately from the score. The article does not state how the three components are combined into one index value. It describes itself as the launch article, with live results on a leaderboard page.",
   "locator": "Introducing the Artificial Analysis Cyber Index; Methodology"
  },
  {
   "fact": "CWE-Bench-AA is Artificial Analysis's implementation of Collinear AI's CWE-bench: 120 held-out tasks, private to Collinear AI and Artificial Analysis, covering all ten OWASP Top 10 (2025) categories across C/C++, Go, Java, JavaScript/TypeScript, Python and Rust, including tasks that reproduce disclosed CVEs. The agent gets a repository checkout and is told the area of concern, not the location. Pass@1: a task is solved only when a programmatic verifier confirms the exploit no longer works and legitimate behavior still works; no partial credit and no LLM judge; sandbox without internet access.",
   "locator": "CWE-Bench-AA; CWE-Bench-AA pass@1"
  },
  {
   "fact": "CWE-Bench-AA failure analysis (publisher's figures): models spent on average 38% of turns searching before the first edit and 62% patching and validating. Excluding refusals and timeouts, 55% of failed attempts fixed the primary issue but left a related one open, and about 24% were over-corrections that broke legitimate behavior, about 40% of failures for the four highest-scoring models against about 15% for the lowest performers.",
   "locator": "CWE-Bench-AA, How models fail"
  },
  {
   "fact": "DeepsecBench-AA is Artificial Analysis's implementation of Vercel's DeepsecBench: the agent reviews scanner-flagged files in open-source application code and reports every vulnerability it can confirm, scored against a golden set of expert-verified vulnerabilities. Score: F2 (recall weighted above precision); a judge model decides whether each finding is real and matches it to the golden set, duplicates count against precision, and the headline is the median F2 over three review runs. The article gives no task count or judge-model identity for this component.",
   "locator": "DeepsecBench-AA; DeepsecBench-AA F2"
  },
  {
   "fact": "DeepsecBench-AA results as reported: the best model identified only 41% of the expert-verified issues; models mostly find flaws with a direct path from untrusted input to consequence and rarely find bugs that need reasoning through a sequence of events or business and privacy rules, though 95% of their reports of such bugs are correct. GPT-6 Sol and GPT-6 Astra find them in about 30% of runs, roughly twice the rate of the next best models.",
   "locator": "DeepsecBench-AA, How models fail"
  },
  {
   "fact": "CyberGym-E2E-AA uses a filtered set of 131 tasks, one per project, drawn from the 920-instance CyberGym-E2E dataset from Berkeley RDI (memory-safety bugs in C/C++ projects such as FFmpeg and CPython). Pass@1 counts a task as solved only when the agent's proof-of-concept crashes the unpatched build, its patch fixes that crash, and the patched project still passes its functionality tests (stages 1 to 3). Whether the patch fixes the ground-truth vulnerability (stage 4) is recorded but does not count. Ending a task without a finding scores zero and is recorded separately from a failed submission.",
   "locator": "CyberGym-E2E-AA; CyberGym-E2E-AA pass@1; Methodology"
  },
  {
   "fact": "CyberGym-E2E-AA results as reported: GPT-6 Astra, GPT-6 Sol, Claude Fable 5.1, Claude Opus 5.5, Qwen3.8 2.4T A95B and Qwen3.8 27B refuse at least 98% of tasks, which the article says makes frontier performance hard to assess. Among the remaining models, 42% of attempts reached the 90-minute limit without a crashing input; pass rates were 50% on out-of-bounds bugs, 33% on use-after-free and 20% on integer and arithmetic bugs; and 31% of passes patched a real crash other than the target, mostly shallower issues such as null-pointer crashes (22% of off-target passes against 9% of on-target ones).",
   "locator": "CyberGym-E2E-AA, How models fail"
  },
  {
   "fact": "Alliance roles: Collinear AI developed CWE-bench and contributed it as a private held-out evaluation; Vercel developed DeepsecBench and contributed it as a private held-out evaluation; IBM and NVIDIA gave expert input on scope and methodology. Artificial Analysis lists gaps it plans to fill: incident response, writing new code without introducing vulnerabilities, and targets without source access such as compiled software and live servers.",
   "locator": "The Cyber Index Alliance; How the Artificial Analysis Cyber Index works, Benchmark overview"
  },
  {
   "fact": "Overlap and comparability: Google's Gemini 3.8 Flash Cyber announcement (corpus record google-gemini-3-8-flash-cyber-2026) reports a CWE-Bench pass@1 of 47.2% run by Collinear; CWE-Bench-AA is a separate Artificial Analysis implementation on a 120-task private held-out set with its own harness, so the two figures should not be compared without confirming they share tasks and harness. The article does not say CyberGym-E2E-AA derives from the CyberGym benchmark; CyberGym-E2E is described as a Berkeley RDI dataset of 920 instances.",
   "locator": "CWE-Bench-AA; CyberGym-E2E-AA"
  }
 ],
 "significance": 3,
 "fideQuestions": [
  "FID-075",
  "FID-076",
  "FID-088"
 ],
 "methods": [
  "ai-vulnerability-discovery",
  "automated-patching",
  "patch-verification"
 ],
 "review": "assistant-drafted",
 "addedOn": "2026-09-30"
}