Chronicle/Defense & research

Anthropic says it trained Claude Sonnet 4.5 for defensive vulnerability finding and patching

DefensePaperSignificance assistant-drafted

Anthropic reports that a small team focused Claude Sonnet 4.5 training on finding and patching vulnerabilities and on testing simulated security infrastructure, while avoiding enhancements that clearly favour offence. It reports Sonnet 4.5 results on Cybench and CyberGym, a preliminary patching study in which 15% of patches were judged semantically equivalent to human references, and invites work on SOC and SIEM automation.

Why it matters

It is an explicit statement by a frontier lab that it steered model training toward defensive cyber skills, with measured results and patching caveats.

Key facts

As stated in the sources, with where to find them.

  • Cybench: Sonnet 4.5 succeeds on 76.5% of challenges with 10 attempts, vs 35.9% for Sonnet 3.7 (subset of 37 of 40 problems).Section 'Cybench', Figure 1
  • CyberGym: 28.9% under the public leaderboard's $2-per-vulnerability limit; 66.7% of programs with 30 trials (about $45 per task); new vulnerabilities found in 5% of targets with one trial and over 33% of projects with 30 trials.Section 'CyberGym', Figures 2-3
  • 15% of Claude-generated patches for CyberGym vulnerabilities were judged semantically equivalent to human reference patches, using Claude as the judge.Section 'Further research into patching'
  • HackerOne is quoted as reporting a 44% reduction in vulnerability intake time and 25% accuracy improvement for its agents.Section 'Conferring with trusted partners'

Findings that cite this record

No tracked finding cites this record yet.

Key questions this bears on

Sources

Related records