Chronicle/Defense & research

Audit finds cybersecurity LLM benchmark scores swing over 80 points with evaluation pipeline choices

DefensePaperSignificance assistant-drafted

Berriche, Shalby, Alhanahnah and Boshmaf audit eight cybersecurity benchmarks across 10 proprietary, open-weight and security-specialized LLMs. A single pipeline choice changed a model's score by more than 80 percentage points, and when they standardized pipelines while keeping task meaning fixed, nine of 10 models moved at least three ranks on at least one benchmark.

Why it matters

Published cyber benchmark rankings may reflect harness and parsing choices as much as model capability.

Key facts

As stated in the sources, with where to find them.

  • Audit of eight cybersecurity benchmarks (including CyberMetric, SecEval, CTI-Bench and AthenaBench), 48,662 questions across 23 tasks, against 10 LLMs; 15 recurring pipeline failure modes identified.Abstract; Introduction; Sections 3-4
  • A single pipeline choice can change a model's score by more than 80 percentage points.Abstract; Introduction
  • Under a harness that standardizes pipeline choices while preserving task semantics, 9 of 10 models shift at least three ranks on at least one benchmark; GPT-5.4 is the exception.Abstract; Table 5

Findings that cite this record

Key questions this bears on

Sources

Related records