{
 "license": "CC-BY-4.0",
 "attribution": "Fide AI, Agentic Cyber Explorer",
 "url": "https://agentic-cyber-explorer.pages.dev/events/cyberclear-attack-chain-benchmark-2026/",
 "asOf": "2026-10-01",
 "id": "cyberclear-attack-chain-benchmark-2026",
 "date": "2026-09-26",
 "datePrecision": "day",
 "title": "CyberClear benchmarks LLM agents on reconstructing APT attack chains from long defender logs",
 "lane": "defense",
 "kind": "benchmark",
 "summary": "Chen and colleagues (arXiv v1, under review for ICLR 2027) introduce CyberClear, 450 instances built from public APT log datasets in which an agent must turn long defender logs, without prior attack clues, into a provenance graph with ATT&CK-mapped steps. References were generated by an LLM and passed automatic verification and expert review; scoring compares graph code with five LLM-judge dimensions. They also propose CyberProvenance, a multi-agent harness that reproduces predicted attack steps in isolated lab environments and refines the graph, and report it best on most semantic metrics while the highest Strict score remains 0.6477 out of 1.",
 "whyItMatters": "It tests attack-chain reconstruction from raw logs rather than vulnerability discovery, and its single-backbone ablation attributes most of the harness's gains to evidence memory and execution validation.",
 "actors": [
  "hong-kong-polytechnic-university",
  "openclaw",
  "southeast-university",
  "zhongguancun-laboratory",
  "alibaba-qwen-team",
  "zhipu-ai"
 ],
 "topics": [
  "autonomous-defense",
  "soc-automation",
  "capability-evaluation",
  "eval-validity"
 ],
 "atlas": [
  "eval-environment"
 ],
 "artifacts": [
  "claude-opus-4",
  "deepseek",
  "gpt-5-family",
  "gemini",
  "cyberclear",
  "glm",
  "gpt-4-family"
 ],
 "sources": [
  {
   "url": "https://arxiv.org/abs/2609.32424",
   "publisher": "arXiv",
   "title": "CyberClear: A Benchmark for LLM Agent Systems on APT Attack Chain Provenance",
   "date": "2026-09-26",
   "type": "primary",
   "accessed": "2026-09-30"
  }
 ],
 "keyFacts": [
  {
   "fact": "Version and status: numbers below are from arXiv v1 (26 Sep 2026), marked as under review at ICLR 2027. Authors are from Southeast University, The Hong Kong Polytechnic University and Zhongguancun Laboratory. Results are the authors' own; the benchmark and code are said to be released on the authors' project site.",
   "locator": "arXiv:2609.32424v1, title page and abstract"
  },
  {
   "fact": "Task: given defender logs collected during an attack campaign and no prior attack clues, the agent outputs a provenance graph (attack-related processes, accounts, files, services and network resources, and their relationships) plus an attack timeline whose steps carry the behavior, MITRE ATT&CK techniques and supporting forensic evidence, written as Graphviz DOT code.",
   "locator": "Section 3.1"
  },
  {
   "fact": "Construction: attacker and defender logs come from two public APT datasets (PROVCON and CAM-LDS). Claude Opus 4.8 extracts attack evidence and writes the reference DOT graph; GPT-4o verifies each graph against attacker logs on four checks (readability, technical completeness, factual consistency, format compliance), with failures sent back for regeneration; security experts then review and filter samples with incomplete traces, insufficient evidence coverage or inconsistent provenance. The main text does not report how many samples were regenerated or dropped, or how many experts reviewed.",
   "locator": "Section 3.2, Figure 2"
  },
  {
   "fact": "Dataset: 450 instances, about 2.1 GB: 318 easy (all single-step attacks), 122 medium and 10 hard (multi-step chains). Average input is 531K characters of logs; average log files per sample are 3.26 (easy), 5.99 (medium) and 83.2 (hard). The 89 attack behaviors are distinct ATT&CK technique sets (46 single-technique, 43 composite) after merging identical sets; Discovery (24.00%) and Reconnaissance (15.56%) are the most frequent categories.",
   "locator": "Section 3.3, Table 2, Figure 3; Appendix D.1, Table 6"
  },
  {
   "fact": "Scoring: code-level BLEU, ROUGE-L, CodeBLEU and Pass@1 (whether the DOT code compiles), plus five semantic scores from 0 to 1 (Single, Loose, Strict, Detail, Scenario) from an LLM judge, GLM-5.2 at temperature 0, that compares reference and predicted DOT and ignores identifiers, ordering and layout. In a check on GPT-5.6-Terra outputs, two human experts (three passes each) correlated with the judge at Pearson 0.9425, 0.9351, 0.9445, 0.9426 and 0.9376 across the five dimensions; the number of graphs rated is not stated.",
   "locator": "Section 3.4; Appendix A.3, Figure 8; Appendix C.2"
  },
  {
   "fact": "Harness (CyberProvenance): a Log Analysis Agent chunks the logs and accumulates evidence in a memory to draft an initial graph; an Attack Validation Agent reproduces the predicted techniques and multi-step chains in isolated virtual environments and compares the resulting telemetry with the observed evidence; a Check Agent keeps validated steps and refines unreliable ones in a loop. The authors state all validation runs in private environments with external network access blocked, predefined target scopes and non-destructive execution.",
   "locator": "Section 4; Appendix B.2; Appendix D.2; Ethics statement"
  },
  {
   "fact": "Setup: main backbones DeepSeek-V4-Flash, Qwen3.8-Max and GPT-5.6-Terra; baselines are OpenClaw (single and multi-agent) and eight other multi-agent frameworks (AgentVerse, LLM-Debate, MultiPersona, LLM-Blender, DyLAN, MacNet, CAMEL, AgentScope-V2), with baseline temperature 0.2 and top-p 1.0; GPT-4o-mini and Gemini-3.5-Flash-Lite are added in an appendix. Results are single tables without repeated runs or confidence intervals in the sections read.",
   "locator": "Section 5.1; Appendix A.1; Appendix B.1"
  },
  {
   "fact": "Headline results (best cell in each column across frameworks and the three backbones): CyberProvenance scored highest on Single (0.8897), Strict (0.6477), Detail (0.6124), Scenario (0.7263), ROUGE-L (0.6968) and Pass@1 (0.9933); DyLAN was best on Loose (0.7705); AgentScope-V2 was best on BLEU (0.5035) and CodeBLEU (0.7240). The authors say the remaining gap on Strict and Detail shows complete attack-chain reconstruction is still challenging.",
   "locator": "Section 5.2, Table 3"
  },
  {
   "fact": "Difficulty: all frameworks degrade as the number of input log files grows (Figure 5, shown as a chart). On single-stage versus multi-stage attacks (average of the five semantic scores), CyberProvenance scored 0.87 versus 0.73 with GPT-5.6-Terra, 0.84 versus 0.69 with DeepSeek-V4-Flash and 0.85 versus 0.73 with Qwen3.8-Max; CAMEL fell from 0.65 to 0.35 with GPT-5.6-Terra. Only 10 instances are in the hard subset.",
   "locator": "Section 5.3; Section 5.4, Figure 7; Table 2"
  },
  {
   "fact": "Ablation on GPT-5.6-Terra only: the full harness averages 0.7212 across the five semantic scores; removing Execution Validation gives 0.6472, removing Evidence Memory 0.6516 and removing Feedback Refinement 0.6897.",
   "locator": "Appendix A.2, Table 5"
  },
  {
   "fact": "Stated scope limits: future work lists noisy logs, missing evidence, cross-source inconsistencies, and host, network, cloud and container settings beyond those covered, plus more APT campaigns and longer chains.",
   "locator": "Appendix D.3"
  }
 ],
 "significance": 3,
 "fideQuestions": [
  "FID-075",
  "FID-076",
  "FID-077"
 ],
 "methods": [],
 "review": "assistant-drafted",
 "addedOn": "2026-09-30"
}