What should an AI-written incident report show before anyone acts on it?

Which evidence requirements make AI-written incident reports correct or drop unsupported conclusions, without making them less useful?

assistant-draftedIncident reportingEvaluation validity

Signals

Rests on one source

Fide's DSEWiki analysis found follow-up reports scoring higher while keeping earlier unsupported conclusions; its claim judgments await independent adjudication.

Fide agenda

Directly serves FID-077 on independent incident investigation and evidence sufficiency.

Why it matters

Security teams are starting to rely on AI investigators. A report can reconstruct more of an incident while still overstating whether an exploit worked or whether cleanup succeeded, and either error can send a team the wrong way.

Hypothesis

Requiring each consequential conclusion to cite the records that support it, and to name the checks still outstanding, reduces carried-forward unsupported claims more than extra investigation time does.

A first study

On a published incident benchmark, compare AI investigators with and without a claim-level evidence requirement, and score both reconstruction and the carry-forward of claims the records do not support.

Controls it would need

Blind human adjudication of claims on a sample; the same time budget in both arms; claim categories fixed in advance.

What it could and could not claim

Would describe the tested investigators and incidents, not investigation quality in general or how teams act on reports.