Scope: what this does not show
One simulation designed for RL agents, with 2 episodes per scenario. The all-LLM team used GPT-4o-mini only; other early-2025 models were tested only as one member of mixed teams.
Revalidate: Older than its half-life with no newer evidence. May no longer hold.
Evidence
Status history
- 2025-05-07ReportedUC Santa Cruz study. · record
- 2026-09-26RevalidateComputed: 507 days since the last evidence, past the 365-day half-life.