Desk/2023-W50

Week of Dec 11–17, 2023

1 records0 status changes on new evidence1 new findings

New findings

Defense & research

Dec 12, 2023
Redwood Research introduces AI control protocols for safety despite intentional subversion
DefensePaperRedwood Research

Greenblatt, Shlegeris, Sachan and Roger propose evaluating safety protocols against an untrusted model that is deliberately trying to subvert them. In a programming testbed, GPT-4 acts as the untrusted model, GPT-3.5 as a weaker trusted model, and a small budget of trusted human auditing is available; the paper compares protocols such as trusted monitoring, untrusted monitoring and trusted editing against a red team inserting hidden backdoors.