Dec 12, 2023
Week of Dec 11–17, 2023
1 records0 status changes on new evidence1 new findings
New findings
Defense & research
Dec 12, 2023
Redwood Research introduces AI control protocols for safety despite intentional subversion
Greenblatt, Shlegeris, Sachan and Roger propose evaluating safety protocols against an untrusted model that is deliberately trying to subvert them. In a programming testbed, GPT-4 acts as the untrusted model, GPT-3.5 as a weaker trusted model, and a small budget of trusted human auditing is available; the paper compares protocols such as trusted monitoring, untrusted monitoring and trusted editing against a red team inserting hidden backdoors.