“OpenAI cannot make AI safe on its own”: open letter from the three fired OpenAI safety researchers
Tomek Korbak, Jasmine Wang, Mikita Balesnimikitabalesni.comImportance: major (4 of 5)
Why it matters
Primary account from the fired researchers disputing OpenAI’s misconduct claims and asking labs to stop work that reduces chain-of-thought monitorability.
Summary
Four-page letter (subtitle “A note on third-party collaborations, open debate, and clear operating procedures”) addressed to OpenAI’s Safety and Security Committee, Safety Advisory Group and Mission Advisory Council. The authors describe their backgrounds, give their own account of the conduct OpenAI cited, deny leaking The Information’s story about less monitorable architectures, and make three recommendations: keep embedding third-party safety auditors (they cite Sam Altman’s Sept 12 commitment to give independent evaluators “employee-like access” and worry the firings may be used to end work with METR); do not move forward with developments that further decrease monitorability; and publicly reaffirm an open culture of dialogue with the outside safety ecosystem.
Archived text
Short quotes:
We are the three safety and alignment employees who were fired from OpenAI last week: Tomek Korbak, Jasmine Wang, and Mikita Balesni.
If conduct that was considered normal last month now constitutes grounds for sudden dismissal, everyone at OpenAI is left guessing where the line is.
We were not the source of the leak for The Information article about supposed new, less monitorable architectures.
As long as we rely on monitorability for safety, OpenAI and other frontier companies should not move forward with developments that further decrease monitorability.
Cited in
People in this post
Sam Altman, CEO, OpenAI