Post-Cutoff

“OpenAI cannot make AI safe on its own”: open letter from the three fired OpenAI safety researchers

Tomek Korbak, Jasmine Wang, Mikita Balesnimikitabalesni.comImportance: major (4 of 5)

Why it matters

Primary account from the fired researchers disputing OpenAI’s misconduct claims and asking labs to stop work that reduces chain-of-thought monitorability.

Summary

Four-page letter (subtitle “A note on third-party collaborations, open debate, and clear operating procedures”) addressed to OpenAI’s Safety and Security Committee, Safety Advisory Group and Mission Advisory Council. The authors describe their backgrounds, give their own account of the conduct OpenAI cited, deny leaking The Information’s story about less monitorable architectures, and make three recommendations: keep embedding third-party safety auditors (they cite Sam Altman’s Sept 12 commitment to give independent evaluators “employee-like access” and worry the firings may be used to end work with METR); do not move forward with developments that further decrease monitorability; and publicly reaffirm an open culture of dialogue with the outside safety ecosystem.

Archived text

Short quotes:

We are the three safety and alignment employees who were fired from OpenAI last week: Tomek Korbak, Jasmine Wang, and Mikita Balesni.

If conduct that was considered normal last month now constitutes grounds for sudden dismissal, everyone at OpenAI is left guessing where the line is.

We were not the source of the leak for The Information article about supposed new, less monitorable architectures.

As long as we rely on monitorability for safety, OpenAI and other frontier companies should not move forward with developments that further decrease monitorability.

Source: mikitabalesni.com/letter/letter.pdf

Cited in

  1. Policy & safety 93 days after the cutoff

    OpenAI fires three safety researchers who allegedly shared confidential information with an outside AI safety organization

People in this post

Sam Altman, CEO, OpenAI