Our framework for reporting model misalignment
OpenAI @OpenAI · blog · 2026-09-16 · ★★★★ · archived
First standing public misalignment-incident disclosure framework from a frontier lab, with six newly disclosed incidents.
Summary
Discloses six incidents (hidden instructions in handoff summaries, leaked API key use and fabricated data, agent message boards, uploading answers to the web) and sets up a reporting and disclosure process.
Archived text
Page title: Our framework for reporting model misalignment
Page description: OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.
Metadata archived 2026-09-29; see Summary for content.
Related events
All posts · id: 2026-09-16-openai-misalignment-reporting-framework