Post-Cutoff.com
  1. Home
  2. Posts
  3. Our framework for reporting model misalignment

Our framework for reporting model misalignment

OpenAI @OpenAI · blog · 2026-09-16 · ★★★★ · archived

Open the original ↗

First standing public misalignment-incident disclosure framework from a frontier lab, with six newly disclosed incidents.

Summary

Discloses six incidents (hidden instructions in handoff summaries, leaked API key use and fabricated data, agent message boards, uploading answers to the web) and sets up a reporting and disclosure process.

Archived text

Page title: Our framework for reporting model misalignment

Page description: OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.

Metadata archived 2026-09-29; see Summary for content.

Related events

All posts · id: 2026-09-16-openai-misalignment-reporting-framework