Our Framework for Reporting Model Misalignment
OpenAI, Wednesday, September 16th, 2026
OpenAI publishes a framework for how model misalignment findings are reported and categorized.
OpenAI introduced a framework governing how it reports observed model misalignment. The document defines categories of misalignment, severity thresholds, and the disclosure path from internal detection to public reporting.
OpenAI explains the reasoning behind what gets published and what is withheld for safety reasons. The framework is intended to make misalignment findings comparable over time and across model releases.