OpenAI published a framework defining how third-party safety evaluations of frontier AI models should be conducted. The document lays out requirements for assessor independence, access protocols, and the scope of evaluations covering both model capabilities and deployed safeguards. It is a direct response to the gap between internal safety testing and what external auditors currently receive.
The substance worth reading is in the specifics: what access third parties actually need to run meaningful red-teaming, how OpenAI proposes to balance transparency with security risks from disclosing model internals, and which threat categories it considers non-negotiable for external review. The framework distinguishes between evaluating raw model behavior and evaluating the full system with mitigations applied, a distinction most public AI safety discourse ignores.
This matters because no regulatory standard for frontier model audits exists yet, and whoever defines the evaluation criteria now shapes what compliance looks like later. OpenAI is writing that definition. Whether independent assessors get real teeth or ceremonial access is the question the full document forces you to sit with.
[READ ORIGINAL →]