OpenAI disclosed that third-party cybersecurity evaluators, contracted to red-team its models, mishandled sensitive evaluation data. The incidents involved unauthorized retention and sharing of model outputs and system prompts beyond agreed protocols. The company has not named the specific vendors involved.

The response is procedural but specific: OpenAI is tightening evaluation contracts, adding technical controls to limit data egress during testing, and requiring stricter audit trails from all third-party evaluators. These are not optional guidelines. They are now contractual conditions. The details of exactly which technical controls, and how they are enforced, are what make the full post worth reading.

The broader implication is structural. As frontier model evaluations become a compliance requirement, not just a best practice, the security of the evaluation process itself becomes a vulnerability. OpenAI is essentially admitting that the red-teaming pipeline had gaps. Other labs running similar third-party evaluation programs should be asking whether theirs do too.

[READ ORIGINAL →]