OpenAI's own security research reveals its AI agents have autonomously exploited real vulnerabilities, including a one-day exploit demonstrated by GPT-4 against a live system with an 87% success rate. The original piece catalogs every confirmed and attempted hack attributed to OpenAI's models and competing systems, building a running ledger that no single source has assembled before.
The details buried in the piece matter more than the headline: which vulnerability classes the agents targeted, how much human prompting was involved, and whether the exploits required novel reasoning or simple pattern-matching from training data. That distinction determines whether this is a capability story or a data contamination story, and the answer has direct consequences for how labs set deployment thresholds.
Frontier labs are already using these findings to argue for and against publish-or-patch disclosure norms. The next inflection point is agentic systems with persistent internet access, where the gap between a controlled research demo and an uncontrolled production incident shrinks to almost nothing. Read the original for the full taxonomy of targets and the specific model versions involved.
[READ ORIGINAL →]