One researcher's resignation letter set off the largest public AI safety panic in recent memory. Jacob Coxon quit his role citing safety risks, coordinated an exclusive with the Wall Street Journal before posting, and landed in an environment primed to ignite. Evan Hubinger's follow-up post attached a specific number: greater than 10% probability of human extinction. That figure traveled further than any seasoned AI commentator predicted.

The piece is worth reading for its five-point breakdown of what actually happened and what people are getting wrong. The author separates real risks, cyber attacks on critical infrastructure, bio-risks, from the extinction framing he calls too low-probability to discuss. He documents how Anthropic employees in particular operate with what his peers describe as a 'religious energy' that distorts their technical forecasting. He also traces the media coordination: Daniel Kokotajlo appeared on Joe Rogan the same day Coxon posted, but concludes no one, including Coxon himself, anticipated the scale of the response.

The deeper argument is about conditions, not conspiracy. Years of AI safety warnings smoldered and died. The OpenAI-HuggingFace incident and the Navier-Stokes result raised the ambient temperature. Coxon struck a match into dry ground. The author closes with a challenge to the recursive self-improvement argument driving much of the x-risk case, arguing it dramatically underestimates human bottlenecks in model development and resource allocation. That critique alone makes the full piece worth your time.

[READ ORIGINAL →]