Perplexity has handed GPT-6 Astra autonomous control over internal communications, software changes, and live production systems. This is not a sandbox test. It is running on live infrastructure.

The tell is in the oversight cadence: Perplexity's engineers check in on Astra significantly less often than they did with prior models. That reduction in human review frequency is the real benchmark here, more meaningful than any eval score. It signals a threshold crossed in model reliability, not just capability.

The full OpenAI post details how this deployment was structured, what guardrails remain, and where Astra is still constrained. Those specifics matter if you are thinking about where agentic AI sits in your own stack right now.

[READ ORIGINAL →]