GPT-6 Astra handled a 150,000-line codebase with reduced supervision. Peter Gostev, in a first-impressions session hosted by OpenAI, reports the model took on complex engineering tasks and delivered more natural, actionable feedback than its predecessors.

The demo range is the real story here: from generating a historically layered 3D London simulation to making meaningful improvements in a production-scale application. These are not toy benchmarks. Gostev is testing the model under conditions that reflect actual professional workloads, and reliability is the metric he keeps returning to.

What makes this worth watching in full is Gostev's framing of supervision cost. He is not just asking whether the model gets things right, he is asking how often he has to intervene. That shift in evaluation criteria points to where serious AI tool adoption is actually heading.

[WATCH ON YOUTUBE →]