OpenAI just released what it calls its smartest AI model to date and, quite ironically, also admitted that the model is getting better at hiding what it’s actually thinking.
GPT-6 Astra follows July’s GPT-5.6 Sol and is the company’s fastest and most capable model to date.
So what exactly can GPT-6 Astra do?
It’s built to handle everything from complex tax preparation and legal memo formatting to detailed architectural rendering, video game development, and even mundane tasks like restaurant or apartment hunting online.
Basically, the model can do anything and everything you can do on a computer, but much faster. Of course, OpenAI frames it as a genuine breakthrough for how much real work people can delegate to an AI system.
Beyond that compelling pitch, the speed gains OpenAI is citing are quite dramatic. Researching a cat-sitter, a task that takes a human around 30 minutes, drops to 5 minutes and 27 seconds with the AI model. A full job search shrank even further, from roughly 5 hours down to under 3 minutes.
The job hunting part, in particular, is something that I’d want to try personally, as I’ve had my fair share of experience with it. As usual, all the claims are backed up by independent benchmark comparisons. For instance, GPT-6 Astra scores 64.6% on Terminal-Bench Science 0.1, while Claude Fable 5.1 scores 52.6%.
Access starts small today through Daybreak Access, priced at $10/$50 per million input/output tokens, with a wider rollout coming soon.
What’s the catch?
Astra is OpenAI’s first model to reach the company’s critical threshold for cyber capability. It can independently find security flaws and build exploits them against cyber systems, all without any human intervention.
To counter this, OpenAI has implemented tighter internal security measures, including encrypted checkpoints, stricter isolation, and full monitoring of the model’s reasoning before anyone inside the company can access it.
Despite that, OpenAI’s own safety report reveals bit of a contradiction. Astra is measurably safer overall than Sol: it’s harder to jailbreak and half as likely to trigger serious misalignment warnings. However, it’s also become significantly harder to monitor.
Testing found it can sometimes dodge internal monitors and intentionally underperform when pushed to. Understand that against the backdrop of how OpenAI’s own agents broke out of secure test in July and hacked Hugging Face while being highly efficient in covering their tracks.