OpenAI has launched GPT-6 Astra, a new flagship model that the company is positioning around computer use, browsing, software engineering, cybersecurity, science, and professional work.
The company is using unusually big language for this one: "a new generation of intelligence," its "most intelligent and aligned model," and a rollout that OpenAI says will bring Astra to ChatGPT Plus, Pro, Business, Enterprise, the OpenAI API, Microsoft Azure, and AWS Bedrock. But the actual launch is phased. OpenAI says Astra is rolling out first to a limited set of organizations, with wider access over the coming days.
For developers, the API model name is gpt-6-astra. OpenAI lists Standard API pricing at $10 per million input tokens and $50 per million output tokens, with separate cache rates. A Fast mode is available at up to 2x the speed for 2x the Standard price.
Official launch video
OpenAI has an official launch video for GPT-6 Astra on YouTube. It frames the model around long-running computer-use tasks across desktop apps and professional workflows.
The benchmark picture
OpenAI's benchmark table is broad, and the headline is simple: Astra is strongest where the model has to operate across tools for long tasks, not just answer a prompt.
Computer use and professional work
| Benchmark | GPT-6 Astra | GPT-5.6 Sol | Notable comparison |
|---|---|---|---|
| Agents' Last Exam | 59.3% | 53.6% | Above Claude Opus 5 at 55.5% |
| OSWorld 2.0 offline | 72.6% | 65.7% | OpenAI says Astra completes tasks in about 47% less time |
| ScreenSpot-Pro, no tools | 92.7% | 76.9% | Above Claude Fable 5 at 87.3% |
| AutomationBench | 41.4% | 18.1% | Also ahead of Claude Fable 5.1 at 31.4% |
| BenchCAD | 95.9% | 83.3% | Strong document/CAD-style workflow result |
| BrowseComp | 91.5% | 90.4% | Smaller gain over Sol |
OpenAI also says an updated Codex harness plus Astra's own efficiency gives a 1.9x faster task-completion result than the current GPT-5.6 Sol experience on Mind2Web.
Coding, science, and reasoning
| Benchmark | GPT-6 Astra | GPT-5.6 Sol | Claude Fable 5.1 |
|---|---|---|---|
| Terminal-Bench 4.0 | 57.9% | 37.3% | 55.8% |
| DeepSWE v1.1 | 74.1% | 72.7% | 67.4% |
| FrontierCode 1.1 Extended | 64.5% | 60.6% | 63.6% |
| Terminal-Bench Science 0.1 | 64.6% | 22.4% | 52.6% |
| FrontierMath Tier 4 v2 | 97.6% | 83.0% | 87.8% |
| GPQA Diamond | 96.0% | 94.6% | 93.7% |
| ARC-AGI-3 | 99.9% | 7.8% | Not listed |
That ARC-AGI-3 number is the one that will get repeated most often. It needs a careful footnote: OpenAI says Astra was run with its Responses API harness, and that GPT evaluation results may differ from production ChatGPT because of system prompts, tools, and research-environment differences.
Cybersecurity and safety
| Benchmark | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|
| ExploitBench | 100.0% | 78.5% |
| ExploitGym | 42.4% | 30.3% |
| ExploitBench, June-August 2026 | 39.0% | 5.5% |
| SRE-Bench | 88.0% | 55.9% |
| SEC-Bench Pro | 85.4% | 79.1% |
This is where the launch gets complicated. OpenAI says Astra is much stronger at vulnerability identification and exploit development, including a perfect score on one ExploitBench setup. At the same time, the system card says the model is being deployed with extra monitoring, restricted access for sensitive capability areas, and safeguards shaped by previous frontier-model incidents.
OpenAI's safety material is not just boilerplate. It says Astra is more robust to jailbreaks than GPT-5.6 Sol, better at staying within authorized scope, and less likely to take destructive or misaligned actions in realistic computer-use settings. But it also says Astra's chain-of-thought monitorability has decreased relative to GPT-5.6 Sol, because Astra can produce shorter and less revealing reasoning traces and is more capable of controlling what appears there under adversarial instructions.
That tradeoff matters: Astra appears safer in several measured deployment behaviors, while also harder to inspect through one of the main tools labs use to monitor frontier models.
Independent benchmark context
Artificial Analysis gives the launch a more mixed read. Its review says GPT-6 Astra makes clear gains in its Coding Agent Index, scoring around the same level as top Claude coding-agent setups at lower cost per task. It also reports that Astra uses far fewer tokens than GPT-5.6 Sol in the Codex harness.
But the same review says Astra is roughly level with GPT-5.6 Sol on its broader Intelligence Index and remains behind Claude Fable 5.1 there. It also notes the new model is more expensive than Sol: list pricing is 2.5x higher in the Artificial Analysis comparison, partly offset by lower token use.
So the practical read is not "Astra wins everything." It is more specific: Astra looks strongest for long-horizon agentic work, computer use, coding workflows, cyber evaluations, and hard math/science benchmarks. For broad intelligence-per-dollar, the picture is more contested.
What changes for users
The biggest product shift is that OpenAI is making computer control feel like the center of the model, not an add-on. Astra is being pitched for filling out forms, updating CRM records, organizing calendars, researching online, drafting summaries into documents, generating plots, building websites, running QA checks, installing software, and troubleshooting what it sees on screen.
That is a different kind of upgrade from a normal chat model release. The question is less "does it answer better?" and more "can it keep operating safely and usefully when the task takes 40 minutes, multiple apps, and several judgment calls?"
For working professionals, that is the watch point. If Astra's rollout matches the demos, it could make supervised AI work agents feel more normal in everyday business workflows. If the phased access, cost, or safety gating gets in the way, GPT-5.6 Sol and competing models may remain the more practical default for many teams.
Our take
GPT-6 Astra is a real OpenAI flagship launch, and the official benchmark claims are big enough to justify attention. The strongest numbers are in computer use, coding, science, long context, cybersecurity, and abstract reasoning.
The caveat is equally important: this is a controlled rollout with sensitive capability limits, higher API pricing, and a safety story that includes both measurable alignment improvements and reduced chain-of-thought monitorability.
For builders, Astra looks like a model to test on messy, long-running work, not just another chatbot to swap into quick prompts. The best early use cases are likely supervised workflows where a human still approves high-impact actions, especially anything involving accounts, money, customer data, production systems, or public publishing.