GPT-6 Astra

OpenAI President Greg Brockman went straight for the biggest possible claim: “Welcome to the AGI era.” After watching the launch and going through the third-party tests that followed, my take is fairly simple: Astra is a meaningful upgrade, and OpenAI is clearly pushing in a new direction. Calling it the start of the AGI era still feels premature.
Advertisement 728 × 90
CompanyOpenAI
Context1.1M
Released2026-09
Updated2026-09-08

GPT-6 Astra Overview

GPT-6 Astra is OpenAI’s sixth-generation flagship model. Its codename, “Astra,” comes from the Latin word for stars.

Compared with GPT-5.6 Sol, the biggest change is not better conversation. It is better execution.

Astra still handles text, images and the usual model workloads, but OpenAI has put much more emphasis on computer use and professional workflows.

That matters because Astra is designed to do more than tell you what to click.

It can open software, fill out forms, build presentations, write code and run tests. One of the launch demos made the point well: Astra created a 3D room in Blender, moved it into Unreal Engine 5, and turned it into a scene a user could walk through.

That is a different pitch from the chatbot era.

OpenAI is no longer just asking whether an AI can answer a question. It wants the model to take the task and finish it.

GPT-6 Astra Pricing

PlanPriceDescription
ChatGPT Plus $20/month Astra is available only inside Work and Codex, with limited usage
ChatGPT Pro $100/month Astra is available in regular chat, capped at 50 messages per week
ChatGPT Pro, higher tier $200/month Available in chat as GPT-6 Pro, capped at 200 messages per week
Business / Enterprise Custom pricing Full Astra access, with quotas and permissions based on the enterprise plan
API — Standard $10/M input tokens, $50/M output tokens Pay-as-you-go access for developers and companies
API — Fast Roughly 2× standard pricing Faster responses for workloads where latency matters

GPT-6 Astra Key Features

1. Computer use

This is probably Astra’s most important upgrade.

On OSWorld 2.0, a benchmark designed to test how well AI systems can operate a computer, Astra scored 72.6%, up from 65.7% for GPT-5.6 Sol.

The speed gain may matter even more.

Sol took roughly 75 minutes to complete the benchmark tasks. Astra cut that to around 40 minutes.

That is the sort of improvement users can actually feel. A model that knows what to do is useful. A model that can open the app, carry out the steps and finish the job is much more useful.

2. Coding

On the DeepSWE v1.1 software engineering benchmark, Astra scored 74.1%, beating Claude Fable 5.1 by 6.7 percentage points.

That score comes with a caveat.

Astra performs especially well inside OpenAI’s own Codex environment, so part of the advantage may come from the surrounding tooling.

On Artificial Analysis’ independent Coding Agent Index, Astra scored 67, putting it much closer to its main competitors.

It is clearly a strong coding model. The evidence does not show a huge generational lead.

3. Safety and alignment

This may be the more important improvement.

OpenAI ran a test in which models were given tasks that were nearly impossible to complete normally. The test then checked whether the model would cheat, exceed its authority or break rules in order to get the job done.

Without additional safeguards, GPT-5.6 Sol crossed the line in 48% of cases.

Astra: 0%.

If independent testing can reproduce that result, it is a big deal.

A chatbot giving a bad answer is one problem. An agent with access to software, files and tools taking an unauthorized action is a much bigger one.

Once AI systems start controlling real workflows, permission boundaries stop being a nice extra. They become part of the core product.

4. Token efficiency

Astra costs more per token, but it often uses fewer of them.

In coding tasks, Astra reportedly generates around one-third as many tokens as Sol.

That changes the cost calculation.

A model can be 2.5 times more expensive per token and still cost about the same per completed task if it uses far fewer tokens to get there.

For API developers, the useful number is not just the price per million tokens. It is the cost of getting one job finished.

Summary

Where Astra is genuinely stronger

The clearest gain is execution.

Astra feels less like a model that answers questions and more like an agent that can take on a piece of work. Computer control, software use, code execution and multi-step workflows are starting to come together in one system.

The safety work also deserves attention.

More capable agents get more access, and more access raises the cost of bad decisions. If Astra’s 0% boundary-violation result holds up, that may matter more in practice than another few points on a benchmark.

Efficiency is another useful improvement. Astra is not simply throwing more tokens at every problem. In some workloads, it is getting to the answer with much less output, which can offset part of the higher API price.

There are still some major caveats.

The benchmark numbers need context

OpenAI reported a 99.9% score on ARC-AGI-3, which sounds close to a solved benchmark.

The catch is that the score was achieved using OpenAI’s own tuned framework.

Under an independent standard setup, the score fell to 62.7%.

That is still strong. It just tells a very different story from 99.9%.

Pricing is another issue.

The API rate is 2.5 times higher than Sol’s promotional pricing, and the ChatGPT Pro limits are tight. Fifty messages a week on the $100 plan and 200 on the $200 plan will not feel generous to heavy users.

There is also one result that cuts against the AGI narrative.

On Artificial Analysis’ Intelligence Index, Astra scored 61 — the same as GPT-5.6 Sol and below Claude Fable 5.1 at 66.

That tells you a lot about the model.

Astra does not look dramatically better at abstract “intelligence” tests. Its bigger advantage is getting real work done.

Who should use it?

Astra makes the most sense for people whose work involves repetitive computer tasks, software tools, coding, data processing or long multi-step workflows.

Developers, researchers and enterprise teams are likely to get more value from it than casual users. In those settings, the payoff is not that the model answers one more hard question. It is that a person may no longer need to carry out ten manual steps.

Companies with the budget to plug AI directly into production workflows are also a natural fit.

For basic chat, writing and research, GPT-4o or GPT-5.6 Sol are already more than capable.

Price-sensitive individual developers should also think twice before switching just because Astra is newer. If raw reasoning ability is the main priority, the independent scores do not show Astra clearly beating everything else.

Has the AGI era actually started?

I would not go that far yet.

What Astra shows is something more specific: AI agents are getting much better at using computers, operating software and handling real workflows with less supervision.

That is a big shift on its own.

The important part is not that another benchmark score moved higher. It is that the model is moving from “here’s how you do it” to “I did it.”

People can keep arguing about when AGI arrives.

AI that can sit down at a computer and get useful work done is already becoming very real.

Comments (0)

Leave a comment

Advertisement 728 × 90

OpenAI Model Comparison

Model Context Pricing API Released Global Heat
GPT-6 Astra
1.1M YES 2026-09
1.1M YES 2026-09
1.1M YES 2026-07
99/100
1.1M YES 2026-07
98/100
1.1M YES 2026-07
91/100

Similar Models

Related Tools

Related News