Models & Assistants

GPT-6 Astra is now the default model for ChatGPT Work and Codex: here's what enterprise teams need to know

OpenAI's September 9 post sets GPT-6 Astra as the default for Work and Codex, with benchmark scores, case studies, and an admin-enable requirement.

models assistants category

On 9 September 2026, OpenAI published a dedicated enterprise post positioning GPT-6 Astra as the new default model for ChatGPT Work and Codex. The post goes well beyond the main launch announcement from 3 September 2026, adding Terminal-Bench scores, real customer case studies, and the governance detail that matters most to anyone managing an enterprise workspace: access is off by default and must be explicitly enabled by an admin.

Here is what the post actually says and what it means for your organisation.

The headline benchmark: Terminal-Bench 4.0

OpenAI’s chosen showcase metric for enterprise work is Terminal-Bench 4.0, a test that evaluates agents on complex terminal-based tasks including software engineering, system configuration, and data analysis.

GPT-6 Astra scores 57.9% on Terminal-Bench 4.0. For context, GPT-5.6 Sol scores 37.3% and Claude Fable 5.1 scores 55.8%. Astra beats Fable 5.1 on this benchmark while costing approximately 63% less per task than Fable 5.1 and around 9% less than Sol.

That cost-per-task figure matters in practice. A model that scores marginally better but runs significantly cheaper is not a marginal upgrade for teams running large agent workloads. It is a meaningful operational difference.

That said, it is worth noting that independent benchmarking by Artificial Analysis tells a more nuanced story. On their broader Intelligence Index, Fable 5.1 scores 66 against Astra’s 61, and Fable 5.1 leads on the Coding Agent Index too (70 to 67). OpenAI’s chosen benchmarks favour Astra’s strengths. The honest picture is that Astra leads on terminal and computer-use tasks, not across every dimension.

What customers are already doing with it

The 9 September post includes several concrete examples from early rollout, which started on 3 September 2026 for enterprises in OpenAI’s Trusted Access Programme.

OpenAI’s own engineering team used Astra inside a test environment to find and fix a memory-allocation bottleneck in Codex sessions. The result was 25x lower turn latency and roughly 30% higher throughput after switching allocators. That is not a benchmark number. That is a real infrastructure diagnosis completed by the model.

On the productivity side, OpenAI reports that Astra can complete Financial Modeling World Cup challenges using computer use approximately four times faster than the winning human competitor. The framing is not that analysts are replaced. It is that they spend less time building the model and more time interpreting results.

Cognition (Devin) integrated Astra into their harness on launch day, noting improvements in computer use, writing, and codebase understanding. Databricks reported Astra achieves state-of-the-art results on their OfficeQA Pro and Pro V2 benchmarks, with significantly better cost per task than GPT-5.6 Sol and a step up in data reasoning and document understanding.

These are not cherry-picked toy examples. They reflect the kinds of tasks enterprise teams are actually running.

The admin detail you need to act on

Here is the part most likely to cause confusion in a real organisation.

Enterprise access is off by default. In ChatGPT Enterprise and Edu workspaces, GPT-6 Astra does not appear automatically. Workspace owners must enable it through the model access controls in workspace settings, either for the entire workspace or for specific roles.

Two things that will not work as you might expect:

  • The two-week admin preview process that applies to other models is not available for Astra at launch. Access will not switch on automatically two weeks after launch.
  • Existing Early Model Access settings do not carry over. Having Early Model Access turned on does not grant members access to Astra.

Admins need to take a deliberate action. If your team is waiting for Astra to appear and it has not, this is almost certainly why.

Who gets access and on which plan

The access structure is layered:

  • Pro $100, Pro $200, Business Premium seats: Full existing Work and Codex allowance applies to Astra.
  • Plus and Business Standard seats: Limited Astra usage is included within Work and Codex only. It does not appear in the chat model picker. Additional usage is available through optional credits.
  • Free and Go plans: No access to Astra.
  • API: Available as gpt-6-astra at $10 per million input tokens and $50 per million output tokens (cached input at $1 per million). Fast mode delivers up to 2x the speed at 2x the price. Astra is also available through Microsoft Azure and Amazon Bedrock.

What is new inside Codex specifically

OpenAI made two notable changes to how Astra operates inside Codex that go beyond raw model capability.

First, Astra can keep persistent notes across context windows rather than compressing everything into a rolling summary at each compaction. Earlier context windows stay searchable, so the model can retrieve a specific test output or requirement from hours back in a session. For long-running engineering tasks, this is a practical improvement in reliability.

Second, OpenAI updated the Codex harness alongside the model to improve the speed of computer use. Combined, these changes produce 1.9x faster task completion compared to GPT-5.6 Sol on the Mind2Web benchmark. Astra also requires Codex CLI version 0.153.0 or newer, so check your tooling before rolling out.

A note on safety

OpenAI publishes safety benchmark results alongside capability numbers. On circumvention of a deliberately evadable Codex Auto-Review denial, Astra scores 0.00% versus 0.29% for Sol. The internal computer-use safety benchmark shows 2.4% for Astra against Sol’s 22.0%. The internal hallucination benchmark drops from 12.2% (Sol) to 4.2% (Astra).

There is one caveat OpenAI flags directly: Astra’s written reasoning is harder to monitor than Sol’s when the model is explicitly prompted to evade monitoring. The explanation is that Astra solves problems in fewer written steps, which means there is less visible reasoning to inspect. For organisations with compliance requirements around AI auditability, that is worth factoring into your governance decisions.

Zero Data Retention is supported for eligible API customers, and OpenAI is testing Private Safety Processing as a way to strengthen safety monitoring while preserving customer privacy.

What you should do now

If you manage a ChatGPT Enterprise or Edu workspace and want your team to access Astra, go to workspace settings and enable it explicitly through model access controls. Do not wait for it to appear on its own.

If you are on a Plus or Business Standard plan, Astra is available in Work and Codex with usage limits. It will not show up in the chat model picker, which catches people out.

If you are an API customer, gpt-6-astra is live. Check the pricing page and factor in the Fast mode option if latency matters more than cost in your workload.

The capability improvements are real and the benchmark numbers are strong on the tasks Astra is built for. The access controls are equally real, and they require deliberate action rather than passive rollout. Get your admin settings sorted before you start evaluating what the model can do for your team.