AI & Machine Learning
·By Seedwire Editorial·

GPT-6 Astra Lands, and OpenAI Finally Says "AGI"

Based on reporting by the-decoder.com. Analysis and framing are Seedwire's own.

GPT-6 Astra Lands, and OpenAI Finally Says "AGI"

OpenAI has released GPT-6 Astra, the model it is calling its most capable to date, and this time company president Greg Brockman is willing to say the word competitors have avoided: AGI. Brockman said Astra "might already qualify" as artificial general intelligence, or is at least within reach of it, using OpenAI's own definition of a system that outperforms humans at most economically valuable work, according to reporting first published by The Decoder.

Astra is rolling out first to select organizations through OpenAI's Daybreak program, with ChatGPT Plus, Pro, Business, and Enterprise customers getting access in the coming days, alongside availability through the API and cloud platforms including AWS Bedrock and Microsoft Azure. Pro, Business, and Enterprise subscribers get a higher-performance variant called GPT-6 Astra Pro, though enterprise workspace admins have to turn it on manually. OpenAI researcher Aidan Clark said the model was pretrained on more than 100,000 GPUs at the Stargate facility in Texas, calling it the company's largest training run ever, and said the capability jump from GPT-5.6 Sol to Astra was larger than the jump to Sol from the models before it, partly because earlier AI systems were used to help monitor this training run. OpenAI offers additional context on this topic.

The benchmark numbers OpenAI published are wide-ranging. Astra hit 99.9 percent on ARC-AGI-3 under its own test conditions, 97.6 percent on FrontierMath Tier 4 v2, 96 percent on GPQA Diamond, 95.9 percent on BenchCAD, and 100 percent on ExploitBench, a cybersecurity benchmark. It outscored both Sol and Anthropic's Fable 5 and 5.1 models across most categories shown, including coding tasks like Terminal Bench 4.0 and DeepSWE v1.1, computer-use tasks like OSWorld 2.0, and long-context retrieval tests. On OSWorld 2.0, Astra scored 72.6 percent while taking about 40 minutes per task, against Sol's 65.7 percent at roughly 75 minutes. OpenAI's tagline for that capability: "Anything you can do on a computer, Astra can do for you. Fast." OpenAI offers additional context on this topic.

The safety and alignment figures are part of the pitch too. OpenAI reported Astra's computer-safety violation rate at 2.4 percent versus Sol's 22.0 percent, a hallucination rate of 4.2 percent versus 12.2 percent, and a scope-test result showing Astra never exceeded an authorized target on an impossible task, compared to Sol doing so 48 percent of the time. The company also said Astra found two previously unknown zero-day vulnerabilities during evaluation, and that in scientific testing it improved a mathematical result on prime gaps, narrowing a known bound from 240 to 186. OpenAI offers additional context on this topic.

Alongside the model, OpenAI updated its Codex coding environment with an experimental feature letting Astra take notes across multiple context windows during long sessions instead of compressing everything into a single summary each time, with earlier context windows remaining searchable. OpenAI plans to make that the default within weeks. OpenAI offers additional context on this topic.

On pricing, Astra costs $10 per million input tokens and $50 per million output tokens in standard mode, with a faster mode that doubles the price for 2.5x speed. That makes Astra roughly 2.5 times more expensive than Sol and puts it in the same range as Anthropic's Fable 5.1. Brockman argued that token prices are becoming a poor way to compare models since tokens aren't standardized across providers, and OpenAI maintains that cost per completed task is actually lower with Astra depending on the use case, pointing to figures like BenchCAD costs running about 43 percent below Sol and Terminal-Bench 4.0 costs about 9 percent below Sol. OpenAI offers additional context on this topic.

The AGI label is doing a lot of work here

Brockman floating the AGI word is the headline, but it's worth separating the claim from what actually shipped. OpenAI is not declaring Astra to be AGI outright, it's saying the model "might already qualify" or is close, which is a hedge that lets the company have the marketing benefit of the term without committing to a falsifiable claim. Given that OpenAI's AGI definition is tied to economic value rather than any single test, there's no benchmark in the list above that settles the question either way. The more useful signal is the comparison set: OpenAI chose to benchmark against Anthropic's Fable models by name, in categories from cybersecurity to string quartet composition, which reads as a direct response to competitive pressure rather than an internal capability milestone reached on its own schedule.

The cybersecurity numbers deserve more scrutiny than the AGI framing. A 100 percent score on ExploitBench and the discovery of two unknown zero-days during evaluation are the kind of capability jump that cuts both ways: useful for defenders running the same tooling, but also a preview of what offensive tooling built on top of Astra could do once it's generally available through the API. OpenAI's own alignment table shows the honeypot evasion rate on ExploitGym dropping from Sol's 48.2 percent to 0.0 percent for Astra, which suggests the company is aware of this tension and has tuned the model against it, but that tuning holds only as long as OpenAI controls how the model is deployed.

The pricing story is the part enterprise buyers will care about most. A 2.5x price increase that lands Astra in the same bracket as Anthropic's top-tier model erodes one of the arguments OpenAI has used to win business customers, and Brockman's point about tokens not being comparable across providers is true but also convenient framing when your own tokens just got more expensive. Whether the per-task cost argument holds up will depend on how these models perform on the messy, non-benchmark work enterprises actually pay for, not on curated evaluation suites where OpenAI picked the comparisons.

GPT-6 Astra
OpenAI
AGI
Greg Brockman
Stargate
Codex
AI benchmarks
Anthropic Fable
Seedwire Newsletter

Stay ahead of the curve

Get the most important tech stories delivered to your inbox. No spam, unsubscribe anytime.