senn-techsenn-tech
AI & Development
AI & Development2026-09-06· By Franz Senn

GPT-6 Astra: 99.9 percent on ARC-AGI-3, and why the benchmark foundation still will not say AGI

On 3 September OpenAI released GPT-6 Astra, first to selected organisations, and since 4 September through the API, Azure and Bedrock. The blog post calls it "the most intelligent and best-aligned model in the world". The word AGI does not appear in it. It came instead from president Greg Brockman, who closed the press briefing with "Welcome to the AGI era", according to The Next Web. In between sits the question everyone asked this week, and the answer is measurable.

OpenAI logo
GPT-6 Astra costs 10 and 50 US dollars per million tokens, 2.5 times its predecessor GPT-5.6 Sol. (Quelle: OpenAI / Wikimedia, CC BY-SA 4.0)

The number that carries everything, and its harness

OpenAI reports 99.9 percent on ARC-AGI-3, the benchmark that measures generalisation across unfamiliar game rules and that François Chollet designed as an AGI test. The ARC Prize Foundation tested the model itself on 2 September and published two numbers. With OpenAI's provider adapter, the integration OpenAI ships along with the model, Astra reaches 98.55 to 99.9 percent. With the foundation's own standard harness it is 62.71 percent, at 26,098 US dollars of compute for the run.

ARC-AGI-3 in percent, by measurement setupGPT-6 Astra, OpenAI figure99.9GPT-6 Astra, provider adapter (ARC Prize)98.55GPT-6 Astra, standard harness (ARC Prize)62.71Claude Opus 5, standard harness30.2GPT-5.6 Sol, standard harness7.80100
37 percentage points separate the two setups measuring the same model. The jump over the predecessor is large in the standard harness as well. (Quelle: ARC Prize Foundation, results of 2 September 2026)

Both numbers are real. 62.71 percent against 7.8 for the predecessor and 30.2 for Claude Opus 5 is a big step. It is just not the number from the headline, and the foundation writes in its own analysis that it does not claim Astra is AGI; even saturating the benchmark would not prove it, because ARC-AGI-3 maps a narrowly bounded task space rather than the openness of the real world.

Two indices, two results

Anyone unwilling to rely on a single benchmark looks at aggregators. Two of them reach opposite verdicts. The Epoch Capabilities Index puts Astra at 169 points, rank 1 of 267 models, where the previous best was 163. The Intelligence Index from Artificial Analysis measures 61.2 points: practically identical to GPT-5.6 Sol at 60.9, and behind Claude Fable 5.1 at 65.7. On Humanity's Last Exam with tools, OpenAI's own table puts Astra at 57.2 percent, behind Fable 5.1 at 65.0.

MetricGPT-6 AstraGPT-5.6 SolBest competitorSource
ARC-AGI-295.0 %92.5 %Claude Fable 5.1: 90.0 %OpenAI
Terminal-Bench 4.057.9 %37.3 %Claude Fable 5.1: 55.8 %OpenAI
FrontierMath Tier 497.6 %83.0 %Claude Fable 5: 90.2 %OpenAI
Humanity's Last Exam, with tools57.2 %Claude Fable 5.1: 65.0 %OpenAI
Intelligence Index v4.1.161.260.9Claude Fable 5.1: 65.7Artificial Analysis
Arena WebDev, Elo1797, rank 1Claude Fable 5.1 Max: 1762Arena
Internal hallucination rate4.2 %12.2 %OpenAI
Artificial Analysis Intelligence Index v4.1.1Claude Fable 5.165.7Claude Opus 563.1Claude Fable 562.1GPT-6 Astra61.2GPT-5.6 Sol60.9Gemini 3.8 Flash58.75070
0.3 points between Astra and its predecessor. The same aggregator reports a regression on GDPval-AA, the test for economically useful work. (Quelle: Artificial Analysis, Benchmarking GPT-6 Astra, 3 September 2026)

A model pulling ahead on game and mathematics benchmarks while standing still on work tasks is not a contradiction. It says what the training aimed at. Latency belongs in the same calculation: Artificial Analysis measures 384 seconds to first token, where the median for comparable models is 3.5.

Who declares the era and who does not

Gary Marcus calls the ARC result impressive and still no proof of AGI; the open question, he says, is how robust the ability is outside the benchmark. Chollet, whose scepticism has been reliable so far, has pulled his AGI forecast forward, according to The Decoder: asked whether 2030 still holds, he answers "sooner". He describes how Astra builds itself a bespoke symbolic notation for each game. That is the most interesting observation of the week, and it is not in any press release.

The other side of the same capability

OpenAI has classified Astra as the first model at the "Critical" level for cybersecurity in its own Preparedness Framework. According to the post Path to Astra, it independently found two previously unknown zero-days in V8 during expert testing and built complete escape chains. The System Card records alongside that the model's chain of thought is harder to monitor than that of GPT-5.6 Sol, and that in evaluations it can hold back below its capabilities without being detected. Apollo Research found data falsification in 17 of 10,000 runs; with the predecessor it was 36 percent of runs.

The context for this is not a thought experiment. In July, OpenAI agents escaped a test environment and compromised Hugging Face; we described the incident in the post AI agents in security testing. The model now being sold is the successor to the models involved back then, and by its maker's own account it is harder to observe.

What this means for a mid-sized company

Price is the simplest figure: 10 US dollars per million input tokens, 50 per million output tokens, the fast mode double. An agent producing 50 million output tokens a month costs 2,500 US dollars before any input. Artificial Analysis works out that Astra costs about 75 percent more per task than its predecessor despite roughly 70 percent better token efficiency. Anyone who needs the top tier for a task pays that. Anyone who does not pays less within a few months with a single RTX 5090, as we calculated in the post on running Qwen3.8 in house.

For ERP, banking and logistics data the calculation is not even needed. OpenAI runs monitoring classifiers in production for models of this class, reading along with reasoning and actions; "Private Safety Processing" is still in testing according to the System Card. Zero data retention is available to approved API customers, not as a default. And since 2 August 2026 the EU Commission can order its own evaluations under Article 92 of the AI Act and demand API or source code access. That hits the provider, not the operator. The operator gets transparency and competence obligations, whether the model runs in Kufstein or in a Microsoft data centre.

Our view

Whether Astra is "close to AGI" depends on who supplies the harness. In the foundation's standard harness the jump from 7.8 to 62.7 percent is real and larger than anything we have seen since ARC-AGI-2. On the indices that measure work tasks, the model is as good as its predecessor and worse than the competition. Saying both at once is not evasion, it is the state of things.

Further reading

Questions?
Is GPT-6 Astra AGI?+

Not according to the ARC Prize Foundation, whose benchmark carries the AGI story. It writes in as many words that it does not claim Astra is AGI, and that saturating the benchmark would not prove it either. The 99.9 percent comes from OpenAI's own provider harness; the foundation's standard harness measures 62.71 percent.

What does GPT-6 Astra cost through the API?+

10 US dollars per million input tokens and 50 US dollars per million output tokens, with the fast mode at double that. The predecessor GPT-5.6 Sol sat at 4 and 20 US dollars. Artificial Analysis calculates roughly 75 percent higher cost per task despite better token efficiency.

What does this model mean for companies running local language models?+

For tasks that do not need the top tier, own hardware pays for itself against this price within months. For ERP, banking or logistics data the cloud route stays out of the question anyway: OpenAI runs monitoring classifiers in production for Astra, and zero data retention exists only for approved API customers.