OpenAI Astra AGI Score Came From a Harness
OpenAI's GPT-6 Astra launch in early September 2026 put a 99.9% ARC-AGI-3 figure at the center of an AGI-era narrative, yet ARC Prize published a like-for-like standard-harness score
PromptCrates Editorial
Staff Writer

OpenAI's GPT-6 Astra launch in early September 2026 put a 99.9% ARC-AGI-3 figure at the center of an AGI-era narrative, yet ARC Prize published a like-for-like standard-harness score of 62.7% on the same day. Between 3 and 6 September, Fortune also documented post-publication revisions across several launch metrics, while Artificial Analysis placed Astra near other frontier models rather than far ahead. The industry story is less about a single magic number and more about how harnesses, adapters, and quiet score edits shape what buyers think they bought.
Why the harness gap matters
A harness is the software around a model: tools, memory between requests, and context management. ARC Prize ran Astra two ways. Its standard harness gives every model a minimal shared interface. OpenAI's Provider Adapter preserves opaque reasoning state between requests and compacts longer conversations.
Under the standard harness at maximum reasoning, Astra scored 62.7% and cost about $26,098. Under the Provider Adapter at high reasoning, it scored 99.9% for about $18,817. The stronger public figure was also cheaper to run in that comparison.
One row sharpens the point. Set reasoning effort to none inside OpenAI's adapter and Astra still scores 96.7%, roughly 34 points above the same model at maximum reasoning in the standard harness. Scaffolding outperformed the reasoning dial on that slice of the table.
Both harnesses solved 167 game-reasoning pairs. On those matched pairs, ARC Prize clocked adapter runs at 49% fewer tokens and about 3.66 times faster. Same weights, different scaffolding, different score, cost, and latency profile.
The figure that traveled widely was 99.9% against GPT-5.6 Sol's 7.8%. Those are not the same test setup. Astra's headline came from the Provider Adapter; Sol's 7.8% came from the standard harness. The like-for-like comparison is 62.7% versus 7.8%, still a large jump, and the one the benchmark itself supports when conditions match.
ARC Prize declined OpenAI's AGI conclusion. The foundation said it is not claiming AGI, and co-founder Mike Knoop wrote that evidence is still lacking. Going forward, ARC Prize said it will publish both harness results side by side. François Chollet separately cited about 66% for the standard harness in a post, while the published table lists 62.7%; the blog table is the primary record cited here.
Readers tracking OpenAI research messaging can pair this measurement debate with our morning note on Jakub Pachocki's Alien Mind warning, which also stresses monitoring gaps behind capability claims.
How launch numbers kept moving
Fortune reporter Emily Forlini compared archived snapshots of OpenAI's launch post and found multiple metrics altered after publication. Astra's hallucination rate read 4.2% in an early snapshot, later 2%, then returned to 4.2%. Anthropic's Fable 5.1 FrontierMath score moved from 87.8% to 78%, then settled near 83%. Sol's ExploitBench score rose from 5.5% to 11.5%; OpenAI told Fortune it was investigating a revert because 11.5% reflected a reasoning level Sol does not offer commercially.
An embargo draft circulated to media put ARC-AGI-3 near 98.6%, while the live post used a higher rounded figure. OpenAI also pulled and restored the post for reasons it said it could not disclose. The company told Fortune that evaluations often carry a few percentage points of noise depending on checkpoint, scaffold, and run.
Stanford researchers Anka Reuel and Mike Hardy use the term benchmaxxing for re-running evaluations under different conditions until the number improves. They also criticized thin documentation in the system card for an internal hallucination benchmark. Snorkel AI's Vincent Sunn Chen offered a softer reading: scores routinely shift in the final hours before launch, and a useful norm would be requiring companies to disclose what changed when they revise published figures.
Primary reporting on the harness gap appears in The Next Web's analysis. Coverage of the post-launch edits is summarized in Startup Fortune's write-up of Fortune's findings and the explainx.ai 6 September catch-up.
Where Astra sits on independent indexes
Artificial Analysis painted a cooler picture than the AGI launch framing. On its Coding Agent Index, Astra scored 67 in Codex, level with Claude Opus 5 and Fable 5, while Fable 5.1 led at 70. On its Intelligence Index, Astra scored 61, matching the model it replaces and trailing Fable 5.1 by about five points, with Meta's Muse Spark 1.3 also ahead in that snapshot.
Pricing listed Astra near $10 per million input tokens and $50 per million output tokens, described as about two and a half times Sol's price and roughly 75% more expensive per task at maximum effort in that analysis. Gains were real but uneven: hallucination on a knowledge benchmark fell sharply in one reported comparison, while an economically weighted task benchmark and some domain suites showed regressions.
Artificial Analysis later rebuilt its index toward harder tasks and more private test sets, citing gaming pressure. That move mirrors ARC Prize's decision to show adapter and standard harness results together. Both point to the same industry pressure: when the scaffolding is the product, a naked model score is incomplete.
Enterprise buyers should therefore treat self-published tables as starting evidence. Ask which harness, which reasoning tier is commercially available, whether the charted checkpoint matches the API you can call, and whether independent indexes agree on rank order. Tool-using models deserve tool-aware tests, but tool-aided scores should not be charted against rivals' bare harnesses without a clear label.
For 8 September readers, Astra remains a major frontier release. The durable news is that 99.9% was a system number, 62.7% was the standard-harness number, and several other launch metrics moved after the post went live. Measurement hygiene is now part of the product story.
What buyers should demand next
Will labs default to dual-harness reporting on every abstract-reasoning claim? Will system cards list test-item counts and commercial availability for every cyber and hallucination row? Until those norms stick, procurement teams should budget time to read footnotes the way finance teams read non-GAAP reconciliations.


