1 week ago
Astra's headline score obscures what the ARC-AGI-3 test measured
OpenAI showed Astra getting a very high score on a test of unfamiliar problems.
However, that score used extra software that remembered parts of earlier attempts and organized the model's context.
When ARC Prize looked at the model under the regular testing conditions, its best score was about 62.7 percent.
This means the two scores measured different setups.
Astra still did something impressive: it used fewer actions than most human testers on nearly all the levels it completed.
That suggests it may have learned efficient ways to understand new environments.
Greg Brockman said he personally believes Astra reached artificial general intelligence, or AGI.
But there is no agreed test that can prove when a system has reached AGI.
OpenAI reported Astra scoring 99.9 percent on ARC-AGI-3 with a provider adapter harness.
ARC Prize said Astra's best score under standard shared-test conditions was 62.7 percent.
The harness preserved reasoning across requests and automatically summarized long runs.
Astra used fewer actions than the median human tester on 96 percent of completed levels.
Greg Brockman said he believed the model had crossed an AGI threshold, while acknowledging no agreed industry test exists.
- Who
- OpenAI's Astra, Greg Brockman, and ARC Prize, which runs ARC-AGI-3.
- What
- A dispute over whether Astra's 99.9 percent score accurately represents its performance on ARC-AGI-3.
- Where
- On the ARC-AGI-3 benchmark and its shared testing environment.
- When
- The issue followed OpenAI's GPT-6 Astra briefing and ARC Prize's subsequent analysis.
- Why
- The 99.9 percent result used a harness that preserved reasoning between requests, making it different from standard model-only comparisons.
Benchmark Comparability Critics
AGI-Progress Supporters
Meaning of the 99.9 percent score
Benchmark Comparability Critics
The score should not be compared with ordinary model results because the harness carried reasoning between attempts and managed context automatically.
AGI-Progress Supporters
The harness may represent a useful practical configuration, even if it measures more than the model acting alone.
What Astra's result demonstrates
Benchmark Comparability Critics
The strongest comparable result is 62.7 percent, so the 99.9 percent headline does not describe the model by itself.
AGI-Progress Supporters
Astra's efficiency is significant: it used fewer actions than the median human on 96 percent of completed levels, suggesting precise internal modeling rather than brute force.
Whether AGI has arrived
Benchmark Comparability Critics
Brockman's claim is a belief rather than a falsifiable benchmark finding because there is no agreed test defining AGI.
AGI-Progress Supporters
Brockman believes the AGI threshold has been crossed, and future observers may identify this model and moment as significant.
Key facts
- Headline score
- 99.9 percent, reported with a provider adapter harness.
- Standard-condition score
- 62.7 percent, identified by ARC Prize as Astra's best observed score on the shared test.
- Other circulated score
- 98.55 percent, or 98.6 percent when rounded, also produced with the harness at maximum reasoning settings.
- Human-efficiency comparison
- Astra used fewer actions than the median human tester on 96 percent of levels it completed.
- Action reduction
- Astra averaged 51.7 percent fewer actions than the median human tester.
- Human baseline
- The comparison used roughly 500 general participants playing the same games.
- AGI testing
- The article says there is no industry-agreed test separating a highly capable model from AGI.
Quotes
Greg Brockman
OpenAI executive who ended the Astra briefing
“Welcome to the AGI era.”
wionews.com









