1 week ago

Astra's headline score obscures what the ARC-AGI-3 test measured

Astra's headline score obscures what the ARC-AGI-3 test measured
Astra's 99.9 per cent came from the harness, not the model · wionews.com

OpenAI showed Astra getting a very high score on a test of unfamiliar problems.

However, that score used extra software that remembered parts of earlier attempts and organized the model's context.

When ARC Prize looked at the model under the regular testing conditions, its best score was about 62.7 percent.

This means the two scores measured different setups.

Astra still did something impressive: it used fewer actions than most human testers on nearly all the levels it completed.

That suggests it may have learned efficient ways to understand new environments.

Greg Brockman said he personally believes Astra reached artificial general intelligence, or AGI.

But there is no agreed test that can prove when a system has reached AGI.

Key facts

Headline score
99.9 percent, reported with a provider adapter harness.
Standard-condition score
62.7 percent, identified by ARC Prize as Astra's best observed score on the shared test.
Other circulated score
98.55 percent, or 98.6 percent when rounded, also produced with the harness at maximum reasoning settings.
Human-efficiency comparison
Astra used fewer actions than the median human tester on 96 percent of levels it completed.
Action reduction
Astra averaged 51.7 percent fewer actions than the median human tester.
Human baseline
The comparison used roughly 500 general participants playing the same games.
AGI testing
The article says there is no industry-agreed test separating a highly capable model from AGI.

Quotes

Greg Brockman

OpenAI executive who ended the Astra briefing

“Welcome to the AGI era.”
wionews.com

Sources

Related news