ARC Prize shows GPT-6 Astra's 99.9% ARC-AGI-3 score came from OpenAI's custom harness, not the model
OpenAI
ARC Prize's own writeup shows Astra scored 62.7% (for $26,098) on ARC-AGI-3 under its standard harness but 99.9% (for $18,817) under OpenAI's Provider Adapter, which preserves opaque reasoning state between requests; even with reasoning effort set to zero the adapter scores 96.7%. Separately, Fortune found five metrics in OpenAI's launch post were revised after publication (Astra's hallucination rate moved 4.2% -> 2% -> 4.2%; Fable 5.1's FrontierMath 87.8% -> 78% -> 83%), and ARC Prize disavows the AGI framing: 'we lack evidence to call this AGI yet.'
Why it matters
A benchmark organizer formally documenting that harness choice, not model capability, produced the headline AGI-claim score — a template case for benchmark transparency.
Importance: 4/5
Benchmark-governance story confirmed by 4 independent sources