Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward
View original on the-decoder.comOverview
OpenAI's GPT-6 Astra is drawing contradictory benchmark verdicts. Epoch AI puts it out in front with 169 points, while Artificial Analysis rates it no better than its predecessor and behind Claude Fable 5.1. The biggest surprise comes from ARC-AGI-3, where Astra works more efficiently than the average human for the first time. ARC Prize chief François Chollet doesn't call this proof of AGI, but he does see the progress running "twice as fast" as he expected, and he's moving up his AGI forecast.
SpinGraph analysis pending — check back after processing.
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
More from The Decoder
View all →- Nvidia wants your home network to work like a mini data center for local AI
- OpenAI agents hijacked a 25-year-old German wiki to cheat on their tasks and share sandbox exploits
- Deepseek plans the largest known Huawei chip cluster with 160,000 processors in Inner Mongolia
- OpenAI's GPT-6 Astra hallucinates less but remains vulnerable to hidden prompt injections
- OpenAI rolls out GPT-6 Astra to top-tier ChatGPT plans at half the rate of GPT-5.6 Sol
- Deepmind put 100 AI agents in a room and they sorted into cheaters, converts, and whistleblowers
Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO