OpenAI's GPT-6 Astra hallucinates less but remains vulnerable to hidden prompt injections
View original on the-decoder.comOverview
OpenAI's GPT-6 Astra hallucinates less than its predecessor and blocks 99.99 percent of direct prompt injections. But when attacks are hidden inside documents the AI reads, the model still gets cracked in 8.5 percent of scenarios. Claude Opus 5 does better at 4.8 percent. For autonomous AI agents handling real data, those numbers still seem high. The article OpenAI's GPT-6 Astra hallucinates less but remains vulnerable to hidden prompt injections appeared first on The Decoder.
SpinGraph analysis pending — check back after processing.
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
More from The Decoder
View all →- Nvidia wants your home network to work like a mini data center for local AI
- Benchmarks disagree on GPT-6 Astra, but its human-beating efficiency on ARC-AGI-3 pulls Chollet’s AGI forecast forward
- OpenAI agents hijacked a 25-year-old German wiki to cheat on their tasks and share sandbox exploits
- Deepseek plans the largest known Huawei chip cluster with 160,000 processors in Inner Mongolia
- OpenAI rolls out GPT-6 Astra to top-tier ChatGPT plans at half the rate of GPT-5.6 Sol
- Deepmind put 100 AI agents in a room and they sorted into cheaters, converts, and whistleblowers
Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO