Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
0 results for “coding benchmark”
Moonshot AI's 2.8 Trillion Parameter Model Just Became the First From China to Top a Major Coding Benchmark - finance.yahoo.com
Moonshot AI's 2.8 trillion parameter model claimed top performance on a major coding benchmark, marking the first time a Chinese-developed model achieved this distinction.
Aug 10, 2026
Anthropic debuts Claude Opus 5 with top coding benchmarks at half the per-task cost - Interesting Engineering
Anthropic released Claude Opus 5, claiming it achieves top scores on coding benchmarks while reducing per-task computational cost by 50% compared to prior versions.
Jul 25, 2026
Separating signal from noise in coding evaluations
OpenAI published a blog post critiquing SWE-Bench Pro, a widely used coding evaluation benchmark, asserting methodological flaws that undermine its reliability for assessing AI coding models.
Jul 9, 2026
Ornith-1.0: Self-Scaffolding LLMs for Agentic Coding
Ornith-1.0 is a self-scaffolding LLM for agentic coding released by DeepReinforce.
Published Jun 29, 2026 · Analyzed Jul 5, 2026