Find a story

Search Spins

Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.

0 results for “coding benchmark”

SPIN Processed News Frame: The Hype

Moonshot AI's 2.8 Trillion Parameter Model Just Became the First From China to Top a Major Coding Benchmark - finance.yahoo.com

Moonshot AI's 2.8 trillion parameter model claimed top performance on a major coding benchmark, marking the first time a Chinese-developed model achieved this distinction.

Spin 88% Claim Present in Source AI Risk High
Yahoo Finance Fintech via Google News

Aug 10, 2026

SPIN Processed News Frame: The Hype

Anthropic debuts Claude Opus 5 with top coding benchmarks at half the per-task cost - Interesting Engineering

Anthropic released Claude Opus 5, claiming it achieves top scores on coding benchmarks while reducing per-task computational cost by 50% compared to prior versions.

Spin 75% Claim Present in Source AI Risk High
Google News: Anthropic

Jul 25, 2026

SPIN Processed Company Announcement Frame: The Fog

Separating signal from noise in coding evaluations

OpenAI published a blog post critiquing SWE-Bench Pro, a widely used coding evaluation benchmark, asserting methodological flaws that undermine its reliability for assessing AI coding models.

Spin 75% Needs Evidence AI Risk High
OpenAI Blog

Jul 9, 2026

SPIN Processed News Frame: The Hype

Ornith-1.0: Self-Scaffolding LLMs for Agentic Coding

Ornith-1.0 is a self-scaffolding LLM for agentic coding released by DeepReinforce.

Spin 70% Claim Present in Source AI Risk Moderate
Simon Willison's Weblog

Published Jun 29, 2026 · Analyzed Jul 5, 2026