Find a story

Search Spins

Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.

0 results for “SWE-Bench”

SPIN Processed News Frame: The Fog

Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires

A Hacker News discussion thread titled 'Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires' contains user comments critiquing AI software engineering evaluation methodologies, but no original reporting, data, or verifiable claims are presented in the provided content.

Spin 0% Needs Evidence
Hacker News Front Page

Published Sep 11, 2026 · Analyzed Sep 14, 2026

SPIN Processed News Frame: The Cushion

OpenAI says it found widespread task issues in SWE-Bench Pro, estimates ~30% of tasks are broken, and retracts its earlier recommendation to adopt the benchmark (OpenAI)

OpenAI publicly retracted its prior recommendation to adopt SWE-Bench Pro as a benchmark after identifying widespread task failures—estimating ~30% of tasks are broken—triggering scrutiny over benchmark validity and AI evaluation rigor.

Spin 65% Claim Present in Source AI Risk Moderate
Techmeme

Jul 9, 2026

SPIN Processed Company Announcement Frame: The Fog

Separating signal from noise in coding evaluations

OpenAI published a blog post critiquing SWE-Bench Pro, a widely used coding evaluation benchmark, asserting methodological flaws that undermine its reliability for assessing AI coding models.

Spin 75% Needs Evidence AI Risk High
OpenAI Blog

Jul 9, 2026