SPIN Unprocessed
Source Tomasz Tunguz tomtunguz.com Analyst
July 28, 2026 ai_technology saas

Aftermarket Harnesses

View original on tomtunguz.com

Overview

The harness now moves the coding benchmark more than the model does. Endor Labs' Agent Security League found GPT-5.5 scored 61.5% functional correctness in Codex & 87.2% in Cursor, & Claude Opus 4.7 scored 87.2% in Claude Code & 91.1% in Cursor. Input tokens are 86-98% of OpenRouter volume, so the harness controls most of the bill through cache discipline. First-party co-design buys real cache hit rates, but a third-party harness can match them.

SpinGraph analysis pending — check back after processing.

Ask AI about this story

Opens with the SpinGraph .md URL and structured context — one click, prompt included.

More from Tomasz Tunguz

View all →

Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO