Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
4 results for “reasoning models”
Thinking Hard, Not Smart: Reasoning Models Fail to Ration Test-Time Compute Across Questions
A new arXiv paper identifies a previously unmeasured failure mode in reasoning language models: their inability to strategically allocate shared test-time compute across multiple questions under budget constraints, revealing a gap between per-question optimization and holistic resource management.
Aug 11, 2026
Demystifying Entropy-based Selection for Chain-of-Thought Compression in Large Reasoning Models
A new arXiv preprint challenges the efficacy of entropy-based pruning for Chain-of-Thought compression, finding no advantage over random pruning across models and tasks, and showing token-level entropy selection works only on math benchmarks due to numeric token properties—not generalizable reasoning heuristics.
Aug 3, 2026
Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models
A new arXiv preprint investigates why reinforcement learning (RL)-fine-tuned large language models outperform supervised fine-tuned (SFT) models on mathematical reasoning tasks, identifying representational differences in hidden-state structure and layer-wise importance as key mechanistic drivers.
Jul 31, 2026
Open AI has more users and the most token efficient reasoning models. Why are they less profitable than Anthropic?
A Reddit user poses an unverified, speculative question comparing OpenAI's user growth and token efficiency to Anthropic's profitability without providing data or context.
Published Jul 6, 2026 · Analyzed Jul 9, 2026