Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
0 results for “inference cost”
Dual-Flow Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation
Researchers propose Dual-Flow Transformers, a novel architecture that decouples prompt prefill and autoregressive decode computation to reduce cumulative inference cost without increasing prefill overhead.
Aug 14, 2026
Loss-Aware Feature-Map Pruning in Convolutional Neural Networks Using Multi-Armed Bandits
A new research paper introduces a feature-map pruning method for CNNs using multi-armed bandit algorithms to selectively remove redundant convolutional channels while preserving model accuracy and reducing compute.
Jul 28, 2026
Google's "Frozen v2" chip reportedly bakes Gemini's architecture directly into silicon for efficiency gains
Google is reportedly developing a custom server chip called 'Frozen v2' that hardcodes Gemini's architecture into silicon, aiming for 6–10× efficiency gains over current TPUs by 2028 to reduce inference costs and gain competitive pricing leverage.
Jul 21, 2026
OpenAI Halves Inference Costs With Software Alone: GPUs Drop to Hundreds - Tech Times
OpenAI claims to have reduced AI inference costs by 50% using only software optimizations, enabling deployment on cheaper hardware like sub-$1,000 GPUs.
Published Jul 3, 2026 · Analyzed Jul 6, 2026
OpenAI Discovers New Way to Cut Inference Costs in Half - The Information
OpenAI claims to have developed a novel method that reduces AI inference costs by 50%, potentially improving model deployment economics and scalability.
Published Jun 30, 2026 · Analyzed Jul 4, 2026