Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
0 results for “inference costs”
Correlation-Aware Structured Pruning for Large Language Models
Researchers propose a new structured pruning method for LLMs that models correlations between model units to improve accuracy-efficiency trade-offs during inference cost reduction.
Sep 22, 2026
Gartner Predicts AI Inference Costs Per Agentic Workflow Will Increase More Than Fivefold Through 2028 - Gartner
Gartner forecasts that the cost of running AI inference for agentic workflows will rise by over 500% between now and 2028, signaling a major economic headwind for operational AI deployment.
Aug 18, 2026
Google's "Frozen v2" chip reportedly bakes Gemini's architecture directly into silicon for efficiency gains
Google is reportedly developing a custom server chip called 'Frozen v2' that hardcodes Gemini's architecture into silicon, aiming for 6–10× efficiency gains over current TPUs by 2028 to reduce inference costs and gain competitive pricing leverage.
Jul 21, 2026
OpenAI Halves Inference Costs With Software Alone: GPUs Drop to Hundreds - Tech Times
OpenAI claims to have reduced AI inference costs by 50% using only software optimizations, enabling deployment on cheaper hardware like sub-$1,000 GPUs.
Published Jul 3, 2026 · Analyzed Jul 6, 2026
OpenAI Discovers New Way to Cut Inference Costs in Half - The Information
OpenAI claims to have developed a novel method that reduces AI inference costs by 50%, potentially improving model deployment economics and scalability.
Published Jun 30, 2026 · Analyzed Jul 4, 2026