Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
4 results for “LLM serving”
A Year in LLM Serving: Workload Evolution, Caching and Load-Balancing
Researchers released a one-year production trace of LLM serving traffic from Chutes to enable more realistic benchmarking and system design, addressing gaps in scale, duration, and granularity of prior workload studies.
Aug 17, 2026
Netflix Details Its In-House LLM Serving Platform with Triton and vLLM
Netflix shared internal engineering insights on deploying LLM inference at scale using Triton and vLLM, revealing technical trade-offs in model serving but not announcing a new product, policy, or external offering.
Jul 27, 2026
FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads
FineServe is a newly released, real-world dataset capturing fine-grained LLM serving workloads from a global commercial marketplace, designed to improve benchmarking and systems design for multi-model LLM deployment.
Jul 23, 2026
Kara: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression
Kara is a new sliding-window KV cache compression method for reasoning LLMs that improves decoding throughput and reduces memory overhead by selectively preserving flexible-sized semantic chunks of the key-value cache during inference.
Published Jul 3, 2026 · Analyzed Jul 6, 2026