Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
3 results for “tokens/s”
Qwen 3.8 27B available on Cerebras at 1500 tokens/s
A community forum post announces the availability of the Qwen 3.8 27B large language model on Cerebras hardware, highlighting a throughput metric of 1500 tokens/s, with no substantive reporting or verification.
Sep 4, 2026
Same demo, two failures on DeepSeek V4 Pro 0813, then V4 Flash finished it
A Reddit user reports two failed attempts to run a specific demo on DeepSeek V4 Pro 0813, while the same demo succeeded on V4 Flash — highlighting potential reliability or completion issues with the Pro variant despite high token generation speed.
Aug 15, 2026
Running Gemma 4 26B at 5 tokens/sec on a 13-year-old Xeon with no GPU
A user reports running Google's Gemma 4 26B model at 5 tokens/sec on a 13-year-old Intel Xeon CPU without GPU acceleration, highlighting low-resource inference feasibility.
Jul 15, 2026