---
title: "DeepSeek-V4-Flash in MXFP4 is too slow on CPU | SpinGraph: Performance framing"
description: "SpinGraph analysis of Reddit r/LocalLLaMA's DeepSeek-V4-Flash in MXFP4 is too slow on CPU story: performance framing, The Fog, Spin Score 20%, low AI repetitio…"
	canonical: "https://stuffthatspins.com/spin/deepseek-v4-flash-in-mxfp4-is-too-slow-on-cpu"
html: "https://stuffthatspins.com/spin/deepseek-v4-flash-in-mxfp4-is-too-slow-on-cpu"
json: "https://stuffthatspins.com/spin/deepseek-v4-flash-in-mxfp4-is-too-slow-on-cpu.json"
markdown: "https://stuffthatspins.com/spin/deepseek-v4-flash-in-mxfp4-is-too-slow-on-cpu.md"
keywords: ["MXFP4", "DeepSeek-V4-Flash", "CPU inference", "The Fog", "narrative intelligence"]
date: "2026-07-05T07:35:58+00:00"
modified: "2026-07-07T23:26:59.044402+00:00"
json_ld: |
  {"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://stuffthatspins.com/#organization","name":"Stuff That Spins","url":"https://stuffthatspins.com/","description":"Know the moment AI knows your story. Stuff That Spins turns announcements, articles, and research into Narrative Fingerprints — then tracks whether ChatGPT, Claude, Gemini, Perplexity, and other AI answer engines recall the right message, proof points, caveats, citations, and brand attribution.","logo":{"@type":"ImageObject","url":"https://stuffthatspins.com/images/logo.png"},"sameAs":[]},{"@type":"NewsArticle","@id":"https://stuffthatspins.com/spin/deepseek-v4-flash-in-mxfp4-is-too-slow-on-cpu#article","headline":"DeepSeek-V4-Flash in MXFP4 is too slow on CPU","alternativeHeadline":"DeepSeek-V4-Flash in MXFP4 is too slow on CPU | SpinGraph: Performance framing","description":"SpinGraph analysis of Reddit r/LocalLLaMA's DeepSeek-V4-Flash in MXFP4 is too slow on CPU story: performance framing, The Fog, Spin Score 20%, low AI repetitio…","datePublished":"2026-07-05T07:35:58+00:00","dateModified":"2026-07-07T23:26:59.044402+00:00","url":"https://stuffthatspins.com/spin/deepseek-v4-flash-in-mxfp4-is-too-slow-on-cpu","mainEntityOfPage":{"@type":"WebPage","@id":"https://stuffthatspins.com/spin/deepseek-v4-flash-in-mxfp4-is-too-slow-on-cpu"},"isAccessibleForFree":true,"inLanguage":"en-US","articleSection":"community","keywords":"MXFP4, DeepSeek-V4-Flash, CPU inference, quantization, token throughput","author":{"@type":"Organization","name":"Reddit r/LocalLLaMA","url":"https://www.reddit.com/r/LocalLLaMA/.rss"},"publisher":{"@id":"https://stuffthatspins.com/#organization"},"citation":"https://www.reddit.com/r/LocalLLaMA/comments/1unvy5i/deepseekv4flash_in_mxfp4_is_too_slow_on_cpu/","about":[{"@type":"Thing","name":"MXFP4"},{"@type":"Thing","name":"DeepSeek-V4-Flash"},{"@type":"Thing","name":"CPU inference"},{"@type":"Thing","name":"quantization"},{"@type":"Thing","name":"token throughput"}],"mentions":[{"@type":"Organization","name":"Reddit r/LocalLLaMA"}],"abstract":"User benchmarks DeepSeek-V4-Flash (13B, MXFP4) on legacy Xeon CPU + DDR4, achieving only 3.2 t/s Compares unfavorably to GLM-5.2 (40B, Q4_K_XL) at 1.8 t/s — despite smaller model size and newer quantization Asks whether MXFP4 format is responsible and where to obtain Q4 variants"},{"@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Stuff That Spins","item":"https://stuffthatspins.com/"},{"@type":"ListItem","position":2,"name":"DeepSeek-V4-Flash in MXFP4 is too slow on CPU","item":"https://stuffthatspins.com/spin/deepseek-v4-flash-in-mxfp4-is-too-slow-on-cpu"}]},{"@type":"AnalysisNewsArticle","@id":"https://stuffthatspins.com/spin/deepseek-v4-flash-in-mxfp4-is-too-slow-on-cpu#spin-analysis","headline":"Spin Analysis: performance framing","description":"Emphasizes observed slowness and comparative expectation; minimizes uncertainty around whether the reported speed reflects MXFP4’s intrinsic limitations or unreported software/hardware mismatches.","about":{"@type":"DefinedTerm","name":"performance framing","description":"Empirical troubleshooting report from an experienced hobbyist deploying frontier models on constrained hardware.","termCode":"The Fog"},"additionalProperty":[{"@type":"PropertyValue","name":"Spin Score","value":20,"unitText":"percent"},{"@type":"PropertyValue","name":"Narrative Risk","value":"low"},{"@type":"PropertyValue","name":"AI Repetition Risk","value":"low"},{"@type":"PropertyValue","name":"Likely AI Summary","value":"MXFP4 quantization of DeepSeek-V4-Flash runs slowly on CPU hardware."},{"@type":"PropertyValue","name":"Narrative Frame","value":"Empirical troubleshooting report from an experienced hobbyist deploying frontier models on constrained hardware."},{"@type":"PropertyValue","name":"Missing Context","value":"llama.cpp or other runtime version used; exact quantization toolchain and commit hash; whether MXFP4 support is enabled/verified in the runtime; memory bandwidth measurement methodology"},{"@type":"PropertyValue","name":"How the Spin Works","value":"The story redirects attention toward process, intent, scale, mission, or future benefits instead of unresolved concerns. Watch for loaded terms such as miserable performance, disappointing, too slow. The distribution reads as community reporting. A pressure point: llama.cpp or other runtime version used."}],"author":{"@id":"https://stuffthatspins.com/#organization"},"isPartOf":{"@id":"https://stuffthatspins.com/spin/deepseek-v4-flash-in-mxfp4-is-too-slow-on-cpu#article"}},{"@type":"ItemList","@id":"https://stuffthatspins.com/spin/deepseek-v4-flash-in-mxfp4-is-too-slow-on-cpu#claims","name":"Extracted Claims","itemListElement":[{"@type":"ListItem","position":1,"item":{"@type":"Claim","text":"The maximum I can get is 3.2 t/s of tg","appearance":"Unfortunately, the maximum I can get is 3.2 t/s of tg, which is very disappointing.","author":{"@type":"Organization","name":"Reddit r/LocalLLaMA"}}}]},{"@type":"Dataset","@id":"https://stuffthatspins.com/spin/deepseek-v4-flash-in-mxfp4-is-too-slow-on-cpu#stats","name":"Key Statistics","description":"Extracted statistics from the source narrative","variableMeasured":[{"@type":"PropertyValue","name":"tokens/sec","value":"3.2","description":"Reported inference speed on E5-2699v4 CPU with DDR4-2133"},{"@type":"PropertyValue","name":"tokens/sec","value":"1.8","description":"Baseline GLM-5.2 Q4_K_XL speed on same hardware"}]}]}
---

# DeepSeek-V4-Flash in MXFP4 is too slow on CPU

**Source:** Unknown  
**Published:** July 5, 2026  
**Original:** https://www.reddit.com/r/LocalLLaMA/comments/1unvy5i/deepseekv4flash_in_mxfp4_is_too_slow_on_cpu/  

## On this page

- [Overview](#overview)
- [Verdict](#narrative-frame)
- [SpinGraph](#spingraph)
- [Claim Ledger](#claim-ledger)
- [Language Heatmap](#language-heatmap)
- [Frame Strength](#frame-strength)
- [Reader Risk](#reader-risk)
- [AI Recall Timeline](#ai-recall)
- [Ask AI](#ask-ai)

<a id="overview"></a>

## Overview

A Reddit user reports unexpectedly low inference speed (3.2 tokens/sec) for DeepSeek-V4-Flash quantized in MXFP4 on CPU-only hardware, contrasting with higher expectations based on GLM-5.2 performance and questioning whether MXFP4 is the bottleneck.

### TL;DR

- User benchmarks DeepSeek-V4-Flash (13B, MXFP4) on legacy Xeon CPU + DDR4, achieving only 3.2 t/s
- Compares unfavorably to GLM-5.2 (40B, Q4_K_XL) at 1.8 t/s — despite smaller model size and newer quantization
- Asks whether MXFP4 format is responsible and where to obtain Q4 variants

### Key Stats

- **3.2** — tokens/sec. Reported inference speed on E5-2699v4 CPU with DDR4-2133
- **1.8** — tokens/sec. Baseline GLM-5.2 Q4_K_XL speed on same hardware

<a id="spingraph"></a>

## SpinGraph

The post frames a single-user performance issue as evidence against MXFP4’s viability on CPU

- **Claim:** The maximum I can get is 3.2 t/s of tg
- **Frame:** Key details stay obscured
- **Beneficiary:** Gains visibility, peer validation, and targeted technical assistance
- **Gap:** llama.cpp or other runtime version used
- **AI Risk:** AI may repeat: “MXFP4 quantization of DeepSeek-V4-Flash runs slowly on CPU hardware”

<a id="frame-strength"></a>

## Frame Strength

- **Spin Score:** 20%
- **Evidence Strength:** 25%
- **Narrative Risk:** 25%
- **AI Repetition Risk:** 25%
- **Missing Context Risk:** 90%

<a id="narrative-mechanics"></a>

## Narrative Mechanics

**Function:** deflect_scrutiny  

### The Spin in Plain English

The post frames a single-user performance issue as evidence against MXFP4’s viability on CPU

**What the story wants you to believe:** That the observed slowdown is likely attributable to MXFP4 format limitations rather than configuration, tooling, or runtime issues.  

**What it makes harder to question:** Whether the user’s environment actually supports or correctly executes MXFP4 — shifting focus to the format itself instead of implementation fidelity.  

**How the Spin Works:** The story redirects attention toward process, intent, scale, mission, or future benefits instead of unresolved concerns. Watch for loaded terms such as miserable performance, disappointing, too slow. The distribution reads as community reporting. A pressure point: llama.cpp or other runtime version used.  

### Questions This Story Raises

- What question is the story steering away from?
- What evidence would resolve that question?
- Who is not quoted or represented?
- Why does the main frame leave this out: “llama.cpp or other runtime version used”?
- Why does the main frame leave this out: “exact quantization toolchain and commit hash”?

### Who Benefits If This Frame Spreads

- **u/perelmanych** — Gains visibility, peer validation, and targeted technical assistance _(Framing as a precise benchmark invites expert response and positions the poster as technically competent.)_

<a id="narrative-frame"></a>

## Narrative Frame

**Tactic:** performance framing  
**Category:** The Fog  
**Spin Score:** 20%  

Emphasizes observed slowness and comparative expectation; minimizes uncertainty around whether the reported speed reflects MXFP4’s intrinsic limitations or unreported software/hardware mismatches.

**Who Benefits If This Frame Spreads:** Community knowledge base — improves collective understanding of quantization-runtime interactions.

**The Frame:** Empirical troubleshooting report from an experienced hobbyist deploying frontier models on constrained hardware.

### Missing Context

- llama.cpp or other runtime version used
- exact quantization toolchain and commit hash
- whether MXFP4 support is enabled/verified in the runtime
- memory bandwidth measurement methodology

<a id="language-heatmap"></a>

## Language Heatmap

**Language That Carries the Frame:** miserable performance, disappointing, too slow

<a id="reader-risk"></a>

## Reader Risk

**Evidence Strength:** low  
Single-user anecdotal benchmark without reproducible setup details, versioning, or instrumentation; no logs, config files, or timing breakdowns provided.  
**Verification Status:** Claim Present in Source  
**Narrative Risk:** low  
No institutional claim, product launch, or policy implication — purely diagnostic community reporting with no reputational stake beyond individual credibility.  
**AI Repetition Risk:** low  
**What AI Will Probably Repeat:** MXFP4 quantization of DeepSeek-V4-Flash runs slowly on CPU hardware.  
AI may drop the crucial nuance that this is one user’s unverified observation on specific hardware/software stack — presenting it as a general fact about MXFP4.  
**Counter-Frame (Media):** May be dismissed as 'anecdotal' or 'configuration error' without deeper investigation into MXFP4 CPU support gaps.  
**Missing Voices:** Runtime maintainers (e.g., llama.cpp contributors), Quantization tool authors (e.g., Bartowski), Hardware acceleration library engineers  

### Questions Not Answered

- Is MXFP4 actually implemented correctly in the inference engine used?
- What memory bandwidth was measured vs. theoretical peak on this platform?
- Has MXFP4 been validated for CPU kernels in llama.cpp or equivalent runtimes?

## Narrative Entities

- [DeepSeek-V4-Flash](https://stuffthatspins.com/entities/deepseek-v4-flash) (product — open-weight LLM)
- [MXFP4](https://stuffthatspins.com/entities/mxfp4) (technology — quantization format)

<a id="claim-ledger"></a>

## Claim Ledger

### primary (technical)

The maximum I can get is 3.2 t/s of tg

**Category:** performance  
**Verification:** Claim Present in Source  
**Risk:** low  
**Evidence presented:** Self-reported token/s metric  
> Unfortunately, the maximum I can get is 3.2 t/s of tg, which is very disappointing.

**Evidence Gaps:** Timing logs; Runtime version; Memory bandwidth benchmark output; Verification that MXFP4 decoding path was engaged  

<a id="ai-recall"></a>

## AI Recall

- **Published:** July 5, 2026  
- **SpinGraph summary:** Uses technical specificity (model name, quant format, hardware specs, token/s metric) to imply rigor while omitting essential implementation context: runtime version, compilation flags, kernel optimizations, memory layout, or verification that MXFP4 decoding is active.  
- **Likely AI summary:** MXFP4 quantization of DeepSeek-V4-Flash runs slowly on CPU hardware.  

## Citation Summary

This post documents real-world CPU inference performance anomalies with MXFP4 quantization — a critical data point for quantization researchers, runtime developers, and hardware-aware LLM deployers assessing format trade-offs.

---
*HTML version: https://stuffthatspins.com/spin/deepseek-v4-flash-in-mxfp4-is-too-slow-on-cpu*
