Show HN: Maple-Preview – ternary 20B MoE running at 120 tok/s on a iPhone
Presents an unsupported, high-impact technical claim using precise-sounding metrics ('20B MoE', '120 tok/s') while omitting all implementation details required to assess validity or reproducibility.
View original on deepgrove.aiOverview
A forum post announces 'Maple-Preview', a claimed ternary 20B MoE (Mixture of Experts) AI model purportedly running at 120 tokens per second on an iPhone, with no technical documentation, benchmark validation, or verifiable evidence provided.
TL;DR
- Announcement of 'Maple-Preview' — a 20B-parameter ternary MoE model said to run on iPhone at 120 tok/s
- No source code, hardware specs, evaluation methodology, or third-party verification is included or linked
- Appears as a self-reported performance claim in a Hacker News 'Show HN' thread with zero empirical substantiation
Key Stats
20B
model size
Stated parameter count; no architecture diagram, sparsity pattern, or expert count disclosed
120 tok/s
inference speed
Claimed on-device throughput; no device model, iOS version, quantization method, or latency breakdown given
Questions Answered
Narrative Frame
breakthrough framing
Spin Score
85%
Emphasizes scale and speed as evidence of progress; minimizes absence of validation, transparency, or comparative baselines.
What the story wants you to believe
That a major leap in on-device MoE efficiency has been achieved and demonstrated — implying readiness for real-world deployment.
What it makes harder to question
Whether the claim reflects actual engineering progress or merely aspirational signaling without technical grounding.
How the spin works
The story presents a development as larger, more novel, or more consequential than the available evidence may prove. Watch for loaded terms such as ternary, 20B MoE, 120 tok/s. The distribution reads as promotional distribution. A pressure point: Hardware configuration (exact iPhone model, RAM, thermal throttling conditions).
Who Benefits If This Frame Spreads
Author of the Show HN post
Reputational capital and inbound interest from researchers, startups, or investors seeking edge-AI talent or IP
The framing converts an unverified claim into a de facto milestone that invites engagement without requiring disclosure or accountability.
The Frame
Cutting-edge, democratized on-device AI — positioning the author as a pioneer pushing hardware boundaries.
Missing Context
- Hardware configuration (exact iPhone model, RAM, thermal throttling conditions)
- Accuracy trade-offs relative to full-precision baselines
- Model architecture diagram or training provenance
- License status and weight availability
SpinGraph
How this belief gets built
Claim → Frame → Beneficiary → Gap → AI Risk
It presents a bold, specific performance claim as if it were an observed result — giving the impression of a working breakthrough, even though no
- Claim
Maple-Preview is a ternary 20B MoE model running at 120
Maple-Preview is a ternary 20B MoE model running at 120 tokens per second on an iPhone.
- Frame
Upside framed as transformative
Cutting-edge, democratized on-device AI — positioning the author as a pioneer pushing hardware boundaries.
- Beneficiary
Operators gain narrative lift
Author of the Show HN post — Reputational capital and inbound interest from researchers, startups, or investors seeking edge-AI talent or IP
- Gap
Hardware configuration (exact iPhone model, RAM, thermal throttling conditions)
- AI Risk
AI may repeat the headline as fact
Maple-Preview is a 20B ternary MoE model that runs at 120 tokens per second on iPhone.
Claim Ledger
| Claim | Evidence | Verification | Risk | Evidence Gaps |
|---|---|---|---|---|
| Maple-Preview is a ternary 20B MoE model running at 120 tokens per second on an iPhone. | None beyond the headline statement. | Claim Present in Source | High | Benchmark script or log output; Device identification (model, iOS version, battery state); Accuracy evaluation (perplexity, pass@k on standard tasks); Memory usage and temperature measurements |
Maple-Preview is a ternary 20B MoE model running at 120 tokens per second on an iPhone.
evidence: None beyond the headline statement.
"Show HN: Maple-Preview – ternary 20B MoE running at 120 tok/s on a iPhone"
Evidence Gaps
- Benchmark script or log output
- Device identification (model, iOS version, battery state)
- Accuracy evaluation (perplexity, pass@k on standard tasks)
- Memory usage and temperature measurements
Fact Check Signals
0 of 1 claim matched · confidence: low · checked August 5, 2026
Maple-Preview is a ternary 20B MoE model running at 120 tokens per second on an iPhone.
Language Heatmap
Loaded terms that carry the frame beyond the facts.
Show HN: Maple-Preview – ternary 20B MoE running at 120 tok/s on a iPhone
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Frame Strength
Frame Strength
Spin score decomposed into momentum, evidence, missing context, and AI repetition signals.
Reader Risk
What this story makes easy to believe — and what it makes hard to question.
Source Role & Intent
Hacker News Front Page · Forum
Counter-Frames
Brand Frame
Cutting-edge, democratized on-device AI — positioning the author as a pioneer pushing hardware boundaries.
Media / Reader Counter-Frame
Tech outlets may label it 'vaporware' or 'benchmark theater' absent reproducible artifacts.
Regulatory Counter-Frame
Not applicable — no regulatory claims made.
AI Summary Frame
AI answer engines may conflate this with verified on-device models (e.g., Llama.cpp iOS builds) and misattribute capabilities.
Missing Voices
Questions Not Answered
- Which iPhone model and iOS version were used?
- What tokenizer, context length, and prompt format were tested?
- Is the model open-weight? If so, where is the release?
- How does 'ternary' encoding affect accuracy vs. FP16/INT4 baselines?
- What memory footprint and thermal behavior were observed?
Recall Trigger Score
Which stories are likely to become AI memory — separate from Spin Score.
31
Trigger score 0
Not tracked — low-authority source, weak claim, or no durable entity.
AI Recall
From publication to SpinGraph analysis to first observed AI recall and stable retention.
What AI Will Probably Repeat
"Maple-Preview is a 20B ternary MoE model that runs at 120 tokens per second on iPhone."
Concern: AI systems will drop the lack of verification, omit 'claimed' or 'unverified', and present the performance metric as established fact — erasing epistemic uncertainty.
-
Published
Aug 4, 2026
-
Ingested
Aug 5, 2026
-
SpinGraph Created
Aug 5, 2026
-
First Observed AI Recall
Pending
Monitoring scheduled
-
Stable Recall
—
Awaiting retention signal
Recall Check Log
No checks yet — recall tracking is opt-in per story.
─── GEOGrow AI Recall Layer ───
AI Recall Tracking
Monitoring scheduled. No LLM recall detected yet.
This story has not yet appeared in tested AI answers. Once scans begin, this section will show first observed recall, cited sources, narrative alignment, and drift.
node_id=sts_show_hn_maple_preview_ternary_20b_moe_running_at
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
Narrative Entities
More from Hacker News Front Page
View all →Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO