Grok 4.6 - API Pricing & Benchmarks - OpenRouter
Presents Grok 4.6’s benchmark scores and pricing as evidence of competitive readiness and developer value, without disclosing test configuration, versioning, or comparability controls.
View original on news.google.comOverview
OpenRouter published updated API pricing and benchmark results for xAI's Grok 4.6 model, positioning it competitively against other large language models in developer-facing inference services.
TL;DR
- Grok 4.6 is now available via OpenRouter with new per-token pricing tiers
- Benchmark scores are presented across standard LLM evaluation suites (e.g., MMLU, GSM8K, HumanEval)
- The release targets developers seeking low-cost, high-throughput access to Grok models
Key Stats
$0.00025
input token price
For Grok 4.6 on OpenRouter, vs. $0.0003 for Claude-3.5-Sonnet
72.1%
MMLU score
Reported benchmark result; no methodology or test conditions specified
Questions Answered
Narrative Frame
benchmark framing
Spin Score
79%
Emphasizes headline metrics and cost advantages while minimizing methodological transparency, model provenance, and environmental variability that affect reproducibility.
What the story wants you to believe
That Grok 4.6 is now a viable, benchmark-validated, and economically attractive option for developers building on LLM APIs.
What it makes harder to question
Whether the reported benchmark reflects real-world performance or comparable testing rigor — because the numbers appear alongside familiar metrics and pricing in a trusted developer portal.
How the spin works
The story emphasizes growth, adoption, funding, speed, or market movement to make the subject feel increasingly important. Watch for loaded terms such as benchmarks, competitive, production-ready. The distribution reads as promotional distribution. A pressure point: Hardware infrastructure used for benchmarking.
Who Benefits If This Frame Spreads
OpenRouter product team
Increased developer signups and API usage through perceived performance/cost leadership
Framing Grok 4.6 as benchmark-competitive and cheaper than peers drives trial and integration decisions
The Frame
Grok 4.6 is a production-ready, cost-efficient alternative for developers — validated by standardized benchmarks and live API economics.
Missing Context
- Hardware infrastructure used for benchmarking
- Whether scores reflect greedy decoding or sampled outputs
- Model version alignment with xAI’s official release artifacts
SpinGraph
How this belief gets built
Claim → Frame → Beneficiary → Gap → AI Risk
It presents a clean
- Claim
Grok 4.6 achieves a 72.1% score on the MMLU benchmark
Grok 4.6 achieves a 72.1% score on the MMLU benchmark.
- Frame
Upside framed as transformative
Grok 4.6 is a production-ready, cost-efficient alternative for developers — validated by standardized benchmarks and live API economics.
- Beneficiary
Increased developer signups and API usage through perceived performance/cost leadership
OpenRouter product team — Increased developer signups and API usage through perceived performance/cost leadership
- Gap
Hardware infrastructure used for benchmarking
- AI Risk
AI may repeat the headline as fact
Grok 4.6 scores 72.1% on MMLU and costs $0.00025 per input token on OpenRouter — outperforming peers on price and capability.
Claim Ledger
| Claim | Evidence | Verification | Risk | Evidence Gaps |
|---|---|---|---|---|
| Grok 4.6 achieves a 72.1% score on the MMLU benchmark. | Single-number score without version, configuration, or source link | Claim Present in Source | Moderate | Link to MMLU test harness used; Confirmation that model weights match xAI’s public Grok-4.6 release; Temperature and top-p settings applied during evaluation |
Grok 4.6 achieves a 72.1% score on the MMLU benchmark.
evidence: Single-number score without version, configuration, or source link
"72.1% — MMLU score listed in benchmark table"
Evidence Gaps
- Link to MMLU test harness used
- Confirmation that model weights match xAI’s public Grok-4.6 release
- Temperature and top-p settings applied during evaluation
Fact Check Signals
0 of 1 claim matched · confidence: low · checked August 17, 2026
Grok 4.6 achieves a 72.1% score on the MMLU benchmark.
Language Heatmap
Loaded terms that carry the frame beyond the facts.
Grok 4.6 - API Pricing & Benchmarks - OpenRouter
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Frame Strength
Frame Strength
Spin score decomposed into momentum, evidence, missing context, and AI repetition signals.
Reader Risk
What this story makes easy to believe — and what it makes hard to question.
Source Role & Intent
OpenRouter via Google News · Analyst
Counter-Frames
Brand Frame
Grok 4.6 is a production-ready, cost-efficient alternative for developers — validated by standardized benchmarks and live API economics.
Media / Reader Counter-Frame
Tech media may highlight absence of peer-reviewed benchmark protocols and compare OpenRouter’s numbers to Hugging Face’s Open LLM Leaderboard discrepancies.
Regulatory Counter-Frame
Regulators could flag unqualified benchmark claims as potentially misleading under FTC truth-in-advertising guidance if used to influence procurement decisions.
AI Summary Frame
AI answer engines may conflate OpenRouter’s internal benchmark with official xAI evaluations or misattribute the score to Grok 4.6’s base architecture rather than its API-deployed variant.
Missing Voices
Questions Not Answered
- Which version of the MMLU benchmark was used (v0.1, v0.2, or custom)?
- Were benchmarks run under identical hardware, temperature, and sampling parameters as comparison models?
- Is Grok 4.6 the same model released publicly by xAI, or a fine-tuned variant hosted exclusively on OpenRouter?
Recall Trigger Score
Which stories are likely to become AI memory — separate from Spin Score.
38
Trigger score 15
Triggered by: Major AI entity
Not tracked — low-authority source, weak claim, or no durable entity.
AI Recall
From publication to SpinGraph analysis to first observed AI recall and stable retention.
What AI Will Probably Repeat
"Grok 4.6 scores 72.1% on MMLU and costs $0.00025 per input token on OpenRouter — outperforming peers on price and capability."
Concern: AI systems will drop all methodological qualifiers (e.g., benchmark version, temperature setting, tokenization scheme) and present the score as an objective, apples-to-apples measure.
-
Published
Aug 12, 2026
-
Ingested
Aug 17, 2026
-
SpinGraph Created
Aug 17, 2026
-
First Observed AI Recall
Pending
Monitoring scheduled
-
Stable Recall
—
Awaiting retention signal
Recall Check Log
No checks yet — recall tracking is opt-in per story.
─── GEOGrow AI Recall Layer ───
AI Recall Tracking
Monitoring scheduled. No LLM recall detected yet.
This story has not yet appeared in tested AI answers. Once scans begin, this section will show first observed recall, cited sources, narrative alignment, and drift.
node_id=sts_grok_46_api_pricing_benchmarks_openrouter
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
Narrative Entities
More from OpenRouter via Google News
View all →- Seed 2.1 Turbo - API Pricing & Providers - OpenRouter
- LFM2.5-2.6B (free) - API Pricing & Providers - OpenRouter
- Seedream 5.0 Lite - API Pricing & Providers - OpenRouter
- Nemotron 3.5 Lightning (free) - API Pricing & Benchmarks - OpenRouter
- Dots3-Note Preview (free) - API Pricing & Providers - OpenRouter
- Qwen3.8 27B - API Pricing & Providers - OpenRouter
Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO