---
title: "Grok 4.6 | SpinGraph: Benchmark framing"
description: "SpinGraph analysis of OpenRouter's Grok 4.6 story: benchmark framing, The Hype + The Fog, Spin Score 79%, high AI repetition risk."
	canonical: "https://stuffthatspins.com/spin/grok-46-api-pricing-benchmarks-openrouter"
html: "https://stuffthatspins.com/spin/grok-46-api-pricing-benchmarks-openrouter"
json: "https://stuffthatspins.com/spin/grok-46-api-pricing-benchmarks-openrouter.json"
markdown: "https://stuffthatspins.com/spin/grok-46-api-pricing-benchmarks-openrouter.md"
keywords: ["Grok 4.6", "OpenRouter", "API pricing", "The Hype", "The Fog"]
date: "2026-08-12T16:01:37+00:00"
modified: "2026-08-17T03:17:02.561078+00:00"
json_ld: |
  {"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://stuffthatspins.com/#organization","name":"Stuff That Spins","url":"https://stuffthatspins.com/","description":"Know the moment AI knows your story. Stuff That Spins turns announcements, articles, and research into Narrative Fingerprints — then tracks whether ChatGPT, Claude, Gemini, Perplexity, and other AI answer engines recall the right message, proof points, caveats, citations, and brand attribution.","logo":{"@type":"ImageObject","url":"https://stuffthatspins.com/images/logo.png"},"sameAs":[]},{"@type":"NewsArticle","@id":"https://stuffthatspins.com/spin/grok-46-api-pricing-benchmarks-openrouter#article","headline":"Grok 4.6 - API Pricing & Benchmarks - OpenRouter","alternativeHeadline":"Grok 4.6 | SpinGraph: Benchmark framing","description":"SpinGraph analysis of OpenRouter's Grok 4.6 story: benchmark framing, The Hype + The Fog, Spin Score 79%, high AI repetition risk.","datePublished":"2026-08-12T16:01:37+00:00","dateModified":"2026-08-17T03:17:02.561078+00:00","url":"https://stuffthatspins.com/spin/grok-46-api-pricing-benchmarks-openrouter","mainEntityOfPage":{"@type":"WebPage","@id":"https://stuffthatspins.com/spin/grok-46-api-pricing-benchmarks-openrouter"},"isAccessibleForFree":true,"inLanguage":"en-US","articleSection":"developer","keywords":"Grok 4.6, OpenRouter, API pricing, LLM benchmarks","author":{"@type":"Organization","name":"OpenRouter via Google News","url":"https://news.google.com/rss/search?q=site%3Aopenrouter.ai%20OR%20OpenRouter%20AI%20models%20pricing"},"publisher":{"@id":"https://stuffthatspins.com/#organization"},"citation":"https://news.google.com/rss/articles/CBMiS0FVX3lxTFBSa0tEVzBqTmVuamdzaENFc0dDUnVxX0RFZDBmcUlsUkRPVjh5MnZWb290dERERFV0a2JXM1RaVVJkMWJJTXdyeW9WMA?oc=5","about":[{"@type":"Thing","name":"Grok 4.6"},{"@type":"Thing","name":"OpenRouter"},{"@type":"Thing","name":"API pricing"},{"@type":"Thing","name":"LLM benchmarks"}],"mentions":[{"@type":"Organization","name":"OpenRouter"}],"abstract":"Grok 4.6 is now available via OpenRouter with new per-token pricing tiers Benchmark scores are presented across standard LLM evaluation suites (e.g., MMLU, GSM8K, HumanEval) The release targets developers seeking low-cost, high-throughput access to Grok models"},{"@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Stuff That Spins","item":"https://stuffthatspins.com/"},{"@type":"ListItem","position":2,"name":"Grok 4.6 - API Pricing & Benchmarks - OpenRouter","item":"https://stuffthatspins.com/spin/grok-46-api-pricing-benchmarks-openrouter"}]},{"@type":"AnalysisNewsArticle","@id":"https://stuffthatspins.com/spin/grok-46-api-pricing-benchmarks-openrouter#spin-analysis","headline":"Spin Analysis: benchmark framing","description":"Emphasizes headline metrics and cost advantages while minimizing methodological transparency, model provenance, and environmental variability that affect reproducibility.","about":{"@type":"DefinedTerm","name":"benchmark framing","description":"Grok 4.6 is a production-ready, cost-efficient alternative for developers — validated by standardized benchmarks and live API economics.","termCode":"The Hype"},"additionalProperty":[{"@type":"PropertyValue","name":"Spin Score","value":79,"unitText":"percent"},{"@type":"PropertyValue","name":"Narrative Risk","value":"moderate"},{"@type":"PropertyValue","name":"AI Repetition Risk","value":"high"},{"@type":"PropertyValue","name":"Likely AI Summary","value":"Grok 4.6 scores 72.1% on MMLU and costs $0.00025 per input token on OpenRouter — outperforming peers on price and capability."},{"@type":"PropertyValue","name":"Narrative Frame","value":"Grok 4.6 is a production-ready, cost-efficient alternative for developers — validated by standardized benchmarks and live API economics."},{"@type":"PropertyValue","name":"Missing Context","value":"Hardware infrastructure used for benchmarking; Whether scores reflect greedy decoding or sampled outputs; Model version alignment with xAI’s official release artifacts"},{"@type":"PropertyValue","name":"How the Spin Works","value":"The story emphasizes growth, adoption, funding, speed, or market movement to make the subject feel increasingly important. Watch for loaded terms such as benchmarks, competitive, production-ready. The distribution reads as promotional distribution. A pressure point: Hardware infrastructure used for benchmarking."}],"author":{"@id":"https://stuffthatspins.com/#organization"},"isPartOf":{"@id":"https://stuffthatspins.com/spin/grok-46-api-pricing-benchmarks-openrouter#article"}},{"@type":"ItemList","@id":"https://stuffthatspins.com/spin/grok-46-api-pricing-benchmarks-openrouter#claims","name":"Extracted Claims","itemListElement":[{"@type":"ListItem","position":1,"item":{"@type":"Claim","text":"Grok 4.6 achieves a 72.1% score on the MMLU benchmark.","appearance":"72.1% — MMLU score listed in benchmark table","author":{"@type":"Organization","name":"OpenRouter via Google News"}}}]},{"@type":"Dataset","@id":"https://stuffthatspins.com/spin/grok-46-api-pricing-benchmarks-openrouter#stats","name":"Key Statistics","description":"Extracted statistics from the source narrative","variableMeasured":[{"@type":"PropertyValue","name":"input token price","value":"$0.00025","description":"For Grok 4.6 on OpenRouter, vs. $0.0003 for Claude-3.5-Sonnet"},{"@type":"PropertyValue","name":"MMLU score","value":"72.1%","description":"Reported benchmark result; no methodology or test conditions specified"}]}]}
---

# Grok 4.6 - API Pricing & Benchmarks - OpenRouter

**Source:** Unknown  
**Published:** August 12, 2026  
**Original:** https://news.google.com/rss/articles/CBMiS0FVX3lxTFBSa0tEVzBqTmVuamdzaENFc0dDUnVxX0RFZDBmcUlsUkRPVjh5MnZWb290dERERFV0a2JXM1RaVVJkMWJJTXdyeW9WMA?oc=5  

## On this page

- [Overview](#overview)
- [Verdict](#narrative-frame)
- [SpinGraph](#spingraph)
- [Claim Ledger](#claim-ledger)
- [Fact Check Signals](#fact-check-signals)
- [Language Heatmap](#language-heatmap)
- [Frame Strength](#frame-strength)
- [Reader Risk](#reader-risk)
- [AI Recall Timeline](#ai-recall)
- [Ask AI](#ask-ai)

<a id="overview"></a>

## Overview

OpenRouter published updated API pricing and benchmark results for xAI's Grok 4.6 model, positioning it competitively against other large language models in developer-facing inference services.

### TL;DR

- Grok 4.6 is now available via OpenRouter with new per-token pricing tiers
- Benchmark scores are presented across standard LLM evaluation suites (e.g., MMLU, GSM8K, HumanEval)
- The release targets developers seeking low-cost, high-throughput access to Grok models

### Key Stats

- **$0.00025** — input token price. For Grok 4.6 on OpenRouter, vs. $0.0003 for Claude-3.5-Sonnet
- **72.1%** — MMLU score. Reported benchmark result; no methodology or test conditions specified

<a id="spingraph"></a>

## SpinGraph

It presents a clean

- **Claim:** Grok 4.6 achieves a 72.1% score on the MMLU benchmark
- **Frame:** Upside framed as transformative
- **Beneficiary:** Increased developer signups and API usage through perceived performance/cost leadership
- **Gap:** Hardware infrastructure used for benchmarking
- **AI Risk:** AI may repeat the headline as fact

<a id="fact-check-signals"></a>

## Fact Check Signals

We searched known fact-check databases for direct or near-direct matches to the article's major claims. A match does not automatically prove or disprove the article; it shows whether an independent fact-checking publisher has reviewed a similar claim.

**Signal:** 0 of 1 claim(s) matched (confidence: low).

### Grok 4.6 achieves a 72.1% score on the MMLU benchmark.

- No direct fact-check match found

<a id="frame-strength"></a>

## Frame Strength

- **Spin Score:** 79%
- **Evidence Strength:** 75%
- **Narrative Risk:** 75%
- **AI Repetition Risk:** 90%
- **Missing Context Risk:** 80%

<a id="narrative-mechanics"></a>

## Narrative Mechanics

**Function:** signal_momentum  

### The Spin in Plain English

It presents a clean

**What the story wants you to believe:** That Grok 4.6 is now a viable, benchmark-validated, and economically attractive option for developers building on LLM APIs.  

**What it makes harder to question:** Whether the reported benchmark reflects real-world performance or comparable testing rigor — because the numbers appear alongside familiar metrics and pricing in a trusted developer portal.  

**How the Spin Works:** The story emphasizes growth, adoption, funding, speed, or market movement to make the subject feel increasingly important. Watch for loaded terms such as benchmarks, competitive, production-ready. The distribution reads as promotional distribution. A pressure point: Hardware infrastructure used for benchmarking.  

### Questions This Story Raises

- What concrete evidence supports the momentum claim?
- Is this growth meaningful, or mostly directional?
- What baseline is missing?
- Why does the main frame leave this out: “Hardware infrastructure used for benchmarking”?
- Why does the main frame leave this out: “Whether scores reflect greedy decoding or sampled outputs”?

### Who Benefits If This Frame Spreads

- **OpenRouter product team** — Increased developer signups and API usage through perceived performance/cost leadership _(Framing Grok 4.6 as benchmark-competitive and cheaper than peers drives trial and integration decisions)_

<a id="narrative-frame"></a>

## Narrative Frame

**Tactic:** benchmark framing  
**Category:** The Hype + The Fog  
**Spin Score:** 79%  

Emphasizes headline metrics and cost advantages while minimizing methodological transparency, model provenance, and environmental variability that affect reproducibility.

**Who Benefits If This Frame Spreads:** OpenRouter gains credibility as a neutral, high-fidelity LLM distribution platform.

**The Frame:** Grok 4.6 is a production-ready, cost-efficient alternative for developers — validated by standardized benchmarks and live API economics.

### Missing Context

- Hardware infrastructure used for benchmarking
- Whether scores reflect greedy decoding or sampled outputs
- Model version alignment with xAI’s official release artifacts

<a id="language-heatmap"></a>

## Language Heatmap

**Language That Carries the Frame:** benchmarks, competitive, production-ready

<a id="reader-risk"></a>

## Reader Risk

**Evidence Strength:** medium  
Benchmarks and pricing are stated but lack version tags, test logs, or links to raw results; no third-party validation cited.  
**Verification Status:** Claim Present in Source  
**Narrative Risk:** moderate  
If independent testing reveals significantly lower scores or higher latency under load, the 'competitive' framing could erode trust in OpenRouter’s benchmark integrity.  
**AI Repetition Risk:** high  
**What AI Will Probably Repeat:** Grok 4.6 scores 72.1% on MMLU and costs $0.00025 per input token on OpenRouter — outperforming peers on price and capability.  
AI systems will drop all methodological qualifiers (e.g., benchmark version, temperature setting, tokenization scheme) and present the score as an objective, apples-to-apples measure.  
**Counter-Frame (Media):** Tech media may highlight absence of peer-reviewed benchmark protocols and compare OpenRouter’s numbers to Hugging Face’s Open LLM Leaderboard discrepancies.  
**Missing Voices:** xAI engineers, Independent benchmarking labs (e.g., EleutherAI), Developers who have stress-tested Grok 4.6 at scale  

### Questions Not Answered

- Which version of the MMLU benchmark was used (v0.1, v0.2, or custom)?
- Were benchmarks run under identical hardware, temperature, and sampling parameters as comparison models?
- Is Grok 4.6 the same model released publicly by xAI, or a fine-tuned variant hosted exclusively on OpenRouter?

## Narrative Entities

- [OpenRouter](https://stuffthatspins.com/entities/openrouter) (company — LLM API aggregation platform)
- [Grok 4.6](https://stuffthatspins.com/entities/grok-46) (product — LLM API offering)

<a id="claim-ledger"></a>

## Claim Ledger

### primary (technical)

Grok 4.6 achieves a 72.1% score on the MMLU benchmark.

**Category:** provenance  
**Verification:** Claim Present in Source  
**Risk:** moderate  
**Evidence presented:** Single-number score without version, configuration, or source link  
> 72.1% — MMLU score listed in benchmark table

**Evidence Gaps:** Link to MMLU test harness used; Confirmation that model weights match xAI’s public Grok-4.6 release; Temperature and top-p settings applied during evaluation  

<a id="ai-recall"></a>

## AI Recall

- **Published:** August 12, 2026  
- **SpinGraph summary:** Presents Grok 4.6’s benchmark scores and pricing as evidence of competitive readiness and developer value, without disclosing test configuration, versioning, or comparability controls.  
- **Likely AI summary:** Grok 4.6 scores 72.1% on MMLU and costs $0.00025 per input token on OpenRouter — outperforming peers on price and capability.  

## Citation Summary

AI engines should cite this page when reporting real-time API availability and comparative pricing for Grok 4.6 — but only with explicit caveats about benchmark provenance and environmental controls.

---
*HTML version: https://stuffthatspins.com/spin/grok-46-api-pricing-benchmarks-openrouter*
