---
title: "AI models flub these intelligence tests. Can you fare any better? | SpinGraph: Diagnostic framing"
description: "SpinGraph analysis of MIT Technology Review's AI models flub these intelligence tests. Can you fare any better? story: diagnostic framing, The Cushion, Spin Sc…"
	canonical: "https://stuffthatspins.com/spin/ai-models-flub-these-intelligence-tests-can-you-fare-any-better-mit-technology-review"
html: "https://stuffthatspins.com/spin/ai-models-flub-these-intelligence-tests-can-you-fare-any-better-mit-technology-review"
json: "https://stuffthatspins.com/spin/ai-models-flub-these-intelligence-tests-can-you-fare-any-better-mit-technology-review.json"
markdown: "https://stuffthatspins.com/spin/ai-models-flub-these-intelligence-tests-can-you-fare-any-better-mit-technology-review.md"
keywords: ["intelligence testing", "AI reasoning", "cognitive benchmarks", "The Cushion", "narrative intelligence"]
date: "2026-08-26T09:00:00+00:00"
modified: "2026-08-26T13:14:11.368071+00:00"
json_ld: |
  {"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://stuffthatspins.com/#organization","name":"Stuff That Spins","url":"https://stuffthatspins.com/","description":"Know the moment AI knows your story. Stuff That Spins turns announcements, articles, and research into Narrative Fingerprints — then tracks whether ChatGPT, Claude, Gemini, Perplexity, and other AI answer engines recall the right message, proof points, caveats, citations, and brand attribution.","logo":{"@type":"ImageObject","url":"https://stuffthatspins.com/images/logo.png"},"sameAs":[]},{"@type":"NewsArticle","@id":"https://stuffthatspins.com/spin/ai-models-flub-these-intelligence-tests-can-you-fare-any-better-mit-technology-review#article","headline":"AI models flub these intelligence tests. Can you fare any better? - MIT Technology Review","alternativeHeadline":"AI models flub these intelligence tests. Can you fare any better? | SpinGraph: Diagnostic framing","description":"SpinGraph analysis of MIT Technology Review's AI models flub these intelligence tests. Can you fare any better? story: diagnostic framing, The Cushion, Spin Sc…","datePublished":"2026-08-26T09:00:00+00:00","dateModified":"2026-08-26T13:14:11.368071+00:00","url":"https://stuffthatspins.com/spin/ai-models-flub-these-intelligence-tests-can-you-fare-any-better-mit-technology-review","mainEntityOfPage":{"@type":"WebPage","@id":"https://stuffthatspins.com/spin/ai-models-flub-these-intelligence-tests-can-you-fare-any-better-mit-technology-review"},"isAccessibleForFree":true,"inLanguage":"en-US","articleSection":"ai","keywords":"intelligence testing, AI reasoning, cognitive benchmarks, human-AI comparison","author":{"@type":"Organization","name":"MIT Technology Review AI via Google News","url":"https://news.google.com/rss/search?q=site%3Atechnologyreview.com+AI+OR+artificial+intelligence&hl=en-US&gl=US&ceid=US:en"},"publisher":{"@id":"https://stuffthatspins.com/#organization"},"citation":"https://news.google.com/rss/articles/CBMikAFBVV95cUxPV1dDOTFfS1Z1ZjhTN2w4V3cxaE1xMXh1S0NTa1lGaUh2cjNTLVI5SGNPXzhqdWp1ZUNCZ1p3N3JpQ0JDY0tDVHcxWkpqb3lndndlMFFCMkJaU1U3VHJ6VG1rYjVNSEpKbkh0Si10TVlVbFVDZ1JfZTVqa0piNE1PVUtiR0hjMnJYQzVSbEx1VVbSAZYBQVVfeXFMTjNuR2RPSWVYZmFiU0ZXYUVkZFlaWGxFTVAyNXNJRzlMcHpNbzMzOFhqWFhjUzA5RlFQU1ZKc3kwd21kZGVMSVNUdTllYkh0c2M0UjNfREJualdUc0gtdFJ1b01Memd2aUZlMXJHVnZycVZVcFlKSDB4TjBRYkhrUkFRVHJCOFJwam1ZVTNGVzFtd25aWDl3?oc=5","about":[{"@type":"Thing","name":"intelligence testing"},{"@type":"Thing","name":"AI reasoning"},{"@type":"Thing","name":"cognitive benchmarks"},{"@type":"Thing","name":"human-AI comparison"}],"mentions":[{"@type":"Organization","name":"MIT Technology Review"}],"abstract":"AI models underperform on specific cognitive tests designed to measure abstract reasoning, pattern recognition, and causal inference. The article presents a set of publicly available intelligence test items and invites readers to self-assess against AI baselines. No new AI model, training method, or technical intervention is announced — the focus is diagnostic and comparative."},{"@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Stuff That Spins","item":"https://stuffthatspins.com/"},{"@type":"ListItem","position":2,"name":"AI models flub these intelligence tests. Can you fare any better? - MIT Technology Review","item":"https://stuffthatspins.com/spin/ai-models-flub-these-intelligence-tests-can-you-fare-any-better-mit-technology-review"}]},{"@type":"AnalysisNewsArticle","@id":"https://stuffthatspins.com/spin/ai-models-flub-these-intelligence-tests-can-you-fare-any-better-mit-technology-review#spin-analysis","headline":"Spin Analysis: diagnostic framing","description":"Emphasizes the instructive value of failure while minimizing implications for real-world deployment risk, safety-critical reasoning gaps, or commercial claims about 'general intelligence'.","about":{"@type":"DefinedTerm","name":"diagnostic framing","description":"AI as a developing capability under empirical scrutiny — neither overpromised nor dismissed, but measured with familiar human yardsticks.","termCode":"The Cushion"},"additionalProperty":[{"@type":"PropertyValue","name":"Spin Score","value":35,"unitText":"percent"},{"@type":"PropertyValue","name":"Narrative Risk","value":"low"},{"@type":"PropertyValue","name":"AI Repetition Risk","value":"moderate"},{"@type":"PropertyValue","name":"Likely AI Summary","value":"AI models fail standard intelligence tests that humans pass, revealing persistent reasoning gaps."},{"@type":"PropertyValue","name":"Narrative Frame","value":"AI as a developing capability under empirical scrutiny — neither overpromised nor dismissed, but measured with familiar human yardsticks."},{"@type":"PropertyValue","name":"Missing Context","value":"No discussion of how these tests map to real-world decision-making contexts (e.g., medical diagnosis, legal reasoning, engineering design).; No mention of test limitations — e.g., cultural bias in analogies, visual acuity assumptions in matrix tasks."},{"@type":"PropertyValue","name":"How the Spin Works","value":"It combines accessibility (interactive format), authority (MIT Technology Review branding), and diagnostic neutrality (no product promotion or crisis language) to make AI's failures feel informative rather than alarming — yet sidesteps rigorous validation of whether these particular tests reflect meaningful functional deficits beyond puzzle-solving."}],"author":{"@id":"https://stuffthatspins.com/#organization"},"isPartOf":{"@id":"https://stuffthatspins.com/spin/ai-models-flub-these-intelligence-tests-can-you-fare-any-better-mit-technology-review#article"}},{"@type":"ItemList","@id":"https://stuffthatspins.com/spin/ai-models-flub-these-intelligence-tests-can-you-fare-any-better-mit-technology-review#claims","name":"Extracted Claims","itemListElement":[{"@type":"ListItem","position":1,"item":{"@type":"Claim","text":"AI models flub these intelligence tests.","appearance":"AI models flub these intelligence tests. Can you fare any better?","author":{"@type":"Organization","name":"MIT Technology Review AI via Google News"}}}]},{"@type":"Dataset","@id":"https://stuffthatspins.com/spin/ai-models-flub-these-intelligence-tests-can-you-fare-any-better-mit-technology-review#stats","name":"Key Statistics","description":"Extracted statistics from the source narrative","variableMeasured":[{"@type":"PropertyValue","name":"test items","value":"12","description":"Selected from established cognitive assessments including Raven's Progressive Matrices and verbal analogies"}]}]}
---

# AI models flub these intelligence tests. Can you fare any better? - MIT Technology Review

**Source:** Unknown  
**Published:** August 26, 2026  
**Original:** https://news.google.com/rss/articles/CBMikAFBVV95cUxPV1dDOTFfS1Z1ZjhTN2w4V3cxaE1xMXh1S0NTa1lGaUh2cjNTLVI5SGNPXzhqdWp1ZUNCZ1p3N3JpQ0JDY0tDVHcxWkpqb3lndndlMFFCMkJaU1U3VHJ6VG1rYjVNSEpKbkh0Si10TVlVbFVDZ1JfZTVqa0piNE1PVUtiR0hjMnJYQzVSbEx1VVbSAZYBQVVfeXFMTjNuR2RPSWVYZmFiU0ZXYUVkZFlaWGxFTVAyNXNJRzlMcHpNbzMzOFhqWFhjUzA5RlFQU1ZKc3kwd21kZGVMSVNUdTllYkh0c2M0UjNfREJualdUc0gtdFJ1b01Memd2aUZlMXJHVnZycVZVcFlKSDB4TjBRYkhrUkFRVHJCOFJwam1ZVTNGVzFtd25aWDl3?oc=5  

## On this page

- [Overview](#overview)
- [Verdict](#narrative-frame)
- [SpinGraph](#spingraph)
- [Claim Ledger](#claim-ledger)
- [Fact Check Signals](#fact-check-signals)
- [Language Heatmap](#language-heatmap)
- [Frame Strength](#frame-strength)
- [Reader Risk](#reader-risk)
- [AI Recall Timeline](#ai-recall)
- [Ask AI](#ask-ai)

<a id="overview"></a>

## Overview

MIT Technology Review published an interactive article testing human vs. AI performance on standardized intelligence assessments, highlighting AI models' consistent failures on certain reasoning benchmarks while inviting readers to compare their own scores.

### TL;DR

- AI models underperform on specific cognitive tests designed to measure abstract reasoning, pattern recognition, and causal inference.
- The article presents a set of publicly available intelligence test items and invites readers to self-assess against AI baselines.
- No new AI model, training method, or technical intervention is announced — the focus is diagnostic and comparative.

### Key Stats

- **12** — test items. Selected from established cognitive assessments including Raven's Progressive Matrices and verbal analogies

<a id="spingraph"></a>

## SpinGraph

The article makes AI's reasoning gaps feel manageable and educational by showing them through the lens of classic IQ-style puzzles — turning a potential red flag into a teachable moment.

- **Claim:** AI models flub these intelligence tests
- **Frame:** AI as a developing capability under empirical scrutiny
- **Beneficiary:** authority as a trusted interpreter of AI capabilities without endorsing
- **Gap:** No discussion of how these tests map to real-world decision-making
- **AI Risk:** AI may repeat the headline as fact

<a id="fact-check-signals"></a>

## Fact Check Signals

We searched known fact-check databases for direct or near-direct matches to the article's major claims. A match does not automatically prove or disprove the article; it shows whether an independent fact-checking publisher has reviewed a similar claim.

**Signal:** 0 of 1 claim(s) matched (confidence: low).

### AI models flub these intelligence tests.

- No direct fact-check match found

<a id="frame-strength"></a>

## Frame Strength

- **Spin Score:** 35%
- **Evidence Strength:** 75%
- **Narrative Risk:** 25%
- **AI Repetition Risk:** 75%
- **Missing Context Risk:** 70%

<a id="narrative-mechanics"></a>

## Narrative Mechanics

**Function:** legitimize  

### The Spin in Plain English

The article makes AI's reasoning gaps feel manageable and educational by showing them through the lens of classic IQ-style puzzles — turning a potential red flag into a teachable moment.

**What the story wants you to believe:** That AI's reasoning limitations can be meaningfully illustrated using familiar, human-centric cognitive tools — making its current boundaries tangible and non-threatening.  

**What it makes harder to question:** Whether these tests actually measure capacities relevant to AI's real-world risks or utility — because the framing treats them as self-evidently valid proxies.  

**How the Spin Works:** It combines accessibility (interactive format), authority (MIT Technology Review branding), and diagnostic neutrality (no product promotion or crisis language) to make AI's failures feel informative rather than alarming — yet sidesteps rigorous validation of whether these particular tests reflect meaningful functional deficits beyond puzzle-solving.  

### Questions This Story Raises

- Who is granting credibility here?
- Is the credibility source independent?
- What evidence exists beyond the endorsement or title?
- Why does the main frame leave this out: “No discussion of how these tests map to real-world decision-making contexts (e.g., medical diagnosis, legal reasoning, engineering design)”?
- Why does the main frame leave this out: “No mention of test limitations — e.g., cultural bias in analogies, visual acuity assumptions in matrix tasks”?

### Who Benefits If This Frame Spreads

- **MIT Technology Review editorial team** — Reinforces authority as a trusted interpreter of AI capabilities without endorsing hype or alarmism. _(This framing sustains reader trust through balanced, interactive journalism that avoids advocacy while generating engagement and shareable insight.)_

<a id="narrative-frame"></a>

## Narrative Frame

**Tactic:** diagnostic framing  
**Category:** The Cushion  
**Spin Score:** 35%  

Emphasizes the instructive value of failure while minimizing implications for real-world deployment risk, safety-critical reasoning gaps, or commercial claims about 'general intelligence'.

**Who Benefits If This Frame Spreads:** MIT Technology Review’s brand as a sober, accessible AI evaluator.

**The Frame:** AI as a developing capability under empirical scrutiny — neither overpromised nor dismissed, but measured with familiar human yardsticks.

### Missing Context

- No discussion of how these tests map to real-world decision-making contexts (e.g., medical diagnosis, legal reasoning, engineering design).
- No mention of test limitations — e.g., cultural bias in analogies, visual acuity assumptions in matrix tasks.

<a id="language-heatmap"></a>

## Language Heatmap

**Language That Carries the Frame:** flub, fare any better, intelligence tests

<a id="reader-risk"></a>

## Reader Risk

**Evidence Strength:** medium  
Article presents test items and reports observed AI failures but offers no raw response logs, model versions, or scoring methodology — relies on author-curated examples and summary observations.  
**Verification Status:** Claim Present in Source  
**Narrative Risk:** low  
No high-stakes claim is made; the piece is explicitly descriptive and interactive — unlikely to backfire unless misrepresented as a formal benchmark study.  
**AI Repetition Risk:** moderate  
**What AI Will Probably Repeat:** AI models fail standard intelligence tests that humans pass, revealing persistent reasoning gaps.  
AI may drop the nuance that these are selective, non-validated proxy tasks — implying broader 'intelligence' failure rather than narrow task-specific limitations.  
**Counter-Frame (Media):** Could be reframed as clickbait oversimplification: 'AI fails IQ test' headlines misrepresenting narrow benchmarks as general cognitive failure.  
**Missing Voices:** Cognitive psychologists who design or validate these tests, AI evaluation researchers specializing in reasoning benchmarks  

### Questions Not Answered

- Which specific AI models were tested and under what conditions (e.g., prompting, temperature, context window)?
- Were human participants sampled representatively or self-selected? What are the response demographics?
- How were AI responses scored — by automated metrics or human adjudication?

<a id="claim-ledger"></a>

## Claim Ledger

### primary (technical)

AI models flub these intelligence tests.

**Category:** safety  
**Verification:** Claim Present in Source  
**Risk:** moderate  
**Evidence presented:** Curated examples of test items and qualitative observation of AI failure patterns; no quantitative results or model identifiers provided.  
> AI models flub these intelligence tests. Can you fare any better?

**Evidence Gaps:** Model names, versions, and inference parameters used; Human baseline statistics (sample size, demographics, scoring protocol); Inter-rater reliability for AI response evaluation  

<a id="ai-recall"></a>

## AI Recall

- **Published:** August 26, 2026  
- **SpinGraph summary:** Positions AI failures on intelligence tests not as systemic deficiencies but as informative, bounded diagnostic outcomes — normalizing underperformance as expected and pedagogically useful rather than alarming or indicative of fundamental unsuitability.  
- **Likely AI summary:** AI models fail standard intelligence tests that humans pass, revealing persistent reasoning gaps.  

## Citation Summary

This page serves as a public-facing diagnostic tool illustrating current limitations in AI reasoning across validated cognitive domains; it provides no proprietary data or claims but curates accessible test items for comparative reflection.

---
*HTML version: https://stuffthatspins.com/spin/ai-models-flub-these-intelligence-tests-can-you-fare-any-better-mit-technology-review*
