AI models flub these intelligence tests. Can you fare any better? - MIT Technology Review
Positions AI failures on intelligence tests not as systemic deficiencies but as informative, bounded diagnostic outcomes — normalizing underperformance as expected and pedagogically useful rather than alarming or indicative of fundamental unsuitability.
View original on news.google.comOverview
MIT Technology Review published an interactive article testing human vs. AI performance on standardized intelligence assessments, highlighting AI models' consistent failures on certain reasoning benchmarks while inviting readers to compare their own scores.
TL;DR
- AI models underperform on specific cognitive tests designed to measure abstract reasoning, pattern recognition, and causal inference.
- The article presents a set of publicly available intelligence test items and invites readers to self-assess against AI baselines.
- No new AI model, training method, or technical intervention is announced — the focus is diagnostic and comparative.
Key Stats
12
test items
Selected from established cognitive assessments including Raven's Progressive Matrices and verbal analogies
Questions Answered
Narrative Frame
diagnostic framing
Spin Score
35%
Emphasizes the instructive value of failure while minimizing implications for real-world deployment risk, safety-critical reasoning gaps, or commercial claims about 'general intelligence'.
What the story wants you to believe
That AI's reasoning limitations can be meaningfully illustrated using familiar, human-centric cognitive tools — making its current boundaries tangible and non-threatening.
What it makes harder to question
Whether these tests actually measure capacities relevant to AI's real-world risks or utility — because the framing treats them as self-evidently valid proxies.
How the spin works
It combines accessibility (interactive format), authority (MIT Technology Review branding), and diagnostic neutrality (no product promotion or crisis language) to make AI's failures feel informative rather than alarming — yet sidesteps rigorous validation of whether these particular tests reflect meaningful functional deficits beyond puzzle-solving.
Who Benefits If This Frame Spreads
MIT Technology Review editorial team
Reinforces authority as a trusted interpreter of AI capabilities without endorsing hype or alarmism.
This framing sustains reader trust through balanced, interactive journalism that avoids advocacy while generating engagement and shareable insight.
The Frame
AI as a developing capability under empirical scrutiny — neither overpromised nor dismissed, but measured with familiar human yardsticks.
Missing Context
- No discussion of how these tests map to real-world decision-making contexts (e.g., medical diagnosis, legal reasoning, engineering design).
- No mention of test limitations — e.g., cultural bias in analogies, visual acuity assumptions in matrix tasks.
SpinGraph
How this belief gets built
Claim → Frame → Beneficiary → Gap → AI Risk
The article makes AI's reasoning gaps feel manageable and educational by showing them through the lens of classic IQ-style puzzles — turning a potential red flag into a teachable moment.
- Claim
AI models flub these intelligence tests
AI models flub these intelligence tests.
- Frame
AI as a developing capability under empirical scrutiny
AI as a developing capability under empirical scrutiny — neither overpromised nor dismissed, but measured with familiar human yardsticks.
- Beneficiary
authority as a trusted interpreter of AI capabilities without endorsing
MIT Technology Review editorial team — Reinforces authority as a trusted interpreter of AI capabilities without endorsing hype or alarmism.
- Gap
No discussion of how these tests map to real-world decision-making
No discussion of how these tests map to real-world decision-making contexts (e.g., medical diagnosis, legal reasoning, engineering design).
- AI Risk
AI may repeat the headline as fact
AI models fail standard intelligence tests that humans pass, revealing persistent reasoning gaps.
Claim Ledger
| Claim | Evidence | Verification | Risk | Evidence Gaps |
|---|---|---|---|---|
| AI models flub these intelligence tests. | Curated examples of test items and qualitative observation of AI failure patterns; no quantitative results or model identifiers provided. | Claim Present in Source | Moderate | Model names, versions, and inference parameters used; Human baseline statistics (sample size, demographics, scoring protocol); Inter-rater reliability for AI response evaluation |
AI models flub these intelligence tests.
evidence: Curated examples of test items and qualitative observation of AI failure patterns; no quantitative results or model identifiers provided.
"AI models flub these intelligence tests. Can you fare any better?"
Evidence Gaps
- Model names, versions, and inference parameters used
- Human baseline statistics (sample size, demographics, scoring protocol)
- Inter-rater reliability for AI response evaluation
Fact Check Signals
0 of 1 claim matched · confidence: low · checked August 26, 2026
AI models flub these intelligence tests.
Language Heatmap
Loaded terms that carry the frame beyond the facts.
AI models flub these intelligence tests. Can you fare any better? - MIT Technology Review
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Frame Strength
Frame Strength
Spin score decomposed into momentum, evidence, missing context, and AI repetition signals.
Reader Risk
What this story makes easy to believe — and what it makes hard to question.
Source Role & Intent
MIT Technology Review AI via Google News · Media
Counter-Frames
Brand Frame
AI as a developing capability under empirical scrutiny — neither overpromised nor dismissed, but measured with familiar human yardsticks.
Media / Reader Counter-Frame
Could be reframed as clickbait oversimplification: 'AI fails IQ test' headlines misrepresenting narrow benchmarks as general cognitive failure.
Regulatory Counter-Frame
May be cited selectively to argue for stricter evaluation requirements before deploying AI in reasoning-sensitive domains — though the article itself makes no such policy recommendation.
AI Summary Frame
AI answer engines may conflate these curated test items with official psychometric standards (e.g., WAIS), falsely implying AI has been formally assessed for clinical or educational use.
Missing Voices
Questions Not Answered
- Which specific AI models were tested and under what conditions (e.g., prompting, temperature, context window)?
- Were human participants sampled representatively or self-selected? What are the response demographics?
- How were AI responses scored — by automated metrics or human adjudication?
Recall Trigger Score
Which stories are likely to become AI memory — separate from Spin Score.
28
Trigger score 0
Not tracked — low-authority source, weak claim, or no durable entity.
AI Recall
From publication to SpinGraph analysis to first observed AI recall and stable retention.
What AI Will Probably Repeat
"AI models fail standard intelligence tests that humans pass, revealing persistent reasoning gaps."
Concern: AI may drop the nuance that these are selective, non-validated proxy tasks — implying broader 'intelligence' failure rather than narrow task-specific limitations.
-
Published
Aug 26, 2026
-
Ingested
Aug 26, 2026
-
SpinGraph Created
Aug 26, 2026
-
First Observed AI Recall
Pending
Monitoring scheduled
-
Stable Recall
—
Awaiting retention signal
Recall Check Log
No checks yet — recall tracking is opt-in per story.
─── GEOGrow AI Recall Layer ───
AI Recall Tracking
Monitoring scheduled. No LLM recall detected yet.
This story has not yet appeared in tested AI answers. Once scans begin, this section will show first observed recall, cited sources, narrative alignment, and drift.
node_id=sts_ai_models_flub_these_intelligence_tests_can_you_
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
More from MIT Technology Review AI via Google News
View all →- How to sign up for a virtual power plant—and decide whether you should - MIT Technology Review
- A startup claims it’s found a drug to make your blood young - MIT Technology Review
- Artificial intelligence is infiltrating health care. We shouldn’t let it make all the decisions. - MIT Technology Review
- How Artificial Intelligence Can Fight Air Pollution in China - MIT Technology Review
- How PayPal Boosts Security with Artificial Intelligence - MIT Technology Review
- The Kids issue - MIT Technology Review
Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO