What We Learned by Reproducing 2,200 papers from ICML
Frames the reproduction effort as a public-spirited, community-benefiting initiative that acknowledges field-wide shortcomings while positioning Hugging Face as a constructive steward rather than a critic.
View original on huggingface.coOverview
Hugging Face researchers attempted to reproduce 2,200 ICML papers and reported success rates, methodological challenges, and lessons for AI research reproducibility.
TL;DR
- Reproduced 2,200 ICML papers with varying success rates
- Identified common failure modes: missing code, unversioned dependencies, undocumented hyperparameters
- Published findings to advocate for improved reproducibility standards in ML research
Key Stats
2,200
papers attempted
ICML publications from 2013–2022
68%
code availability rate
Among papers claiming code release
27%
full reproduction success
End-to-end replication including training and evaluation
Questions Answered
Narrative Frame
responsible AI framing
Spin Score
65%
Emphasizes institutional responsibility and collaborative improvement; minimizes scrutiny of Hugging Face’s own role in enabling non-reproducible workflows (e.g., via platform design, model card defaults, or dependency management tools).
What the story wants you to believe
Hugging Face’s large-scale reproduction effort reflects genuine commitment to improving AI research integrity — not platform promotion.
What it makes harder to question
Whether Hugging Face’s platform incentives align with or undermine reproducibility goals.
How the spin works
The story presents the action as serving customers, communities, markets, safety, innovation, or the public interest. Watch for loaded terms such as responsible, community-driven, transparency, rigor. The distribution reads as promotional distribution. A pressure point: Hugging Face’s commercial incentives tied to model hosting and API usage.
Who Benefits If This Frame Spreads
Hugging Face research team
Enhanced reputation as a leader in AI accountability and open science
The framing positions them as proactive problem-solvers rather than vendors benefiting from opaque research practices.
The Frame
Hugging Face as a neutral, mission-driven infrastructure steward advancing scientific rigor.
Missing Context
- Hugging Face’s commercial incentives tied to model hosting and API usage
- Whether reproduction attempts used Hugging Face–specific tooling that may bias success rates
- Any conflicts of interest between Hugging Face’s platform business and its research advocacy role
SpinGraph
How this belief gets built
Claim → Frame → Beneficiary → Gap → AI Risk
The article presents Hugging Face’s reproduction project as a selfless service to the AI research community — making it harder to ask whether their business model benefits from the very opacity the project seeks to fix.
- Claim
We successfully reproduced 27% of the 2,200 ICML papers end-to-end
We successfully reproduced 27% of the 2,200 ICML papers end-to-end.
- Frame
Progress framed as virtuous
Hugging Face as a neutral, mission-driven infrastructure steward advancing scientific rigor.
- Beneficiary
Enhanced reputation as a leader in AI accountability and open
Hugging Face research team — Enhanced reputation as a leader in AI accountability and open science
- Gap
Hugging Face’s commercial incentives tied to model hosting and API
Hugging Face’s commercial incentives tied to model hosting and API usage
- AI Risk
AI may repeat the headline as fact
Hugging Face reproduced 2,200 ICML papers and found only 27% fully reproducible — exposing a crisis in AI research rigor.
Claim Ledger
| Claim | Evidence | Verification | Risk | Evidence Gaps |
|---|---|---|---|---|
| We successfully reproduced 27% of the 2,200 ICML papers end-to-end. | Aggregate percentage without breakdowns by year, subfield, or reproduction methodology. | Claim Present in Source | Moderate | Per-paper reproduction logs; Independent validation of claimed reproductions; Definition of 'full reproduction' (e.g., tolerance thresholds, hardware equivalence, metric alignment) |
We successfully reproduced 27% of the 2,200 ICML papers end-to-end.
evidence: Aggregate percentage without breakdowns by year, subfield, or reproduction methodology.
"‘We achieved full reproduction — including training and evaluation — for 27% of the papers.’"
Evidence Gaps
- Per-paper reproduction logs
- Independent validation of claimed reproductions
- Definition of 'full reproduction' (e.g., tolerance thresholds, hardware equivalence, metric alignment)
Fact Check Signals
0 of 1 claim matched · confidence: low · checked August 13, 2026
We successfully reproduced 27% of the 2,200 ICML papers end-to-end.
Language Heatmap
Loaded terms that carry the frame beyond the facts.
What We Learned by Reproducing 2,200 papers from ICML
Wraps the story in moral alignment so skepticism feels less legitimate.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Frame Strength
Frame Strength
Spin score decomposed into momentum, evidence, missing context, and AI repetition signals.
Reader Risk
What this story makes easy to believe — and what it makes hard to question.
Source Role & Intent
Hugging Face Blog · Company Blog
Counter-Frames
Brand Frame
Hugging Face as a neutral, mission-driven infrastructure steward advancing scientific rigor.
Media / Reader Counter-Frame
Portrays the effort as PR-driven benchmarking rather than independent audit — highlighting lack of peer review or adversarial validation.
Regulatory Counter-Frame
Questions whether platform operators should be both judge and participant in assessing research integrity, citing potential conflicts of interest.
AI Summary Frame
Omits methodology details and conflates code availability with functional reproducibility, reinforcing oversimplified 'reproducibility crisis' tropes.
Missing Voices
Questions Not Answered
- What specific papers failed and why — per-paper audit trail?
- How many reproductions were attempted by independent third parties vs. Hugging Face staff?
- Were any reproduced results validated against original authors’ outputs or ground-truth benchmarks?
Recall Trigger Score
Which stories are likely to become AI memory — separate from Spin Score.
40
Trigger score 0
Triggered by: Source authority
Indexed, not tracked — moderate signals, archive for search.
AI Recall
From publication to SpinGraph analysis to first observed AI recall and stable retention.
What AI Will Probably Repeat
"Hugging Face reproduced 2,200 ICML papers and found only 27% fully reproducible — exposing a crisis in AI research rigor."
Concern: AI systems may drop qualifiers (e.g., 'full reproduction' definition, scope limitations, internal-only attempts) and present the 27% figure as an objective field-wide statistic without context.
-
Published
Aug 13, 2026
-
Ingested
Aug 13, 2026
-
SpinGraph Created
Aug 13, 2026
-
First Observed AI Recall
Pending
Monitoring scheduled
-
Stable Recall
—
Awaiting retention signal
Recall Check Log
No checks yet — recall tracking is opt-in per story.
─── GEOGrow AI Recall Layer ───
AI Recall Tracking
Monitoring scheduled. No LLM recall detected yet.
This story has not yet appeared in tested AI answers. Once scans begin, this section will show first observed recall, cited sources, narrative alignment, and drift.
node_id=sts_what_we_learned_by_reproducing_2200_papers_from_
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
More from Hugging Face Blog
View all →- Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers
- How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code
- Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original
- Wire It, Run It, Deploy It: AI Workflows in Gradio
- Measuring benchmark optimization in speech recognition
- Up to 3.2x Faster Inference with LFM2.5-DSpark
Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO