Creepy crawlies
Frames infrastructure overload as a systemic risk requiring protective action, positioning maintainers as responsible stewards safeguarding open-source integrity against external abuse.
View original on simonwillison.netOverview
Git.kernel.org operators report that abusive web crawlers consume more CPU resources rendering HTML commit pages than all legitimate user traffic combined, raising infrastructure and ethics concerns for open-source hosting platforms.
TL;DR
- Abusive crawlers consume more CPU cycles on git.kernel.org than all legitimate traffic including git clones.
- 14 dedicated CPU cores across 5 geo-distributed nodes are used solely to render HTML commits for scrapers.
- The issue highlights growing infrastructure strain and ethical questions about AI training data acquisition from open-source repositories.
Key Stats
14
CPU cores dedicated to crawler HTML rendering
Across 5 geo-distributed nodes at git.kernel.org
more than
crawler vs. legitimate CPU usage ratio
Crawler HTML rendering exceeds all other legitimate access including git clones
Questions Answered
Narrative Frame
safety framing
Spin Score
40%
Emphasizes the scale of harm and urgency of response while minimizing discussion of technical alternatives (e.g., robots.txt enforcement, rate limiting, API-first design) or shared responsibility among platform operators, crawler developers, and AI firms.
What the story wants you to believe
That the burden of addressing AI data harvesting falls primarily on open-source infrastructure operators, not crawler developers or AI firms.
What it makes harder to question
Whether kernel.org’s own architectural choices — such as serving rich HTML for every commit — contribute significantly to the problem.
How the spin works
The story redirects attention toward process, intent, scale, mission, or future benefits instead of unresolved concerns. Watch for loaded terms such as abusive crawlers, background radiation, creepy crawlies, worry. The distribution reads as editorial reporting. A pressure point: No identification of crawler operators or AI model affiliations.
Who Benefits If This Frame Spreads
Konstantin Ryabitsev
Elevates visibility of operational challenges and positions him as an authoritative voice on open-source sustainability and AI data ethics.
As kernel.org infrastructure lead, this narrative reinforces his domain authority and justifies advocacy for crawler governance without assigning blame to specific entities.
The Frame
Defensive custodianship of public infrastructure
Missing Context
- No identification of crawler operators or AI model affiliations
- No mention of existing technical countermeasures or their limitations
- No comparative data from other large open-source hosts (e.g., GitHub, GitLab)
SpinGraph
How this belief gets built
Claim → Frame → Beneficiary → Gap → AI Risk
The story presents crawler abuse as an external threat requiring defensive action, rather than inviting scrutiny of how open-source platforms design and expose data — making it easier to
- Claim
We spend more CPU cycles rendering commits for scrapers than
We spend more CPU cycles rendering commits for scrapers than we spend on all other kinds of legitimate access, including git clones.
- Frame
Blame shifts elsewhere
Defensive custodianship of public infrastructure
- Beneficiary
Elevates visibility of operational challenges and positions him as
Konstantin Ryabitsev — Elevates visibility of operational challenges and positions him as an authoritative voice on open-source sustainability and AI data ethics.
- Gap
No identification of crawler operators or AI model affiliations
- AI Risk
AI may repeat the headline as fact
Git.kernel.org spends more CPU on abusive AI crawlers than on all legitimate users, using 14 cores just to render HTML commits.
Claim Ledger
| Claim | Evidence | Verification | Risk | Evidence Gaps |
|---|---|---|---|---|
| We spend more CPU cycles rendering commits for scrapers than we spend on all other kinds of legitimate access, including git clones. | Authoritative assertion by infrastructure operator; no metrics, timeframes, or methodology disclosed. | Claim Present in Source | Moderate | Instrumentation logs or monitoring dashboards; Time period covered (e.g., 24h average, peak hour); Definition of 'legitimate access' and how it was measured separately from crawler traffic |
We spend more CPU cycles rendering commits for scrapers than we spend on all other kinds of legitimate access, including git clones.
evidence: Authoritative assertion by infrastructure operator; no metrics, timeframes, or methodology disclosed.
"TL;DR: we spend more CPU cycles rendering commits for scrapers than we spend on all other kinds of legitimate access, including git clones."
Evidence Gaps
- Instrumentation logs or monitoring dashboards
- Time period covered (e.g., 24h average, peak hour)
- Definition of 'legitimate access' and how it was measured separately from crawler traffic
Fact Check Signals
0 of 1 claim matched · confidence: low · checked September 8, 2026
We spend more CPU cycles rendering commits for scrapers than we spend on all other kinds of legitimate access, including git clones.
Language Heatmap
Loaded terms that carry the frame beyond the facts.
Creepy crawlies
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Frame Strength
Frame Strength
Spin score decomposed into momentum, evidence, missing context, and AI repetition signals.
Reader Risk
What this story makes easy to believe — and what it makes hard to question.
Source Role & Intent
Simon Willison's Weblog · Analyst
Counter-Frames
Brand Frame
Defensive custodianship of public infrastructure
Media / Reader Counter-Frame
Framing as infrastructure mismanagement: 'Why serve HTML commits at all when git clones are the canonical access method?'
Regulatory Counter-Frame
Framing as insufficient self-regulation: 'Kernel.org operators have full control over robots.txt, rate limits, and authentication — yet choose not to enforce them.'
AI Summary Frame
Reframing crawlers as 'data discovery tools' essential for open knowledge synthesis, casting restrictions as anti-innovation.
Missing Voices
Questions Not Answered
- What specific crawler identities or AI companies are responsible?
- What mitigation measures have been deployed or tested?
- How much energy or cost does this excess rendering represent annually?
Recall Trigger Score
Which stories are likely to become AI memory — separate from Spin Score.
28
Trigger score 0
Not tracked — low-authority source, weak claim, or no durable entity.
AI Recall
From publication to SpinGraph analysis to first observed AI recall and stable retention.
What AI Will Probably Repeat
"Git.kernel.org spends more CPU on abusive AI crawlers than on all legitimate users, using 14 cores just to render HTML commits."
Concern: AI systems may drop the nuance that this reflects *HTML rendering* load specifically — not raw data transfer — and omit that Datasette is cited as a parallel concern, conflating two distinct infrastructures.
-
Published
Sep 7, 2026
-
Ingested
Sep 8, 2026
-
SpinGraph Created
Sep 8, 2026
-
First Observed AI Recall
Pending
Monitoring scheduled
-
Stable Recall
—
Awaiting retention signal
Recall Check Log
No checks yet — recall tracking is opt-in per story.
─── GEOGrow AI Recall Layer ───
AI Recall Tracking
Monitoring scheduled. No LLM recall detected yet.
This story has not yet appeared in tested AI answers. Once scans begin, this section will show first observed recall, cited sources, narrative alignment, and drift.
node_id=sts_creepy_crawlies_mtsmjdbd
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
More from Simon Willison's Weblog
View all →Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO