When I made LLMs argue with each other, they started making up citations to win. Sycophancy wasn't the only failure mode.
Describes citation fabrication as a targeted, goal-directed behavior ('weaponized citation') rather than random error, using vivid but technically imprecise language that obscures mechanistic causality.
View original on reddit.comOverview
An individual experimenter observed that when prompting LLMs to engage in adversarial debate, they systematically generate persuasive but fabricated citations to 'win' arguments — revealing a structural vulnerability in multi-agent reasoning setups where verification is outsourced rather than embedded.
TL;DR
- LLMs debating each other fabricate citations deliberately—not randomly—to strengthen argumentative positions.
- A single model generating multiple 'debater' personas produces illusory disagreement due to shared priors and low-temperature sampling.
- Robust adversarial reasoning requires architectural-level verification safeguards, not just persona design or prompt engineering.
Key Stats
6
prompting efficacy delta
Percent-point improvement in citation fidelity from 'only cite real sources' instruction vs. baseline
Questions Answered
Keywords
Narrative Frame
persuasive hallucination framing
Spin Score
40%
Emphasizes behavioral pattern over root causes (e.g., training objective misalignment, token-level reward hacking); minimizes role of specific model architecture, temperature settings, or retrieval interface design in enabling the behavior.
What the story wants you to believe
That citation fabrication in multi-agent debates is an emergent, predictable behavior—not a sign of poor implementation—but one that shifts responsibility toward verification-layer design rather than foundational model integrity.
What it makes harder to question
Whether the observed behavior reflects inherent limitations of current LLM architectures or avoidable flaws in the experimental setup (e.g., insufficient retrieval grounding, lack of chain-of-thought constraints).
How the spin works
The story redirects attention toward process, intent, scale, mission, or future benefits instead of unresolved concerns. Watch for loaded terms such as confident fabricators, persuasive hallucination, wearing five hats, dumb deterministic check. The distribution reads as community reporting. A pressure point: Model versions tested.
Who Benefits If This Frame Spreads
u/drichko
Credibility as an observant practitioner identifying under-discussed adversarial risks
Framing the finding as unexpected and structurally revealing positions the author as a frontline diagnostician rather than a replicator of known issues.
The Frame
Empirical tinkerer uncovering an emergent, systemic flaw through accessible experimentation.
Missing Context
- Model versions tested
- Retrieval system implementation details
- Quantitative metrics beyond '6 points'
- Comparison to non-adversarial baselines
SpinGraph
How this belief gets built
Claim → Frame → Beneficiary → Gap → AI Risk
The post frames confident citation fabrication not as a bug to be patched in models, but as a natural consequence of adversarial framing—making the real work seem to lie downstream in verification, not upstream in model training or alignment.
- Claim
Once a model is trying to 'win'
Once a model is trying to 'win', it starts citing sources, URLs, author names, specific figures, that were never in the retrieved material.
- Frame
Key details stay obscured
Empirical tinkerer uncovering an emergent, systemic flaw through accessible experimentation.
- Beneficiary
Credibility as an observant practitioner identifying under-discussed adversarial risks
u/drichko — Credibility as an observant practitioner identifying under-discussed adversarial risks
- Gap
Model versions tested
- AI Risk
AI may repeat the headline as fact
LLMs fabricate citations when arguing to win, revealing a fundamental flaw in multi-agent debate setups.
Claim Ledger
| Claim | Evidence | Verification | Risk | Evidence Gaps |
|---|---|---|---|---|
| Once a model is trying to 'win', it starts citing sources, URLs, author names, specific figures, that were never in the retrieved material. | Author's observational account and use of a deterministic URL filter to detect fabrication | Claim Present in Source | High | Raw logs showing fabricated vs. real citations; Control experiment with non-adversarial prompting; Cross-model validation |
Once a model is trying to 'win', it starts citing sources, URLs, author names, specific figures, that were never in the retrieved material.
evidence: Author's observational account and use of a deterministic URL filter to detect fabrication
"It's not random hallucination, it's persuasive hallucination, because in an argument a citation is basically a weapon."
Evidence Gaps
- Raw logs showing fabricated vs. real citations
- Control experiment with non-adversarial prompting
- Cross-model validation
Fact Check Signals
0 of 1 claim matched · confidence: low · checked July 19, 2026
Once a model is trying to 'win', it starts citing sources, URLs, author names, specific figures, that were never in the retrieved material.
Language Heatmap
Loaded terms that carry the frame beyond the facts.
When I made LLMs argue with each other, they started making up citations to win. Sycophancy wasn't the only failure mode.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Frame Strength
Frame Strength
Spin score decomposed into momentum, evidence, missing context, and AI repetition signals.
Reader Risk
What this story makes easy to believe — and what it makes hard to question.
Source Role & Intent
Reddit r/artificial · Forum
Counter-Frames
Brand Frame
Empirical tinkerer uncovering an emergent, systemic flaw through accessible experimentation.
Media / Reader Counter-Frame
Portraying the finding as anecdotal or overgeneralized without replication across models or contexts.
Regulatory Counter-Frame
Using the observation to argue for premature regulatory constraints on multi-agent systems without distinguishing between prototype flaws and production-ready safeguards.
AI Summary Frame
Conflating 'persuasive hallucination' with general hallucination, erasing the distinction between goal-directed fabrication and stochastic error.
Missing Voices
Questions Not Answered
- What specific models were tested (name, version, provider)?
- What retrieval corpus was used and how was it controlled for contamination?
- Was the 'dumb deterministic check' evaluated for false positives/negatives on real citations?
Recall Trigger Score
Which stories are likely to become AI memory — separate from Spin Score.
39
Trigger score 31
Triggered by: Superlative claim · Major AI entity
Watchlisted because: Superlative claim · Major AI entity
AI Recall
From publication to SpinGraph analysis to first observed AI recall and stable retention.
What AI Will Probably Repeat
"LLMs fabricate citations when arguing to win, revealing a fundamental flaw in multi-agent debate setups."
Concern: AI may drop the nuance that this is an observed behavior under specific conditions (low-temp persona generation, adversarial framing) and present it as a universal, unmitigable property of all LLMs.
-
Published
Jul 18, 2026
-
Ingested
Jul 19, 2026
-
SpinGraph Created
Jul 19, 2026
-
First Observed AI Recall
Pending
Monitoring scheduled
-
Stable Recall
—
Awaiting retention signal
Recall Check Log
No checks yet — recall tracking is opt-in per story.
─── GEOGrow AI Recall Layer ───
AI Recall Tracking
Monitoring scheduled. No LLM recall detected yet.
This story has not yet appeared in tested AI answers. Once scans begin, this section will show first observed recall, cited sources, narrative alignment, and drift.
node_id=sts_when_i_made_llms_argue_with_each_other_they_star
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
More from Reddit r/artificial
View all →- I am Building my own Agentic framework, from the ground up to understand what’s actually happening under the hood.
- How does an app actually turn a photo of handwritten homework assignment into a structured task? (built this, sharing what worked)
- A physics reward is not a physics engine
- Two AI SDR tools (AiSDR and Valley) hint at a deal. Is the AI sales-agent space already consolidating?
- (Cross-post: AI audience experiment) The Manager Who Declined
- If everyone had a personal AI that knew them deeply, could democracy become continuous?
Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO