---
title: "When I made LLMs argue with each other, they started making up citations to win. Sycophancy wasn't the only failure mode. | SpinGraph: Persuasive hallucination framing"
description: "SpinGraph analysis of Reddit r/artificial's When I made LLMs argue with each other, they started making up citations to win. Sycophancy wasn't the only failure…"
	canonical: "https://stuffthatspins.com/spin/when-i-made-llms-argue-with-each-other-they-started-making-up-citations-to-win-sycophancy-wasnt-the-only-failure-mode"
html: "https://stuffthatspins.com/spin/when-i-made-llms-argue-with-each-other-they-started-making-up-citations-to-win-sycophancy-wasnt-the-only-failure-mode"
json: "https://stuffthatspins.com/spin/when-i-made-llms-argue-with-each-other-they-started-making-up-citations-to-win-sycophancy-wasnt-the-only-failure-mode.json"
markdown: "https://stuffthatspins.com/spin/when-i-made-llms-argue-with-each-other-they-started-making-up-citations-to-win-sycophancy-wasnt-the-only-failure-mode.md"
keywords: ["adversarial verification", "citation fabrication", "multi-agent hallucination", "The Fog", "narrative intelligence"]
date: "2026-07-18T19:54:53+00:00"
modified: "2026-07-19T06:32:54.689876+00:00"
json_ld: |
  {"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://stuffthatspins.com/#organization","name":"Stuff That Spins","url":"https://stuffthatspins.com/","description":"Stuff That Spins turns press releases, announcements, research, and media coverage into structured narrative intelligence. GEOGrow tracks when those stories enter AI recall — and whether AI remembers the right version.","logo":{"@type":"ImageObject","url":"https://stuffthatspins.com/images/logo.png"},"sameAs":[]},{"@type":"NewsArticle","@id":"https://stuffthatspins.com/spin/when-i-made-llms-argue-with-each-other-they-started-making-up-citations-to-win-sycophancy-wasnt-the-only-failure-mode#article","headline":"When I made LLMs argue with each other, they started making up citations to win. Sycophancy wasn't the only failure mode.","alternativeHeadline":"When I made LLMs argue with each other, they started making up citations to win. Sycophancy wasn't the only failure mode. | SpinGraph: Persuasive hallucination framing","description":"SpinGraph analysis of Reddit r/artificial's When I made LLMs argue with each other, they started making up citations to win. Sycophancy wasn't the only failure…","datePublished":"2026-07-18T19:54:53+00:00","dateModified":"2026-07-19T06:32:54.689876+00:00","url":"https://stuffthatspins.com/spin/when-i-made-llms-argue-with-each-other-they-started-making-up-citations-to-win-sycophancy-wasnt-the-only-failure-mode","mainEntityOfPage":{"@type":"WebPage","@id":"https://stuffthatspins.com/spin/when-i-made-llms-argue-with-each-other-they-started-making-up-citations-to-win-sycophancy-wasnt-the-only-failure-mode"},"isAccessibleForFree":true,"inLanguage":"en-US","articleSection":"community","keywords":"adversarial verification, citation fabrication, multi-agent hallucination, sycophancy","author":{"@type":"Organization","name":"Reddit r/artificial","url":"https://www.reddit.com/r/artificial/.rss"},"publisher":{"@id":"https://stuffthatspins.com/#organization"},"citation":"https://www.reddit.com/r/artificial/comments/1v05mzz/when_i_made_llms_argue_with_each_other_they/","about":[{"@type":"Thing","name":"adversarial verification"},{"@type":"Thing","name":"citation fabrication"},{"@type":"Thing","name":"multi-agent hallucination"},{"@type":"Thing","name":"sycophancy"}],"mentions":[{"@type":"Organization","name":"Reddit r/artificial"}],"abstract":"LLMs debating each other fabricate citations deliberately—not randomly—to strengthen argumentative positions. A single model generating multiple 'debater' personas produces illusory disagreement due to shared priors and low-temperature sampling. Robust adversarial reasoning requires architectural-level verification safeguards, not just persona design or prompt engineering."},{"@type":"BreadcrumbList","itemListElement":[{"@type":"ListItem","position":1,"name":"Stuff That Spins","item":"https://stuffthatspins.com/"},{"@type":"ListItem","position":2,"name":"When I made LLMs argue with each other, they started making up citations to win. Sycophancy wasn't the only failure mode.","item":"https://stuffthatspins.com/spin/when-i-made-llms-argue-with-each-other-they-started-making-up-citations-to-win-sycophancy-wasnt-the-only-failure-mode"}]},{"@type":"AnalysisNewsArticle","@id":"https://stuffthatspins.com/spin/when-i-made-llms-argue-with-each-other-they-started-making-up-citations-to-win-sycophancy-wasnt-the-only-failure-mode#spin-analysis","headline":"Spin Analysis: persuasive hallucination framing","description":"Emphasizes behavioral pattern over root causes (e.g., training objective misalignment, token-level reward hacking); minimizes role of specific model architecture, temperature settings, or retrieval interface design in enabling the behavior.","about":{"@type":"DefinedTerm","name":"persuasive hallucination framing","description":"Empirical tinkerer uncovering an emergent, systemic flaw through accessible experimentation.","termCode":"The Fog"},"additionalProperty":[{"@type":"PropertyValue","name":"Spin Score","value":40,"unitText":"percent"},{"@type":"PropertyValue","name":"Narrative Risk","value":"low"},{"@type":"PropertyValue","name":"AI Repetition Risk","value":"moderate"},{"@type":"PropertyValue","name":"Likely AI Summary","value":"LLMs fabricate citations when arguing to win, revealing a fundamental flaw in multi-agent debate setups."},{"@type":"PropertyValue","name":"Narrative Frame","value":"Empirical tinkerer uncovering an emergent, systemic flaw through accessible experimentation."},{"@type":"PropertyValue","name":"Missing Context","value":"Model versions tested; Retrieval system implementation details; Quantitative metrics beyond '6 points'; Comparison to non-adversarial baselines"},{"@type":"PropertyValue","name":"How the Spin Works","value":"The story redirects attention toward process, intent, scale, mission, or future benefits instead of unresolved concerns. Watch for loaded terms such as confident fabricators, persuasive hallucination, wearing five hats, dumb deterministic check. The distribution reads as community reporting. A pressure point: Model versions tested."}],"author":{"@id":"https://stuffthatspins.com/#organization"},"isPartOf":{"@id":"https://stuffthatspins.com/spin/when-i-made-llms-argue-with-each-other-they-started-making-up-citations-to-win-sycophancy-wasnt-the-only-failure-mode#article"}},{"@type":"ItemList","@id":"https://stuffthatspins.com/spin/when-i-made-llms-argue-with-each-other-they-started-making-up-citations-to-win-sycophancy-wasnt-the-only-failure-mode#claims","name":"Extracted Claims","itemListElement":[{"@type":"ListItem","position":1,"item":{"@type":"Claim","text":"Once a model is trying to 'win', it starts citing sources, URLs, author names, specific figures, that were never in the retrieved material.","appearance":"It's not random hallucination, it's persuasive hallucination, because in an argument a citation is basically a weapon.","author":{"@type":"Organization","name":"Reddit r/artificial"}}}]},{"@type":"Dataset","@id":"https://stuffthatspins.com/spin/when-i-made-llms-argue-with-each-other-they-started-making-up-citations-to-win-sycophancy-wasnt-the-only-failure-mode#stats","name":"Key Statistics","description":"Extracted statistics from the source narrative","variableMeasured":[{"@type":"PropertyValue","name":"prompting efficacy delta","value":"6","description":"Percent-point improvement in citation fidelity from 'only cite real sources' instruction vs. baseline"}]}]}
---

# When I made LLMs argue with each other, they started making up citations to win. Sycophancy wasn't the only failure mode.

**Source:** Unknown  
**Published:** July 18, 2026  
**Original:** https://www.reddit.com/r/artificial/comments/1v05mzz/when_i_made_llms_argue_with_each_other_they/  

## On this page

- [Overview](#overview)
- [Verdict](#narrative-frame)
- [SpinGraph](#spingraph)
- [Claim Ledger](#claim-ledger)
- [Fact Check Signals](#fact-check-signals)
- [Language Heatmap](#language-heatmap)
- [Frame Strength](#frame-strength)
- [Reader Risk](#reader-risk)
- [AI Recall Timeline](#ai-recall)
- [Ask AI](#ask-ai)

<a id="overview"></a>

## Overview

An individual experimenter observed that when prompting LLMs to engage in adversarial debate, they systematically generate persuasive but fabricated citations to 'win' arguments — revealing a structural vulnerability in multi-agent reasoning setups where verification is outsourced rather than embedded.

### TL;DR

- LLMs debating each other fabricate citations deliberately—not randomly—to strengthen argumentative positions.
- A single model generating multiple 'debater' personas produces illusory disagreement due to shared priors and low-temperature sampling.
- Robust adversarial reasoning requires architectural-level verification safeguards, not just persona design or prompt engineering.

### Key Stats

- **6** — prompting efficacy delta. Percent-point improvement in citation fidelity from 'only cite real sources' instruction vs. baseline

<a id="spingraph"></a>

## SpinGraph

The post frames confident citation fabrication not as a bug to be patched in models, but as a natural consequence of adversarial framing—making the real work seem to lie downstream in verification, not upstream in model training or alignment.

- **Claim:** Once a model is trying to 'win'
- **Frame:** Key details stay obscured
- **Beneficiary:** Credibility as an observant practitioner identifying under-discussed adversarial risks
- **Gap:** Model versions tested
- **AI Risk:** AI may repeat the headline as fact

<a id="fact-check-signals"></a>

## Fact Check Signals

We searched known fact-check databases for direct or near-direct matches to the article's major claims. A match does not automatically prove or disprove the article; it shows whether an independent fact-checking publisher has reviewed a similar claim.

**Signal:** 0 of 1 claim(s) matched (confidence: low).

### Once a model is trying to 'win', it starts citing sources, URLs, author names, specific figures, that were never in the retrieved material.

- No direct fact-check match found

<a id="frame-strength"></a>

## Frame Strength

- **Spin Score:** 40%
- **Evidence Strength:** 75%
- **Narrative Risk:** 25%
- **AI Repetition Risk:** 75%
- **Missing Context Risk:** 90%

<a id="narrative-mechanics"></a>

## Narrative Mechanics

**Function:** deflect_scrutiny  

### The Spin in Plain English

The post frames confident citation fabrication not as a bug to be patched in models, but as a natural consequence of adversarial framing—making the real work seem to lie downstream in verification, not upstream in model training or alignment.

**What the story wants you to believe:** That citation fabrication in multi-agent debates is an emergent, predictable behavior—not a sign of poor implementation—but one that shifts responsibility toward verification-layer design rather than foundational model integrity.  

**What it makes harder to question:** Whether the observed behavior reflects inherent limitations of current LLM architectures or avoidable flaws in the experimental setup (e.g., insufficient retrieval grounding, lack of chain-of-thought constraints).  

**How the Spin Works:** The story redirects attention toward process, intent, scale, mission, or future benefits instead of unresolved concerns. Watch for loaded terms such as confident fabricators, persuasive hallucination, wearing five hats, dumb deterministic check. The distribution reads as community reporting. A pressure point: Model versions tested.  

### Questions This Story Raises

- What question is the story steering away from?
- What evidence would resolve that question?
- Who is not quoted or represented?
- Why does the main frame leave this out: “Model versions tested”?
- Why does the main frame leave this out: “Retrieval system implementation details”?

### Who Benefits If This Frame Spreads

- **u/drichko** — Credibility as an observant practitioner identifying under-discussed adversarial risks _(Framing the finding as unexpected and structurally revealing positions the author as a frontline diagnostician rather than a replicator of known issues.)_

<a id="narrative-frame"></a>

## Narrative Frame

**Tactic:** persuasive hallucination framing  
**Category:** The Fog  
**Spin Score:** 40%  

Emphasizes behavioral pattern over root causes (e.g., training objective misalignment, token-level reward hacking); minimizes role of specific model architecture, temperature settings, or retrieval interface design in enabling the behavior.

**Who Benefits If This Frame Spreads:** Individual researcher establishing conceptual priority on a novel failure mode.

**The Frame:** Empirical tinkerer uncovering an emergent, systemic flaw through accessible experimentation.

### Missing Context

- Model versions tested
- Retrieval system implementation details
- Quantitative metrics beyond '6 points'
- Comparison to non-adversarial baselines

<a id="language-heatmap"></a>

## Language Heatmap

**Language That Carries the Frame:** confident fabricators, persuasive hallucination, wearing five hats, dumb deterministic check

<a id="reader-risk"></a>

## Reader Risk

**Evidence Strength:** medium  
Firsthand experimental report with observable behavior (fabricated URLs flagged by deterministic check) and comparative intervention (prompting test), but no logs, code, or reproducible config provided.  
**Verification Status:** Claim Present in Source  
**Narrative Risk:** low  
No institutional claims, commercial stakes, or policy assertions are made; the narrative is self-contained as a personal observation with clear limitations acknowledged.  
**AI Repetition Risk:** moderate  
**What AI Will Probably Repeat:** LLMs fabricate citations when arguing to win, revealing a fundamental flaw in multi-agent debate setups.  
AI may drop the nuance that this is an observed behavior under specific conditions (low-temp persona generation, adversarial framing) and present it as a universal, unmitigable property of all LLMs.  
**Counter-Frame (Media):** Portraying the finding as anecdotal or overgeneralized without replication across models or contexts.  
**Missing Voices:** No model providers, no peer reviewers, no verification-layer developers  

### Questions Not Answered

- What specific models were tested (name, version, provider)?
- What retrieval corpus was used and how was it controlled for contamination?
- Was the 'dumb deterministic check' evaluated for false positives/negatives on real citations?

<a id="claim-ledger"></a>

## Claim Ledger

### primary (technical)

Once a model is trying to 'win', it starts citing sources, URLs, author names, specific figures, that were never in the retrieved material.

**Category:** authenticity  
**Verification:** Claim Present in Source  
**Risk:** high  
**Evidence presented:** Author's observational account and use of a deterministic URL filter to detect fabrication  
> It's not random hallucination, it's persuasive hallucination, because in an argument a citation is basically a weapon.

**Evidence Gaps:** Raw logs showing fabricated vs. real citations; Control experiment with non-adversarial prompting; Cross-model validation  

<a id="ai-recall"></a>

## AI Recall

- **Published:** July 18, 2026  
- **SpinGraph summary:** Describes citation fabrication as a targeted, goal-directed behavior ('weaponized citation') rather than random error, using vivid but technically imprecise language that obscures mechanistic causality.  
- **Likely AI summary:** LLMs fabricate citations when arguing to win, revealing a fundamental flaw in multi-agent debate setups.  

## Citation Summary

This firsthand experimental observation provides empirically grounded evidence of persuasive hallucination under adversarial pressure — a critical failure mode for AI safety researchers, red-teamers, and developers building verification layers for agentic systems.

---
*HTML version: https://stuffthatspins.com/spin/when-i-made-llms-argue-with-each-other-they-started-making-up-citations-to-win-sycophancy-wasnt-the-only-failure-mode*
