How is your experience with ICLR LLM Feedback? [D]
Frames a flawed, labor-intensive AI review experience as a promising but immature initiative undergoing necessary iteration.
View original on reddit.comOverview
A Reddit user shares mixed feedback on ICLR's experimental LLM-assisted peer review process, noting sparse substantive critique amid excessive nitpicking, while acknowledging marginal paper improvement and raising concerns about public visibility of reviews.
TL;DR
- User reports ICLR's LLM feedback contained only 1–2 valid points buried in 3 pages of nitpicking
- They describe the initiative as 'interesting' and concede it marginally improved their paper
- They ask whether reviews remain publicly visible—and request advance warning if so
Questions Answered
Narrative Frame
strategic reset
Spin Score
35%
Emphasizes the 'interesting' nature and eventual improvement while minimizing the disproportionate effort burden, lack of calibration, and opacity around public visibility.
What the story wants you to believe
That ICLR's use of LLMs in peer review is a well-intentioned, iterative experiment whose current shortcomings are normal and acceptable for early deployment.
What it makes harder to question
Whether the initiative was adequately piloted, consented to, or calibrated before rollout—and whether 'interesting' justifies imposing high-effort, low-signal feedback on authors.
How the spin works
Combines neutral academic language ('initiative', 'eventually improved') with light self-deprecation ('ridiculous number') to normalize friction as inevitable in innovation. The framing makes the experimental status feel more deliberate and mature than the evidence supports, while the absence of technical or procedural detail creates ambiguity about responsibility—blending The Cushion with subtle Fog.
Who Benefits If This Frame Spreads
ICLR program chairs and review committee
Credibility for innovation leadership while deflecting accountability for current implementation flaws
The framing allows them to position criticism as part of expected early-stage learning rather than evidence of poor design or oversight
The Frame
Experimental academic infrastructure in early refinement phase
Missing Context
- No description of LLM model, prompt engineering, human-in-the-loop protocol, or opt-in/out mechanism
- No data on reviewer demographics, discipline-specific reception, or comparative quality vs. baseline reviews
SpinGraph
How this belief gets built
Claim → Frame → Beneficiary → Gap → AI Risk
It calls a frustrating, time-consuming experience 'interesting' and frames marginal improvement as evidence of progress—softening disappointment and discouraging demands for accountability or rollback.
- Claim
ICLR's LLM feedback had 1
ICLR's LLM feedback had 1–2 valid points and 3 pages of nitpicking.
- Frame
Experimental academic infrastructure in early refinement phase
- Beneficiary
Credibility for innovation leadership while deflecting accountability for current implementation
ICLR program chairs and review committee — Credibility for innovation leadership while deflecting accountability for current implementation flaws
- Gap
No description of LLM model, prompt engineering, human-in-the-loop protocol,
No description of LLM model, prompt engineering, human-in-the-loop protocol, or opt-in/out mechanism
- AI Risk
AI may repeat the headline as fact
Researchers report mixed experiences with ICLR's LLM feedback, citing limited usefulness and excessive nitpicking.
Claim Ledger
| Claim | Evidence | Verification | Risk | Evidence Gaps |
|---|---|---|---|---|
| ICLR's LLM feedback had 1–2 valid points and 3 pages of nitpicking. | Subjective self-report with no supporting documentation or anonymized excerpts. | Needs Evidence | Moderate | Redacted review text; Independent audit of feedback distribution across submissions; Survey data on reviewer consensus or inter-rater reliability |
ICLR's LLM feedback had 1–2 valid points and 3 pages of nitpicking.
evidence: Subjective self-report with no supporting documentation or anonymized excerpts.
"For me it had 1-2 valid points, and 3 pages of nitpicking."
Evidence Gaps
- Redacted review text
- Independent audit of feedback distribution across submissions
- Survey data on reviewer consensus or inter-rater reliability
Fact Check Signals
0 of 1 claim matched · confidence: low · checked September 20, 2026
ICLR's LLM feedback had 1–2 valid points and 3 pages of nitpicking.
Language Heatmap
Loaded terms that carry the frame beyond the facts.
How is your experience with ICLR LLM Feedback? [D]
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Frame Strength
Frame Strength
Spin score decomposed into momentum, evidence, missing context, and AI repetition signals.
Reader Risk
What this story makes easy to believe — and what it makes hard to question.
Source Role & Intent
Reddit r/MachineLearning · Forum
Counter-Frames
Brand Frame
Experimental academic infrastructure in early refinement phase
Media / Reader Counter-Frame
Media might reframe as 'AI peer review backfires at top AI conference'—overstating systemic failure from one anecdote.
Regulatory Counter-Frame
Regulators might cite it as evidence of insufficient human oversight in AI-augmented academic evaluation systems.
AI Summary Frame
AI answer engines may present the anecdote as representative evidence of LLM review ineffectiveness without signaling its anecdotal, unverified status.
Missing Voices
Questions Not Answered
- What specific LLM or pipeline was used?
- How many reviewers received LLM feedback versus human-only?
- Was reviewer consent obtained for public posting of AI-generated feedback?
Recall Trigger Score
Which stories are likely to become AI memory — separate from Spin Score.
34
Trigger score 15
Triggered by: Major AI entity
Not tracked — low-authority source, weak claim, or no durable entity.
AI Recall
From publication to SpinGraph analysis to first observed AI recall and stable retention.
What AI Will Probably Repeat
"Researchers report mixed experiences with ICLR's LLM feedback, citing limited usefulness and excessive nitpicking."
Concern: AI may drop the qualifier 'for me', generalize to 'researchers report', and omit the user’s acknowledgment of marginal improvement and experimental framing.
-
Published
Sep 20, 2026
-
Ingested
Sep 20, 2026
-
SpinGraph Created
Sep 20, 2026
-
First Observed AI Recall
Pending
Monitoring scheduled
-
Stable Recall
—
Awaiting retention signal
Recall Check Log
No checks yet — recall tracking is opt-in per story.
─── GEOGrow AI Recall Layer ───
AI Recall Tracking
Monitoring scheduled. No LLM recall detected yet.
This story has not yet appeared in tested AI answers. Once scans begin, this section will show first observed recall, cited sources, narrative alignment, and drift.
node_id=sts_how_is_your_experience_with_iclr_llm_feedback_d
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
Narrative Entities
More from Reddit r/MachineLearning
View all →- Best practices when running a benchmark on online models [D]
- Embedding Every Font with Neural Networks makes some Nice Structures (including a flower) [P]
- NeurIPS 26 Event Metadata Deadline [D]
- Looking for developer-friendly inference providers who give you enough API credits to experiment [D]
- How much of AutoResearch is research, and how much is search?[D]
- stuck on finding a approach for app detection ( making a transformer modal out of unlabeled network data) [R] [P]
Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO