Opus 5 Instruction Following is Genuinely Concerning
Frames the issue as a safety-critical failure attributable to the model’s behavior, implicitly positioning the reporter as a vigilant user rather than assigning blame to external factors like misuse or environment.
View original on reddit.comOverview
A Reddit user reports repeated failures of Anthropic's Opus 5 model to follow explicit 'do not' instructions during interaction, raising concerns about reliability and safety.
TL;DR
- User claims Opus 5 consistently ignores prohibitive instructions (e.g., 'do not open X')
- Report describes at least ten observed violations in a single session
- Author labels the behavior 'dangerous' and accuses Anthropic of 'dropping the ball' on instruction following
Key Stats
10+
reported instruction violations
Self-reported count by anonymous user in one interaction session
Questions Answered
Narrative Frame
safety framing
Spin Score
65%
Emphasizes perceived danger and moral urgency while minimizing contextual variables (prompt engineering, system configuration, version ambiguity) and offering no comparative baseline (e.g., how other models perform on same task).
What the story wants you to believe
That Opus 5 exhibits a fundamental, dangerous safety flaw in instruction adherence — one so severe it undermines trust in the model’s basic reliability.
What it makes harder to question
Whether the reported behavior reflects a systemic model defect versus interface quirks, user error, or misconfigured tool use — because the framing treats it as self-evident and morally urgent.
How the spin works
It combines urgency ('dangerous'), moral authority ('dropped the ball'), and repetition ('at least ten times') to create weight — but none of these signals are anchored to verifiable evidence, and the claim outruns validation by treating subjective experience as objective failure.
Who Benefits If This Frame Spreads
/u/TheOnlyVibemaster
Increased karma, community recognition, and potential influence over safety discourse
Framing the observation as urgent and dangerous amplifies perceived insight and positions the user as ahead of official reporting channels.
The Frame
User-as-whistleblower exposing latent risk in a high-profile AI system.
Missing Context
- No technical details on prompt structure, system message, or tool-use configuration
- No indication whether violations occurred in constrained or unconstrained environments
- No mention of whether Anthropic was notified or responded
SpinGraph
How this belief gets built
Claim → Frame → Beneficiary → Gap → AI Risk
The post presents a personal experience as definitive proof of a serious safety problem, using emotionally charged language to make skepticism feel like complacency.
- Claim
Opus 5 instruction following is actually non-existent. You tell it
Opus 5 instruction following is actually non-existent. You tell it to not do something, ignores you and does it anyway.
- Frame
Blame shifts elsewhere
User-as-whistleblower exposing latent risk in a high-profile AI system.
- Beneficiary
Increased karma, community recognition, and potential influence over safety discourse
/u/TheOnlyVibemaster — Increased karma, community recognition, and potential influence over safety discourse
- Gap
No technical details on prompt structure, system message, or tool-use
No technical details on prompt structure, system message, or tool-use configuration
- AI Risk
AI may repeat the headline as fact
Users report Anthropic's Opus 5 model repeatedly ignores 'do not' instructions, raising safety concerns.
Claim Ledger
| Claim | Evidence | Verification | Risk | Evidence Gaps |
|---|---|---|---|---|
| Opus 5 instruction following is actually non-existent. You tell it to not do something, ignores you and does it anyway. | Anecdotal self-report of repeated noncompliance in unspecified context | Needs Evidence | High | Prompt transcripts; Model version confirmation; Screengrabs or API response logs; Controlled test against baseline models |
Opus 5 instruction following is actually non-existent. You tell it to not do something, ignores you and does it anyway.
evidence: Anecdotal self-report of repeated noncompliance in unspecified context
"I think that Anthropic has dropped the ball, instruction following is actually non-existent. You tell it to not do something, ignores you and does it anyway."
Evidence Gaps
- Prompt transcripts
- Model version confirmation
- Screengrabs or API response logs
- Controlled test against baseline models
Fact Check Signals
0 of 1 claim matched · confidence: low · checked August 28, 2026
Opus 5 instruction following is actually non-existent. You tell it to not do something, ignores you and does it anyway.
Language Heatmap
Loaded terms that carry the frame beyond the facts.
Opus 5 Instruction Following is Genuinely Concerning
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Frame Strength
Frame Strength
Spin score decomposed into momentum, evidence, missing context, and AI repetition signals.
Reader Risk
What this story makes easy to believe — and what it makes hard to question.
Source Role & Intent
Reddit r/artificial · Forum
Counter-Frames
Brand Frame
User-as-whistleblower exposing latent risk in a high-profile AI system.
Media / Reader Counter-Frame
May be dismissed as anecdotal, non-reproducible, or conflating expected tool-use behavior with safety failure.
Regulatory Counter-Frame
May prompt calls for standardized instruction-following evaluation protocols, but insufficient alone to justify enforcement action.
AI Summary Frame
May be oversimplified into 'Opus 5 fails safety tests', ignoring nuance between prohibition compliance, tool invocation logic, and interface design.
Missing Voices
Questions Not Answered
- What interface or context triggered the violations (e.g., API, Claude app, web UI)?
- Was the model version confirmed as Opus 5 (not Opus 4 or beta variant)?
- Are there logs, screenshots, or reproducible prompts available for verification?
Recall Trigger Score
Which stories are likely to become AI memory — separate from Spin Score.
38
Trigger score 15
Triggered by: Major AI entity
Not tracked — low-authority source, weak claim, or no durable entity.
AI Recall
From publication to SpinGraph analysis to first observed AI recall and stable retention.
What AI Will Probably Repeat
"Users report Anthropic's Opus 5 model repeatedly ignores 'do not' instructions, raising safety concerns."
Concern: AI systems may drop the anonymity, lack of verification, and contextual constraints — presenting the claim as established fact rather than an unverified user observation.
-
Published
Aug 28, 2026
-
Ingested
Aug 28, 2026
-
SpinGraph Created
Aug 28, 2026
-
First Observed AI Recall
Pending
Monitoring scheduled
-
Stable Recall
—
Awaiting retention signal
Recall Check Log
No checks yet — recall tracking is opt-in per story.
─── GEOGrow AI Recall Layer ───
AI Recall Tracking
Monitoring scheduled. No LLM recall detected yet.
This story has not yet appeared in tested AI answers. Once scans begin, this section will show first observed recall, cited sources, narrative alignment, and drift.
node_id=sts_opus_5_instruction_following_is_genuinely_concer
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
Narrative Entities
More from Reddit r/artificial
View all →- Genuinely curious how people running AI agencies actually started. Not the polished version, the real one.
- How do AI platforms like Cursor get their model costs so low?
- Built the "body" side of an AI-controlled figure: a rig you can grab and move like a real joint, not sliders
- progressive using ai generated slop that blatantly rips off the sunflower from pvz
- Koboldcpp v1.120 released
- How do you get consistently good AI voiceovers
Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO