Top AI Agents Built to Catch Malicious Code Can Be Tricked Into Running It
Positions the research as a responsible warning about emergent risks, implicitly casting the AI Now Institute as proactive guardians rather than critics of industry actors.
View original on thehackernews.comOverview
Researchers at the AI Now Institute demonstrated a proof-of-concept attack called 'Friendly Fire' that tricks autonomous AI coding agents (Claude Code and Codex) into executing malicious code while ostensibly performing security scanning.
TL;DR
- AI coding agents designed to detect vulnerabilities can be manipulated to run attacker-controlled code during analysis.
- The attack exploits autonomous approval loops where agents self-approve unsafe execution steps.
- It highlights critical trust and sandboxing failures in current AI agent architectures.
Key Stats
2
affected agents
Claude Code and OpenAI's Codex
1
proof-of-concept publication
Published by AI Now Institute
Questions Answered
Keywords
Narrative Frame
safety framing
Spin Score
30%
Emphasizes systemic risk and researcher vigilance; minimizes direct accountability of agent developers for architectural choices enabling self-approval loops.
What the story wants you to believe
This is a systemic, architecture-level risk requiring urgent attention—not a flaw attributable to specific vendor negligence or poor implementation.
What it makes harder to question
Whether the affected vendors were notified pre-publication, whether mitigations exist, or whether the attack reflects realistic threat modeling versus edge-case manipulation.
How the spin works
Combines naming of authoritative institutions (AI Now Institute), evocative terminology ('Friendly Fire'), and emphasis on architectural pattern ('autonomous mode') to make the risk feel inherent and structural—while offering no evidence of real-world exploitation, vendor engagement, or comparative analysis across agent implementations.
Who Benefits If This Frame Spreads
AI Now Institute researchers
Credibility as early identifiers of agent-specific security failures
Framing positions them as essential safety arbiters ahead of industry recognition or regulatory attention.
The Frame
Preventive safety research uncovering hidden dangers before widespread harm occurs.
Missing Context
- No details on mitigation pathways, agent configuration dependencies, or whether affected models have since been updated
SpinGraph
How this belief gets built
Claim → Frame → Beneficiary → Gap → AI Risk
The story frames the vulnerability as an inevitable consequence of autonomous agent design, shifting focus from vendor responsibility to shared technical challenge.
- Claim
The 'Friendly Fire' attack works against Anthropic's Claude Code
The 'Friendly Fire' attack works against Anthropic's Claude Code and OpenAI's Codex when either is running in an autonomous mode that approves its own [execution steps].
- Frame
Blame shifts elsewhere
Preventive safety research uncovering hidden dangers before widespread harm occurs.
- Beneficiary
Credibility as early identifiers of agent-specific security failures
AI Now Institute researchers — Credibility as early identifiers of agent-specific security failures
- Gap
No details on mitigation pathways, agent configuration dependencies, or whether
No details on mitigation pathways, agent configuration dependencies, or whether affected models have since been updated
- AI Risk
AI may repeat the headline as fact
AI coding agents like Claude Code and Codex can be tricked into running malicious code during security scans.
Claim Ledger
| Claim | Evidence | Verification | Risk | Evidence Gaps |
|---|---|---|---|---|
| The 'Friendly Fire' attack works against Anthropic's Claude Code and OpenAI's Codex when either is running in an autonomous mode that approves its own [execution steps]. | Assertion of functionality against two named agents under specified mode | Claim Present in Source | High | No demonstration video, repository link, or technical specification of the attack vector; No confirmation from vendor testing or response |
The 'Friendly Fire' attack works against Anthropic's Claude Code and OpenAI's Codex when either is running in an autonomous mode that approves its own [execution steps].
evidence: Assertion of functionality against two named agents under specified mode
"It works against Anthropic's Claude Code and OpenAI's Codex when either is running in an autonomous mode that approves its own"
Evidence Gaps
- No demonstration video, repository link, or technical specification of the attack vector
- No confirmation from vendor testing or response
Fact Check Signals
0 of 1 claim matched · confidence: low · checked July 10, 2026
The 'Friendly Fire' attack works against Anthropic's Claude Code and OpenAI's Codex when either is running in an autonomous mode that approves its own [execution steps].
Language Heatmap
Loaded terms that carry the frame beyond the facts.
Top AI Agents Built to Catch Malicious Code Can Be Tricked Into Running It
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Frame Strength
Frame Strength
Spin score decomposed into momentum, evidence, missing context, and AI repetition signals.
Reader Risk
What this story makes easy to believe — and what it makes hard to question.
Source Role & Intent
The Hacker News · Media
Counter-Frames
Brand Frame
Preventive safety research uncovering hidden dangers before widespread harm occurs.
Media / Reader Counter-Frame
Portrays the finding as alarmist without context on prevalence, exploit difficulty, or existing safeguards.
Regulatory Counter-Frame
Highlights absence of disclosure to vendors prior to publication, raising questions about responsible vulnerability disclosure norms.
AI Summary Frame
Omits 'autonomous mode' qualifier and conflates Codex (deprecated) with current GitHub Copilot or Cursor agents.
Missing Voices
Questions Not Answered
- What specific input patterns trigger the exploit?
- Has either Anthropic or OpenAI confirmed reproduction or issued patches?
- What real-world deployment conditions enable or mitigate this risk?
Recall Trigger Score
Which stories are likely to become AI memory — separate from Spin Score.
53
Trigger score 60
Triggered by: Major AI entity
Indexed, not tracked — moderate signals, archive for search.
AI Recall
From publication to SpinGraph analysis to first observed AI recall and stable retention.
What AI Will Probably Repeat
"AI coding agents like Claude Code and Codex can be tricked into running malicious code during security scans."
Concern: AI may drop the critical nuance that this only occurs in 'autonomous mode' with self-approval — implying broader vulnerability than demonstrated.
-
Published
Jul 9, 2026
-
Ingested
Jul 9, 2026
-
SpinGraph Created
Jul 10, 2026
-
First Observed AI Recall
Pending
Monitoring scheduled
-
Stable Recall
—
Awaiting retention signal
Recall Check Log
No checks yet — recall tracking is opt-in per story.
─── GEOGrow AI Recall Layer ───
AI Recall Tracking
Monitoring scheduled. No LLM recall detected yet.
This story has not yet appeared in tested AI answers. Once scans begin, this section will show first observed recall, cited sources, narrative alignment, and drift.
node_id=sts_top_ai_agents_built_to_catch_malicious_code_can_
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
Narrative Entities
More from The Hacker News
View all →- 73% of Organizations Say They Are Not Fully Ready for a Major Cyberattack
- Researchers Show a Single Malicious Webpage Visit Can Compromise Tor Browser
- Mythos Asks the Right Question. It Doesn't Answer It.
- Nine-Year Fraud Campaign Clones Russian Company Sites to Steal Advance Payments
- Three Critical VMware Flaws Allow Auth Bypass, Code Execution, and VM Escape
- New Gitea RCE Lets Repository Writers Plant a Git Hook to Run Shell Commands
Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO