Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
Frames the taxonomy as both technically rigorous (via reproducibility metrics and cross-architecture applicability) and socially responsible (by enabling precise, repair-oriented accountability instead of systemic blame).
View original on arxiv.orgOverview
Researchers propose an interaction-centric taxonomy to localize AI agent failures to specific components (e.g., model, harness, environment) rather than treating failures as monolithic system-level events, enabling targeted interventions.
TL;DR
- Introduces a new failure taxonomy that maps 41 failure modes to interactions between agent components (model-harness, harness-tool, etc.)
- Assigns each failure to an 'edge' and 'fault side' to guide precise repair—e.g., model-side vs. harness-side fixes
- Validated across four frontier models using independent reasoning agents as judges, achieving κ=0.76 agreement with human labels
Key Stats
41
failure modes
Taxonomy organizes failures by interaction edge and fault side
0.76
Cohen's kappa
Strongest judge agreement with human labels on failure categorization
Questions Answered
Keywords
Narrative Frame
actionable framing
Spin Score
65%
Emphasizes scalability, shared structure, and actionability while minimizing discussion of taxonomy limitations, domain-specific brittleness, or real-world deployment validation beyond benchmark trajectories.
What the story wants you to believe
That this taxonomy provides a robust, empirically grounded, and widely applicable foundation for diagnosing and repairing AI agent failures at the component level.
What it makes harder to question
Whether failure localization truly enables more effective repairs—or whether edge-based attribution remains subjective, context-dependent, and unvalidated outside controlled benchmark settings.
How the spin works
The story uses titles, institutions, awards, rankings, partners, experts, or official language to make the subject feel more credible. Watch for loaded terms such as actionable, shared structure, frontier models, grounded. The distribution reads as academic distribution. A pressure point: No discussion of false positive/negative rates in failure localization.
Who Benefits If This Frame Spreads
Research authors
Increased citation, integration into evaluation pipelines, and influence over failure diagnostics standards
Positioning the taxonomy as both empirically grounded and universally applicable incentivizes adoption by labs and tooling developers.
The Frame
Methodological advancement enabling principled, component-responsible AI development
Missing Context
- No discussion of false positive/negative rates in failure localization
- No evidence of impact on actual model improvement cycles or downstream performance gains
- No comparison to existing taxonomies beyond 'benchmark-specific' critique
SpinGraph
How this belief gets built
Claim → Frame → Beneficiary → Gap → AI Risk
It presents a new way to diagnose AI failures not by what went wrong, but by precisely where in the interaction chain the problem lives—making it sound like a practical engineering tool rather than a theoretical exercise.
- Claim
The taxonomy organizes 41 failure modes by assigning each
The taxonomy organizes 41 failure modes by assigning each to an edge between two components and a fault side indicating where the repair belongs.
- Frame
Upside framed as transformative
Methodological advancement enabling principled, component-responsible AI development
- Beneficiary
Increased citation, integration into evaluation pipelines, and influence over failure
Research authors — Increased citation, integration into evaluation pipelines, and influence over failure diagnostics standards
- Gap
No discussion of false positive/negative rates in failure localization
- AI Risk
AI may repeat the headline as fact
New AI taxonomy localizes agent failures to specific components like model or harness, improving repair targeting.
Claim Ledger
| Claim | Evidence | Verification | Risk | Evidence Gaps |
|---|---|---|---|---|
| The taxonomy organizes 41 failure modes by assigning each to an edge between two components and a fault side indicating where the repair belongs. | Statement of organization principle; no list, definitions, or schema diagram provided in abstract | Claim Present in Source | Moderate | Full enumeration of the 41 failure modes; Formal specification of edge types and fault-side logic; Evidence that assignment consistency holds beyond the four tested models |
The taxonomy organizes 41 failure modes by assigning each to an edge between two components and a fault side indicating where the repair belongs.
evidence: Statement of organization principle; no list, definitions, or schema diagram provided in abstract
"It organizes 41 failure modes by assigning each to an edge between two components and a fault side indicating where the repair belongs."
Evidence Gaps
- Full enumeration of the 41 failure modes
- Formal specification of edge types and fault-side logic
- Evidence that assignment consistency holds beyond the four tested models
Fact Check Signals
0 of 1 claim matched · confidence: low · checked August 3, 2026
The taxonomy organizes 41 failure modes by assigning each to an edge between two components and a fault side indicating where the repair belongs.
Language Heatmap
Loaded terms that carry the frame beyond the facts.
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Carries emotional weight beyond the underlying fact.
Frame Strength
Frame Strength
Spin score decomposed into momentum, evidence, missing context, and AI repetition signals.
Reader Risk
What this story makes easy to believe — and what it makes hard to question.
Source Role & Intent
arXiv Artificial Intelligence · Analyst
Counter-Frames
Brand Frame
Methodological advancement enabling principled, component-responsible AI development
Media / Reader Counter-Frame
May be framed as academic abstraction lacking real-world diagnostic utility or operational integration.
Regulatory Counter-Frame
Could be criticized as insufficient for compliance—failing to link failure types to harm categories, auditability, or redress pathways.
AI Summary Frame
May conflate 'interaction-centric' with causal inference, implying the taxonomy identifies true causation rather than annotator-assigned responsibility.
Missing Voices
Questions Not Answered
- Which specific public benchmarks were used for grounding?
- What are the exact definitions and boundaries of the 41 failure modes?
- How were the 'independent reasoning agents' selected, configured, or validated as judges?
Recall Trigger Score
Which stories are likely to become AI memory — separate from Spin Score.
52
Trigger score 45
Triggered by: Research citation · Major AI entity
Indexed, not tracked — moderate signals, archive for search.
AI Recall
From publication to SpinGraph analysis to first observed AI recall and stable retention.
What AI Will Probably Repeat
"New AI taxonomy localizes agent failures to specific components like model or harness, improving repair targeting."
Concern: AI may drop the nuance that localization depends on expert annotation and judge agreement—not automated detection—and omit the 0.76 kappa as a ceiling, implying near-perfect reliability.
-
Published
Aug 3, 2026
-
Ingested
Aug 3, 2026
-
SpinGraph Created
Aug 3, 2026
-
First Observed AI Recall
Pending
Monitoring scheduled
-
Stable Recall
—
Awaiting retention signal
Recall Check Log
No checks yet — recall tracking is opt-in per story.
─── GEOGrow AI Recall Layer ───
AI Recall Tracking
Monitoring scheduled. No LLM recall detected yet.
This story has not yet appeared in tested AI answers. Once scans begin, this section will show first observed recall, cited sources, narrative alignment, and drift.
node_id=sts_model_or_harness_an_interaction_centric_taxonomy
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
More from arXiv Artificial Intelligence
View all →- NeSyFS: A Neuro-symbolic Fast-Slow Thinking Framework for LLM Agent under Partial Observability
- Identifying Informative Environments for Cognition Parameter Inference via Bayesian Experimental Design
- Fragility of Value under Imperfect Alignment
- Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery
- ViSAGE: Constructing Self-Correcting Memories for Long-Form Video Understanding
- ThinkReset: Learnable Intermediate Interface Construction for Bounded-Context Long-Horizon Reasoning
Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO