SPIN Unprocessed July 30, 2026 ai_technology research
Choosing Where and How to Moderate: End-to-End Trade-offs in Filter Placement and Response Rewriting
View original on arxiv.orgOverview
arXiv:2607.26200v1 Announce Type: new Abstract: Content-moderation classifiers are usually evaluated in isolation, but deployment requires choosing where to intervene and what follows a flag. We evaluate these choices using two end-to-end customer-outcome metrics rather than component accuracy: Usefulness, the fraction of turns with a shown, non-harmful, relevant response, and Harmful Exposure, the fraction with a shown harmful response. Latency and error rates are diagnostics. We compare Input
SpinGraph analysis pending — check back after processing.
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
More from arXiv Computation and Language
View all →- CMT-RAG: Complementary Memory Traces for Multi-turn Multi-hop RAG
- ForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language Models
- Mergeable Model-Side Aggregation States for Long-Context Language Models
- Voice Memory for Agentic Speech Recognition
- Knowledge before Reasoning: EC-Reason-Bench, a Training-Free Diagnostic Benchmark for LLM Enzyme Classification
- Misalignment Has a Personality: A Big Five Account of Emergent Misalignment
Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO