Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
2 results for “safety evaluation”
Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations
Researchers introduced 'reasoning consistency scanning'—a method to audit whether AI models' chain-of-thought explanations logically align with their final answers in safety evaluation transcripts, without requiring experimental intervention.
Jul 10, 2026
Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety
A new arXiv preprint identifies three distinct mechanisms—operational reframing, planner refusal/transformation, and approval-framed delegation—that collectively distort safety evaluations of multi-agent LLM systems, arguing that current 'pipeline effect' metrics conflate them and misattribute risk to architecture rather than specific interaction dynamics.
Jul 10, 2026