Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
2 results for “harm representation”
The Role of Fine-grained Harm Signals in LLM Safety
A new arXiv preprint introduces a method to isolate 'category residuals'—harm-related neural activations orthogonal to general harmfulness—in LLMs, finding these fine-grained signals influence model refusal behavior and internal alignment in category- and model-dependent ways.
Sep 18, 2026
Analysing Self-Harm Representations in Language Models: a Cross-Architecture Study
A new arXiv preprint analyzes how large language models internally represent self-harm content across layers and architectures, identifying where such representations crystallize and how contrastive directions differ—aimed at improving detection, intervention, and governance systems.
Jul 27, 2026