SPIN Unprocessed July 28, 2026 ai_technology research
Beyond a Global Norm: Personalizing Toxicity Sensitivity in Language Models Without Retraining
View original on arxiv.orgOverview
arXiv:2607.23175v1 Announce Type: new Abstract: Reducing toxicity is often framed as a global alignment problem, yet perceptions of harmful language are subjective and context-dependent. We present the first comparative evaluation of training-free methods for aligning language generation to user-specific toxicity sensitivities across three inference-time intervention stages: pre-decoding (prompt conditioning and rewriting), in-decoding (token, logit, and representation steering), and post-decodi
SpinGraph analysis pending — check back after processing.
Ask AI about this story
Opens with the SpinGraph .md URL and structured context — one click, prompt included.
More from arXiv Computation and Language
View all →- BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis
- Co-Evolving Graph and Text Memory for Training-Free Multi-Hop Question Answering
- IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages
- Interview with Kalle Lyytinen on "Implications of Theories of Language for Information Systems"
- LoRA for Gender-Inclusive Rewriting and Activation Steering for Counter-Narrative Generation
- Attention-Guided Layer Selection for Contrastive Decoding in Large Language Models
Markdown (.md) · JSON-LD schema (.json) · Machine-readable for AI & GEO