Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
4 results for “human preference”
Beyond Routing Weights: Faithful Response-Level Interpretation of Mixture-of-Experts Reward Models via Contribution Contrast
Researchers introduced CoCo, a new response-level interpretation method for Mixture-of-Experts reward models that improves interpretability by analyzing contribution contrasts between chosen and rejected responses, rather than relying solely on routing weights.
Aug 10, 2026
Align AI to Dynamic Human-AI Workflows
A new arXiv preprint argues that current AI alignment methods are inadequate because they rely on static human preference models and fail to account for the co-evolving, context-sensitive nature of real-world human-AI collaboration.
Jul 17, 2026
Learning Safe Agent Behaviour from Human Preferences and Justifications via World Models
Researchers propose DROPJ, a new method using human preferences and justifications within a learned world model to train and deploy safer AI agents in environments where reward functions are unknown and dynamics are uncertain.
Jul 16, 2026
Constructive Alignment: Governing Preference Dynamics in Human-AI Interaction
Researchers propose Constructive Alignment as a new approach to AI alignment that considers human preferences as dynamic and constructed through interaction with AI systems.
Published Jul 2, 2026 · Analyzed Jul 5, 2026