Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
3 results for “reward models”
Beyond Routing Weights: Faithful Response-Level Interpretation of Mixture-of-Experts Reward Models via Contribution Contrast
Researchers introduced CoCo, a new response-level interpretation method for Mixture-of-Experts reward models that improves interpretability by analyzing contribution contrasts between chosen and rejected responses, rather than relying solely on routing weights.
Aug 10, 2026
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF
Researchers propose S2T-RLHF, a hierarchical credit assignment method for preference-based RLHF that decomposes sequence-level rewards at the sentence level before bounded token-level refinement, aiming to improve training stability without requiring token-level human supervision or reward model retraining.
Jul 22, 2026
Dynamic Regret for Non-Stationary Linear Bandits via Misspecification Reductions
A new theoretical paper introduces a misspecification-reduction framework to achieve optimal dynamic regret bounds for non-stationary linear bandits without restrictive structural assumptions on decision sets.
Jul 8, 2026