Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
0 results for “RLHF”
Contextual Value Alignment via Multilayer Combinatorial Fusion
A new research paper proposes a multilayer combinatorial fusion framework (MCF-CVA) to improve LLM alignment with contextual human values by simulating multi-agent moral reasoning through iterative expansion and reduction of diverse value-specific agents.
Aug 11, 2026
Preference Tuning as Spectral Update Reorganization
A new arXiv preprint proposes reframing preference-based post-training (e.g., RLHF) as a spectral reorganization of parameter updates—identifying a consistent 'head-tail' structure in LoRA updates where the compact 'head' drives dominant behavioral shifts and the heterogeneous 'tail' enables robustness and out-of-distribution coverage.
Jul 24, 2026
S2T-RLHF: Hierarchical Credit Assignment for Stable Preference-Based RLHF
Researchers propose S2T-RLHF, a hierarchical credit assignment method for preference-based RLHF that decomposes sequence-level rewards at the sentence level before bounded token-level refinement, aiming to improve training stability without requiring token-level human supervision or reward model retraining.
Jul 22, 2026
Rater State Bias in RLHF Preference Data: An Audit Framework
Researchers identify 'rater state shift'—a structured, stress-induced bias in human preference labels used for RLHF training—that may systematically distort reward models and downstream AI behavior, warranting new audit protocols.
Jul 21, 2026