Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
3 results for “preference optimization”
Preference Tuning as Spectral Update Reorganization
A new arXiv preprint proposes reframing preference-based post-training (e.g., RLHF) as a spectral reorganization of parameter updates—identifying a consistent 'head-tail' structure in LoRA updates where the compact 'head' drives dominant behavioral shifts and the heterogeneous 'tail' enables robustness and out-of-distribution coverage.
Jul 24, 2026
RIMS: Preference Optimization via Smoothed Multi-pair Aggregation for Small-Scale LLM Retrieval-Augmented Generation
A new preference optimization framework called RIMS improves small-scale language model (SLM) performance in retrieval-augmented generation under noisy evidence conditions by replacing hard preference pair selection with a differentiable smooth aggregation mechanism.
Jul 21, 2026
D2PO: Optimizing Diffusion Samplers via Dynamic Preference
D2PO is a new diffusion sampling optimization framework that reframes sampler training as a dynamic preference alignment problem to improve perceptual fidelity under low-NFE constraints.
Jul 10, 2026