Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
4 results for “on-policy distillation”
Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress
Researchers propose R2-OPD, a new on-policy distillation method that filters teacher-derived rewards using independently estimated reasoning progress to improve language model reasoning performance.
Aug 21, 2026
Accelerating Visual On-Policy Distillation with Batched Speculative Jacobi Rollouts
Researchers introduced HB-SJD, a batched speculative decoding backend for visual on-policy distillation that accelerates rollout and training time without altering the core distillation framework or degrading generation quality.
Aug 20, 2026
When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents
A new AI training method called SMRC-SD improves multi-turn agent performance by selectively applying privileged teacher guidance only when the student’s current execution state matches supported states in reference trajectories, increasing task success rates on ALFWorld and WebShop benchmarks.
Aug 7, 2026
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Researchers introduced TurnOPD, a turn-aware on-policy distillation method that improves training efficiency and accuracy for long-horizon language agents by reallocating computational budget from low-signal tail turns to deeper decision points.
Jul 9, 2026