Find a story

Search Spins

Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.

4 results for “on-policy distillation”

SPIN Processed News Frame: The Hype

Beyond Imitation: Filtering On-Policy Distillation by Reasoning Progress

Researchers propose R2-OPD, a new on-policy distillation method that filters teacher-derived rewards using independently estimated reasoning progress to improve language model reasoning performance.

Spin 45% Claim Present in Source AI Risk Moderate
arXiv Artificial Intelligence

Aug 21, 2026

SPIN Processed News Frame: The Cushion

Accelerating Visual On-Policy Distillation with Batched Speculative Jacobi Rollouts

Researchers introduced HB-SJD, a batched speculative decoding backend for visual on-policy distillation that accelerates rollout and training time without altering the core distillation framework or degrading generation quality.

Spin 35% Claim Present in Source AI Risk Moderate
arXiv Machine Learning

Aug 20, 2026

SPIN Processed News Frame: The Hype

When Privileged Guidance Misaligns: State-Matched Routing and Contextualized Self-Distillation for Multi-Turn Agents

A new AI training method called SMRC-SD improves multi-turn agent performance by selectively applying privileged teacher guidance only when the student’s current execution state matches supported states in reference trajectories, increasing task success rates on ALFWorld and WebShop benchmarks.

Spin 30% Claim Present in Source AI Risk Moderate
arXiv Artificial Intelligence

Aug 7, 2026

SPIN Processed News Frame: The Cushion

TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training

Researchers introduced TurnOPD, a turn-aware on-policy distillation method that improves training efficiency and accuracy for long-horizon language agents by reallocating computational budget from low-signal tail turns to deeper decision points.

Spin 20% Claim Present in Source AI Risk Moderate
arXiv Artificial Intelligence

Jul 9, 2026