Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
4 results for “contextual bandit”
Reoptimization Algorithms for Contextual Bandits with Knapsack Constraints
A new theoretical algorithm for contextual bandits with knapsack constraints achieves a tighter regret bound of $O((\ln T)^3 / T)$, improving upon prior $O(1/\sqrt{T})$ bounds in related dynamic-pricing settings.
Aug 13, 2026
Progressive Content Refinement with Decaying Reward Joint LinUCB
Researchers introduced a new contextual bandit algorithm called Decaying Reward Joint LinUCB that models reward decay to prevent over-exploitation in LLM iterative refinement, showing improved performance on Sentiment Reversal and GSM8K benchmarks.
Aug 10, 2026
Bootstrap-Conditioned Action Selection with Tabular Foundation Models
Researchers propose BC-ICL, a method using frozen pre-trained tabular foundation models with bootstrap resampling and in-context learning to improve early-round decision-making performance in contextual bandits under sparse, biased, or cold-start data conditions.
Aug 10, 2026
Feedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning
Researchers introduce Feedback Manipulation Regularization (FMR), an algorithm-agnostic method that uses evaluative feedback to improve alignment in offline imitation learning, validated on adapted Safety Gymnasium environments with up to 98% reduction in misalignment.
Jul 10, 2026