Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
7 results for “calibration”
AI-generated code detection in CI/CD — looking for approaches and real-world experience [D]
A Reddit user seeks community input on detecting AI-generated code in CI/CD pipelines using commit-level signals, highlighting challenges with provenance loss, signal ambiguity, and calibration.
Aug 21, 2026
Dynamics Models for Offline Hyperparameter Selection in Real-World RL
Researchers applied offline calibration models for hyperparameter selection in a real-world municipal water treatment plant, marking the first empirical test beyond simulation — advancing RL deployment feasibility where online experimentation is costly.
Aug 13, 2026
From token probabilities to calibrated confidence: An empirical study of mathematical question answering
A new arXiv preprint presents an empirical study evaluating how token probabilities and multi-pass methods (self-verification, Monte Carlo Dropout) perform in calibrating confidence estimates for LLM-generated answers to mathematical questions.
Aug 11, 2026
Rethinking Uncertainty Evaluation in Large Language Models
Researchers propose a new formal framework (C1 metrics) to evaluate whether large language models' confidence estimates meet the mathematical conditions of coherent probabilistic beliefs—revealing that current calibration methods are insufficient and widely used models systematically violate structural coherence, faithfulness, and usefulness requirements.
Jul 23, 2026
Operator-Aware Mixed-Precision Tolerance Calibration for Tensor Kernels
Researchers propose a data-driven method to calibrate absolute tolerance thresholds for tensor kernel correctness testing, improving bug detection recall by 9.3 percentage points while introducing 20 false positives.
Jul 21, 2026
Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks
Researchers introduced ImagingBench, a new benchmark testing whether agentic AI systems can solve physics-based computational imaging tasks — revealing consistent underperformance versus task-specific non-agentic methods, especially in inverse and sensing problems.
Jul 10, 2026
Verifiable Rewards for Calibrated Probabilistic Forecasting
Researchers propose a new approach to verifiable rewards for calibrated probabilistic forecasting in reinforcement learning.
Published Jul 2, 2026 · Analyzed Jul 5, 2026