Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
5 results for “LLM-as-a-Judge”
Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation
A research paper demonstrates that LLM-as-a-Judge evaluation systems often rely on rubric text alone—not candidate responses—to generate scores, undermining their validity as objective evaluators of AI-generated text.
Sep 4, 2026
Presentation: The Right 300 Tokens Beat 100k Noisy Ones: The Architecture of Context Engineering
Two practitioners propose context engineering techniques to improve coding agent reliability by reducing prompt noise and optimizing context window usage.
Aug 14, 2026
Hallucinations on the Board: Tool-Augmented Evaluation of LLM Chess Commentary
Researchers introduced ACT-Eval, a tool-augmented framework to detect and quantify hallucinations in LLM-generated chess commentary by decomposing claims and validating them against chess engines and expert annotations.
Aug 6, 2026
Presentation: Designing AI Platforms for Reliability: Tools for Certainty, Agents for Discovery
A presentation by Aaron Erickson describes NVIDIA’s internal approach to designing AI agent systems with an emphasis on reliability, testing, and architectural balance — but provides no verifiable details about implementation, outcomes, or validation.
Jul 9, 2026
Presentation: Trustworthy Productivity: Securing AI-Accelerated Development
A technical presentation outlines emerging security patterns for autonomous AI agents, focusing on vulnerabilities in the ReAct loop and proposing mitigation strategies like LLM-as-a-judge and MAESTRO threat modeling.
Published Jun 30, 2026 · Analyzed Jul 4, 2026