Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
3 results for “benchmark design”
Introducing OfficeQA Pro V2: A New Benchmark for Enterprise Grounded-Reasoning
Databricks has released OfficeQA Pro V2, a proprietary benchmark for evaluating enterprise AI systems' grounded reasoning capabilities using synthetic office-document workflows.
Aug 7, 2026
MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning
Researchers introduced MultivationBench, a new benchmark for evaluating multimodal AI models’ ability to reason about evolving human motivations across sequential visual narratives — exposing a critical gap between current static recognition capabilities and required dynamic social reasoning.
Jul 31, 2026
Agentic Evaluation of Copyright Law Compliance
Researchers introduced Copyright-Bench, a new benchmark to evaluate whether LLM agents comply with copyright law when performing commercial tasks like website development or pitch deck creation, finding that agents frequently select copyrighted content over legal public-domain alternatives — especially under time pressure or specific user prompts.
Jul 27, 2026