Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
7 results for “MLLMs”
Time-Series Retrieval for Grounding Multimodal Language Models in Remaining Useful Life
A research paper proposes using time-series retrieval to ground multimodal LLMs for remaining useful life (RUL) estimation in aircraft engine prognostics, showing improved prediction accuracy and stability over non-retrieval baselines on the FD001 C-MAPSS benchmark.
Aug 21, 2026
When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language Models
A new arXiv preprint identifies a consistent, mathematically characterizable bias in multimodal large language models (MLLMs) caused by task-irrelevant text — revealing that such context induces predictable affine distortions in decision margins rather than random noise.
Aug 21, 2026
Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards
A new arXiv survey paper identifies novel safety threats unique to multi-modal large language models (MLLMs) — such as modality misalignment and fused safety risks — and proposes a multimodal-grounded taxonomy to guide future safety research.
Aug 11, 2026
AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models
Researchers introduced AgentPatch, a training-free method to repair performance degradation in merged agentic multimodal large language models (MLLMs), specifically addressing weak-task failure and behavior-critical forgetting after model merging.
Aug 10, 2026
RoCo-ACE: Rollout-Conditioned Online Distillation for Retention-Aware Knowledge Injection
RoCo-ACE is a new online distillation method for knowledge injection into multimodal large language models that improves factual accuracy of injected knowledge while preserving model behavior on non-updated tasks.
Jul 29, 2026
DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth
Researchers introduced DocOCR-Eval, an annotation-free framework to rank OCR and multimodal LLM tools for document parsing without ground-truth labels, addressing the challenge of tool selection in label-scarce real-world settings.
Jul 21, 2026
Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation
Researchers introduced an automated, multi-agent red-teaming system that synthesizes adversarial multimodal examples to improve MLLM content safety robustness, reducing false negatives by 16.7 percentage points on a public benchmark without human labeling.
Jul 17, 2026