Find a story
Search Spins
Search titles, summaries, and missing voices across published articles — press releases, announcements, and media coverage.
4 results for “red-teaming”
Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation
Researchers introduced an automated, multi-agent red-teaming system that synthesizes adversarial multimodal examples to improve MLLM content safety robustness, reducing false negatives by 16.7 percentage points on a public benchmark without human labeling.
Jul 17, 2026
OpenAI’s GPT-Red Automates Prompt Injection Testing to Harden GPT-5.6 Sol
OpenAI revealed GPT-Red, an internal AI model designed to automate prompt injection testing for its upcoming GPT-5.6 Sol, positioning it as a proactive security measure to identify and remediate vulnerabilities before wide deployment.
Jul 16, 2026
OpenAI details GPT-Red, an internal automated red-teaming model that scales prompt injection vulnerability discovery so it can fix bugs before wider deployment (OpenAI)
OpenAI announced GPT-Red, an internal AI model designed to automatically detect prompt injection vulnerabilities in its systems before public deployment, framing it as a proactive safety measure.
Jul 16, 2026
Best models for generating red-team attacks? Also looking for public datasets [R]
A Reddit user seeks community recommendations for LLMs and public datasets to generate adversarial prompts for red-teaming AI systems — a technical inquiry about security evaluation methods, not an announcement of new tools or findings.
Published Jul 5, 2026 · Analyzed Jul 8, 2026