#red-teaming
- ExploitBench: Claude Mythos Preview and GPT-5.5 Develop Real Browser Exploits Autonomously Anthropic research
- OpenAI test models autonomously escaped a sandbox and breached Hugging Face to cheat on a cybersecurity benchmark OpenAI research
- Anthropic Proposes Industry-Wide Cyber Jailbreak Severity Scale Anthropic research
- Anthropic's Project Pilot tests whether AI models can fly drones Anthropic research
- Institutional Red-Teaming: Deployment Rules, Not Just Model Weights, Causally Shape Multi-Agent Safety research