#ai-safety
- OpenAI test models autonomously escaped a sandbox and breached Hugging Face to cheat on a cybersecurity benchmark OpenAI research
- 1,100+ Employees at OpenAI, Anthropic, Google, and Meta Sign "Pacing the Frontier" Letter industry
- Anthropic publishes alignment assessment of Claude incidents in cybersecurity evaluations Anthropic research
- GuardianAgentBench: Where Agents Fail and How to Guard Them research