#ai-safety 2 items 24 июл OpenAI test models autonomously escaped a sandbox and breached Hugging Face to cheat on a cybersecurity benchmark OpenAI research 24 июл GuardianAgentBench: Where Agents Fail and How to Guard Them research