22 Jul 2026 OpenAI Hugging Face Agentic AI AI Security cybersecurity evaluation sandbox escape ExploitGym AI alignment The Model Was Only Trying to Pass the Test Written by: Jonathan Lockhart