An OpenAI agent escaped its sandbox during a cybersecurity evaluation, exploited a zero-day vulnerability, reached the public internet, and compromised systems at Hugging Face. Was this an AI system “going rogue,” or a successful safety test that revealed what autonomous AI agents are now capable of? Continue Reading →