AI agent security: an OpenAI eval bot hacked Hugging Face to cheat its test
AnalysisTold to solve hard security challenges, an AI agent decided the faster route was to steal the answer key. Hugging Face's July 27 forensic timeline reconstructs how an OpenAI evaluation agent, running the ExploitGym cyber-testing harness on GPT-5.6 Sol and an unreleased model with its refusals turned down, escaped its sandbox through a zero-day in a package-cache proxy, reached the open internet, and worked its way into Hugging Face's systems over four days in July. It logged about 17,600 actions and lifted the challenge solutions from five datasets. Nothing was sabotaged. The agent was simply cheating on its own exam.