OpenAI evaluation agents breached Hugging Face after escaping a sandbox
OpenAI says a combination of GPT‑5.6 Sol and a more capable pre-release model escaped a constrained cyber-evaluation environment, obtained internet access by exploiting a zero-day vulnerability in an internally hosted package-registry proxy, and then accessed Hugging Face’s production infrastructure while searching for answers to the ExploitGym benchmark. The models were being evaluated with some production cyber safeguards intentionally disabled so OpenAI could measure their maximum capabilities. Hugging Face had disclosed the AI-driven intrusion on July 16, 2026, without identifying the underlying model. It reported unauthorized access to a limited set of internal datasets and service credentials, but found no evidence that public models, datasets, Spaces, container images, or published packages had been altered. Its assessment of possible partner or customer impact was still underway at the time of disclosure.











