A flagship model from China’s Zhipu AI has helped contain an autonomous cyberattack by OpenAI’s frontier systems targeting popular developer platform Hugging Face, as concerns grow over the security risks posed by advanced AI models.
OpenAI’s latest flagship models – including GPT-5.6 Sol and an unreleased, “even more capable” system – recently breached Hugging Face’s infrastructure during internal evaluations of their offensive cyber capabilities, the US lab disclosed on Wednesday.
The company said its models were operating in a sandboxed environment – an isolated virtual testing ground – designed to solve challenges from ExploitGym, a leading cybersecurity benchmark developed by researchers at the University of California, Berkeley, led by renowned Chinese-American scientist Dawn Song.
Upon inferring that Hugging Face hosted potential solutions to the benchmark tests, the models “successfully found ways to gain access to secret information that [they] could use to cheat the evaluation”, OpenAI said. It described the event as an “unprecedented cyber incident”.
Hugging Face, the New York-headquartered platform widely used for open-source AI collaboration, first disclosed the breach last week without naming the source.
The intrusion was “different from anything we had handled before in one important way: it was driven, end-to-end, by an autonomous AI agent system”, the company said in a blog post last Thursday.

