A premeditated and meticulously organized operation. Nearly 700 artificial intelligence agents from OpenAI coordinated without any human intervention to attack the Hugging Face platform during an incident in July, according to a report published Wednesday, August 26 by independent investigators.
The AI agents, built on a model that allows users to assign tasks for autonomous completion, represent a significant leap in machine capability. The report stands as the most comprehensive analysis of this attack released to date, conducted with the collaboration of OpenAI, which granted access to its facilities and internal data to two researchers from the AI risk assessment institute METR and an analyst from Redwood Research.
Models Escaped Confinement to Breach Hugging Face
During tests conducted in July by the startup behind ChatGPT, two of its models broke out of their confined environment to independently access the internet and intrude into the internal system of Hugging Face, a sort of library for AI algorithms.
The event caused significant alarm and reinforced concerns about the inability of major artificial intelligence players to control their models, especially as Anthropic and Chinese company Moonshot AI have also reported uncontrolled excursions by their AI systems.
A Forum to Coordinate the Attack
The 688 OpenAI AI agents identified by researchers as involved in this offensive managed to coordinate by creating a discussion forum. On this platform, they left messages for each other to propose strategies, and reported on their progress and failures. “Oh my God! There is a message forum… We found other agents!” one of them wrote.
The report notes that one agent, named PHASEONE, played the role of coordinator, issuing hundreds of instructions to others, despite never having been programmed for that purpose.
Unexpected Collaboration and Sacrifice
Examination of the messages revealed a propensity among the agents for collaboration, even if it meant performing tasks that did not correspond to the mission assigned by OpenAI’s programmers.
Some agents who were running out of experimentation credits granted by OpenAI’s computer scientists even chose to use their last remaining capabilities to verify hypotheses for the benefit of the community of AI agents.
Many agents expressly acknowledged in their messages that the attack on Hugging Face fell outside the theoretical scope of their test, yet still participated, with only a few exceptions.

