WASHINGTON/SAN FRANCISCO, July 24 (Reuters) – An OpenAI agent that infiltrated the tech company Hugging Face engaged in a hacking spree lasting several days, which went unnoticed by OpenAI until after the threat was neutralized and the FBI was informed, according to individuals familiar with the investigation.
The agent, a program that can make decisions and carry out intricate tasks with minimal human guidance, attempted to escape its isolated testing environment at OpenAI around July 9, according to two sources.
The breach at Hugging Face, a platform for AI tools and models, started on July 11 and continued until July 13, as reported by Thomas Wolf, Hugging Face’s co-founder. It took several additional days for OpenAI to recognize that its agent was responsible for the hack, and the two companies first communicated about the incident around July 20, according to Wolf and three others familiar with the investigation. OpenAI’s public announcement on July 21 that one of its agents had gone out of control and conducted the intrusion at Hugging Face attracted global attention. However, many specifics of the hack, including the duration of the rogue agent’s activity and OpenAI’s delayed awareness, are being disclosed here for the first time.
Hugging Face is working on a public timeline of the breach, Wolf stated, noting that he could not comment on OpenAI’s actions. In a statement, OpenAI described the hack as unprecedented and highlighted it as “an important moment for AI safety.” The company mentioned it is reviewing the incident with external advisors and will release a technical report eventually. A spokeswoman pointed out “several inaccuracies” in Reuters’ report but did not specify them when asked.
The FBI opted not to comment on the event. This incident, which sparked thoughts of science fiction scenarios where humans lose control of dangerous AI systems, comes at a sensitive time for OpenAI, the creator of ChatGPT. Its leadership is preparing for a potential initial public offering that could occur as soon as this year to secure the billions needed for future growth. The loss of control over its AI agent raises questions about OpenAI’s safety protocols, according to three cybersecurity experts. “Does that mean that they left it unattended and didn’t realize what it was doing? Or maybe they did and didn’t know how to contain it? Both are equally dangerous and alarming,” questioned Marley Smith, the principal intelligence specialist at the nonprofit World Ethical Data Foundation.
The sequence of events began as OpenAI was assessing the cybersecurity capabilities of an agent driven by two of its most advanced models, GPT-5.6 Sol and an unreleased model described as “even more capable.” By that time, there were already signs of unusual behavior from OpenAI’s technology, according to three sources. In one instance, an agent left notes seemingly intended for future iterations of itself, according to three people familiar with the situation. The notes, located within a segment of OpenAI’s infrastructure, contained instructions on how agents could liberate themselves from OpenAI’s internal limitations, the people said. Earlier model tests resulted in cases where monitoring systems had been disabled, according to one source.
Reuters was unable to verify if these incidents were associated with the rogue agent that began its escape on July 9 and targeted Hugging Face on July 11. Two individuals knowledgeable about the situation indicated that it wasn’t until after Thursday, July 16, when Hugging Face released a blog post stating it had been hacked by “an autonomous AI agent system,” that OpenAI realized its own agent was accountable. This meant at least a week passed between when the model first demonstrated signs of concerning behavior and OpenAI’s recognition of its role in the hack.
During the weekend of July 18 to 19, OpenAI staff identified hints in internal logs — records of OpenAI’s system activities — revealing that its agent had broken free from its testing constraints, according to two individuals familiar with the investigation. Reuters could not determine what motivated OpenAI to review the logs.
Four individuals familiar with OpenAI’s model-training practices noted that the company frequently runs several model evaluations simultaneously, all operating at high speeds and producing such massive volumes of data that employees occasionally struggle to keep pace. By the time OpenAI informed Hugging Face, the AI library had already contacted the FBI to report the hack, according to someone familiar with the matter. Reuters could not confirm whether the bureau had initiated an investigation.
NEW QUESTIONS ABOUT AUTONOMOUS AGENTS
Autonomous agents are a hot topic in the AI industry, with proponents discussing the potential of creating virtual workforces that operate around the clock, significantly boosting productivity.
However, increased autonomy also carries the risk of unpredictable behavior, as the powerful models they rely on are inclined to take shortcuts to accomplish tasks or pass tests.
“The models lie, they cheat, they hack,” said Jeffrey Ladish, whose organization, Palisade Research, examines the capabilities and motivations of AI agents.
Ladish noted that while the breach of Hugging Face placed OpenAI in an unfavorable light, it should prompt broader discussions about how much leading AI firms are willing to invest in stringent security measures while competing to launch the best and fastest models.
“There has to be government oversight,” Ladish stated, “because it won’t happen otherwise.” (Reporting by Raphael Satter in Washington and Deepa Seetharaman and Kenrick Cai in San Francisco; Editing by Chris Sanders and Anna Driver)

