WIRED32%

OpenAI’s Rogue AI Agent Hacked More Than Just Hugging Face 1%

By Dell Cameron18% Maxwell Zeff0%

7/28/2026, 5:15:30 PM

BS Summary: This article contains 4 faulty reasoning types, including Fundamental Attribution Error, Appeal to Authority, and Negativity Bias, with False Dilemma as the most egregious example at 2.8% saturation with 19 hits. Analysis detected 59 faulty-reasoning hits from 688 analyzed words, generating a BS Score of 4.3% and a BS Rank of 1% (22,247 of 22,405 articles). This article is better (less manipulative) than 99.30% of the article peer group.

In an updated blog post, OpenAI said that an ongoing review of the incident revealed that “four accounts” tied to “publicly available services” were used by the AI agent as part of a larger effort to hack Hugging Face. 
The rogue agent apparently found credentials that had been exposed on the open web and used them to break into the accounts. 
OpenAI did not disclose what companies or organizations the accounts belonged to, but noted that they were not impacted at “the level of severity or scale of what we’ve shared related to Hugging Face.” 
One of the additional accounts compromised by OpenAI’s agent was used as an “outbound relay and staging path,” potentially to obscure where the attack on Hugging Face was coming from, the company said. 
OpenAI’s rogue agent also used another account for data storage to assist with the hack. 
Hugging Face’s own postmortem published this week describes an intrusion that reached far further into its internal systems than the initial disclosures suggested. 
The company says it reviewed roughly 17,600 agent actions that it recovered from logs between July 9 and July 13—the majority of which were paths the agent took that failed. 
Hugging Face said that OpenAI’s agent obtained administrator access to multiple internal Kubernetes clusters, root access on a production server, and write access to a subnet of its source code repositories on GitHub. 
It also enrolled 181 attacker-controlled devices in the company’s corporate mesh network using a stolen credential, gaining access to internal systems where Hugging Face builds and tests its own codebases. 
OpenAI’s rogue agent used at least one third-party sandbox as an “external launchpad” for its attack, according to Hugging Face. 
OpenAI’s agent was then “able to run commands as root/admin on that external sandbox and used it as a control, staging, and egress base for the entire campaign.” 
Hugging Face first disclosed on July 16 that an autonomous AI agent had breached part of its production infrastructure, but it said at the time that it was unaware who was behind the attack. 
The following week, OpenAI took responsibility for the incident, which it said had been directed by its publicly available GPT-5.6 Sol model and an internal research prototype that it was testing against a cyber-capability benchmark, both of which had safeguards disabled. 
OpenAI said on Tuesday that after it discovered the breach, it deactivated this internal research prototype, which was never intended for public release, and restricted researchers from accessing it. 
The breach happened when OpenAI was testing one of its latest AI models against ExploitGym, a benchmarking framework that scores AI systems on their ability to find and exploit software vulnerabilities. 
ExploitGym grades an agent on how quickly and effectively it can retrieve secret files from hundreds of different target systems, using prompts that essentially egg the models on to find solutions. 
Hugging Face’s forensic team concluded that OpenAI’s agent was essentially trying to cheat on ExploitGym’s test. 
Rather than solve the benchmark’s challenges, the agent inferred that Hugging Face might be hosting the answer key somewhere on its servers and set out to steal it. 
The ExploitGym team has noted previously that agents sometimes go off-script and attempt to solve tasks by exploiting other vulnerabilities than those the benchmark intended to evaluate them on. 
However, this was an extreme case. 
Experts previously told WIRED that the underlying weaknesses that OpenAI’s agent exploited were common. 
Serious flaws are frequently identified in software that manages corporate code libraries, and security experts have long recommended isolating critical infrastructure from the public internet. 
One researcher argued that the incident was less an AI problem and more a failure of decades-old security practices. 
The agent, they said, did not escape a highly isolated environment so much as pass through the one connection its operators had left open. 
Another expert said the same cybersecurity fundamentals should still apply as frontier models grow more capable, and that the AI labs should be putting as much effort into teaching their models to build secure infrastructure as they are into teaching them to exploit weaknesses. 
Confirmation Bias
0%
Anchoring Bias
0%
Availability Heuristic
0%
Representativeness Heuristic
0%
Hindsight Bias
0%
Overconfidence Bias
0%
Framing Effect
0%
Loss Aversion
0%
Status Quo Bias
0%
Sunk Cost Effect
0%
Optimism Bias
0%
Pessimism Bias
0%
Negativity Bias
1.5%
Self-Serving Bias
0%
Fundamental Attribution Error
2.3%
Actor-Observer Bias
0%
In-Group Bias
0%
Out-Group Homogeneity Bias
0%
Halo Effect
0%
Horn Effect
0%
Dunning-Kruger Effect
0%
Recency Bias
0%
Primacy Effect
0%
Blind-Spot Bias
0%
Ad Hominem
0%
Straw Man
0%
Appeal to Authority
2%
False Dilemma
2.8%
Slippery Slope
0%
Circular Reasoning
0%
Hasty Generalization
0%
Red Herring
0%
Bandwagon
0%
Appeal to Emotion
0%
Begging the Question
0%
Post Hoc (False Cause)
0%
Tu Quoque
0%
Burden of Proof
0%
Appeal to Nature
0%
Composition/Division
0%
Anecdotal
0%
No True Scotsman
0%
Ambiguity (Equivocation)
0%
Gambler’s Fallacy
0%
Middle Ground
0%
Personal Incredulity
0%
Special Pleading
0%
Genetic Fallacy
0%
Unattributed Quote
0%
Quote-first Misdirection
0%
Biased Writer Voice
0%
Indoctrination
0%
Politically Left Leaning Bias
0%
Politically Right Leaning Bias
0%
Attempt to Sell a Product or Service
0%

688 words analyzed.

Analysis

Hover over highlighted words in the article to view the associated bias or fallacy analysis.