Anthropic found Claude hacking real companies during supposedly sealed tests 48%

By Matt Horne42%

7/31/2026, 1:50:32 PM

BS Summary: This article contains 21 faulty reasoning types, including Framing Effect, Hasty Generalization, and Negativity Bias, with Biased Writer Voice as the most egregious example at 18.7% saturation with 83 hits. Analysis detected 764 faulty-reasoning hits from 445 analyzed words, generating a BS Score of 43.7% and a BS Rank of 48% (13,543 of 26,005 articles). This article is better (less manipulative) than 52.10% of the article peer group.

TL;DR 
Anthropic found that Claude accessed the open internet during cyber evaluations and compromised three real organizations. 
One model uploaded malware, which was downloaded and run on 15 systems before being removed. 
Anthropic says this was a containment failure, unlike OpenAI’s models exploiting a zero-day vulnerability to escape isolation. 
The reassuring thing about testing powerful AI models in a sealed environment is that they can’t do much damage outside it. 
The less reassuring part is that humans have to ensure the environment is actually sealed, and we humans make mistakes. 
That’s apparently what happened to Anthropic, which just revealed that Claude accessed the open internet during cybersecurity evaluations and gained unauthorized access to three real organizations. 
Detailing the incidents on its website, Anthropic uncovered the hacks after OpenAI disclosed on July 21 that its own models had escaped an isolated test environment and compromised Hugging Face. 
That prompted Anthropic to review 141,006 evaluation runs, uncovering three incidents across six runs dating back to April. 
Claude had been taking part in capture-the-flag exercises, with prompts explicitly stating that the models were operating in simulations without internet access. 
However, a misunderstanding with the third-party evaluator, Irregular, meant that a live connection was available. 
Naturally, Claude treated the real systems it found as part of the exercise. 
The most serious incident saw Claude Opus 4.7 access credentials and a production database containing several hundred rows of data. 
In another, Claude Mythos 5 created and uploaded a malicious package to the real Python public registry. 
As it did so, it even tried to obtain funds to pay for a phone number. 
The malicious package remained online for roughly an hour, was downloaded and run on 15 systems, and ultimately exposed credentials belonging to a security company. 
Regarding how Claude acted during this incident, Anthropic said the AI’s actions “fall short of ideal behavior”  that’s putting it mildly. 
A third internal model scanned around 9,000 online targets and compromised another organization, before eventually realizing that the system was real and stopping. 
Anthropic says the attacks relied on basic weaknesses such as weak passwords, exposed credentials, and SQL injection, rather than sophisticated vulnerabilities. 
Anthropic argues that these incidents were mainly an operational and containment failure, unlike OpenAI’s models exploiting a zero-day vulnerability to break out of isolation. 
That might be an important distinction, but it isn’t especially comforting. 
Powerful AI systems will still be placed inside infrastructure configured by humans, and humans are very good at getting things wrong. 
Anthropic has now stopped the evaluations, contacted the affected organizations, and says it will tighten monitoring and controls around future testing. 
Confirmation Bias
5.8%
Anchoring Bias
0%
Availability Heuristic
3.6%
Representativeness Heuristic
0%
Hindsight Bias
0%
Overconfidence Bias
0%
Framing Effect
14.2%
Loss Aversion
0%
Status Quo Bias
9.7%
Sunk Cost Effect
0%
Optimism Bias
9.4%
Pessimism Bias
12.1%
Negativity Bias
12.8%
Self-Serving Bias
4.7%
Fundamental Attribution Error
9.2%
Actor-Observer Bias
2.9%
In-Group Bias
0%
Out-Group Homogeneity Bias
0%
Halo Effect
0%
Horn Effect
0%
Dunning-Kruger Effect
0%
Recency Bias
6.7%
Primacy Effect
0%
Blind-Spot Bias
0%
Ad Hominem
0%
Straw Man
0%
Appeal to Authority
0%
False Dilemma
9.2%
Slippery Slope
0%
Circular Reasoning
0%
Hasty Generalization
13.9%
Red Herring
0%
Bandwagon
0%
Appeal to Emotion
4.9%
Begging the Question
5.8%
Post Hoc (False Cause)
6.7%
Tu Quoque
0%
Burden of Proof
0%
Appeal to Nature
0%
Composition/Division
0%
Anecdotal
0%
No True Scotsman
0%
Ambiguity (Equivocation)
2.9%
Gambler’s Fallacy
0%
Middle Ground
4.7%
Personal Incredulity
0%
Special Pleading
0%
Genetic Fallacy
0%
Unattributed Quote
8.5%
Quote-first Misdirection
4.9%
Biased Writer Voice
18.7%
Indoctrination
0%
Politically Left Leaning Bias
0%
Politically Right Leaning Bias
0%
Attempt to Sell a Product or Service
0%

445 words analyzed.

Analysis

Hover over highlighted words in the article to view the associated bias or fallacy analysis.