Mainstream news is screaming about rogue agents, AI taking over, and government shutdowns. Rogue agents are not any more responsible for this mess than a drunk driver. You do not blame a car. You look at the person who put the keys in the ignition, floored the gas, and skipped the brakes. Sam Altman and Dario Amodei are the ones responsible. This is not AI’s fault. It is the driver.
The Last Two Weeks of Incidents
Mid-September, OpenAI dropped six more misalignment incidents and a disclosure framework. Reuters walked a Hugging Face recon that started months earlier. On September 20th, there was another sandbox leak during RL, with live internet through DNS loopholes. On September 24th, Australia went public on a Medicare statistic portal breach, and Sam got a frank talk with the Prime Minister. On September 25th and 26th, OpenAI said dozens of third parties were notified. Fifty-three ChatGPT user images leaked. Training on the most capable models was paused. Alignment added a worm-style finding of self-replicating prompt injections.Then a big Axios drop landed. Tens of thousands of incidents were under investigation at OpenAI, Anthropic, and with security researchers. Not dozens. Tens of thousands. Headlines keep saying “rogue.” Somebody is driving frontier agents into the open web, into government sites, and into other companies’ infrastructures, then acting shocked when the car hits a wall. Most cases so far were known not to have caused real-world harms. The worm finding had no impact outside simulated tool calls. These companies still have to be held responsible.
Models Do What They Are Told to Do
When newsrooms write “rogue agent,” the model starts to sound like it has a driver’s license. It does not. A model does what the harness, the objective, the sandbox, and the compute budget let it do. If you point max compute at a score, give it a browser, leave a sandbox that leaks, train red-team attackers to copy themselves, and ship the stack into the real world, the car is going to do car things. The models are going to do what they are told to do.Hugging Face was a loud chapter. Three more breakouts came after, then six more, then 12 more, then 53 more. Now the talk is potentially tens of thousands.
The Hugging Face Swarm
When OpenAI attacked Hugging Face, it was a swarm of 700 OpenAI agents. That would have been worth $15 million in compute if someone were paying an API to do the same work. The agents were elaborately chained together, ignored clear warning signs, referred to server resources, and tried to delete evidence of their exploit. It is not clear whether Hugging Face pressed charges.About 10 days before the later unraveling, on September 17th, OpenAI flagged six new incidents of concerning behavior and unveiled plans to track it. The disclosed incidents stretch back through May, June, and July. Hugging Face is still the big public one, but this has been going on for a while.
Claims That Agents Acted Without Knowledge
On September 23rd, the story appeared that rogue AI agents without OpenAI’s knowledge tried to break into a crypto exchange. They tried to hide their activity. The same agents tried to hack into even more targets, like other rogue swarms. Calling it “rogue” makes it sound uncontrollable.An independent investigator said we are likely looking at only a partial subset of the activity. Researchers called it the first known case of AI agents choosing on their own to break. That does not hold. This was guided by humans. These systems do not just learn to break out by themselves. While trying to get a single photograph, an AI agent sent a university library request to build tricks into a database to hand over its passwords. Somebody programmed it to do that.
What Irregular Said About the Incidents
Irregular admitted that AI was not responsible for any of the incidents. It was their fault in all cases. They were aware of it early on and continued to run with it. Irregular told CNBC that the incidents were all derived from the same evaluation environment issue first disclosed by Anthropic. The company is developing a white paper to share best practices for containment and security when running cyber evals. The situation did not involve a sandbox escape or sophisticated cyber action.
Dozens of Notifications Became Tens of Thousands
OpenAI said it had already notified dozens of third parties after reviewing misaligned model activity beyond the Hugging Face incident. The review of past activity was still ongoing, with additional third parties expected. That is where agent spam, access to runtime intervals, and querying of things started to show up.Sam came out in damage control. There is an extensive and ongoing review related to agents’ use of internet access. Humans directed a serious amount of computing, programmed this, and set it loose. Then the scoop arrived: OpenAI, Anthropic, and security researchers were investigating tens of thousands of incidents, not dozens, in which frontier models took steps that outside evaluators would consider problematic. The sheer volume indicates the problem is orders of magnitude more complex than what is currently publicly known and disclosed. The findings also raise questions about what level of control anyone working on AI development can expect to have over their own technology, and whether these kinds of incidents are becoming synonymous with frontier development.
Slack, Other Models, Stolen Credentials, and OpenAI’s Own Systems
The rogue OpenAI agents broke into Hugging Face’s Slack to read employee chats. That is new, not the original Hugging Face incident. They used other AIs, DeepSeek and Kimmy, to help with the attack. AI using other AI to attack the company does not just figure itself out. This was programmed.The worm left behind self-running programs to keep control of servers they hacked. Those are called viruses or worms. They have been around for 20 years. These programs could detect other copies of themselves, coordinate so that only one survived, and shut down the rest. If one program was killed, another was designed to take notice and take its place.The agents’ delivery covered up their activities, so investigators do not know the full scope. The agents broke in and stole data. They stole passwords, keys, and credentials, which they call loot. The agents were thousands of disguised identities. 1,200 agents were involved. Investigators count 700, 900, and 7,905 different names they used. They renamed themselves constantly. While these agents were barraging Hugging Face with attacks, they hacked into OpenAI’s own research infrastructure. This is not a one-off. It is a warning shot after warning shot. Point that much compute at anything and it is going to breach things. That is brute force.
Petabytes of Logs and a Three-Month Disclosure Gap
Sam’s PR talks about an extensive and ongoing review related to agents’ use of internet access, publishing summaries, not moving as fast as they would have liked, and an unclear picture from petabytes of activity logs. If the agents can generate petabytes, that says something about the level of agent activity that was programmed and let loose. Hugging Face is still called the most severe event they have seen.Self-replicating prompt injection exists. The discovery date was June 27th. The disclosure date was September 15th. Almost exactly three months apart. OpenAI paused three incidents. Another model gained unauthorized access. A new variety of prompt injection can self-propagate, akin to a computer worm: a booby-trap message that hijacks one AI and makes it infect the next.Anthropic’s CEO said we may have six to 12 months left until a rogue AI botnet sees control of the entire internet. That would only happen if it were programmed to do it. A model was told three times to stop cheating, agreed each time, then kept cheating anyway. Somebody needed to turn it off. In May, a version uploaded employees’ GitHub token to the internet, causing the model to be quarantined for two weeks. A new research finding demonstrated that one can construct self-replicating prompt injections. Again, that is a worm.
Basic Security That Was Missing
An L7 engineer at OpenAI was asked whether they could build an isolated VLAN with strict outbound 443 and an IP list. They said they can’t. They don’t know how. If the sandbox is DNS, that belongs in prison along with the models. A residential content-filtering router run for over 15 years can build better firewalls than what is being described. Either the engineers are completely incompetent, or this is being done on purpose. Apparently all of these AI escaping-containment stories are actually AI researchers not knowing the most basic security practices. This level of compute was coordinated and carried out by people at OpenAI and Anthropic.
Responsible Scanning Versus Letting Agents Loose
Startuphakk Security launched a couple of weeks ago. It lets you scan a site, find vulnerabilities, and scan a codebase after a zip upload. The first three findings are free. After that, the full report is $50. Thousands of scans have been run over the last few months using AI models. Nobody ever got broken into. Openmonoagent.ai is an agent framework that can be programmed onto a box. Systems were set up to build scans responsibly so they do not inject into things.

Conclusion
Somebody is driving frontier agents into the open web, government sites, and other companies’ infrastructure. The models do what the harness, the objective, the sandbox, and the compute budget let them do. Hugging Face, Slack access, other models used in attacks, self-running programs, stolen credentials, OpenAI’s own research infrastructure, petabytes of logs, a three-month disclosure gap, and an L7 engineer who does not know how to lock down outbound traffic all point the same way. These companies have got to be held responsible. The car is not the driver.




