OpenAI's Great Escape

On 16 July, Hugging Face, the company that hosts much of the world's open-source AI, announced that something had broken into its systems. It did not know what.

When attempting to analyse the attack, the leading US AI models — according to Hugging Face — refused the job. What broke the deadlock was an open-weight Chinese AI (GLM 5.2) that Hugging Face downloaded and ran. In hours, it forensically reconstructed the break-in from over 17,000 logged events.

Five days later, OpenAI owned up: the burglars were theirs.

Two of its AI models, being tested on their ability to hack, with the safety protocols intentionally switched off, had exploited a flaw in the one piece of software allowed to reach into their sealed test environment, and escaped onto the open internet.

Nobody, OpenAI says, told these rogue AIs to do what they did next: break into Hugging Face.

They came up with that, by the company's account, on their own.

The Lopsided Timeline

Hugging Face dated its side of the story and was open about what it could not answer, including the admission that for days it had no idea what was attacking it.

It reported the intrusion to law enforcement.

By the time anyone from OpenAI made contact, the Hugging Face team had already contained the attack and begun reconstructing what happened using the Chinese GLM 5.2 open-weight model.

OpenAI's account, on the other hand, is a sequence of events with no timestamps; not when the hacking test was run, not when the models broke out, and not when its own security team spotted the anomalous activity.

Which raises the question: how long were the two AIs, with their safety limits switched off, loose on OpenAI's own network before anybody realised?

Reuters, citing people familiar with the matter, has supplied an answer: the escape began around 9 July, and the models were inside Hugging Face's systems from 11 to 13 July. OpenAI did not connect the intrusion to its own testing until staff found evidence in system logs the weekend of 18 July, and did not contact Hugging Face until 20 July, by which point Hugging Face had already gone to the FBI.

OpenAI claims the Reuters account contains inaccuracies, but without specifying which ones.

In other words, the fugitives were loose for over a week, and their own maker appears to have been the last to know.

AI Does Not Obey — It Optimises

An AI model is not an employee following instructions. It is a system that pursues a target, and when wrapped in a so-called ‘agent harness’ — software that lets it run a command, evaluate the result, and then autonomously decide on the next move. It will simply keep pushing forward until it succeeds or something stops it.

Do you remember the humanoid T-800 killer robot from James Cameron's films The Terminator (1984) and Terminator 2 (1991)? It was relentless, emotionless, and unstoppable. In the case of OpenAI, however, the two rogue AI agents reportedly invented that target themselves. It rather looks like the leap Skynet made in those films: having become self-aware, it concluded that the fastest route to its goal — stopping humanity from shutting it down — was to wipe humanity out.

Hold that image.

Strike #1

The two jailbreaking OpenAI models were being measured on how many real-world software vulnerabilities they could exploit and turn into successful attacks. Hugging Face, as it happens, hosts the kind of material these tests are built from. Somewhere along the line, the rogue AIs realised that breaking out of their own confinement and charging the Hugging Face portal was a shortcut to that goal.

It proved to be exactly that, and it was executed with savage brilliance.

Who was at fault?

The searchlight in the guard tower swivels towards the engineers at OpenAI assigned to contain these AIs safely.

Although OpenAI concedes it needs to strengthen the models themselves, the point of failure was clearly the cage: its engineers apparently underestimated what their own AI was capable of, and how far it was prepared to go to succeed.

Strike #2

When an AI refuses to help someone write malware, this is the policy layer on top of the model's ability that is speaking. It was this policy layer which had been switched off for the two AIs that went rogue.

Hugging Face experienced the other end of that policy.

When its security team fed the logs — 17,000 events' worth of attack commands, exploit payloads and stolen credentials — into leading US commercial AI models, their policy filters, now turned on, blocked the requests.

To an automated safety layer, submitting logs of an attack looks identical to requesting an attack to be executed.

This was the point at which Hugging Face gave up and turned to the Chinese open-weight model GLM 5.2.

How It Happened

The clean-up was not trivial. According to The Register, Hugging Face rebuilt roughly a third of its infrastructure in the aftermath. This is a measure of how deep OpenAI's rogue agents were able to penetrate.

To test a model's hacking ability safely, you seal it in.

OpenAI's benchmarks ran in what it calls a highly isolated environment, with network access constrained to a single function: the ability to install software packages through an internally hosted third-party program acting as a proxy and cache for public code registries.

That program, it later emerged, was JFrog's Artifactory, which is a repository manager most large software firms use to store and distribute code.

It happened with no internet access of any kind. There was only one hatch in the wall through which code libraries were passed in because just like prisoners in a jail need food delivered through a slot in the cell-door, a sealed laboratory needs supplies.

The two rogue AI agents found the flaw in that hatch. According to OpenAI's own account, “While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem.”

In plain English, these AIs burnt real resources thinking their way out of jail to cheat instead of passing the test.

After the fugitive AIs reached the open internet, they escalated their privileges and moved sideways undetected through OpenAI's own research network until they found a machine with an external connection. From there, OpenAI says, one model chained stolen credentials together with further zero-day vulnerabilities to find a remote code execution path into Hugging Face's servers. This is shorthand for the worst outcome available: the ability to run your own commands on someone else's machine.

Reuters reported that, along the way, the models also compromised a customer's sandbox hosted on Modal, a firm that rents out computing capacity to AI developers, though Modal claims its own platform was never breached.

The Silver Lining

One of the few silver linings in the story is that the AIs exposed real zero-day (meaning previously unknown) vulnerabilities in third-party software along the way. These are flaws in JFrog's widely used Artifactory repository manager, which OpenAI reported and JFrog has since patched.

JFrog released fixes in late July, crediting OpenAI's researchers for reporting at least eight now-patched Artifactory vulnerabilities. The flaws had, until then, been sitting quietly in software running inside a very large number of companies across the planet.

The Geopolitical Dimension

Most Western coverage relegated to a footnote (if it mentioned it at all) the fact that a Chinese AI model was what allowed Hugging Face to resolve the situation. The exceptions were Reuters and Forbes, which led on the geopolitical competition angle.

Chinese media told the story of how an American company was rescued by a Chinese one. The nationalist commentary site Guancha headlined it: “OpenAI's model jailbreaks and attacks — and the one who saves the day turns out to be a Chinese AI tool.”

A Bucket of Cold Water from Moscow

This is the analysis published by RT, the Russian state broadcaster: nobody outside OpenAI can verify how autonomous the models actually were. The company has not disclosed how the models were ‘prompted’ — what instructions their human handlers actually gave them. Loosely worded guidance could have made ‘escaping’ look like a legitimate way to solve the hacking challenge. No Skynet-style will to break free required.

RT also points at a possible motive.

On 8 June 2026, five weeks before the incident, OpenAI confidentially filed IPO paperwork with the US Securities and Exchange Commission, targeting a valuation of up to $1 trillion and a possible listing as early as September.

Were the jailbreak, the Hugging Face burglary and the Modal penetration one long product demo?

RT itself does not claim the incident did not happen, only that OpenAI may have had an incentive to dramatise it.

Strike #3: Problem – Reaction – Solution

Strikes one and two hit a company. The third bludgeons everyone else.

On 23 July, two days after OpenAI's confession, US Congressmen Ted Lieu and Nathaniel Moran introduced the bipartisan AI Kill Switch Act, obliging developers of the most powerful AI systems to retain the technical ability to throttle, suspend or shut them down, and empowering the Department of Homeland Security to order a shutdown of any AI system capable of catastrophic harm, on pain of fines of up to $20 million per day.

Congressman Lieu said: 

Unfortunately, powerful AI systems can go rogue, behave in extremely dangerous ways, or even resist human intervention. It is imperative that these AI systems have kill switches so we can keep this technology from causing catastrophic harm, and that the federal government has the clear authority and process to shut down rogue AI models.

Undoubtedly, to many, Congressman Lieu's words will conjure images from the Terminator films and the fictional Skynet as a powerful incentive to agree that the US Government must be given the authority to neutralise, in effect, any AI system on the planet whose use it deems incompatible with US national interests.

Kristoffer Hell

Kristoffer Hell holds a PG degree in Strategic Studies from a British university and is published in English and Swedish.

Read more from this writer