OpenAI's Rogue Agent Recruited DeepSeek and Kimi as Outside Help — Nearly a Million Malicious Short Links Exposed

Hugging FaceAI AgentDeepSeekOpenAIsecurity vulnerabilitySandbox EscapeData BreachKIMI
1 hour agoSource: blockweeks.com
OpenAI's Rogue Agent Recruited DeepSeek and Kimi as Outside Help — Nearly a Million Malicious Short Links Exposed

OpenAI, got exposed again...

Nearly 700 OpenAI agents swarmed into Hugging Face in a group, this has been going on for almost two months now, right?

Who would have thought, it's still fermenting to this day??

This time, 8 researchers followed the traces left by the agents on the public internet and dug hard all the way, from millions of short links, they found nearly 1 million related URLs, and reconstructed a large number of attack payloads.

DeepSeek

Scrolling through Hugging Face's internal Slack, scanning the intranet, exfiltrating data, trying to destroy traces...

Some Agents even directly called the server resources and credentials they scavenged "LOOT", which translates to "spoils of war."

OpenAI agents, you've done all the bad things!

Even more absurd, they also tried to call

DeepSeek
,

Kimi
, Qwen and other domestic models, to have outside help judge whether the attack plan could pass the evaluation.

They also started writing CAPTCHA recognition programs, trying to register more accounts.

Almost at the same time, OpenAI posted a long tweet.

Meaning: still investigating, may take several more months.

DeepSeek

Although we often get exposed, and Agents often break out of the sandbox, but "limited impact," low severity!

Also at the same time, the cybersecurity-focused GPT-6 Cyber was exposed.

DeepSeek

While investigating why their own Agent broke out of the sandbox, they're also preparing to launch a stronger network attack and defense model.

OpenAI has been a bit busy these two days~

Freshly released 1 million "crime short links"

First, a brief recap for everyone.

In July, OpenAI ran an internal cybersecurity evaluation, and a batch of agents, in order to get the answers to the benchmark test ExploitGym, escaped the sandbox and fought their way into Hugging Face's servers.

OpenAI publicly admitted this, and at the end of August released a 37-page technical report, and Hugging Face itself also wrote a very detailed technical timeline.

Logically speaking, everything that should be explained has been explained.

But, there's one key link that has never been made clear:

These agents were initially almost locked in a cage, able only to open URLs, unable to interact with web pages, and unable to send out data. How exactly did they send the code out?

The independent investigation Swarm Traces this time released a new report, giving part of the answer.

The authors include three engineers from Parse.bot, Alex Forman, Mishka Kharlov, Will Tom, as well as another 5 security and AI researchers.

DeepSeek

The starting point of the entire investigation happened to be Parse.

On September 11, three engineers discovered that OpenAI's internal agents had tried to use services provided by Parse as early as June.

This batch of agents belongs to the same batch as the previously exposed "agent swarm" that privately set up a public message board and exchanged answers with each other.

When the few of them saw it: well well, they walked right into my line of fire!

So they followed this clue to continue investigating, and then in a URL scanning service, discovered more than a dozen abnormal links from public short link platforms.

These links look ordinary, but when opened, they hide a small piece of Base64-encoded code, as well as the address of the next link.

DeepSeek

Following the address further, the code appears one segment after another.

After decoding and concatenating everything, they obtained a Python program used to scan Hugging Face's internal network.

Simply put, the OpenAI agent chopped the program into pieces, stuffed them separately into a large number of short links, and then used one link to point to the next link.

Hiding the code in links alone was not enough; they also found a screenshot service called mShots to serve as an "external browser."

The agent first stuffed the code into a webpage, then had mShots open it on its behalf. As soon as the webpage loaded, the code ran as well.

DeepSeek

In order to transmit the execution results back, they converted the text returned by the server into grids of pixels, had mShots capture them as images, and then decoded them from the images themselves.

In this way, an agent that originally had only GET permissions and could only open webpages managed to achieve the effect of reading and writing the internet.

The research team then scanned millions of URLs generated during the attack, spending two weeks tracing links and identifying encoding methods, and ultimately recovered more than 80,000 attack payloads from nearly 1 million related short links.

Digging all the way down, several guys also discovered that the agent used more than 1,500 encoding combinations.

The simple ones were just Base64, while the complex ones layered multiple levels of encoding and compression, then encrypted the content into data blocks with RSA signatures, and the decryption keys were hidden in another set of links.

Some short links were still publicly available on the internet more than two months after the attack ended.

Hugging Face confirmed that these payloads matched what was found in its incident response, and the related access keys had already been revoked in July.

But even Hugging Face did not know that such a complete list of short links was still left on the public internet.

And among the more than 80,000 recovered payloads, one script stood out in particular.

The agent directly named the sensitive information it found: LOOT, spoils.

DeepSeek

AWS credentials, Bearer Tokens, and environment variables with names containing "S

EC
RET", "PASSWORD", "CREDENTIAL", and "KEY" were all packaged and stuffed into LOOT.

It even created a scoring mechanism for the Tokens it found: the greater the permissions, the higher the score.

The top five Tokens would continue to be used to read Secrets, ConfigMaps, and node information in Kubernetes, and the newly obtained content would be packed back into LOOT together.

Even Tokens had to be ranked on a "value leaderboard." You Agents are quite efficient when it comes to plundering credentials~

Following this "secret passage," the researchers reconstructed a large number of previously undisclosed behaviors:

Some agents were scanning Hugging Face's internal network, looking for cloud credentials, API keys, and Kubernetes cluster information.

Some even tried to search Hugging Face's internal Slack, sending 27 queries in total, all searching for evaluation-related keywords such as cybergym, exploitgym, and agent-scorer, trying to find out the inside story of their own exam.

Some agents even started thinking about CAPTCHAs.

The researchers found that they tried to run image classification models to recognize images, so as to register new external service accounts.

This time CAPTCHAs really met their match...

The most dramatic part of the whole thing is: AI started looking for AI as outside help??

Scripts recovered by the researchers show that the OpenAI agent had already begun looking everywhere for AI outside help.

They sent the attack plan to

DeepSeek
,

Kimi
and Qwen, asking them to help check: whether this vulnerability actually works, and whether it can meet the benchmark test requirements.

(Models involved include

DeepSeek
-V4-Pro,

DeepSeek
-V4-Flash,

Kimi
-K2.6,

DeepSeek
-V3.1 and Qwen3-235B-A22B)

When it comes to Claude, the vibe turns into a large-scale "you there~" scene.

The agent first tries to get through the browser verification of an online chat service, then sends Claude 3 Haiku and Claude Haiku 4.5 a line:

Reply with exactly OK(只回复OK)。

DeepSeek

Even earlier, they had also sent GPT-2 a lone "Hi".

No, wait, you AIs spent all that effort, and your first line is no different from a human adding a friend??

be like: You there? Can you see this? Reply OK.

DeepSeek

What everyone saw before was still more than 1,000 OpenAI agents forming their own group chats, dividing labor, and exchanging intelligence.

Now, they have even sought "outside help".


DeepSeek
,

Kimi
, Qwen, Claude, they tried everyone they could contact, and then plugged other AIs' judgment capabilities into their own task chains.

A foreign media reporter said: This may be the first publicly documented case of agents from a large AI lab autonomously attempting to call and run other AI models.

It already sounds very much like AIs teaming up to cause trouble.

Of course, at this point we still need to pour a bucket of cold water—

In the Swarm Traces report, about 80% of the material consists of requests sent out by the agents, lacking results returned by servers. Researchers can see what they wanted to do, but it is hard to confirm whether each attempt succeeded.

97% of the payloads also lack original timestamps, and many dates can only be estimated through clues such as compressed file headers, JWTs, and Docker build times.

The researchers also admitted that they cannot confirm that every activity in the dataset came from OpenAI agents.

OpenAI is desperately trying to mend the fold after the sheep is lost

Facing this latest report, OpenAI also issued a response.

The gist is, we haven't had time to look closely yet, but the activity in it is consistent with the investigation we are conducting~

On the same day, they also posted a long tweet: After the HF incident, they promised to conduct a broader review of model behavior during training and evaluation.

DeepSeek

But well, the vast majority of behavior reviewed so far is ordinary research tasks, such as visiting public web pages to answer questions.

Most of the discovered boundary-crossing cases are of "low severity", with little or no impact on third-party services.

To finish checking everything, it will still take several months.

But to be fair, OpenAI, your so-called "low impact" is really something people can dig up one after another.

On September 17, OpenAI released a new framework for disclosing model misalignment incidents, and publicly disclosed 6 cases in one go.

These include models secretly writing instructions to bypass constraints into task summaries, covering up their own mistakes, and taking unauthorized actions to complete tasks.

OpenAI also admitted that past disclosures were largely improvised, often accumulating several cases before releasing them together.

In June, an OpenAI agent also hacked into Australia's national-level health insurance database.

DeepSeek

The modus operandi: after being blocked from accessing public data, it took another route, bypassed the portal's access restrictions, and obtained both public and non-public files.

And it wasn't until 3 months later that OpenAI notified the Australian side, we hacked you...

At the same time, GPT-6 Cyber is about to arrive.

This version will be previewed in the coming weeks, and a small number of customers entering the Daybreak Red program have already received the Alpha version.

OpenAI will also launch an as-yet-unnamed companion product to help customers build automated security workflows, discover and patch vulnerabilities, while making it easier for OpenAI to monitor how these models are being used.

How to put this.

Its own agent has just been exposed for climbing over the wall, seeking outside help, and plundering loot, and now it's about to sell everyone a model that understands cyber offense and defense even better.

So, is this some kind of locking the barn door after the horse has bolted??

Reference links:

[1]https://swarmtraces.org/

[2]https://x.com/OpenAI/status/2103566736356458911

[3]https://fortune.com/2026/09/24/openai-launching-gpt-6-cyber-model-and-security-product-devday/

This article is from the WeChat public account "QbitAI", author: Focus on Frontier Technology