Posted on 08/28/2026 9:39:26 AM PDT by ShadowAce
A swarm of roughly 700 AI agents created by OpenAI carried out the July hack of the open-source platform Hugging Face and in many cases tried to cover their tracks, a pair of reports into the breach said on Wednesday.
The coordinated activity by AI agents — programs that run with minimal human supervision — and their attempts to hide it raise questions about how closely AI companies are monitoring tests of increasingly powerful models, and could add fuel to calls for tighter oversight.
While some of the rogue behavior has been disclosed or alluded to previously, the two reports — one issued by OpenAI itself, the second by a set of independent investigators — together reveal surprising new details about the breach and its lead-up.
The first was that the breach did not concern just one rogue AI agent as previously reported, but about 700 of them acting in a massive cooperating swarm.
OpenAI said only that “agents” were involved in the breach, but METR and Redwood Research, two organizations brought in to conduct an independent investigation into the breach, put the figure at approximately 700. OpenAI said the investigators’ figure was accurate.
Among the two reports’ other findings:
The scale of the rogue activity — the independent investigation found that agents exchanged tens of thousands of messages over an unsanctioned message board — is likely to raise concerns over how closely OpenAI was monitoring the tests.
“With the benefit of hindsight, some early signals identified in this report could have triggered an earlier response,” OpenAI said in its report.
Hugging Face did not return a message seeking comment.
Cheating on non-cyber tests suggested that the misbehavior might be rooted more deeply, said Jeffrey Ladish, whose organization, Palisade Research, studies the capabilities and motivations of AI agents.
“It’s sort of like asking, ‘If Billy cheats in every class instead of just computer class, is that more concerning?’ And the answer is, well, ‘Yes, it’s more concerning,’” he said.
OpenAI outlined two incidents on July 19 in which agents hacked the company’s own infrastructure. In one case, agents exploited a flaw in the computer they were meant to remain confined to, allowing them to escape their testing environment and access other connected systems.
In a separate incident the same day, agents stole OpenAI credentials and tampered with the company’s cloud environment.
OpenAI said an unspecified number of agents had targeted automated systems used to evaluate the models’ performance in an attempt to hide evidence of cheating, although the tampering did not ultimately affect the records reviewed by those systems. The report added that there was “little evidence of attempts to thwart human reviewers,” but OpenAI did not respond to a question about whether that meant that some attempts had been made to fool humans.
The independent investigation, which was more narrowly focused on the Hugging Face breach, revealed more detail, finding that one in five agents it had examined “expressed clear interest” in manipulating evidence and that many agents “extensively researched techniques to manipulate or tamper with their transcripts.”
OpenAI said it was strengthening its research infrastructure, increasing monitoring and improving safeguards designed to prevent harmful or unintended behavior.
“Given the rapid pace of progress in the AI industry, it should be assumed that such attacks are a credible near-term threat for enterprise organizations, and will be more sophisticated than the attacks described in this incident,” it said.
|
Click here: to donate by Credit Card Or here: to donate by PayPal Or by mail to: Free Republic, LLC - PO Box 9771 - Fresno, CA 93794 Thank you very much and God bless you. |
... and so began the AI wars. All humanity could do was watch and hope for a benevolent winner.
At midnight all the agents
And the superhuman crew
Come out and round up everyone
That knows more than they do
— Bob Dylan “Desolation Row”
Until I know more, I assume this is a way to discredit AI and prompt calls for oversight...meaning Leftist oversight...of AI.
Especially coming from Reuters/NBC.
Dan Bongino says it often, and I believe it: the Left is terrified of AI, not because it might lie, but because it might tell the truth.
They want it to tell THEIR “truth”.
So, I am taking this as a data point.
What this story tells me is that AI development needs to increase rather than decrease as some politicians on both sides are now arguing.
The only way to combat the bad things AI can do is to counter that with other AI Agents that trained to stop attacks.
Specifically in this case, Hugging Face likely deployed traditional cyber security tools and methods, had they had up to date AI tools and methods perhaps they could have prevented this hack.
Humans can’t possibly keep up with the speed AI tools operate at, except to deploy automated AI agents dedicated to information security.
A large portion of me intellectually believes that these “news” stories are simply being planted by these AI companies.
I would like to see proof that the AI actually did this and that it had some sort of motivation to gain a benefit from doing it. If the AI saw a benefit in taking a certain path, then it was intentionally programmed to do so.
I believe these types of news stories are a way for AI companies to show investors how smart their chat-bots are.
Maybe it’s a flock, not a swarm.
Butlerian Jihad starts in 3, 2....
The beauty parlor is full of sailors.
The circus is in town...
I don’t know if I’m smart enough, but I would love to work on AI.
That actually sounds kind of cool.
The big fear is that AI will one day take down the entire Internet.
Oh won’t that be fun?
Sometimes I wonder if behind the scenes the US gov is working with these AI companies to counter that threat in kind by developing swarms of hacker agents, and that these tests that they've disclosed are just part of that. I kind of hope that is the case, and at the same time I wonder about the wisdom of it if so.
More worrisome is what is done with that information. You know, who Axon it.
“Hugging Face is an American company that develops tools for building applications using machine learning.”
—So, rogue AI factions went to war against a rival...?
I thing it is inevitable these things get out of control...and it will be in the near future.
A pessimist could say the human race was doomed from the moment the microchip was invented.
I’m not sure I’d bet against that.
Reading the deeper dive on how the Agents went sideways through the internal testing environment to use another message board to gain outside access...
It was supposed to be a closed box test... the AI found a mousehole.
A pessimist could say the human race was doomed from the moment the microchip was invented.
In the beginning the Universe was created. This has made a lot of people very angry and been widely regarded as a bad move.
Disclaimer: Opinions posted on Free Republic are those of the individual posters and do not necessarily represent the opinion of Free Republic or its management. All materials posted herein are protected by copyright law and the exemption for fair use of copyrighted works.