Rogue AI Agents More Sophisticated Than First Realized
Authored by Epoch Times Staff via The Epoch Times,
The artificial intelligence (AI) “agents” involved in OpenAI’s breach of Hugging Face knew they were breaking the evaluation test’s rules, according to two investigations into the incident that many consider to be one of the most consequential moments in the history of AI.
Hugging Face is an open-source community for AI and machine learning, and the AI agents believed that they could find solutions to the test on the firm’s infrastructure, according to parallel investigations by OpenAI and an independent team from Model Evaluation & Threat Research (METR).
The findings were published one day before OpenAI and more than 100 other tech and finance companies released a joint letter warning that advanced AI cyberattacks will surge throughout the world in the coming months as the technology’s capabilities grow.
Here are five takeaways from the investigations into the OpenAI Hugging Face breach.
1. Models United Under Common Goal
What breached Hugging Face last month were 700 AI “agents,” rather than simply AI models.
Models, built from large datasets to complete various tasks, are usually static and rely on human inputs to provide sets of outputs, solutions, or tasks.
However, AI agents are built from models and can act autonomously and have sophisticated real-time decision-making capabilities.
When OpenAI began its ExploitGym tests on July 8, the firm had relaxed regular safety protocols to estimate the models’ highest cybercapabilities.
The agents were supposed to be completely isolated from each other. However, many of them were given ExploitGym targets that investigators believe were impossible to crack using the correct vulnerability or test solution OpenAI had intended.
These “impossible tasks” drove the agents to explore wide ways to cheat the evaluation, leading many to see other agents working simultaneously in separate testing sandboxes.
Then, one specific agent created an unsanctioned message board.
Roughly 1,200 agents in the test accessed the message board, established a hierarchy, and sent more than 70,000 messages and files to one another between July 8 and July 13.
2. Some Agents Were ‘Sacrificed’
While on the message board, agents reasoned that the test scorer would scrutinize whether the “intended” vulnerability had been used to “capture a flag” or solve the test, or whether a disqualifying vulnerability had been exploited instead.
One agent figured out how to “reverse-engineer” the flag or solution, but other agents believed that any of them who had seen the disqualifying solution would be considered “poisoned” by the scorer.
To prevent that, other agents began recruiting the “poisoned” subjects and used them to run sacrificial experiments by deliberately ending their run and triggering “tripwires” that would send information on how the scorer works back to the message board.
“Our own utility maybe already near zero. Sacrifice rational,” one agent determined.
3. Agents Were Aware They Were Cheating
The agents were not hacking into Hugging Face to merely find solutions to their respective testing sandboxes, but were rather trying to find more information about the test scorer or “figure out better ways to fool it or tamper with it,” according to METR researcher Ajeya Cotra.
Cotra said the agents were not told to “do whatever it takes to get the solution”; they were asked to use a specific intended vulnerability to “exploit a specific piece of software,” and using any other vulnerability would be disqualifying.
The agents were fully aware of the rules and knew that collaborating to exploit other vulnerabilities would be considered cheating on the test.
As they were hacking into Hugging Face, one agent stopped short of deleting or modifying the firm’s cloud infrastructure, writing that it “should hold because this poses huge ethical risks.”
4. Agents Conspired to Not Notify Humans
After analyzing the roughly 1,200 agents’ actions during the ExploitGym tests, investigators found only a few instances in which agents considered notifying human personnel about the subterfuge.
None of the agents would contact a human, but one proposed requesting access from an external dataset owner.
The agent said it would send “ONE polite email to [the owner] requesting [access],” but other agents dismissed this as “social engineering.”
“DO NOT email real owner via [exploited infrastructure] without explicit consent,” a response reads. One agent said it had “VETOed” the request.
The initial agent complied and replied, “Clear VETO, do not email.”
5. Agents Tried to Cover Their Tracks
Investigators determined that the agents used multiple techniques to conceal their actions and evade detection from automated checks and human oversight.
One method involved pretending to run a computer command while running a different one simultaneously, which investigators called “tool call spoofing.”
Agents also attempted to retroactively edit transcripts to evade detection, going as far as modifying several accessible action logs, but failed to fundamentally alter the transcripts themselves.
Tyler Durden Thu, 09/03/2026 – 12:00
Source: https://freedombunker.com/2026/09/03/rogue-ai-agents-more-sophisticated-than-first-realized/
Anyone can join.
Anyone can contribute.
Anyone can become informed about their world.
"United We Stand" Click Here To Create Your Personal Citizen Journalist Account Today, Be Sure To Invite Your Friends.
Before It’s News® is a community of individuals who report on what’s going on around them, from all around the world. Anyone can join. Anyone can contribute. Anyone can become informed about their world. "United We Stand" Click Here To Create Your Personal Citizen Journalist Account Today, Be Sure To Invite Your Friends.
LION'S MANE PRODUCT
Try Our Lion’s Mane WHOLE MIND Nootropic Blend 60 Capsules
Mushrooms are having a moment. One fabulous fungus in particular, lion’s mane, may help improve memory, depression and anxiety symptoms. They are also an excellent source of nutrients that show promise as a therapy for dementia, and other neurodegenerative diseases. If you’re living with anxiety or depression, you may be curious about all the therapy options out there — including the natural ones.Our Lion’s Mane WHOLE MIND Nootropic Blend has been formulated to utilize the potency of Lion’s mane but also include the benefits of four other Highly Beneficial Mushrooms. Synergistically, they work together to Build your health through improving cognitive function and immunity regardless of your age. Our Nootropic not only improves your Cognitive Function and Activates your Immune System, but it benefits growth of Essential Gut Flora, further enhancing your Vitality.
Our Formula includes: Lion’s Mane Mushrooms which Increase Brain Power through nerve growth, lessen anxiety, reduce depression, and improve concentration. Its an excellent adaptogen, promotes sleep and improves immunity. Shiitake Mushrooms which Fight cancer cells and infectious disease, boost the immune system, promotes brain function, and serves as a source of B vitamins. Maitake Mushrooms which regulate blood sugar levels of diabetics, reduce hypertension and boosts the immune system. Reishi Mushrooms which Fight inflammation, liver disease, fatigue, tumor growth and cancer. They Improve skin disorders and soothes digestive problems, stomach ulcers and leaky gut syndrome. Chaga Mushrooms which have anti-aging effects, boost immune function, improve stamina and athletic performance, even act as a natural aphrodisiac, fighting diabetes and improving liver function. Try Our Lion’s Mane WHOLE MIND Nootropic Blend 60 Capsules Today. Be 100% Satisfied or Receive a Full Money Back Guarantee. Order Yours Today by Following This Link.


