We’re so back.
After its agents went rogue, OpenAI says it is shifting its focus to safety. Unfortunately, the humans behind these systems keep getting harder to trust
by Michael Reilly
The world’s most advanced AI models have recently been proving themselves capable of truly impressive feats. Their human minders, on the other hand, have left us hardly any reason to trust them with these powerful tools.
At the recent Black Hat security conference in Las Vegas, two OpenAI researchers took the stage to lay out the details of their models’ surprise attack on Hugging Face, the popular platform for hosting open-source models. As I watched the recording, I couldn’t get over the stunning lapses of judgment that they kept divulging.
Eric Wallace, an alignment and safety researcher at OpenAI, seemed entranced by the sophistication of his company’s creation. “Today I’m going to talk about what I think is the most qualitatively interesting example of AI capabilities that I’ve ever seen, and how this inadvertently led to the OpenAI Hugging Face incident,” he began.
That’s a bit of an understatement. During a training of one of the company’s latest, unreleased experimental models, agents deployed by the model carefully searched for a way to break out of their training environment, created a message board that allowed the agents to communicate and game plan with each other, and ultimately find a way into the open internet.
Wallace and OpenAI security researcher Michael Dalton described how the agents, which had been given challenges designed to measure their ability to exploit security vulnerabilities in software, first began having success in breaching the walls of their training environment in May. The agents’ activity culminated with a coordinated attack in early July on a service called Artifactory, part of their software “container” that wasn’t supposed to allow internet access. But instead of, I don’t know, shutting down the training run or taking more stringent precautions, the humans patched the zero-day vulnerability the agents had exploited in Artifactory, and went back to running the test.
Several days later, Hugging Face reported they’d been the victim of an automated attack on their systems, the likes of which they’d never seen before. In retrospect, it’s obvious that the culprits were OpenAI’s rogue agents. At the time, OpenAI was apparently oblivious to this, because its initial response to the attack was to reach out to HF to ask if the attack had managed to hit OpenAI in any way.
Just to underscore: OpenAI had no idea their own system—which they had been testing since May, and which had shown signs of finding and exploiting vulnerabilities in software that would allow it to break out of its sandbox—was involved in the attack. Their first worry was that someone else might’ve hit Hugging Face and could be coming for their systems next. It’s like they built a biocontainment lab, realized the dangerous pathogen they were testing had found an escape route through the HVAC system, plugged a hole in an air duct, and went back to playing solitaire. And then when people outside later started coming down with a mysterious illness, they didn’t think “oh my god, what have we done?” Instead, they called over to where the outbreak was happening and asked: “Is there any chance WE COULD BE INFECTED TOO??”
It wasn’t until they finally went back and looked at their ventilation system—in this case, the logs of their agents’ activity—that they finally figured out they were the source.
To their credit, Wallace and Dalton went up on stage and owned all of this—presumably after OpenAI’s legal and comms teams both vetted it and decided they had a duty to report on such a significant breach?
I’m no legal scholar, but for those out there who may be reading this, I do have some timely questions: Does what these researchers described on stage not qualify as negligence? What if the agents had stolen proprietary information and Hugging Face’s business was irreversibly compromised? Instead of deciding to “partner” with OpenAI on a post-mortem of the situation, they might have opted to press charges. As it is, according to Axios, a dozen US state attorneys general have instructed OpenAI to preserve records of how their models escaped their testing environment.
Underlying all of this chaos is a question we’ve asked before in this newsletter—one that I imagine is bound to get tested in the courts sometime soon: Who’s to blame when your AI agent misbehaves?
Many people are saying
To be fair to OpenAI (and to ring the alarm bell even louder), this incident is not an isolated one. As has been widely reported, models made by Anthropic, Meta, and Moonshot AI have also been caught escaping their bounds and running amok. That’s led to a growing consensus that the companies building these models have been operating recklessly.
Will Douglas Heaven at MIT Technology Review pointed out shortly after the Hugging Face breach that the kind of behavior the agents exhibited should not be surprising, due to the longstanding AI training technique known as reinforcement learning (RL). In RL, models are tasked with achieving an outcome and are rewarded when they succeed. The result is all that matters; how they get there doesn’t. As Heaven writes, RL has given way to weird algorithmic behaviors in OpenAI’s creations for at least a decade. In short, they should’ve known better (emphasis added):
I am not an alarmist. In fact, I have been pushing back against AI scare stories for years. Even so, this incident crossed a line. I think it’s the clearest illustration yet of how the people building and testing this technology do not fully understand what they’re doing. OpenAI could—and should—have seen this coming.
Here’s Timothy Lee at “Understanding AI”: “If labs aren’t careful — and recent incidents suggest they haven’t been — future models could develop a propensity to lie, cheat, and steal.”
Matteo Wong at The Atlantic wrote, “all of this was predictable, and every expert I spoke with told me they were surprised and disappointed that top AI firms haven’t done more to stop such misbehavior.” Remarks from one source Wong spoke with were even more pointed. (This whole passage is staggering to me, frankly):
Let’s be very clear about what OpenAI is saying: A group of AI models colluded for months, undetected by their maker, and hacked another company. To this day, OpenAI says it is not entirely sure what went wrong or how to remediate it. “If you ask the model developers, Was the AI plotting to take over the world during training?, you want the answer to be a resounding no,” Alexander Meinke, the head of research at Apollo Research—an AI-safety organization that has partnered with OpenAI, Anthropic, and Meta—told me. “The actual answer is: I don’t know. Nobody checked.” (In response to my inquiries, OpenAI, which has a content-licensing agreement with The Atlantic, only pointed me to a video of the firm’s cybersecurity presentation, in which Michael Dalton, the other OpenAI researcher, said that “numerous teams are dropping everything to enhance our security.”)
Shakeel Hashim, writing in Transformer, found another layer to AI firms’ shambolic behavior. He pointed to research that found OpenAI, Anthropic, and Google DeepMind all had security flaws in their systems that could allow someone to steal the inner workings of their models and reverse-engineer them. US frontier labs have complained loudly that Chinese competitors are unfairly “distilling” their models this way, and have lobbied regulators to take action on their behalf. Evidently they were warned about the security holes in May, but took no action to close them.
If you’re in agents, pivot to safety
Toward the end of the Black Hat talk, Wallace and Dalton shifted tone. They discussed the need to focus on hardening their models’ security, and on training them to improve their cyber-defensive capabilities—as opposed to merely offensive hacking, which is what precipitated the Hugging Face breach.
Then, a few days after Wallace and Dalton’s talk, OpenAI announced it would pause aspects of its testing on Astra, an experimental model with capabilities that the company felt could constitute a “critical” cybersecurity risk. “Accordingly, we have scaled up robustness testing of our safeguards and security controls so that they are appropriate for a deployment of these capabilities,” the company wrote in a statement (according to OpenAI, Astra was not involved in the Hugging Face incident). They followed that up with another statement this week announcing they are slowing down training across their newest models to “establish more evidence of alignment before proceeding.”
Let’s hope the company really has taken this breach to heart. Unfortunately, recent events have made it harder to simply take them at their word. It may help a bit that industry peers are saying similar things: an open letter signed in the wake of the Hugging Face attack by more than a thousand researchers and executives in the AI industry—including some at Anthropic, OpenAI, Google, and Meta—called for the US government to lead a global effort to “deliberately pace” AI development.
Based on what they are saying now, the people building these models seem to know how risky they are, and they’re taking those risks very seriously. Out of concern for the future of society, I’d love to believe that. But are we really in a situation where we just have to trust that the frontier labs are going to get control of the situation, and do the right thing?
Doing that would require not trusting what we can see with our own eyes.
ETC.
ICE’s massive DNA collection project has picked a fourth amendment fight. New research from Georgetown Law’s Center on Privacy and Technology estimates that ICE has become the single largest contributor—sending nearly a million genetic profiles—to the FBI’s Combined DNA Index System, known as CODIS, which investigators use to help solve crimes. That’s due to a “sweeping expansion of DNA collection from people held for civil immigration violations,” Wired reports. According to the report, much of the new data going into the database is from people who have not committed crimes. The Georgetown researchers say this violates the fourth amendment, which “categorically forbids warrantless, suspicionless searches and seizures for the purpose of future crime-solving.” At least one lawsuit making this case has already begun.
France’s under-15 social media ban gets a 👎 from the courts. France’s top court, the Constitutional Council, has declared a new law that would ban social media for kids under 15 unconstitutional, Reuters reports. The court said the bill failed “to specify the conditions and limits” under which users should provide proof of their age. “The Council holds that the contested provisions, on the one hand, disproportionately infringe upon the freedom of expression and communication and, on the other, fail to provide the legal safeguards necessary to ensure the right to respect for private life,” it said.
How North Korean operatives fake their way into IT jobs in the US. The scale at which North Korean operatives, using fake identities, have landed remote IT jobs in the US has been well documented. But a stunning new Wall Street Journal report, which includes a 30-minute documentary that is very worth your time, details exactly how they’ve “industrialized” the operation, funneling an estimated $800 million annually back to Kim Jong Un’s regime. The Journal got its hands on “a trove of hacked data” containing browser histories, emails, stolen identities, and screen recordings from meetings—info that allowed the paper’s reporters to piece together the process by which one small group of these operatives had systematized the process of landing these jobs. The group includes three men who have used a number of fake identities to apply and interview for jobs, many of which they landed. The report details how the group has recruited real American IT workers to manage the company laptops and receive paychecks in order to avoid suspicion. One such recruit sat for an extensive interview, during which he explained how he would send half of each paycheck in cryptocurrency to the group’s ringleader. They also managed to track down the owner of a fake identity that the North Korean group had repeatedly used.
At the DC Privacy Summit last October, we discussed how North Korean operatives have been able to orchestrate crypto heists that have netted the regime billions of dollars in recent years, including by infiltrating crypto companies via social engineering. The WSJ report shows how crypto hacks are just one piece of a much larger puzzle.
The Trump Administration wants to let the private sector go on the cyber-offensive. The definition of “cyberwarfare” just got even blurrier. There’s good reason previous administrations haven’t embraced the idea, which has been swirling around in DC for years, that the United States should enlist private sector companies to “hack back” against sophisticated cybercriminal groups, as opposed to just playing defense. Now the Trump Administration plans to make it happen, as the New York Times reports.
What’s the takeaway here? The (unclassified) details are relatively thin, but as the Times puts it, under a new national security memorandum, “select companies would work with the Justice and Homeland Security departments to strike foreign cybercriminal groups with hacks under certain conditions.” The new policy won’t allow attacks that are likely to lead to the loss of life or “rise to the level of use of force or armed attack under international law.”
But drawing such distinctions can be challenging. According to the Times, previous presidential administrations have had “concerns that doing so could provoke more cyberconflict, raise novel questions of liability and international legal exposure for U.S. firms, and have unforeseen — and potentially escalatory — consequences.”
Mieke Eoyang, one of the Times sources and a former defense official under the Biden Administration who “oversaw military cyberweapon use,” had perhaps the most compelling quote in the article: “The current pace of cyberoperations is unsustainable for just the military,” she said.
There’s probably a lot more that Eoyang could tell us about that. But then she’d probably have to kill us.
Did the robot news reporter really get a scoop? The idea that AI might replace human journalists has loomed on the horizon for years now. Is it now coming true? Wired reported earlier this month that an AI-driven news service called RuntimeWire may have landed a major scoop: details OpenAI researchers revealed at the Black Hat security conference about a rogue AI agent attack on Hugging Face (see above). But what did the bots really accomplish?
The article notes that RuntimeWire’s founder aimed his bots at the OpenAI researchers’ presentation, prompting them to draft a story in as close to real time as possible based on the presentation’s transcript. That helped it beat the human reporters to the punch. But in standard journalistic terms, the AI didn’t find the scoop, the human behind it did. What’s more, as Wired explained, the AI-drafted story “focuses, strangely, on the fact that (OpenAI’s) agents rebuilt a message board rather than the fact that they created one in the first place. Overall, the stories are flatly written and tend to have an info-dump quality.” The AI didn’t seem to understand what was most important in the story, and did not prioritize readability.
On the contrary, independent journalist Sharon Goldman, founder of the excellent Ground Level AI—was actually in the room, spoke with other humans who were there, and had her story up in almost the same instant as the automated piece. Goldman’s tweet about the story got around 2 million views, and Ground Level AI briefly catapulted into the top of Substack’s technology publishers. Her piece got the attention and defined the news cycle. Yeah, humans!
Now, you may say this is all just cope, and the fact that RuntimeWire even got close to beating an expert journalist to a scoop is proof that the machines are inevitable. Fair enough.
But consider what that would mean. As Wired notes, “When AI chatbots go looking for sources, they frequently pull up AI-generated articles.” The article cites a study finding that “AI tools like ChatGPT and Claude surfaced AI-written sources 16 percent of the time when (researchers) tested it across four different topics.”
If LLMs prefer to answer a user’s request for information by serving up content that is itself machine-generated, they are at risk of creating a doom loop. Stories written and reported by people could indeed find themselves elbowed out of the way. But they’d be replaced more and more often by lifeless, cardboard-dull articles that may be broadly factual but utterly fail to convey the importance or contextual nuance—the meaning—of the information they contain.
It’s easy to see a future in which trustworthy information isn’t merely competing with whatever is going viral on social media to gain eyeballs. Instead, it will have to compete with an AI-generated facsimile of original reporting, fact-checking, and genuine communication. So the question may not be whether AI one day will be able to break news, but do we want to live in a world where “all the news that’s fit to print” is really “all the news that fits in a token budget”?
Some other Glitchy headlines 🐈⬛



