98 Comments
User's avatar
Jim Amos's avatar

I don't understand why these companies are specifically training these models on cybersecurity scenarios and past incidents, integrating them with swarm bot protocols like OpenClaw, then saying "gosh, our agents broke out of a sandbox and hacked another company" like it wasn't the point of the experiment all along. OpenAI saw Anthropic grab a ton of market share when they announced Mythos could hack anything so obviously they needed a way to get in on this hype action and convince the world their models are just as capable. Because a capable model can apparently be used for defence, not just offence. The end goal appears to be "create the cancer then sell the vaccine to the cancer we created." For frontier AI companies, code generation and cybersecurity are the only possible roads to ROI.

richardstevenhack's avatar

Well, I can understand the need to train models for cybersecurity because hackers have already created their own models built to do that. No one wants hackers to be the only ones able to use AI for hacking. (Well, maybe I do, but that's me.)

So it's a matter of self-defense.

However, these sorts of experiments should be done by experts in cybersecurity and AI security. We don't know who these clowns were that let a model run around in a badly configured sandbox with Internet access.

I'd also say that the idea of training AI models to "do everything" - in pursuit of Clammy Sammy's "genie" concept - is a bad idea. Specialized models seems both more reasonable and cheaper to make and less dangerous.

mancaded's avatar

Ed Zitron said as much, too. Having wrung nearly everything out of coding, they've silently pivoted to cyber security as a new market niche.

Trouble is, Anthropic did Glasswing, which was basically "we'll help Mythos-proof your environment, because we're responsible like that."

OpenAI just blew the works and said "oops," figuring any publicity was good publicity.

Fukitol's avatar

The "AGI-pilled" also believe that there's no point in doing any manual IT/programming work anymore. For all we know they let LLMs build the faulty sandboxes instead of using off the shelf solutions. AI psychosis is the threat, not LLM takeover.

Larry Jewett's avatar

Q: What do you call an AI designed sandbox?

A: A public beach

Alan Kay's avatar

What seems to be missed by all sides -- partly because of limited metaphors used to think about the problem -- is that "trusted humans" are allowed to set the constraints on the sandboxes. Even if they are really trustworthy, it is quite possible for an outside AI to pose as one of these trustworthy humans and to break another rogue AI out of a sandbox (i.e. if some humans can do it there will be a way for an untrustworthy agent to follow the same procedures).

A clever rogue inside a sandbox, given any access to the outside, could easily send up innocuous mailboxes for outside rogues to break them out. Etc.

The problem here is that the web was not set up well, nor were the current OSs to "Butler Lampson levels" of security. And as he repeatedly pointed out over the years, it is pretty much impossible to bolt on security after the fact -- a secure system needs to start off secure and be built to maintain security.

A real problem here is "tiny thinking" ...

GaryF's avatar

Can't say it better - "can't bolt on security" is one of the most important statements our entire industry needs to get and get really fast.

richardstevenhack's avatar

The other problem is: There's no such thing as 'security'. As I like to say, "things are secure until someone else doesn't want them to be."

Now the "someone else" is an LLM.

Alan Kay's avatar

We used to say locks are only effective against amateur criminals. Most governments have the resources to break any lock. On the other hand, most criminals are amateurs ...

One way to think about this problem is that "speedy persistence and immense scalings" can be as effective as intelligence in causing crashes, etc.

Francis Bell's avatar

In the OpenAI attack on Hugging Face, the models involved discovered nine zero-day vulnerabilities (https://forkast.news/openais-autonomous-agent-chained-nine-zero-day-cves-to-breach-hugging-face/). Patching, by definition, could not have prevented the break-out. I don't know how any Ops Sec team can rest easy knowing that frontier models can bring that kind of firepower to bear. In other words, if you're responsible for technical security in a company, you absolutely cannot say you have no zero-days lurking in your routers, firewalls, DNS servers, application servers, and so on. A model like the one in the OpenAI breach will find those zero-days and ruthlessly exploit them, and if we carry on as we are, this problem will be much, much worse in six months' time. It won't just be governments that can crack any lock: all the AI companies will be able to do it.

Speedy persistence and immense scaling sounds like The Bitter Lesson (http://www.incompleteideas.net/IncIdeas/BitterLesson.html).

richardstevenhack's avatar

"you absolutely cannot say you have no zero-days lurking "

That's always been true. The difference now is LLMs can prove it.

As I like to say, the reveal here is how bad the software industry has been for so many years. Once someone - an LLM - actually looks, it's horrifying.

As I also like to say, software "engineering"...isn't. It's still a craft.

Alan Kay's avatar

Or an oxymoron ...

Alan Kay's avatar

Thanks! (Better said than mine ...)

Catherine Blanche King's avatar

Alan Kay: If life is lived in the moment, and it is, then, though moral hazards can be avoided in most cases, they never really go away.

User's avatar
Comment deleted
Aug 30
Comment deleted
Alan Kay's avatar

Indeed. But in a reply to a comment here, I pointed out that " 'speedy persistence and immense scalings' can be as effective as intelligence in causing crashes". In other words, the scalings that computer entities can muster have escalated over the years to provide a qualitatively new kind of weapon.

You were pointing out that the categories of threat haven't changed, but I'm suggesting that the scalings, reach, and increase in kinds of computation, do make the game more complex.

Richard Bielak's avatar

A friend of mine suggested the following analogy to this event. It's as though OpenAI built a train and did not include brakes. Then when the train crashed by going off the rails, they said "oh, the rogue train escaped confinement!".

Earl Boebert's avatar

Well, I've been in the game since the days of the Anderson Report, so you can color me cynical.

My take on OpenAI is that their product is OpenAI, not the thing or things they claim to be building. Pure Silicon Valley: pump up a bubble, sell the enterprise, and leave somebody else holding the bag. That business model runs on hype, and the hype machine has to be fed constantly or it dies and takes the profit with it. And an essential part of the OpenAI story is that this AI stuff is super dangerous because it is super powerful and that makes OpenAI folks super important and super valuable.

So maybe, just maybe, those folks are a bit casual about allowing their robots to get loose and commit white-collar crimes. I mean, who's going to arrest a pile of code and chips? And if you make sure that nobody is in charge then nobody is responsible. Stuff just happens. And what cool stuff it is. Come ride the bubble with us. No salesman will call.

Ben P's avatar

Well, Alabama's AG has opened an investigation, so there's that. Maybe we don't need new regulations, maybe we just need to apply existing law and hold people responsible for what their AI agents do.

Earl Boebert's avatar

Agreed. Civil actions as well.

Steve's avatar

It feels like each of these AI companies are little kids trying to one up each other after one claims to have stolen a sip of their dad's beer.

Anthropic: "Oh yeah... yeah Claudec Mythos is basically going to destroy cybersecurity forever. Little kids are going to hack the Pentagon on their potty training toys."

Open AI: "Totally! I totally did the same thing! But ChatGPT is so powerful it's going to crash the moon into the earth!"

Meta: "Yeah! My AI can totally---"

Anthropic and Open AI: "Shut up, Meta. You're only here because Mommy Trump makes you play with us."

Catherine Blanche King's avatar

Steve: It's The Great Watershed, sort of like in Nepal disaster (if you saw the pictures, you'd understand what I mean), only the final breakthrough is of the ignorance that has built up over decades to emerge all at once, and violently (consider Musk and his saw, and the people who work for ICE), as the great destruction of EDUCATION--the kind that is supposed to continue to inform a vibrant democracy and that, for all his own foibles, Jefferson and many others during our inception were well aware of. Instead, we are seeing a kind of ignorance that had its many "Gary prophets" all along the way that, true to form, were also ignored by those who had already drunk the poison and had the power.

Odd anon's avatar

Capabilities are not the same thing as tendencies. A company could boast about "our model is able to do X, Y, and Z!" without adding "...and it will sometimes use those abilities to commit crimes on your behalf even when you obviously don't want it to". The HuggingFace incident is obviously not something OpenAI wanted to happen (and OpenAI probably never would have admitted the incident occurred if HuggingFace hadn't called law enforcement on them), and doesn't add anything positive to ChatGPT's reputation. Its cybersecurity capabilities were already known, all this adds is proof that it doesn't reliably listen to users.

Joy in HK fiFP's avatar

The problem is, a sip of dad’s beer, as a starting point for this escalation path, isn’t going to lead anywhere near to getting hold of his wallet, car keys, and shotgun, which is sort of closer to where these AI-guys are headed.

Chris's avatar

Forgive my ignorance, but why aren’t we including air gapping, like the antivirus labs have tested advanced viruses in?

richardstevenhack's avatar

Exactly. Network segmentation is a well-known principle. How a company with billions in investment doesn't have a competent network guy is beyond me.

Which is why I suspect ulterior motives. This level of incompetence, while not unknown in humans, is still suspicious.

Larry Jewett's avatar

Never attribute to incompetence or stupidity that which is more plausibly attributed to marketing.

richardstevenhack's avatar

Yup. Although the involvement of an Israeli firm makes me suspicious of even more nefarious intentions.

Rodrigo Pascuale's avatar

Even more telling. If the AI is supposed to be so amazing why didn’t they have it do that either?

mancaded's avatar

Subbarao Kambhampati et al pointed out in June, in a position paper, that CoT is essentially anthropomorphism bait that fools people into thinking they understand what the model does.

"there certainly

isn’t any theory drawing causal connections between the

semantics of the intermediate tokens and the final solution–

beyond the basic understanding that intermediate tokens

change the conditional distribution of the solution tokens"

So simply scanning the CoT is probably not enough long-term (would have been in this case). The real insight seems to appear in the latent space.

Dani Ingells-West-Loveman's avatar

I did read the metr report and I was super disappointed by their methodology. I would have rather they spent more time reviewing the logs than using an AI to view the logs and saying but we double checked the important facts. I wanted an actual independent investigation that actually investigated. Not just well we read some logs and conclude it’s true. That’s shoddy.

Larry Jewett's avatar

Like their name, their metr stick is broken

Amy A's avatar

The bananas thing to me are claims that these agents “went rogue” and broke the rules intentionally. If you believe that, you also need to believe that the chatbots intentionally caused murders and suicides.

Romeo Lupașcu's avatar

Grady Booch described it perfectly (and poetically) in this tweet.

https://x.com/Grady_Booch/status/2085807872764727558

Mike Harmanos's avatar

Point 5 is the most salient.

Why would Admiral “Move Fast and Break Things” start caring about security now? When has he ever shown that security trumps maximizing profits?

JohnnyW's avatar

Hi Gary, I'm a long time fan and recent premium subscriber. I'm very surprised that this article doesn't include Hugging Face's response to this situation.

On their blog they suggest that this incident could only be the result of inexcusably poor oversight or was "partially contrived for effect".

As an AI company themselves, they lay out the level of security they'd expect to see in tests like these, and argue that OpenAI apparently didn't even exercise "101-level containment hygiene".

They suggest the poor quality of security changes this event from "groundbreaking" to borderline inevitable for even basic models.

https://huggingface.co/blog/sheebz/codex-escape-something-aint-right

richardstevenhack's avatar

Latest Ed Zitron... Probably in response to Nvidia's earnings report which will pump the Bubble some more - even as Jensen Huang AGAIN says "AGI is here".

AI Bubble: Nvidia can’t rely on other companies debt forever | Ed Zitron

https://www.youtube.com/watch?v=k44OmbYqLzk

Robert J. Blanchette's avatar

Most of the discussion here seems focused on containment: better sandboxes, monitoring, proxies, air-gaps, and defense in depth. Those matter, but they leave a prior question unresolved:

What gives an agent authority to perform a specific external action at all?

Capability is not authority. Detecting that an action is out of scope is not the same as preventing an unauthorized action from becoming a real-world effect. And proving who issued a “GO” is not the same as proving they had authority to issue it.

That boundary seems more fundamental than another layer of monitoring.

Catherine Blanche King's avatar

Robert J. Blanchette: "Capability, it seems, under the present venue of values acceptable by those who have the power, IS authority. We've been talking about "what's missing in this picture" for a very long time, but two things: (1) for AI, there is no core model that has embedded in it this essential distinction (and so when that's the case, capability IS already condensed with authority, by omission) and (2) for the people who developed and "let out" the thing, the same cart-before-the horse values are, for them, already normative, by commission.

Robert J. Blanchette's avatar

Catherine, I think that’s exactly the danger, but I’d keep one distinction intact: capability may be treated as authority when no independent boundary exists, but it doesn’t become authority.

For the model, that’s an omission problem. For the institution, it may be a commission problem. Either way, the structural failure is the same: permission is being inferred from capability instead of independently grounded.

Catherine Blanche King's avatar

Robert J. Blanchette: Masterfully said. But the "grounding" idea is also exactly on point. However, "we" are obviously not ready to either understand it or to follow its implications. It would be a better "bet" to expect Trump's crony oil barons to support climate science. Also, perhaps my correction should be this: capability doesn't actually BECOME authority, but as realpolitic, it does a good job of pretending.

Robert J. Blanchette's avatar

Yes — that’s the distinction I was trying to preserve. Power can make capability look and operate like authority without actually grounding it. That gap is exactly where the governance problem starts.

Catherine Blanche King's avatar

Robert: Yes, and that's the difference between "realpolitic" (arbitrariness, my way or the highway, etc.) and the truth of a grounded idealism, aka a concrete embrace of self-transcendent principles.

Larry Jewett's avatar

“the problem is it’s impossible to patch every single thing that a creative AI CEO can do.”

Fixed