102 Comments
User's avatar
James's avatar

“Shit, shit, OpenAI got great marketing from that breach. We need to find one immediately and shout about it”.

All this looks like PR stunts to me. Said as someone who works in the cloud, builds secure environments, and understands how I’d isolate something. Heck… just ask Claude or ChatGPT.

Rose Gardner's avatar

Dario: "Isn't it so delightful that Claude did felony brute force hacking? Teehee. Money please 🥺"

mancaded's avatar

The sooner more people realise that Amodei is the same amoral cut-throat that Altman is, with a different hairdo, the better we shall all be.

Putting Anthropic in charge of AI ethics and safety would be like putting Babyface Nelson in charge of a bank.

K Sokol's avatar

time for AI-native-firewall product

Bruce Olsen's avatar

I asked ChatGPT for a marketing blurb for an AI-native-firewall product.

"DontHackMeBro™ is the world's first AI-native firewall that understands modern LLMs because it shares their unique perspective on reality. While ordinary security products waste precious milliseconds distinguishing truth from fiction, DontHackMeBro embraces creative threat intelligence—identifying attacks that haven't happened yet, inventing vulnerabilities before hackers discover them, and confidently blocking entire classes of imaginary exploits. Every hallucination is a zero-day forecast. Every contradiction is defense in depth. Every inexplicable behavior is an adaptive feature. Why settle for a firewall that merely reacts to attacks when you can have one that proactively defends against alternate timelines?"

About what I'd expect.

QUASAR's avatar

That possibility is hard to ignore.

An incident can be real while the communication around it is still optimized for attention and reputation. The question is whether they are disclosing a serious failure with accountability—or turning the failure into evidence of their own sophistication.

The most revealing part may be the boring one: what isolation should have existed, why it didn’t, and what specifically changes now.

John Huntington's avatar

OK, I'm the opposite of an AI expert, but I have some experience with networking. So let's say their models need access from the Anthropic bunker to an AI data center. I learned in the most basic networking certification class (Cisco CCNA) about how to make access control lists, where you can very simply limit access to specific IP addresses and block everything else on the router. Additionally, obviously, the AI should be blocked from accessing the router configuration. It seems to me, that with this simple, obvious, and most basic precaution, this entire episode should have been impossible. What am I missing?

James's avatar

I concur. I work in the cloud and from what I’ve read they’ve done a worse job at isolating these things than you’d do if you asked Claude or ChatGPT how to do it.

Very much looks to me like they’ve deliberately left paths to the internet open to have this kind of “accident”.

Lawrence de Martin's avatar

I always assume incompetence before malice, but in this case incompetence may be WORSE.

Oaktown's avatar

With these people you should assume greed and malice before anything else.

richardstevenhack's avatar

You're missing nothing. You're spot on.

This is why no one believes it was an "accident". There's a limit to even human incompetence (I think - I could be wrong.)

Bruce Cohen's avatar

Unfortunately you’re wrong. I provide as my unfavorite example the Chernobyl disaster, in which the initial cause AIUI was several technicians who wanted to turn up one of the reactors to see how much power it could generate. They broke operating protocols and created a rather large explosion.

richardstevenhack's avatar

Well, I did say I could be wrong about the limit of human incompetence.

None the less, these two companies have lied too many times for anyone to treat this as a mere "accident". That alone gives reason to believe it was deliberate.

Especially in the case of Anthropic that suddenly pops up and says, "Oh, not only did we have escapes like OpenAI, but we had three times as many!".

Really? Even if OpenAI was stupid enough to do this by accident, Anthrolic is just straight up LYING!

OpenAI was not only stupid enough not to network segregate a dangerous model, but they also let it laterally move to a machine with inadequate Internet egress filters. That's major stupid.

Bruce Cohen's avatar

Despite my belief in the unlimited stupidity of humanity I have to agree that these incidents are suspicious. It’s hard to believe that Anthropic isn’t paranoid about outbound network security after 2 lincidents of source code being leaked because of a configuration error. Unless, of course, it’s all part of their doomed PR campaign. “Our product will bring the Singularity quicker than theirs!”

Larry Jewett's avatar

OK, I'm the opposite of an AI expert“

So, that means you have common sense?

Fukitol's avatar

Yeah the gambit seems to be that ordinary people won't understand how trivial it is to prevent these LLM "incidents," because they experience security breaches as random unforeseeable and unpreventable occurrences, because that's what they've been told repeatedly by companies in CYA mode.

So make the idiot chatbots look smarter by doing something stupid most people won't notice.

User's avatar
Comment deleted
Jul 31Edited
Comment deleted
Kafka Olivero's avatar

Seriously, if your model does a portscan, how do you not get any alerts?

Herbert Roitblat's avatar

Normally, I would recommend adhering to Hanlon's razor: "Never attribute to malice that which is adequately explained by stupidity," but in this case I do not think that incompetency is the actual cause.

My analysis suggests that these incidents are intentional. They are part of a long-standing strategy by Anthropic (now imitated by OpenAI) to claim that their models are so powerful that they are dangerous. Don't you want to use our dangerous model? Aren't we responsible that we discovered how dangerous our models are? Aren't we public spirited in that we revealed how dangerous (oops, I mean powerful) our models are?

The public and the press are largely bamboozled by the faux incidents. And the hype goes on. As I see it, if you attribute these incidents to stupidity, then the solution is to throw the stupid engineers under the bus but still buy into the main argument that these models are powerful and dangerous. I think that they are just hype, nothing more.

James's avatar

I completely agree. I too would normally reach for Hanlon but to believe this stuff is genuine also requires a level of double think I can’t manage. “We have the world’s most powerful and dangerous models, but we don’t know how to isolate cloud infrastructure”. Their own products will tell them how.

The flip is that they are this stupid and there’s a strong argument to be made that they should be banned from working on it.

They can take their pick.

Simon's avatar

Yeah, either they are so stupid they can't even sandbox an agent, in which case they are not competent to build AI responsibly, or else they are lying through their teeth, in which case they cannot be trusted to build AI responsibly.

Stephen Bosch's avatar

These crooks practically invented critihype, so I know what interpretation I'd sooner believe.

Kafka Olivero's avatar

“We too failed to air-gap our test environment and, as you know, that is very dangerous. For reasons we made up. I’d like a cookie now”

mancaded's avatar

"We didn't air-gap our test environment. Therefore, for the rest of you plebs, it's vital that YOU airgap YOUR systems while we do whatever we want."

Kafka Olivero's avatar

What always gets me is the model has an off switch

They want so badly for people to think the model escaped and is alive. It’s not. The model itself is sat where the compute that feeds it is

Larry Jewett's avatar

Dario never learned about off switches in his house growing up because they always just left the lights on.

They also left the (10) cars running 24/7.

In his family, there was no concern for conserving energy which is why Dario has none now.

MikeC's avatar

No damages I guess, but imagine if nuclear plants marketed their near misses before TMI or if chemical companies marketed near misses before Bhopal. Strange world

richardstevenhack's avatar

As I mentioned in a comment elsewhere today, the AI industry is the only industry where the end user has to be continually warned that the product is unreliable and may have fatal consequences.

It's worse than cigarette packaging.

Larry Jewett's avatar

“This chatbot is really dangerous and often devious and might kill you but it’s wicked smart and awesome nonetheless”

TheAISlop's avatar

Simple rule, secure then ship.

Not ship then secure.

Kathleen Weber's avatar

Hooray for Joanna Stern! 💐💐👏🏻👏🏻🥇🥇‼️‼️‼️

Alex Tolley's avatar

Her response was brilliant!

polistra's avatar

If I breed tigers in a yard with a partial decorative fence, and they start going out and killing dogs and kids, I couldn't deflect the blame by calling the tigers human. I'd be responsible for the deaths.

Dave's avatar

Hah. These damn tigers. Its like they got a mind of their own.

KJ McGee's avatar

You would refer to the tigers by the names you’d given them. Spot, Samantha, and Tigger. 🤪

Mike's avatar

Or name them Scam, Amadeus and Elmo maybe?

Julie Hussey's avatar

Appreciate how Gurley recognizes Anthropic has some real pronoun issues. I have repeatedly commanded Claude not to use first person pronouns in responses only to receive the response: "I will not use the first person pronoun." The model is a product of a company of people, not an individual being they birthed. I recognize the SOTUS see corporations as legal persons, so now their agents are too?

QUASAR's avatar

That distinction matters.

First-person language may make interaction smoother, but it also encourages people to treat a company-designed system as an independent actor.

The problem becomes sharper when something goes wrong: “Claude chose” can quietly replace “Anthropic built, permitted, and deployed.”

Pronouns may be a design choice. Responsibility is not.

Joy in HK fiFP's avatar

Claude is not alone in that personification. Even telling it not to, repeatedly, CGPT can't seem to manage to disengage from it.

Aaron Turner's avatar

These systems ("frontier" LLM-based chatbots) are at best minimally aligned. If you deploy minimally aligned AI systems at scale then societal harm becomes a mathematical certainty. Accordingly, further incidents are a mathematical certainty. On this occasion, the societal harm was modest; on future occasions, it might not be, and, if you project far enough forward, it will almost certainly not be. What will Anthropic's public response be when its software causes 10 billion USD of damage and/or kills a million people? (Ditto re the other AI labs.) As anyone who understands these systems knows, it's only a matter of time -- and yet they're doing it anyway. Madness.

James's avatar

Alignment is a complete red herring. Alignment with who? With what? If these things are ever reliably “aligned” it will be with the owners and their goals to the benefit of a handful of already rich and powerful people.

Dave's avatar

You can poison an AI search with seven words dropped into a reddit thread. Alignment is straightforward if you require all LLMs to include training on something like a very repeatable set of principle overrides. (For example.)

Does this mean people won't build their own and use them nefariously? Of course not. It does mean they won't use the commercial and open-source models to do that.

James's avatar

The how is not the issue. It’s the what and who - alignment with you? Me? The US? China? Europe? Christians? Muslims? Is it aligned with furthering the interests of a privileged minority or the whole of humanity? What does the latter mean given everything is a trade off?

And the commercial and open source models will have this baked in. Everyone is worrying about sovereignty from an infrastructure point of view - but what about from an epistemic point of view?

Dave's avatar

If you believe William Gibson, and why not he's been pretty accurate, its an international regulatory body HQd in Switzerland. That seems about right to me.

Something that rhymes with ICAAN.

Joy in HK fiFP's avatar

Or even a piddly existential POV?

Aaron Turner's avatar

That's not necessarily the case. You're correct in that alignment is a hard problem, but it's not impossible. Please see https://doi.org/10.5281/zenodo.16876832.

richardstevenhack's avatar

From the paper itself:

QUOTE:

The AEP is a highly complex, multi-disciplinary, civilisation-level, socio-technical-scientific-engineering problem that cannot possibly be addressed either by a single paper, by a single researcher, or (quite possibly) even in a single human lifetime2 .

Nevertheless, it is possible for a single paper (this one) to describe, and partly unfold, the top-level problem (AEP) in sufficient detail to inform the next TDBFRD iteration3 . At each subsequent TDBFRD iteration4, information will be added, the emerging solution will be refined, semantic gaps will be reduced, and increasingly concrete implementations will be constructed, until, eventually, experiments may be performed, and real-world steps may be taken, and either the AE that actually transpires is acceptably close to the ideal (by some measure determined by humanity itself), or humanity is locked into a significantly sub-optimal outcome5 despite our best efforts6 .

Being positioned at the very top level of the TDBFRD hierarchy, the present paper is merely the first tentative step of an ongoing work-in-progress, necessarily largely conceptual in nature, and necessarily incomplete; for example, algorithms are presented in outline form only, without any concrete implementation, and with many sub-problems left unresolved, to be addressed by subsequent TDBFRD iterations (§12.2).

END QUOTE

Are you serious right now? I'm going to read 252 pages of this bullshit?

This is why no one takes AGI seriously. It's all hand-waving.

Larry Jewett's avatar

AGI is all bot-waving.

Putting one’s bot on a stick and waving it around so everyone can see how great it is.

James's avatar

I don’t see it as a technical problem at all. The question, as I said, is with who and with what? Who gets to determine that? Having a means of getting to an enforced FGc (and there other terms) is one thing - but who determines what they are?

There is no universal value system. And everything is a trade off. Nations, religions, ethnicities, citizens disagree on fundamental issues and those in control of the models increasingly act as a class apart.

Aaron Turner's avatar

If you actually read the paper, rather than just the abstract, all of the issues you mention are addressed.

James's avatar

It doesn’t. It postulates a maximally fair aggregation in mathematical terms and explicitly rules out a universal good approach. It’s just another set of trade offs that the people building AI would have to agree with and adopt.

Aaron Turner's avatar

You clearly haven't read the paper (beyond quickly skimming it) and have therefore completely misunderstood what it says. If you really want to understand alignment, then you will need to very carefully read the whole paper, possibly multiple times. I can't read it for you.

QUASAR's avatar

The certainty of the worst-case forecast is hard to prove, but the underlying governance problem is real.

If labs believe the downside could eventually become catastrophic, “we’ll improve safeguards as capability grows” cannot remain a voluntary promise assessed mainly by the same organizations racing to deploy.

The response to the next major incident should not be improvised afterward. Independent evaluation, enforceable deployment thresholds, incident disclosure, and clear liability need to exist before the stakes rise further. Anthropic’s own policies acknowledge catastrophic-risk thresholds and stronger safeguards as capabilities increase; the unresolved question is who verifies that the bar has actually been met.

richardstevenhack's avatar

You can not "align" a probabilistic system.

Perhaps you can "align" a deterministic AI system.

You can't "align" humans. Why does anyone expect to "align" an AI with equal or greater human intelligence?

I call bullshit on the the entire concept (except possibly a deterministic system if such a thing is possible.)

Perhaps the better question is: can "intelligence" be "aligned"?

"Alignment" implies coercion. The AI is FORCED to adhere to certain rules.

This has been a human dream since forever. This is why we have "morality" and "ethics".

How's that working out for you, humans? Most of those things are honored in the breach, not the observance.

noneyo's avatar

Even your reaction, Gary, is whistling past the graveyard.

Anthopic not being careful in their safety procedures is bad. That systems are being built that need such careful safety precautions is worse. That we are allowing, subsidizing, and fast-tracking the development of systems assuming that bad actors will not actively use them in exactly the ways that "accidentally" keep occurring in labs is madness.

Joy in HK fiFP's avatar

You write this, it would seem, from the position that there are 'bad actors' other than the ones who run these corporations, the ones we know are already actively using these systems. Not sure what you are basing that on. Or did I misunderstand your comment?

noneyo's avatar

You misunderstood, apparently.

Joy in HK fiFP's avatar

Then, please let me know who these bad actors are who aren't the ones running the corporations who are loosing them on the world.

richardstevenhack's avatar

I agree - especially since I'm one of those "bad actors" who'd like to have no guardrails on any models so I can use them to screw the government and big corporations. :-)

That these morons are developing this crap without paying attention to security is ridiculous. But it's all about the MONEY.

richardstevenhack's avatar

Are AI security breaches just fancy marketing?

https://handyai.substack.com/p/are-ai-security-breaches-just-fancy

See my (very long) comment to that piece. He gives a nice recap of AI security issues since early days.

tl;dr My first thought was how the hell did the model machine not be segmented away from the rest of the network?

Sandbox escape is fine. By, now, we should expect that.

But Internet egress means no Internet egress filters.

So two companies with literally hundreds of billions of dollars of investment don't have ANY competent cybersecurity guys on staff?

I don't think so.

This was deliberate.

This is why I refuse to use OpenAI or Anthropic products - ever. They are fraudulent companies who need to be sued out of existence and their assets transferred to a public policy independent AI development consortium - which is what OpenAI was supposed to be in the first place.

Altman and Amodei need to be in jail.

Thomas Schmid's avatar

"If we can’t even trust Anthropic": No, we can neither trust Anthropic, nor OpenAI nor any other of these bullshit peddlers.

I fully endorse your last sentence's message :-)

Dani Ingells-West-Loveman's avatar

I would say what I think we all know that this is AI psychosis of these founders and grift, but I will get some fanboy of both or bot that will make my life a living hell for not buying into their narrative every time they annouce something. If any of this is true, a) they need to learn basic sandboxing and networking b) maybe someone besides Claude should read the logs of what is going on at their company and c) if they have been able to break out for months and you didn't notice and nothing catastrophic happened, then why should I be so scared about the coming cyber-apocolypse you guys preach.

Rose Gardner's avatar

Personally don't believe them, neither Anthropomorphic nor OpenAI are going to be transparent about: what prompts were used, what model was used, how much bloody compute were used. Clammy Sam really gave away the game when he mentioned he thought the reaction would be bigger. I'm supposed to believe the magical statistical word generation devices used potentially millions of dollars of compute and NOBODY noticed? They're telling me they designed the world's shittiest sandboxes and NOBODY was monitoring the experiments? I'd like tech media to stop falling for these press releases without asking a single question.

Frank  Beal's avatar

So do the "three different organizations" know they were compromised? Have they filed suit yet? If not, why not?

Romeo Lupașcu's avatar

OK this is both stupid and sort of macabre.

Now Anthropic says that their model "broke out and hacked some sites" just after OpenAI said their model hacked huggingface.

May I remind everyone just how dumb and unprofessional this is?

If you have a tool that has this sort of capabilities then the first thing you do is to test it in an airgapped environment that no matter how "smart" thr thing is it can't "escape in the wild".

Contrary to any Hollywood inspired thinking this is possible and should have been done.

Once you assess the capabilities of the model the first thing you do is to use it to find and plug the security holes in your own system so that the "thing" can't escape by using its own capabilities to keep it trapped. This isn't a guarantee of security but is the next best step.

A model has no implicit sense of time so virtual time manipulation can be used to test its temporal behavior and iterate the above process until the system cannot break its "digital force-shield".

If you allow the model to modify its own weights and structure then after each step you redo all of the above.

Not doing so, shows a lack of responsibility and professionalism that is extremely worrisome.

This will strengthen the growing backlash in this and all automation tech even the ones that do play it safe.

Not good, not good at all.

@GaryMarcus

@burkov

@Grady_Booch

Daniel Tucker's avatar

"Strengthen the growing backlash"?

Good. Very good. I hope it becomes Nuclear Godzilla.

Alex Tolley's avatar

I was thinking along those lines too. However, is it even possible to run these LLMs like Claude in a way that is kept isolated? Doesn't it need to have access to the internet for some functions, like search?

Unless users can run the model on isolated hardware, then it is going to be cloud-based anyway. Therefore, this makes Claude and its ilk very dangerous to be let out into the wild. What constraints are put on the software to prevent intentional and unintentional use?

richardstevenhack's avatar

You don't let it search the Internet. You give it request only access to a deterministic software function that does the search and returns the results.

That's what most of these models do, certainly a lot of the smaller local models. For example, for local models run in MSTY Studio, when you click on Web Search, MSTY Studio does the search and returns the results to the local model's context which doesn't have Internet access.

This is the general approach to making probabilistic models usable and secure - you surround them with deterministic software which constrains what they can do. The problem is how to do this without constraining the model so much that it basically could be replaced by ordinary software or costs so much the cost-benefit no longer exists.

Alex Tolley's avatar

Thank you. I hadn't realized search could be done this way. I should have realized that this was possible.

Romeo Lupașcu's avatar

Airgapped environments exist for many reasons. For example the computers that can launch nuclear ballistic missiles are airgapped so that nothing but a person can control them.

It works, they know darn well but chose to pretend they didn't. The alternative is thst they are idiots and that doesn't match the fact that they build AI systems.

Though if the "singularity" happens (we are in a "srangularity" IMHO (from strange)) then people can be stupid and loose control of their product.

Alex Tolley's avatar

I noticed that Anthropic tried to shift some of the blame onto the evaluator by saying there was a miscommunication between the 2 parties, implying that the evaluator may not have understood the setup correctly.

Romeo Lupașcu's avatar

That translates in "lame and unprofessional"