93 Comments
User's avatar
Green Duck's avatar

Problem being that transparency is hard. If the protocol is formalized ahead of time, then LLM vendors will train their models to detect when they are being tested, and to behave accordingly. See the Volkswagen Dieselgate scandal, where precisely this happened: https://en.wikipedia.org/wiki/Volkswagen_emissions_scandal. So the tests will necessarily be opaque and adaptive - but one person's "opaque and adaptive" is another person's "unaccountable and corrupt".

nAxis's avatar

We already see this with training to the benchmark as well. This is why ARC-AGI continually need to be versioned because by the time they're released we see a clear incremental process as the models are trained on that cohort's version of the test.

Bill Donahue's avatar

And then of course after that even if an LLM maker was nailed and penalized after getting caught doing such a thing, we'd see that other companies would still try to design their way around testing, like Stellantis and Cummins did with Dodge diesel trucks after Volkswagen was "made an example of" for doing it.

polistra's avatar

Transparent theft is still theft.

nAxis's avatar

Ah yes, taxes.

Romeo Lupașcu's avatar

So... no one told him that LLMs can't be tested enough to declare them safe? Like anyone checked just how much state needs to be verified to even touch the surface of that monster? Is he a smart guy?

Larry Jewett's avatar

Hassabis knows that LLMs not only can’t be deemed but can’t be MADE secure but he works for Google and Google wants the public to believe otherwise.

Who better to convince them?

Catherine Blanche King's avatar

Larry Jewett: I think they want to keep him, however.

richardstevenhack's avatar

Exactly. More succinct than my long-winded comment.

Bonnie Blodgett's avatar

Hassabis believes that AI will kill us all. He would like to put the genie back in the bottle but of course that can't be done. So guardrails it is. Here's hoping the whole AI project blows up when investors lose confidence.

Catherine Blanche King's avatar

Bonnie Blodgett: I haven't read the paper yet, but can you cite your statement about Hassabis' belief that AI will kill us all?

Bonnie Blodgett's avatar

I can't cite my statement's source because my comment was based on my interpretation of the paper and of the interview his biographer gave to Novara Media.

I think Hassabis is hedging his bets. He wants to be good (don't we all?) but above all he wants to be right. So he very eloquently threads the needle, presenting two options and then suggesting that it's not his choice to make but ours, as if we ordinary folks can and must choose between risking the destruction of humanity for the sake of curing disease. Or not.

This is a false choice. Moreover it can't be our choice. How the hell do we know? We're not brilliant scientists. Whereas Hassabis, the most brilliant of them all, not only can but MUST make the choice.

The correct question is not, in this case, could a manmade technology cure disease, but rather, does that possibility justify risking the end of mankind? This is not a bet anyone should take (imo), but humans are irrational. We make decisions based on emotion (envy, greed, etc.) rather than objective data as AI would make them if it truly were rational. Sadly, AI is only as rational as the data it indiscriminately hoovers up.

Anyway, Hassabis bet his own career (his life's work) on AI being a force for good. Putting myself in his shoes, I'd like to believe that I'd use my intelligence (the human kind, not AI) to raise an alarm by ending my own AI research involvement. I would admit that I was wrong.

I am not condemning him any more than I condemn the scientists who allowed hubris and "the fog of war" to cloud their understanding of the potential consequences of making a nuclear bomb. Obviously, that anything justifies dropping such a weapon on innocent people is preposterous. But we humans are, if nothing else, really good at lying, especially to ourselves.

Hope at least some of this makes sense.

Catherine Blanche King's avatar

Bonnie Blodgett: Thank you for your thoughtful reply.

I am reminded of a statement in a political text I read long ago that said (paraphrased) that one irony about keeping up with democracy is that to get one's children to live well in one, you cannot treat them democratically at first. Rather, well-being in a democracy is a long-term developed thing--lots of guidance, good examples, cultural supports, tacit by just watching and copying (as children do anyway), and then overtly educational, and never merely ideological but at the end, and only when one has reached maturity as a slow process of development, one's political ground is not guaranteed good, but has the best chance of being a conscious choice that serves everyone well.

This can be related to Aristotle's statement in his Ethics that having good parents is not only a good thing, it's the only thing . . . the rest of his text suggests that he means that having good and constant parents can create the psychic and intellectual ground for a person's being able to foster well-being and to live rightly with others as is possible under one's circumstances.

I look around and have no wonder at the presence of all the skepticism and pessimism, even nihilism, we see on this blog and elsewhere. And to those who did not have anywhere near "good parents," it's a long and difficult journey to recover what others often take for granted, but then one knows something terribly significant by one's experience, something others do not know.

Further, when I first read and saw Hassabis talk, I thought he was probably naively optimistic but have since changed my mind. There are good intentions, where I think he qualifies as a star about or, again, a genuine article; (but then I used to think SCOTUS was an unbiased institution and still pine for Walter Cronkite); and then there are the limitations of even a fine intelligence who, on principle, cannot see around the corners of human living in history (as Demis seems to understand by the article--which I read fully a while ago), even if one's intentions are sterling. Historical blowback is what cannot be completely known by any of us, even if we all had good intentions.

I came away from what I already knew about DH and then the article with a view that Demis puts a lot of stock in the good OF intelligence, and the good IN intelligence which (and if I understand him correctly) I am with him on that. (At least it doesn't talk about an invisible hand that no one really understood in the first place).

My hope is that he can gather enough similarly trustworthy well-intentioned people together, with enough expertise and applications savvy, to make a difference. I just hope there are enough people who know that some people really are "bad actors" and build protections (as the article also suggests) against such influences and that some of us didn't have good parents and so not everyone understands or endorses what it takes to arrive at (loosely speaking) democratic/republican politics.

Bobby Hardenbrook's avatar

Not to be cynical or anything...but isn't this just google wanting to slow down the frontier model makers since they have fallen out of the race right now? Gemini isn't even in the top 5 right now for frontier coding models.

Catherine Blanche King's avatar

Bobby Hardenbrook: Everything has its potential underside. The point to watch for is the answer to the question: Which is the paramount motivator?--it usually shows up at some point, especially when someone experiences pressure.

I happen to love cynics, except when that's all they know how to do. BTW, watch Trump. He has made a Big Business out of exposing, speaking from, and taking advantage of the underside of everything he is involved with. I never cease to be amazed but have learned to expect it.

nAxis's avatar

What are your thoughts on GLM 5.2 breaking the frontier advantage illusion and the fact that open weight models will notably not be preflight tested which in effect disadvantages US frontier progress? Further, the idea that US frontier models can be pulled at anytime without notice, how can any corp go all in on any US frontier model under those terms?

As in all things when it comes to software infrastructure, ipen source will win. While Mythos is a step up, its capabilities are strongly supported with harness improvements. Routing Claude Code Ultracode effort harness to GLM 5.2 approaches Mythos in infosec today (exceeds it in design and some dev tasks as well). In 6 months it'll 5.x will have surpassed the Mythos of today and it'll still be open weight and still will not be preflight tested...

Mitchell Harper's avatar

China is already considering export controls on all models, so there is no reason to believe that the US instituting a product certification process would not be the starting point of a campaign convincing China to do the same as part of a global treaty on IT product proliferation. One could argue Wassenaar is a historical counterexample that eventually became obsolete due to open source proliferation, but I do think the world is structurally different today to where that example is not binding precedent if enough power brokers feel threatened by biological or radioactive aggression.

The above is the short answer. The problem with this answer is that the US is currently committing the most heinous forms of biological and radiological terrorism in the Gaza Strip and Lebanon right now and all this talk about future threats is a distraction from present annihilation. Targeting systems supported economically by the US are recommending to drop bombs full of depleted uranium or other heavy metals on the most densely populated civilian areas on the planet which will lead to guaranteed cancer deaths and birth deformities like in Fallujah. They are recommending buildings full of asbestos as targets, which will lead to tens of thousands if not hundreds of thousands of guaranteed mesothelioma deaths in the Gaza Strip and everywhere downwind of there as mesothelioma stays in the air for a long time (I grew up not too far from Libby, MT, so I personally understand the fallout from asbestos lasts decades). Despite multiple very clear international conventions, US backed forces are dropping carcinogenic herbicides and white phosphorus all over Southern Lebanon in order to turn the south into a Rafah-like wasteland.

Anthropic is currently hiring a "Safeguards Enforcement Analyst, Radiological & Nuclear Harms" with a listed salary of $245,000 - $285,000. The only way anyone can get paid that much money is due to the exploitation of the global working class and imperialist control over economic flows. Meanwhile, they sell products to a US government that is sending depleted uranium penetrating munitions into so many civilian areas causing incalculable health damage and potentially preventing the safety of human reproduction for decades, right now. I suspect that is "out of scope" for this role.

The question of what is "out of scope" for such checks is the quickest way to understand the contradictions of any preflight checks under an imperialist economic order. "Are we selling products to people causing mass immunocompromise by enforcing a total siege of nutritious food to a population while also preventing the repair of sewage and waste disposal facilities while rats and fleas run rampant through the entire civilian population (I just read a story today about a grandmother who woke up in her "tent" in the Gaza Strip with a bleeding chest from all the rat bites she sustained while sleeping) and total lack of disease surveillance creates a Petri dish to create new and dangerous biological pathogens that threaten the whole world" is likely not in scope for the kinds of checks and processes proposed. This question of scope has troubled me throughout my entire software validation career, and that's why I am ultimately ambivalent in my feeling about the urgency of implementing such processes although I cannot find any reason to fight the imposition of such processes and think they could be a starting point for further analysis.

nAxis's avatar

Firstly excellent reply. I don't think I know enough about the second part of your answer to even comment beyond saying what you're describing sounds terrifying and appalling.

As to the first part I think the point I'm trying to make is that even if China decides on some kind of export control in the future, we are rapidly hitting a point where even if you were to freeze the current generation of models, both closed and open weights, we're at a capability point where we're already in no man's land.

Further the idea that China is going to implement export controls on their AI models, while feasible, historically when it comes to their economic advantage and in particular this AI race, I predict that they're more likely to let their AI industry continue to run as hard as possible to further diminish U.S. frontier advantage. As I said even if they were to implement these controls tomorrow, GLM 5.2 represents a point in time where open weight capability approaches that of mythos today. I think with the right amount of focused harness engineering, it could be used to surpass mythos in terms of capability and without safeguards.

Now of course there's the question of where you run these open weight models but we're already seeing a proliferation of providers who run these open weight models by constructing private endpoints, a la wafer.ai. Now we could argue that wafer is a U.S. company and might also be bound by certain restrictions in the future. Again these restrictions are gonna take a while to proliferate down to that level but even then I could imagine wafer-like businesses being set up in neutral jurisdictions. Given the increased reliance on these types of models, demand should naturally move towards open source and neutral inference providers.

Mitchell Harper's avatar

Full disclosure: I think you are likely correct about the incentives, but structurally, China and the US are both competent at putting people in jail or punishing people who they deem national security threats. So it's hard for me to conclusively say that the US unilaterally implementing such controls is only going to cause competitive disadvantage and that this is the only possible outcome after multiple rounds of economic iteration.

It's so hard for me to think through how to address this issue, but the only iota of clarity I have is that I do think product liability lawsuits will likely be the ultimate driver of the behavior of US firms regardless as long as Congress doesn't neuter the ability of consumers to file suit.

Mitchell Harper's avatar

Another point to consider: how many open weight Chinese models will return meaningful results for queries relating to figures such as the Tiananmen Square Tank Man? There is already a tacit model preflight checklist in China.

nAxis's avatar

That's where LLM fuzzing comes into play... You can get any model to do basically anything and tell you anything during its forward pass, particularly if you also control the inference compute. This is how we're seeing even local models perform significant infiltration and red teaming when paired with exploit frameworks because they've been fuzzed to hell, all inhibitions removed, and given a harness in which to systematically analyze and exploit whatever you put in front of it. So while Mythos is cool and the supposed jailbreak of Fable (which incidentally wasn't specific to Fable), again, taking GLM 5.2 and giving it the right harness for whatever job you want it to do, it'll answer and do anything you want it to do, assuming you've properly fuzzed it (which isn't that difficult).

Mitchell Harper's avatar

It's possible to both restrict the training set and introduce heavy bias for certain parameters to deter fuzzing efficacy if an enforcement regime is sufficiently strict. This introduces a trade-off against "economic value" and is politically complicated, but I suspect product liability lawsuits and IP protection lawsuits will force these practices upon model product vendors regardless, so the only question remaining is if only the judiciary forces this and not the legislature or executive.

Catherine Blanche King's avatar

Mitchell Harper: Even a brief perusal of China's long history can tell an interested person that their leaders have drawn (a) what is excellent from their attractive but varied moral and religious traditions, (b) twisted them around, and (c) applied them to a "winner take all regardless of who dies in the process" mentality and set of political practices. (Just my take on the present situation in China. But be careful about what one likes about China, as is the same with the United States, only in a completely different philosophical context.)

Mitchell Harper's avatar

I mean Japan didn't help to set a good tone for post-revolutionary China given the millions of murders that set a pretty grim moral backdrop for social development and US officials talk openly and often about how they do war games where they would use a nuclear weapon against China if necessary to meet military objectives, so I think China is a pretty moderate country morally in terms of care for human death all things considered. I think Ali Kadri in The Accumulation of Waste: A Political Economy of Systemic Destruction (2023) put the record of the human toll of Chinese and US social development side by side in a way that strikes me as potentially less orientalist than what I glean from your comment.

Not that I want to be praising China for moral development, but if the metric is "disregard for who dies in the process," US imperialism has such a poor record on that front that it's hard to address that poor record adequately to where I don't even understand why we are having this conversation.

Catherine Blanche King's avatar

Michael Harper: And still, we can criticize the U.S. (presently) without getting thrown in jail. One of the problems that is endemic to democracy is that freedom without earning about responsibility leaves the worst among us to use democratic values to "game the system" and thereby gradually weaken its central source of power--in the people.

Another is that when things go well for decades, the people who didn't actually experience contrary political systems (like fascism during WWII) tend to assume it will always a be smooth sailing civil culture and so we relax the hold of a watchful " people" by reducing legitimate "guardrails," which opens to door to someone like Trump or, on blogs like this one, for people to make rightful comparisons that are unfortunately of much less significance than a comparison would be of the policies of the governing political orders under examination; and where one has the principles, but also where some who live here abuse them, and one is clearly waiting for the right time to implement their own.

I have always been involved in education in some way precisely because I think the only way to give support to "the people" in a democracy is through a good education with excellent political history and philosophy, and even though that qualification as fulfilled is still no guarantee.

Catherine Blanche King's avatar

FYI the below link is to a YouTube video of Demis Hassabis sharing the stage with Wendy Hall in interview just a couple of days ago. Long, but extremely informative and forward looking:

YG_85639_72014_WO_EN_tch_VDN_Vid_Res_1920x1080

schwortz's avatar

"Rapid progress". More bullshitting from tech bros to desperately try to prop up the bubble that's soon gonna deflate or burst. Meta's already been indicating they plan to sell off excess compute. Doesn't sound like a great investment if you have excess to sell.

Catherine Blanche King's avatar

schwortz: "More bullshitting," maybe from some of those "tech bros" around Hassabis, but I seriously doubt that Hassabis fits that analysis. But if you think that, on principle, it cannot happen that someone is genuine about what they say then you are not worth listening to because, for only one reason, you are also complicit and so suspect of holding that somewhat lesser view.

My own view is that, with Hassabis' willingness to say and to mean "I don't know," and his general upbeat attitude about everything else, is that he provides less naiveté and unwarranted optimism than he displays an outer (higher and dynamic) attitude that is in open conflict with a range of apparently bottomless lesser, even insidious views.

schwortz's avatar

That's a misrepresentation and strawman of what I said. I did not say that these changes cannot happen. I am merely calling out these claims as bs. If anything I concede that these claims and predictions are perhaps potentially conceivable, like discovering extraterrestrial life, but without further substantial evidence, they aren't worth a damn. Also, if you're going to misrepresent what I said, then you yourself are not worth listening to.

When I see the headline for his own Twitter post make broad bold proclamations like "AGI is probably only a few short years away" contrary to existing evidence of what we know from existing technologies and research and affiliated with DeepMind who in turn is affiliated with a big tech giant like Google, I know there's probably bullshit there. Someone is trying to raise the value of their company's stocks, investments, get public attention, and keep investors convinced that there's progress.

Moreover you didn't even bother to address my point of Meta selling excess compute. If there's so much progress and value, why would there even be a surplus, let alone a desire or need to sell it, of compute? These things simply do not add up.

Catherine Blanche King's avatar

schwortz: The end of your note is rightly a question--not a statement of fact, which is not exactly how your first statement reads: "More bullshitting from tech bros to desperately try to prop up the bubble that's soon gonna deflate or burst" and later: "I know there's probably bullshit there. Someone is trying to raise the value of their company's stocks, investments, get public attention, and keep investors convinced that there's progress." You probably suspect rightly, but you do not know.

You moved the goalpost from your first note to your second note anyway. I agree that what you say (second note) is suggestive, KNOWING what we do from some of their past behaviors as some pretty solid evidence; and it may even be true; but it's hardly knowledge ("I know") though you do moderate it THE SECOND TIME with "probably." Also, if "things don't add up" it's certainly cause to question further, and to be wary and even to keep looking for solid evidence.

In this case, however, I called the statements as you wrote them, and I think you probably misrepresented yourself--by omission--because like the rest of us, we are all a bit more reactionary than we prefer to be in a more civil national environment (if you are in the US like I am), and who are so very tired of being gaslighted and even royally screwed by people we trusted (and needed to trust) and who have left their spines, moral selves, and what's could be left of their self-respect in the toilet.

schwortz's avatar

Yes I admit I may have moved the goalpast a bit in moderating my tone, but even potential con artists deserve some benefit of the doubt. Moreover me characterizing the rhetoric used by supposed experts, be they Hassabis and Hinton, or Altman and Musk, as bullshitting does not suggest that the predictions or technologies being propped up cannot reach the goals or levels promised, however unlikely or absurd such outcomes may be. Nuclear threats posed by alleged nations (ie North Korea and Iran) can still be quite possible and foreseeable, even if they are unlikely and fringe scenarios. Similar logic applies here. It’s conceivable to imagine “human level intelligence” or AGI being achievable, even if current evidence and models show this to be highly unlikely and improbable.

At the risk of moving goalposts, that’s the conundrum that’s also bullshit and not worth acknowledging. Tech salesmen aren’t just propping up alleged success and predictions, they’re propping up predictions and claims that no one can truly definitively prove nor disprove. No one will truly have access to their models, the data used to train them, and all the inputs and puts and methodologies used, at least not all of that data and assuming it’s credible and not altered. At that point you may as well ask someone to prove or disprove the existence of god(s). You can’t due to absence of evidence.

Finally, yes. I am a US citizen and thoroughly exhausted and sick of rhetoric and hype propped up by big tech companies. There haven’t been meaningful results and they keep begging for more, all the while displacing workers, stealing data, and of course siphoning away money that could be better invested elsewhere, including on the incredibly wasteful and costly AI data centers.

So yes, in my mind and for many others, Demis Hassabis claims here among others made by him and other tech bros in the past, are indeed bullshit.

Catherine Blanche King's avatar

schwortz: As I read and watch Hassabis here and elsewhere, and it's not a fine point to make I think, Hassabis doesn't really make knowledge or truth claims--he's carefully speculating, and says so, and I think thoughtfully. I didn't go back and scour his every word, but I don't remember reading or hearing an "I know." (If it matters to you, correct me if I am mistaken in this.) And in my business, which is teaching and writing about philosophy, that kind of clarity matters.

More certainly, I think it pretty clear he wants to bring the best of what he, himself, has experienced in some notable others in order NOT to leave the field to the carnival barkers and self-serving sophists. Unless I am mistaken, again, he doesn't have the power to end the movement as a whole but wants to bring clarity and moral purpose to the argument that we all should do what we can to make it work for all concerned, rather than against us. If true, that's a far piece from being BS.

I also think there are many more people "out there" who have much more power than I or you, but who are genuinely interested in the same thing and that we, out here in blogland, haven't heard of, but that he has access to. And I don't think there is much of a choice here, anyway, though I do share your negative take, I don't think it's a given that everyone's BS-ing, which brings me around to my original note to you about making room for reasonable hope..

Purnima Gauthron's avatar

Gary: Sure, but one doesn't have to be knighted and/or laureated to endorse preflight safety checks. Pre-release model testing is of course necessary but it cannot determine whether a specific action is legally or institutionally admissible. Who is to maintain continuous oversight of deployed models, actions, incidents, and consequences when the FAA is already understaffed.

Demis is not only fully removed from the reality of "sand" (I guess he means silicon when future generation will not be of silicon) but fully misinformed on FINRA.

1/ FINRA is a self-regulatory organization run by Wall Street for Wall Street. Should Wall Street regulate itself?

2/ FINRA does not evaluate, endorse, or approve any products or securities. FINRA's job is to supervise public firms only around market activity, fair conduct, and disclosure of risks. I've been in finance for close to 20 years. You know when someone is grasping at straws when they start throwing out big organization names like NATO, UN, WHO, IMF, FINRA, etc and act like they fully understand and run it.

David Evanoff's avatar

Using FINRA as a benchmark for AI governance is an attempt to regulate tomorrow’s neural architectures with yesterday's compromised, rent-seeking corporate guilds.

John Kalafut PhD's avatar

umm... what?? Yes - there were SOO many Swiss inventors in the late 19th Century submitting 'inventions' that implicated molecular diffusion, quantized energy packets interacting with atoms (photo-electric effect) and extensions of Lorentzian invariance to reframe how we view the concept of time and space. Sure.

And then you think that Einstein and Kassapis may, could have or do act sometimes in ways that are self-serving motives vs. the common good? Like ANY human being EVER in the past 250,000+ years?? It is exactly because individuals can be non-trustworthy, and cabals of people (eg: the robber barons, current tech-bros) that our civilizations established social contract, the common law, and governing mechanisms that (in theory) strive to ensure a level-playing field. Insert Natural law in here for kicks, depending on one's disposition.

Now, we see in our current US and most of the world a return to authoritarianism, but even in that model, there are guardrails and norms that are established, often in complicated spheres of influence/kick-backs etc.. This is precisely what Gary and Kassapis are talking about.

We have seen the 'play' before when we let a few wealthy and privileged set the ground rules, determine even what game we are playing and there are then decades of societal upheaval. And whereas China has a regulatory mindset and frameworks, the "party" is just a larger cabal that is opaque and self-serving. Far from the ideals of Engels, Marx and even Lennon (the Beatle or the philosopher).

Richard Pinch's avatar

Does Hassabis (or anyone else for that matter) think that we have the terminology, theories, techniques, tools, tradecraft and technical workforce to do all these things? If so, where are they?

James Maconochie's avatar

Gary, we have connected a couple of times, starting with scale, and I have always believed we have many similar observations and concerns, so it is genuinely good to see this. Preflight testing is the upstream remedy. Most of what passes for AI governance intervenes after deployment, which is treating the symptom while conceding the disease.

One thing I would push on. The FDA analogy contains a piece the FINRA framing risks losing. Drugs do not simply pass a trial and walk free. There is pharmacovigilance afterward, meaning named humans who remain continuously accountable for what the product does in the world. My worry with formalization is that the test becomes the finish line. A checkpoint gets automated, a box gets ticked, and the humans who actually bear the consequences drift out of the loop exactly when the system scales.

So the question I keep landing on: in the model Hassabis proposes, once a model passes, who remains accountable, and for how long?

Hoping this is the turning point you say it could be. J.

John McIntire's avatar

I am suspicious of “funding would need to be substantial” without independent money because it leads to the companies paying for the tests and then hiding the results they dont like because they paid for them. There has to be an FDIC / FRB funding model — levies on the companies that develop and use the tools. The levies would fund the independent experts whose work would in the public domain and who would not be required to sign NDAs.

Catherine Blanche King's avatar

All: Below FYI is a "gift" article from the NYTimes about what's going on in AI and the China competition.

I'd like to say to those on this site who talk about choosing between the United States and China as they compare the horrific histories of both: If China were in charge politically, we'd all get arrested for saying what we say on this and probably other blogs--as another blogger wrote on this site as a the best prescient warning of the century.

But it occurred to me that those who take the view that they prefer China are reviving the infamous spirit of Neville Chamberlain as he came back from Germany waving a commitment paper signed by Hitler. I hope I am wrong in this as things actually play out, however . . .

https://www.nytimes.com/2026/07/17/business/china-ai-moonshot-kimi.html?unlocked_article_code=1.ylA.UqZ6.m46-xqslWLl1&smid=url-share