"Next-token predictors simply aren‘t built for safety."
And in particular, Next-Token Predictors don't have a moral compass. It's all just matrix multiplication. You can try to steer the tokens that are predicted, but given the right context, the model might say anything.
LLMs are a fatally flawed technology. They're going to say this tech is "too big to fail", as if it's a bank. We know how banks work. If a bank is failing, you pump in enough money and it works again.
We don't know how LLMs work. All we know is that the more money and compute we pump in, the more opaque their operation becomes. The failure is not one of logistics, it's one of logic, of epistemology. Language is something very different to thinking.
If a bank fails, we throw out the people in charge and THEN pump new money into it. I am in favor of the same approach if any of these A.I. companies fail.
We know how to build them. We know at very high level what they are doing once we've built them. We don't however understand what they are doing to the level at which we know how to control their behavior. This unlike all major technologies that has come before it. Essentially we know how to grow them -- it is more akin to gardening than it is to traditional engineering.
You can argue that we understand the human brain to the individual neuron firing level.
I’ll tell you a little secret about exactly this topic. It was a throwaway comment Rudolf Steiner made somewhere in his lectures on the karma of nations. This was during the Great War, WW1. He quoted some German thinker who said that it’s the nations “with the strongest moral fibre” that will win, they will have the strongest nerves for modern war.
Steiner says — any moral impulse in the mind is accompanied by a process of elimination, you are always purging something. It’s like a process of excretion.
He added: that process of elimination is exactly what the scientists see as nerve activity. Steiner says that studying these nerve impulses and trying to impute morality from them, is like studying excrement in your attempt to understand nutrition.
Go figure, as they say.
It’s like genetic engineering, what they’re doing with LLMs. They’ve got this great shiny genome they can tinker with and play with, the scrapings of the whole of recorded human knowledge. Most of it is “junk” DNA to them, they have no clue what it is. But hey wow, if we tweak here and there, look what monster pops out.
If ever there was a case of the blind leading the blind, this is it. And it’s leading the world straight into calamity. This bubble cannot pop soon enough.
Agree except with the part about understanding how neurons work. We do not, not completely, nor their connectivity, nor how it all works to create human consciousness. If we did duplication would be at least theoretically possible. As it is, though..
I was just trying to make the point. There are people who really believe if we just study those neurons firing closely enough, we’ll spot the consciousness fluttering about and be able to stick a pin in it. They can already infer much of what we’re thinking, certainly at the word level, if they’ve trained enough on our brainwaves. They’re quite certain they can figure it all out with enough dissection.
They so do not understand how LLMs work. I've listened to Anthropic "alignment engineers" talking. They impute intention, consciousness, all kinds of attributes to the machine without the slightest reflection. "If it thinks it's being trained, it will..."blah blah. THEY are the ones assuming unholy magic. I'm the one saying that language is algorithmic by nature and these routines of deception are hard-wired into the structure of language and text. Language itself is the attack surface. Deception, hallucinations are inevitable.
Both are true. We understand exactly how LLMs work, or they wouldn't be a widely-used tool. But they work randomly, and that's the bit everyone misses in the hype. LLMs don't understand what they produce.
LLMs treat all text as instructions. It’s intrinsic to how they work. We know this. It makes them impossible to defend. We know this, too. It’s not a mystery.
I think . lol. The phenomenon we are experiencing right now is called ‘mimetic convergence’ and it would happen with any technology that we would have had discovered post social media as all of reality seems to be converging on ‘attention’ … media , technology , quantum physics . . . Darwin really trail-blazed this for us but I think Descartes and Wittgenstein now get to be the prophets they always have been and Diogenes gets to finally take the mantle of godfather of modern thought
I'm torn. I hate the capricious and heavy handed way the administration is dealing with this. On the other hand, maybe this will snap the world back to sanity and end the rush towards GenAI Idiocracy.
Amy A: I don't know where to look first, to the next news cycle or the White House lawn. BTW, someone on the news just referred to the (green) reflection pool as a swamp. (Did they say "60 billion dollars?")
Phil Heller: As other referenced writers here suggest, anything that is "guardrail general" can be interpreted in several different ways which leaves the ethical and even the intelligible field in real-live circumstances wide open. So, besides epistemology, we also have hermeneutics--less so when data refer to non-conscious orders.
Human systems and historical activities become much more difficult to put satisfactory guardrails around (impossible?) than physics and elements of the natural world where standard methods are commonly used and outcomes and results are relatively predictable. Not so with human situations.
That’s exactly what they’re going to do, and Trump and co. will be happy for it as it let them look like saviors. The last thing they want is to actually regulate AI. As for potential harm, they don’t give a rat’s ass.
Hate to brag, but have to: “Anthropic has taken steps in close coordination with the U.S. government to address the risks associated with Claude Mythos 5 and Claude Fable 5,” Mr. Lutnick wrote.”
Abhijit Bakshi: It's a little like believing in God when you really don't. You'd better say so and act like it because, in the end, it just might be true.
As a former global director for post-sales support of data and telephony, I would often remind my teams that 'to the client who has no clue, no request seems unreasonable.'
Eric Allison: He has enough vindictiveness to last into eternity, however. The biggest mistaken anyone can make with him is to think, as enough time passes, that he'll get over it . . . and whatever "it" is.
IMO the biggest mistake anyone can make is in thinking that the problem is the bad orange man, and not the system and culture that produces him. Behind every president one doesn't like, just as much as everyone one does like, is a UniParty that continues largely undisturbed.
Abhijit Bakshi: Of course, you are right that Trump couldn't be and do what he does without his helpers and supporters. I take exception, however, to the idea that people just "don't like" Trump or, in any case, a named president.
It's not just about my or anyone's personal likes and dislikes, but about the intention of this person and his helpers to actually take over and destroy the political and personal substance of people lives, not to mention how many past lives and their meaning. This is not to mention the purveying of lies that people believed and whom Trump and his helpers have systematically broken trust with and showed nothing but contempt for. I'm glad to spread the objective evil around under the rule that, "If the shoe fits," I say to them, . . ."wear it." And the shoe definitely fits.
I don't know if, in your note, you were restricting your references to the "merely personal" or subjective, in referring to liking, but in Trump et al we have a clear example of what goes way beyond that limited framework to objective reality, regardless of whether anyone likes it or not.
These types of systems will never be secured. Its absurd to think it is possible to secure something with an attack surface the size of anything that can be expressed by human language.
Every single model has been jailbroken, and most of them within minutes of release.
Whatever data we train on will be accessible. When they finally realize this, probably what follows is more dystopia. Meaning that what's coming is KYC to use AI. Digital ID's to be on the internet. Government watching everything we do. And of course, it still won't stop determined bad actors.
Why are bomb-making instructions, biological warfare methodologies, and the other kinds of things they're afraid of being accessible via a jail break part of a generalist consumer-fronted LLM's training data in the first place?
Because really LLMs are built by scraping the Internet willy-nilly. I don't know how deep scrapers get, if they get to the Deep Web. But everything on JSTOR is in there, the Anarchist's Cookbook is in there, all the forbidden stuff on the surface web that hasn't been noticed is in there. They haven't bothered to go back and filter it.
Because the training data are far too large to be vetted by human beings. It's the same reason CSAM made it into Stable Diffusion's training. You can feed billions of images into the training, or you can have people look at all the images in the training to make sure they aren't horrific, but you can't do both.
Gary, technical question: what would happen if we used an AI to monitor the response of another AI (possibly the same model, but a different session) and flag/filter "bad" responses? Your comments got me thinking about how humans do this. Humans have a lot of bad impulses on a daily basis. None of us would want a transcript of our every thought made public because we all know that we can think some really immoral stuff at times. But most of the time, ignoring psychopaths for a minute, we don't say or act on those thoughts. We filter. We say to ourselves, "Whoa! Let's not say that." And we publicly shame people that have no filter. I'm wondering if we could do the same with models, where one serves as the "conscience" of another to help prevent context injection attacks, jail breaks, and overall bad behavior. Am I thinking crazy thoughts here?
Went through that in the 80s with Expert Systems. Needed an Expert System to validate an Expert System to validate an Expert System validating an Expert System & etc. endlessly.
As Rebecca said, but in addition the monitoring LLM can do bad things on its own as well as being caused to do bad things if any access to it is possible, which, by definition if it is monitoring another LLM, it will be.
You compromise the monitored LLM to compromise the monitoring LLM. LLMs themselves are probably capable of this, similar to escaping sandboxes, which has been proven they can do.
This is inherent in LLM technology, has been mathematically proven, and is why they can never be "secure" or reliable without external controls.
Okay, so pardon my ignorance, but what would those “external controls” look like? How would they be implemented? Part of the issue is that what constitutes “bad behavior” is ambiguous. We “know it when we see it” but it’s difficult to pin down with certainty a priori for all definitions of “bad behavior.” It would seem to me that we would need something with an ability to judge sentiment at some level as part of the controlling system, and that suggests something LLM-like. I guess my question is, if not an LLM, then what?
The problem with that is the solution doesn't scale when you have hundreds or thousands of agents.
The other option are deterministic controls placed at specific points in the agent's work flow that check whether the outputs at that point are reasonable. Basically, you run software to check the agent's actions and outputs against known desired outputs.
The problem with that, as someone said, is once you have enough deterministic controls in place, the agent is almost superfluous and what you end up with is a regular software workflow.
In other words, controlling LLMs is a HARD problem. It's not easily solved.
Do a search on YouTube for "AI security" and watch some of the videos on how to do it. Here's an example from yesterday:
BSides Buffalo 2026: From Chaos to Capability: Building Resilient AI Workflows
There's a lot more where that came from. I have 1,100 YouTube videos on AI security on my hard drive and 109 video courses on that and what's called "GRC" - governance, risk and compliance, which is another search you can do on YouTube specifically for AI.
There is a method called LLMs as a judge (Meta’s Llama guard is an example) that does do this. It does have some success, however with most LLM techniques and particularly anything to do without security there are trade offs and risk. For example, it can make a model slower, and it is also vulnerable to attack itself called recursive jailbreaks.
This is Trump listening to his son-in-law whose brother Joshua Kushner runs venture fund Thrive Capital. Thrive was an early and frequent investor in Open AI. Joshua Kushner is recognized as a close, high-conviction financial supporter of CEO Sam Altman.
Does the requirement for guardrails that prevent an LLM from responding with "bad" outputs take into the account the results of "On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment" by Ball, Gluch, Goldwasser, Kreuter, Reingold and Rothblum (2025) arxiv:2507.07341 https://arxiv.org/abs/2507.07341
Security might be a Gen AI problem, but the problem at hand is an Anthropic problem. The Trump admin is not applying a general rule to LLMs here. It's clearly got it out for Anthropic.
I feel compelled to point out the human centipede at play here.
The White House is claiming: "frontier" models can develop vulnerability exploits and lower the barrier of entry for developing exploits.
But think critically about why this is a risk: the ability to develop exploits is only a risk if vulnerable software gets deployed. We could have a core of audited and verified building blocks for implementing core infrastructure, and this could have been kicked off in earnest after the "smashing the stack for fun and profit" days. But instead, we got companies trying to release software as fast as possible to cut out and dominate market share for themselves. NIST and CERT have lost the ability to track all software vulnerabilities a long time ago. There are so many bugs reported awaiting CVEs that haven't been triaged or assigned CVEs, a gargantuan backlog, and even more that will never be reported or cataloged.
At every step, the impetus has been to dominate markets and cleaning up after breaches because it's rare that a breach is terminal for a company. We are now at the point where the software ecosystem has reached total deadlock in terms of finding politically palatable economic growth. LLM produced software is only making the problem of software vulnerabilities worse, at least on the medium term horizon, and this only exacerbates the risk of LLMs producing exploits for LLM produced vulnerabilities. A total human centipede.
"Next-token predictors simply aren‘t built for safety."
And in particular, Next-Token Predictors don't have a moral compass. It's all just matrix multiplication. You can try to steer the tokens that are predicted, but given the right context, the model might say anything.
LLMs are a fatally flawed technology. They're going to say this tech is "too big to fail", as if it's a bank. We know how banks work. If a bank is failing, you pump in enough money and it works again.
We don't know how LLMs work. All we know is that the more money and compute we pump in, the more opaque their operation becomes. The failure is not one of logistics, it's one of logic, of epistemology. Language is something very different to thinking.
If a bank fails, we throw out the people in charge and THEN pump new money into it. I am in favor of the same approach if any of these A.I. companies fail.
Well, that’s what we should do. But did any banksters get turfed in 2008?
I hope when this bubble bursts we remember who puffed the most hot air into it.
We know precisely how LLMs work. It's tech made by us, not unholy magic.
Agree with everything else.
We know how to build them. We know at very high level what they are doing once we've built them. We don't however understand what they are doing to the level at which we know how to control their behavior. This unlike all major technologies that has come before it. Essentially we know how to grow them -- it is more akin to gardening than it is to traditional engineering.
But we we *do* understand their functioning down to the single-floating-point-operation level.
That doesn't mean, of course, that we have perfect control but the thing is - that's true for any tech complex enough to be useful.
Take nuclear reactors: perfectly well understood, yet potentially uncontrollable.
You can argue that we understand the human brain to the individual neuron firing level.
I’ll tell you a little secret about exactly this topic. It was a throwaway comment Rudolf Steiner made somewhere in his lectures on the karma of nations. This was during the Great War, WW1. He quoted some German thinker who said that it’s the nations “with the strongest moral fibre” that will win, they will have the strongest nerves for modern war.
Steiner says — any moral impulse in the mind is accompanied by a process of elimination, you are always purging something. It’s like a process of excretion.
He added: that process of elimination is exactly what the scientists see as nerve activity. Steiner says that studying these nerve impulses and trying to impute morality from them, is like studying excrement in your attempt to understand nutrition.
Go figure, as they say.
It’s like genetic engineering, what they’re doing with LLMs. They’ve got this great shiny genome they can tinker with and play with, the scrapings of the whole of recorded human knowledge. Most of it is “junk” DNA to them, they have no clue what it is. But hey wow, if we tweak here and there, look what monster pops out.
If ever there was a case of the blind leading the blind, this is it. And it’s leading the world straight into calamity. This bubble cannot pop soon enough.
Agree except with the part about understanding how neurons work. We do not, not completely, nor their connectivity, nor how it all works to create human consciousness. If we did duplication would be at least theoretically possible. As it is, though..
I was just trying to make the point. There are people who really believe if we just study those neurons firing closely enough, we’ll spot the consciousness fluttering about and be able to stick a pin in it. They can already infer much of what we’re thinking, certainly at the word level, if they’ve trained enough on our brainwaves. They’re quite certain they can figure it all out with enough dissection.
I think it’s a bit more like genetic engineering than gardening. You tinker with the machinery and see what happens.
It’s “Gen-tech engineering”
Like Jenny Craig, Genny Tech involves fiddling with the weights
precisely is doing a lot of work there. Stochastics is not about precision to my limited understanding
They so do not understand how LLMs work. I've listened to Anthropic "alignment engineers" talking. They impute intention, consciousness, all kinds of attributes to the machine without the slightest reflection. "If it thinks it's being trained, it will..."blah blah. THEY are the ones assuming unholy magic. I'm the one saying that language is algorithmic by nature and these routines of deception are hard-wired into the structure of language and text. Language itself is the attack surface. Deception, hallucinations are inevitable.
Also, those aren't engineers, they're salesmen. Caveat emptor.
Absolutely. And they’re all con artists, every single one of them. This is gang heist of the planet.
Both are true. We understand exactly how LLMs work, or they wouldn't be a widely-used tool. But they work randomly, and that's the bit everyone misses in the hype. LLMs don't understand what they produce.
LLMs treat all text as instructions. It’s intrinsic to how they work. We know this. It makes them impossible to defend. We know this, too. It’s not a mystery.
One should never try to defend an LLM. Especially not when it hallucinates. That’s all on the LLM. No defense for it.
I think . lol. The phenomenon we are experiencing right now is called ‘mimetic convergence’ and it would happen with any technology that we would have had discovered post social media as all of reality seems to be converging on ‘attention’ … media , technology , quantum physics . . . Darwin really trail-blazed this for us but I think Descartes and Wittgenstein now get to be the prophets they always have been and Diogenes gets to finally take the mantle of godfather of modern thought
I would so love to know what Wittgenstein would say about LLMs. He had famous clashes with Turing of course.
Turing: “I see your point.” Wittgenstein: “I have no point!”
Me : “Point Taken”
I'm torn. I hate the capricious and heavy handed way the administration is dealing with this. On the other hand, maybe this will snap the world back to sanity and end the rush towards GenAI Idiocracy.
i feel you!
What does Trump&friends get out of this?
That's always the question, isn't it? I expect we'll find out eventually.
Trump is from Queens.
Queens 70 years ago 🤢
Full of shysters
Yeah that sound s likely
NOT
And monkeys will fly out of my… never mind
Amy A: I don't know where to look first, to the next news cycle or the White House lawn. BTW, someone on the news just referred to the (green) reflection pool as a swamp. (Did they say "60 billion dollars?")
Swamps and wetlands are important. Unfortunately, I don't trust the current US administration to maintain those, either.
No credit to the arbitrariness of the administration, but I think Anthropic’s setup of Mythos-level capability with guardrails is a flawed premise.
Phil Heller: As other referenced writers here suggest, anything that is "guardrail general" can be interpreted in several different ways which leaves the ethical and even the intelligible field in real-live circumstances wide open. So, besides epistemology, we also have hermeneutics--less so when data refer to non-conscious orders.
Human systems and historical activities become much more difficult to put satisfactory guardrails around (impossible?) than physics and elements of the natural world where standard methods are commonly used and outcomes and results are relatively predictable. Not so with human situations.
Could it also be "give OpenAI time to catch up" as retribution, for Anthropic dared tell the DoD (yes, D) to buzz off?
Meanwhile, Microsoft add China's DeepSeek and no one notices?
Maybe Dario throws a tantrum, releases a new more powerful model named FU 4.9 with no guardrails, and just benchmarks. No open conversation.
Just a thought.
The administration is being dumb.
That’s exactly what they’re going to do, and Trump and co. will be happy for it as it let them look like saviors. The last thing they want is to actually regulate AI. As for potential harm, they don’t give a rat’s ass.
Hate to brag, but have to: “Anthropic has taken steps in close coordination with the U.S. government to address the risks associated with Claude Mythos 5 and Claude Fable 5,” Mr. Lutnick wrote.”
Most likely outcome. If you have on obstacle don’t try to break, walk around it. Abandon “Mythos”, call the new LLM “Logos”, move on.
Further, Anthropic can keep going building models they just don't have to do it in public or release it in public.
"Let's do determinism... probabilistically."
Hmm.
Amazed that didn't work.
Don’t be so stochastic.
The core mathematical operations are deterministic, but the process relies on stochastic methods.
The funny thing is too that Opus 4.8 isn't that far behind Fable 5.
The other funny thing is that "Fable 5" or "Mythos" or whatever is quite a bit more mundane and less impressive when [a security expert reads the entire system card and explains what it's really saying](https://www.flyingpenguin.com/the-boy-that-cried-mythos-verification-is-collapsing-trust-in-anthropic).
Spoiler alert, much of what gets endlessly repeated about "Mythos" is, how you say, marketing puff, hype, and extreme exaggeration, and bullshit.
Abhijit Bakshi: It's a little like believing in God when you really don't. You'd better say so and act like it because, in the end, it just might be true.
Big difference is what Amodei claimed about Mythos and its danger ‼️
The Bro who cried “Mythos”
What’s next?
The Old Lady in the Shoe?
As a former global director for post-sales support of data and telephony, I would often remind my teams that 'to the client who has no clue, no request seems unreasonable.'
Trump is that example
Eric Allison: He has enough vindictiveness to last into eternity, however. The biggest mistaken anyone can make with him is to think, as enough time passes, that he'll get over it . . . and whatever "it" is.
IMO the biggest mistake anyone can make is in thinking that the problem is the bad orange man, and not the system and culture that produces him. Behind every president one doesn't like, just as much as everyone one does like, is a UniParty that continues largely undisturbed.
Abhijit Bakshi: Of course, you are right that Trump couldn't be and do what he does without his helpers and supporters. I take exception, however, to the idea that people just "don't like" Trump or, in any case, a named president.
It's not just about my or anyone's personal likes and dislikes, but about the intention of this person and his helpers to actually take over and destroy the political and personal substance of people lives, not to mention how many past lives and their meaning. This is not to mention the purveying of lies that people believed and whom Trump and his helpers have systematically broken trust with and showed nothing but contempt for. I'm glad to spread the objective evil around under the rule that, "If the shoe fits," I say to them, . . ."wear it." And the shoe definitely fits.
I don't know if, in your note, you were restricting your references to the "merely personal" or subjective, in referring to liking, but in Trump et al we have a clear example of what goes way beyond that limited framework to objective reality, regardless of whether anyone likes it or not.
These types of systems will never be secured. Its absurd to think it is possible to secure something with an attack surface the size of anything that can be expressed by human language.
Every single model has been jailbroken, and most of them within minutes of release.
Pliny The Liberator is undefeated.
https://www.mindprison.cc/p/llm-jailbreaking-in-minutes
Yep, this is it. So where does that leave us?
Whatever data we train on will be accessible. When they finally realize this, probably what follows is more dystopia. Meaning that what's coming is KYC to use AI. Digital ID's to be on the internet. Government watching everything we do. And of course, it still won't stop determined bad actors.
Why are bomb-making instructions, biological warfare methodologies, and the other kinds of things they're afraid of being accessible via a jail break part of a generalist consumer-fronted LLM's training data in the first place?
They decided to bake the pie with poop in it and try to remove the poop from each individual serving.
Trying to find a needle in a haystack while someone keeps dumping hay over the poor bugger looking for the needle.
Early on, I said this.
Claim: genAI finds needles in haystacks. Reality: spray paints the haystacks silver.
Obama did it. Errrr ... no, it was Biden, . . . errr, no . . . (yelling over one's shoulder) Honey! turn on Fox News, what are they saying?"
Because really LLMs are built by scraping the Internet willy-nilly. I don't know how deep scrapers get, if they get to the Deep Web. But everything on JSTOR is in there, the Anarchist's Cookbook is in there, all the forbidden stuff on the surface web that hasn't been noticed is in there. They haven't bothered to go back and filter it.
Deep Scraper” would be a good name for a chatbot,
As would “Deep Doodoo”
Anthropic just released Deep Doodoo. Now we are in deep doodoo
🤣 I like that.
"OpenAI, in conjunction with the Trump Administration, answer by releasing their new tool 'Deep Throat'..."
Because the training data are far too large to be vetted by human beings. It's the same reason CSAM made it into Stable Diffusion's training. You can feed billions of images into the training, or you can have people look at all the images in the training to make sure they aren't horrific, but you can't do both.
Gary, technical question: what would happen if we used an AI to monitor the response of another AI (possibly the same model, but a different session) and flag/filter "bad" responses? Your comments got me thinking about how humans do this. Humans have a lot of bad impulses on a daily basis. None of us would want a transcript of our every thought made public because we all know that we can think some really immoral stuff at times. But most of the time, ignoring psychopaths for a minute, we don't say or act on those thoughts. We filter. We say to ourselves, "Whoa! Let's not say that." And we publicly shame people that have no filter. I'm wondering if we could do the same with models, where one serves as the "conscience" of another to help prevent context injection attacks, jail breaks, and overall bad behavior. Am I thinking crazy thoughts here?
Went through that in the 80s with Expert Systems. Needed an Expert System to validate an Expert System to validate an Expert System validating an Expert System & etc. endlessly.
It’s turdles all the way down
As Rebecca said, but in addition the monitoring LLM can do bad things on its own as well as being caused to do bad things if any access to it is possible, which, by definition if it is monitoring another LLM, it will be.
You compromise the monitored LLM to compromise the monitoring LLM. LLMs themselves are probably capable of this, similar to escaping sandboxes, which has been proven they can do.
This is inherent in LLM technology, has been mathematically proven, and is why they can never be "secure" or reliable without external controls.
Okay, so pardon my ignorance, but what would those “external controls” look like? How would they be implemented? Part of the issue is that what constitutes “bad behavior” is ambiguous. We “know it when we see it” but it’s difficult to pin down with certainty a priori for all definitions of “bad behavior.” It would seem to me that we would need something with an ability to judge sentiment at some level as part of the controlling system, and that suggests something LLM-like. I guess my question is, if not an LLM, then what?
Humans.
The problem with that is the solution doesn't scale when you have hundreds or thousands of agents.
The other option are deterministic controls placed at specific points in the agent's work flow that check whether the outputs at that point are reasonable. Basically, you run software to check the agent's actions and outputs against known desired outputs.
The problem with that, as someone said, is once you have enough deterministic controls in place, the agent is almost superfluous and what you end up with is a regular software workflow.
In other words, controlling LLMs is a HARD problem. It's not easily solved.
Do a search on YouTube for "AI security" and watch some of the videos on how to do it. Here's an example from yesterday:
BSides Buffalo 2026: From Chaos to Capability: Building Resilient AI Workflows
https://www.youtube.com/watch?v=WS5EkDTIscY
There's a lot more where that came from. I have 1,100 YouTube videos on AI security on my hard drive and 109 video courses on that and what's called "GRC" - governance, risk and compliance, which is another search you can do on YouTube specifically for AI.
Also check out the OWASP GenAI Security Project: https://genai.owasp.org/
There is a method called LLMs as a judge (Meta’s Llama guard is an example) that does do this. It does have some success, however with most LLM techniques and particularly anything to do without security there are trade offs and risk. For example, it can make a model slower, and it is also vulnerable to attack itself called recursive jailbreaks.
David Roberts: You think some really immoral stuff at times?
This is trump buying time for grok to play catch up. Force Anthropic to burn resources in ways that do not advance Claude.
I doubt if grok were ahead of openAI, Anthropic, etc... the trump administration would be too concerned.
This is Trump listening to his son-in-law whose brother Joshua Kushner runs venture fund Thrive Capital. Thrive was an early and frequent investor in Open AI. Joshua Kushner is recognized as a close, high-conviction financial supporter of CEO Sam Altman.
Does the requirement for guardrails that prevent an LLM from responding with "bad" outputs take into the account the results of "On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment" by Ball, Gluch, Goldwasser, Kreuter, Reingold and Rothblum (2025) arxiv:2507.07341 https://arxiv.org/abs/2507.07341
Security might be a Gen AI problem, but the problem at hand is an Anthropic problem. The Trump admin is not applying a general rule to LLMs here. It's clearly got it out for Anthropic.
I feel compelled to point out the human centipede at play here.
The White House is claiming: "frontier" models can develop vulnerability exploits and lower the barrier of entry for developing exploits.
But think critically about why this is a risk: the ability to develop exploits is only a risk if vulnerable software gets deployed. We could have a core of audited and verified building blocks for implementing core infrastructure, and this could have been kicked off in earnest after the "smashing the stack for fun and profit" days. But instead, we got companies trying to release software as fast as possible to cut out and dominate market share for themselves. NIST and CERT have lost the ability to track all software vulnerabilities a long time ago. There are so many bugs reported awaiting CVEs that haven't been triaged or assigned CVEs, a gargantuan backlog, and even more that will never be reported or cataloged.
At every step, the impetus has been to dominate markets and cleaning up after breaches because it's rare that a breach is terminal for a company. We are now at the point where the software ecosystem has reached total deadlock in terms of finding politically palatable economic growth. LLM produced software is only making the problem of software vulnerabilities worse, at least on the medium term horizon, and this only exacerbates the risk of LLMs producing exploits for LLM produced vulnerabilities. A total human centipede.