First, abstracting from all the words in my long analysis on OpenAI Astra yesterday, the overall point was that Astra may not be the breakthrough that OpenAI wants you to think it was.
Having worked at several VC-backed tech startups on a much smaller scale than the AI companies, I can confidently state that the pressure to produce results - or the appearance of them - and for everyone within the company to publicly toe the company line has to be immense at these firms. Way too much money at stake for everyone. Everything they say and claim has to be taken lightly and rigorously verified. Until proven, it's all marketing fluff.
That may depend on who has control. When I worked at a VC-funded biotech tools company creating tools to use gene expression to ascertain drug effects and toxicity, the scientists were very much ensuring that their results were robust and passed muster. The CEO and sales team may not have been so science-centric, but the scientists also attended sales meetings to show the product as well as test client samples, so there was no hype AFAICT.
My sense is that much of the fluffy/sensationlist science "reporting" on the web is based on PR reports manufactured to draw attention. How many breathless reports on some breakthrough or discovery prove a momentary flash that never pans out, prove a mistake (and sometimes manipulated results), and do not have the speculated potential claimed. The phrase "one day this may change [X]" is often a giveaway of hype.
We know the big AI companies do not have profitable business models and are desperate to raise the capital to keep the momentum going as well as cash out via IPOs. The pressure from investors to get out whole and preferably with a big payoff must be excruciating on management. Color me unimpressed when these sorts of stories keep getting trotted out to keep the investors happy.
And it seems, with few exceptions, those who have the most power presently are thinking from the least developmental state and order of human affairs. Theories that have emerged and that well-cover the natural and physical sciences do not and CANNOT explain what's wrong with that picture. (See my latest note in this blogpost.)
I’m referring to the institutional lies. I worked on the inside of big tech, and there would be plenty of debate internally. For scientists to praise their systems for searching the Internet for a solution, are they worthy of the benefit of the doubt?
Calling them liars is more charitable—it at least tacitly lets them live in a dreamworld in which they knew better rather than a more damning one, that they were too ignorant or too stupid to know better.
Which of those offends their fragile egos the least?
In any case, I don’t think their egos ought to be the thing we concern ourselves with. If they’re wrong, the harm is significantly worse.
I'm not sure I completely understand, but my point is simple: If you want someone to listen to you carefully and learn something and perhaps change their behavior, you do not call them a liar, and say that that is charitable. If you think they are hopelessly incorrigible, it's OK, although you run the risk of also turning away readers who recoil at calling people liars without evidence of intentionality on their part, which is generally part of the definition. (I did find it maddening that the mainstream media went for years without calling Trump a liar on the grounds he may not know the truth, and vividly remember the first time I heard it done, on CNN.)
The thing is I don’t expect them to listen to me, nor will they modify their behavior on critiques from the outside. I worked inside high tech corporations for years, and I’m resolute that these people are lying about what they really believe.
Besides, is there a kinder way to say that they’re liars? That they accidentally mislead? That they’re dupes? What magical formula is there for changing their views? The media already plays their rhythm uncritically. Why should I? Does it turn off those who aren’t apologists?
When a scientist brags that these systems solve a math problem, then reveals that it just looked up the solution online, how exactly are you suggesting we describe that? When the systems magically can solve the same set of math problems without any insider knowledge of what was tried and what wasn’t, do you honestly believe they’re telling you the truth? Do you believe you can “nicety” them into telling us what is really happening?
Well said, it's all useless stuff in the real world and considering trillions of investment and trillions more coming. Yeah they generated proofs for certain types of questions but failed on 100k other types. In fact we'll never know because there's no transparency
You make a good case in favor of government research. Privately funded research and startups have incentives in all the wrong places and preclude progress on rare diseases, new antibiotics, or long-term research goals which don't necessarily guarantee short-term profits.
There is a persistent pattern of looking at subjects in a highly reductive way, that OpenAI and others tend to show us.
Mathematics at this level isn’t about finding x=2, but it appears so because with LLMs everything is imagined as boiling down to problem / solution framing.
In fact we can see advanced math as often about uncovering new expressive possibilities, and so what becomes more interesting is the process of how a solution is uncovered than whether the problem is even solved.
Not to mention the real reason for doing this is to create yet another false anecdote, that if LLMs can solve erdos problems etc they can obviously do important things in the enterprise - which is still highly dubious.
Tao's issue parallels the issue in cybersecurity: the AIs can find bugs in software far faster than they can be remediated.
And a lot of those bugs are relatively trivial, to boot. Real bugs, but not necessarily easy to exploit. But they all have to be reviewed to determine which ones are really serious. And this is breaking a lot of software development teams, especially in the under/un-paid open source community.
There is a possibility that these problems had already been solved or that someone has a theory that gave the AI a way to get to the proof. They spend a huge amount of time hiring mathematicians to work on these and they scanned tons of rare books. The thing that upsets me is they usually say our model solved this: here is a counter to the proof, but we don’t get access to the logic it used even when they publish the findings. It also takes humans much longer to disprove their proofs or to come to a consensus so I mean did it really solve it? We won’t know until the proof is accepted by all mathematicians and not just ones that work for these companies.
Yes, if we think that resolving obscure math conjectures will have a major impact on the world this is important. But if all LLMs could do this, it would be a solid confirmation of the value of LLMs, even if it is rain on OpenAI's marketing campaign.
"How good Astra is in open-ended real-world problems very much remains to be seen." Let me be the first to reveal the answer: It won't be good at open-ended real-world problems. I'm open to a wager.
With reference to the cartoon Harari in his latest book Nexus makes the point that the truth is expensive and that lying (and by extension bullshitting) are cheap. AI makes bullshitting even cheaper but it doesn’t reduce the cost of truth.
AI actually increases the cost of truth because it drowns the truth in a sea of bullshit making it all the more difficult to find.
For every one instance of a correct AI math proof or correct AI science result, there are probably thousands of incorrect ones generated. And humans still have to sift through them all to decide which is which. The journals are being inundated with AI generated stuff these days and a lot of garbage is undoubtedly getting through because there are simply not enough human reviewers to cover the deluge.
That was my thought as well. Already we see a big issue in science publishing where the AI output overwhelms the system. It becomes very difficult to wade through it.
The key point is that one can not simply focus on the few cases where LLMs have performed well or even outstandingly well while ignoring everything else
But that is precisely what Altman, Amodei and the others would have us do.
I am also reminded with "GPT4 o3-preview", the first recursive ('thinking' - NOT) model with their strong math and ARC-AGI-1 results (with unlimited compute), both of which did not really materialise on the official release. It was clearly fine-tuned on the results it was creating.
I think one should just assume that they are lying about and concealing as much as possible, given their past displays. Not about the proofs, which I guess are probably true, if mathematicians have already checked them. But everything else is fair game for the most abject mendacity.
For instance, the claim that the cost was something like $2000. First, if your belief that that number is even loosely related to the truth is based on something like "Would AI companies lie to me?", then that question kind of answers itself. Even if it does have some relationship with the truth, you can't discard any kind of misleading element. Spend a million times that much, then present only the few proofs you were able to get, which happened to also be the cheap ones, and give the cost just for them? Sure, they would do that. Spend many times as much, get a bunch of proofs, but only report the cheapest ones? Sure, they would do that. Count only the "relevant" parts of the inference process when assessing cost? Sure, they would do that. Quote the public-facing loss-leader token cost rather than the actual cost to them to run a test? Oh, yes, they would certainly do that.
Did the Anthropic person say how much they needed to spend? Did they describe their testing procedure in detail? Would you trust them if they had? For instance, would you trust a random person on X when they say no Internet access was used, right after they admitted that they had failed to restrict Internet access, several times?
The assertion that the proofs are all new? They have gotten caught in the past passing off existing proofs as new, so don't believe that without proof. Don't assume that a chatbot couldn't plagiarize or take hints from some obscure proof in the many books they scanned, for instance.
Gary, excuse me, I saw a New York Times article about a 900 kg bomb dropped by the US Air Force on a house in Qeshm, far from any military site, where a family was reportedly killed. Could this have happened due to an AI hallucination?
Quite possibly, but on the other hand, "military intelligence" has always been an oxymoron. The same question was raised on the first day when the US launched Tomahawk missiles at a girls' school in Mnab, killing around 170 schoolgirls at a site which had not been a military site for sixteen years. Reporters from the US visited the site and saw absolutely zero military connections anywhere around it.
The US military was just exposed today as requesting in some meeting "new ideas". All the decades of planning the Iran war has failed miserably, so now they're just bombing target lists almost at random.
They can't damage Iranian missile launching sites because most of them are underground in mountains and tunnels and can be repaired within a day or so. Leaving Iran's 16,000 missiles and 40,000 drones fully available.
So they're just hitting civilian targets as a "punishment campaign", which as Professor Robert Pape from the University of Chicago who has studied air campaigns for thirty years has been saying "never works."
The military also considered doing these strikes of Trump's - that he chickened out on Saturday night - to reduce Iran's missile and drone arsenal. Professor Pape has demonstrated that this is next to impossible given the terrain in Iran and that drones are the size of a king-sized bed and can be hidden anywhere in an area larger than the size of California.
Short of using nuclear weapons, or a ground campaign that will require a million or more US/NATO soldiers for ten years, the US has NO military solution to Iran.
"Doing mathematics" is much more than solving certain types of problems that require nothing creative as inputs to the solution. David Mumford, a distinguished mathematician gave a great talk about this two years ago at the Fields Institute.
Ugly proofs of not exactly central problems isn't going to cut it. And LLMs still cannot do arithmetic or reliably reproduce lists. They aren't going to prove any results I will wish I had gotten to first.
I listen to Tao and Penrose a lot. They are good! I am not a mathematician or a physicist, I am an educator with a science bent on psychology and technology.
Having worked at several VC-backed tech startups on a much smaller scale than the AI companies, I can confidently state that the pressure to produce results - or the appearance of them - and for everyone within the company to publicly toe the company line has to be immense at these firms. Way too much money at stake for everyone. Everything they say and claim has to be taken lightly and rigorously verified. Until proven, it's all marketing fluff.
That may depend on who has control. When I worked at a VC-funded biotech tools company creating tools to use gene expression to ascertain drug effects and toxicity, the scientists were very much ensuring that their results were robust and passed muster. The CEO and sales team may not have been so science-centric, but the scientists also attended sales meetings to show the product as well as test client samples, so there was no hype AFAICT.
My sense is that much of the fluffy/sensationlist science "reporting" on the web is based on PR reports manufactured to draw attention. How many breathless reports on some breakthrough or discovery prove a momentary flash that never pans out, prove a mistake (and sometimes manipulated results), and do not have the speculated potential claimed. The phrase "one day this may change [X]" is often a giveaway of hype.
We know the big AI companies do not have profitable business models and are desperate to raise the capital to keep the momentum going as well as cash out via IPOs. The pressure from investors to get out whole and preferably with a big payoff must be excruciating on management. Color me unimpressed when these sorts of stories keep getting trotted out to keep the investors happy.
Alex Tolley: ". . . on who has control." Indeed.
And it seems, with few exceptions, those who have the most power presently are thinking from the least developmental state and order of human affairs. Theories that have emerged and that well-cover the natural and physical sciences do not and CANNOT explain what's wrong with that picture. (See my latest note in this blogpost.)
Yep—why would we believe them when they lie so frequently?
The grift is the game itself—how many rounds can you make it with double or nothing?
Like cancer, they’ll inhale the resources of the system until no one can work on anything in this space.
I returned to school to work on AI back in 2010. DL was about to take off, and I wasn’t impressed by it.
Years of working in elite high tech did not budge me a whisker off. Instead, I found plenty of interesting problems that do matter.
I hope there is a place for the neoLuddite techies such as me.
Are they lies, or do they believe it when they say it, even if the belief is formed
because their salary depends on it? This is an important distinction. If lying,
they are not receptive to data and reason. If they believe, they may be open. When
suddenly someone shifts and becomes an LLM critic, no one says "you were lying all
that time!" No, we realize that they believed it then and have been influenced by data.
But if we call believers who have not yet recanted "liars," they won't listen because they know we are dead wrong about them.
I’m referring to the institutional lies. I worked on the inside of big tech, and there would be plenty of debate internally. For scientists to praise their systems for searching the Internet for a solution, are they worthy of the benefit of the doubt?
Calling them liars is more charitable—it at least tacitly lets them live in a dreamworld in which they knew better rather than a more damning one, that they were too ignorant or too stupid to know better.
Which of those offends their fragile egos the least?
In any case, I don’t think their egos ought to be the thing we concern ourselves with. If they’re wrong, the harm is significantly worse.
I'm not sure I completely understand, but my point is simple: If you want someone to listen to you carefully and learn something and perhaps change their behavior, you do not call them a liar, and say that that is charitable. If you think they are hopelessly incorrigible, it's OK, although you run the risk of also turning away readers who recoil at calling people liars without evidence of intentionality on their part, which is generally part of the definition. (I did find it maddening that the mainstream media went for years without calling Trump a liar on the grounds he may not know the truth, and vividly remember the first time I heard it done, on CNN.)
The thing is I don’t expect them to listen to me, nor will they modify their behavior on critiques from the outside. I worked inside high tech corporations for years, and I’m resolute that these people are lying about what they really believe.
Besides, is there a kinder way to say that they’re liars? That they accidentally mislead? That they’re dupes? What magical formula is there for changing their views? The media already plays their rhythm uncritically. Why should I? Does it turn off those who aren’t apologists?
When a scientist brags that these systems solve a math problem, then reveals that it just looked up the solution online, how exactly are you suggesting we describe that? When the systems magically can solve the same set of math problems without any insider knowledge of what was tried and what wasn’t, do you honestly believe they’re telling you the truth? Do you believe you can “nicety” them into telling us what is really happening?
Well said, it's all useless stuff in the real world and considering trillions of investment and trillions more coming. Yeah they generated proofs for certain types of questions but failed on 100k other types. In fact we'll never know because there's no transparency
You make a good case in favor of government research. Privately funded research and startups have incentives in all the wrong places and preclude progress on rare diseases, new antibiotics, or long-term research goals which don't necessarily guarantee short-term profits.
Stop criticizing my frontier villages! -- Grigory Potemkin
There is a persistent pattern of looking at subjects in a highly reductive way, that OpenAI and others tend to show us.
Mathematics at this level isn’t about finding x=2, but it appears so because with LLMs everything is imagined as boiling down to problem / solution framing.
In fact we can see advanced math as often about uncovering new expressive possibilities, and so what becomes more interesting is the process of how a solution is uncovered than whether the problem is even solved.
Not to mention the real reason for doing this is to create yet another false anecdote, that if LLMs can solve erdos problems etc they can obviously do important things in the enterprise - which is still highly dubious.
Tao's issue parallels the issue in cybersecurity: the AIs can find bugs in software far faster than they can be remediated.
And a lot of those bugs are relatively trivial, to boot. Real bugs, but not necessarily easy to exploit. But they all have to be reviewed to determine which ones are really serious. And this is breaking a lot of software development teams, especially in the under/un-paid open source community.
There is a possibility that these problems had already been solved or that someone has a theory that gave the AI a way to get to the proof. They spend a huge amount of time hiring mathematicians to work on these and they scanned tons of rare books. The thing that upsets me is they usually say our model solved this: here is a counter to the proof, but we don’t get access to the logic it used even when they publish the findings. It also takes humans much longer to disprove their proofs or to come to a consensus so I mean did it really solve it? We won’t know until the proof is accepted by all mathematicians and not just ones that work for these companies.
How does frontier math change the day to day need for grounded reliability without hallucinations.
Do that instead. Make it good.
We have plenty of intelligence, fix the damn product. Watching AI weenies say they can be math weenies is trite.
Yes, if we think that resolving obscure math conjectures will have a major impact on the world this is important. But if all LLMs could do this, it would be a solid confirmation of the value of LLMs, even if it is rain on OpenAI's marketing campaign.
"How good Astra is in open-ended real-world problems very much remains to be seen." Let me be the first to reveal the answer: It won't be good at open-ended real-world problems. I'm open to a wager.
With reference to the cartoon Harari in his latest book Nexus makes the point that the truth is expensive and that lying (and by extension bullshitting) are cheap. AI makes bullshitting even cheaper but it doesn’t reduce the cost of truth.
AI actually increases the cost of truth because it drowns the truth in a sea of bullshit making it all the more difficult to find.
For every one instance of a correct AI math proof or correct AI science result, there are probably thousands of incorrect ones generated. And humans still have to sift through them all to decide which is which. The journals are being inundated with AI generated stuff these days and a lot of garbage is undoubtedly getting through because there are simply not enough human reviewers to cover the deluge.
That was my thought as well. Already we see a big issue in science publishing where the AI output overwhelms the system. It becomes very difficult to wade through it.
The key point is that one can not simply focus on the few cases where LLMs have performed well or even outstandingly well while ignoring everything else
But that is precisely what Altman, Amodei and the others would have us do.
Apres AI, le deluge!
I am also reminded with "GPT4 o3-preview", the first recursive ('thinking' - NOT) model with their strong math and ARC-AGI-1 results (with unlimited compute), both of which did not really materialise on the official release. It was clearly fine-tuned on the results it was creating.
I think one should just assume that they are lying about and concealing as much as possible, given their past displays. Not about the proofs, which I guess are probably true, if mathematicians have already checked them. But everything else is fair game for the most abject mendacity.
For instance, the claim that the cost was something like $2000. First, if your belief that that number is even loosely related to the truth is based on something like "Would AI companies lie to me?", then that question kind of answers itself. Even if it does have some relationship with the truth, you can't discard any kind of misleading element. Spend a million times that much, then present only the few proofs you were able to get, which happened to also be the cheap ones, and give the cost just for them? Sure, they would do that. Spend many times as much, get a bunch of proofs, but only report the cheapest ones? Sure, they would do that. Count only the "relevant" parts of the inference process when assessing cost? Sure, they would do that. Quote the public-facing loss-leader token cost rather than the actual cost to them to run a test? Oh, yes, they would certainly do that.
Did the Anthropic person say how much they needed to spend? Did they describe their testing procedure in detail? Would you trust them if they had? For instance, would you trust a random person on X when they say no Internet access was used, right after they admitted that they had failed to restrict Internet access, several times?
The assertion that the proofs are all new? They have gotten caught in the past passing off existing proofs as new, so don't believe that without proof. Don't assume that a chatbot couldn't plagiarize or take hints from some obscure proof in the many books they scanned, for instance.
Astra-damus?
Or Asta-damnus?
Gary, excuse me, I saw a New York Times article about a 900 kg bomb dropped by the US Air Force on a house in Qeshm, far from any military site, where a family was reportedly killed. Could this have happened due to an AI hallucination?
Quite possibly, but on the other hand, "military intelligence" has always been an oxymoron. The same question was raised on the first day when the US launched Tomahawk missiles at a girls' school in Mnab, killing around 170 schoolgirls at a site which had not been a military site for sixteen years. Reporters from the US visited the site and saw absolutely zero military connections anywhere around it.
The US military was just exposed today as requesting in some meeting "new ideas". All the decades of planning the Iran war has failed miserably, so now they're just bombing target lists almost at random.
They can't damage Iranian missile launching sites because most of them are underground in mountains and tunnels and can be repaired within a day or so. Leaving Iran's 16,000 missiles and 40,000 drones fully available.
So they're just hitting civilian targets as a "punishment campaign", which as Professor Robert Pape from the University of Chicago who has studied air campaigns for thirty years has been saying "never works."
The military also considered doing these strikes of Trump's - that he chickened out on Saturday night - to reduce Iran's missile and drone arsenal. Professor Pape has demonstrated that this is next to impossible given the terrain in Iran and that drones are the size of a king-sized bed and can be hidden anywhere in an area larger than the size of California.
Short of using nuclear weapons, or a ground campaign that will require a million or more US/NATO soldiers for ten years, the US has NO military solution to Iran.
None.
"Doing mathematics" is much more than solving certain types of problems that require nothing creative as inputs to the solution. David Mumford, a distinguished mathematician gave a great talk about this two years ago at the Fields Institute.
http://www.fields.utoronto.ca/talks/How-will-AI-affect-mathematical-research-Online
Ugly proofs of not exactly central problems isn't going to cut it. And LLMs still cannot do arithmetic or reliably reproduce lists. They aren't going to prove any results I will wish I had gotten to first.
Why does no one talk about the 2025 Deep Seek Lean-Prover LLM? *Last year* I did a huge chunk of Lean 4 theorem-proving with that just to learn Lean!
What about other math LLMs like Harmonic’s Aristotle or Axiom? Axiom has the guidance of an incredibly famous mathematician, Ken Ono.
Curious that everyone hangs on Altman & Dario’s every word while ignoring the LLM work of more serious & renowned mathematicians.
Why don’t the real experts doing serious math work, & not just searching for the easily tractable items, get airtime?
I listen to Tao and Penrose a lot. They are good! I am not a mathematician or a physicist, I am an educator with a science bent on psychology and technology.
Why are you so skeptical of Sam I Am's "undertaking of great advantage but nobody to know what it is"?
I do not like green eggs and HAIm.