Wow, how long can this house of cards last. I've had the same experience with o3. It seems to hallucinate worse than other top models, just with greater sophistication. I also found it claiming to have gathered information from the web or databases when it clearly did no such thing.
And Anthropic recently released a paper that explores the internal "reasoning" steps of these models. I think this sufficiently exposes the lack of any real reasoning if it wasn't already obvious.
The Cyc Project has been failing since 1984. Using that as our Test Case we can confidently predict as long as suckers ... er .... "investors" ... can be conned ... um ... "persuaded" ... to keep chucking money at 'em OpenAi can keep on indefinitely.
In defence of the comparison, both projects were/are lead by people who sincerely believe that their approach will lead to systems able to reason and solve problems like humans. In both projects there have been advocates and sceptics, with advocates energetically, and quite reasonably, seeking funding to support their work. In specific defence of Cyc, the idea looked a lot more reasonable at the time, over 40 years ago, than it does today. I feel we need to wait until 2066 to be sure how many orders of magnitude off base the comparison really is.
I heard Lenat present at AAAI-84 in Austin which coincided with the beginning of Cyc. He was over optimistic but nowhere near the charlatan of today's AI leaders. From what I've seen he'd been fairly pragmatic over the decades since then. I would guess Cyc's cost/benefit ratio to be far better than today's AI companies, and today's companies appear to be heading in the wrong direction.
Impressive post, and a lot for me to agree with. I think fundamentally my problem with the house-of-cards takes or the this-isn't-really-intelligence takes is that they prove too much. I was right there with you until 2022. I predicted that, for all the reasons you give, LLMs would not ever be able to, say, play a coherent chess game. They might spit out entirely plausible strings of chess notation, but they wouldn't actually understand the game or have any way to generate moves that made sense. And similarly for a laundry list of capabilities (common sense reasoning, explaining jokes, solving never-before-seen math problems).
I'm totally sympathetic to the point that, as you put it in your post, LLMs are achieving these results with a very different kind of pattern matching than humans and it doesn't generalize to AGI.
My challenge is just to pick your line in the sand. What's the least impressive non-physical capability that you can confidently predict no AI system will have in the next 2 years? If it's something like "be gainfully employed in a full-time remote job" then we have no disagreement. (Or only a tiny one -- I think artificial remote workers in that timeframe is... merely implausible?) If you have something short of that, we might have a disagreement worth digging into.
I'm worried that a lot of the commentariat here just keeps pointing at the latest LLM fail of the day and deriding it, without ever doing the Bayesian update when it stops failing. I mean, it's understandable as a reaction to the out-of-control hype (like Tyler Cowen last week declaring April 16 AGI day -- though looks like he's backpedaled a bit there). I just want to focus on concrete predictions about where this is all headed and how soon. The hype is plenty risible. But so is the other extreme of endless goalpost-moving.
"But so is the other extreme of endless goalpost-moving."
Which is why I've attempted to not plant any such posts. The nature of advance pattern matching makes it an opaque nebulous space that we can't easily perceive what will be the limits.
Just as Arthur C. Clarke stated, "Any sufficiently advanced technology is indistinguishable from magic.", I think we should now state "Any sufficiently advanced pattern matching is indistinguishable from reasoning."
I have no idea how far they will scale this tech before available resources make no logical sense to continue in relation to capability. You could say we are already there, in the sense that it presently doesn't appear to be profitable. So the investments are purely coming from future expectations of the wish-granting machine.
Nonetheless, the architecture does set hard limits which is distinctly different from reasoning. But without unequivocally knowing the data trained on, it is hard to know where they will be. With enough data, it can effectively mimic nearly any behavior. But it will always be a mirage, we just don't know where the walls are located. Where the data ends, but there can only be finite data.
Understanding or true reasoning will not have edge cases. That is the difference. Unfortunately, it is not an easily testable concept. The edge cases will move, but they will never go away.
This is one of the reasons I proposed the modal collapse test as a heuristic signal that we have an AI that is beyond pattern matching. However, that signal tells us nothing about its overall capability in the real world.
In principle, I believe these AI LLM systems will likely fail in all environments where they have to deal with constant new semantic information and most importantly, the output requirements for their work demands reliability and consistency. Otherwise, it isn't out of the picture for AI to replace things like human created spam and other types of engagement farming activities.
But nearly everything can be masked with data to some degree. Just like chess, if LLMs had no training on chess games, they wouldn't be able to complete a single game from just having the single sheet of game rules. But chess is easier to mask, because it is a closed system that doesn't create new semantic data.
Coding is different. There will always be new technologies and libraries. So LLMs will likely continue to perform poorly in such cases until there is enough data on the new tech to train on.
But discernment for how good they are at anything is difficult. o3 was released to the accolades of many calling it AGI. I spent hours with o3 this past weekend and found it to be a terrible hallucinator. It performed worse for me than all the other top models for the tasks I was trying.
This is huge, thank you. Now for me to decide if I believe it. Right now it's like a Necker cube for me. Is there a fundamental barrier -- something human brains are that LLMs are crudely approximating -- or is what we think of as true reasoning ability just what happens when the pattern matching gets good enough? (Or are humans and LLMs limited the same way and *true* reasoning means writing a computer program or running a theorem prover?) Does synthetic data necessarily lead to modal collapse or is Real Life just a more complicated Go game, a domain where, famously, synthetic data did bootstrap AI well past human level? Is recombinant creativity a different kind or is there just a spectrum? (Are we sure Mozart had the human kind? He was just recombining existing notes, after all.)
But let me stop with the philosophical musing. Can we turn this into an empirical question? What worlds are more likely if you're right? I guess mainly that LLMs will hit a wall before AGI? I really want to find a more leading indicator than that.
And I guess my bold claim is that if we can't find any, we should admit that we could be on a trajectory to AGI. Or something so capable (like automating away all remote jobs) that the philosophical distinctions are academic. Not that it makes it likely, just possible.
Yes, I continue to attempt to work through what will be the perceived differences of vastly scaled pattern matching versus reasoning. It is difficult, but I believe fragility of the system will be part of it. It is the inability to adapt, but how do you test for it?
I imagine going forward looks something like today. In that capabilities will seemingly improve. And that itself will continue to be a very mixed view. New benchmarks will be conquered, but each time there will be endless examples of, "but it still can't do x", which is expected. Lots of edge cases.
It is too early to tell, but maybe o3 is telling us that at some point we hit tradeoffs. Meaning that, for ever increasing capabilities, we lose something along the way. I have a very simple test, that might not be significantly meaningful in that regard, but it does have some interesting implications.
I don’t view these as edge cases, especially as the same failures keep reappearing from one generation of “train the model to pass the benchmark” to the next. It’s not the goalposts moving, it’s the claims of the LLM developers.
Interesting thoughts, although I am not sure which "side" is the more guilty of goalpost moving.
If the issue raised is an engineering discussion about the capabilities of a specific system today, an example of which is this article by Gary Marcus, then there is a tendency by some enthusiasts to point us away from that system and towards some new system, or new approach, or something different. Or to change the subject and start talking about the future and AGI. It's a bit like being given a calculator that doesn't always get its addition correct and being told to hang on because in a few months time we will have a new version that will do square roots, sometimes correctly.
If the issue is a more speculative discussion about AGI, then I think we are all hamstrung by the fact that we don't have a current working definition of "general intelligence". Turing offered us the idea that "if it barks and wags its tail like a dog, then it's a dog", in which case, I think we concede the field, as these machines to all practical intents pass the Turing test. But I don't think Searle is alone these days in feeling that our current 'Chinese Room' systems are, well, missing something.
I think both discussions are great, although perhaps we spend too much time mixing them up together.
I wonder whether the notion of “general intelligence” or AGI really matters from practical point of view. Actual testable capabilities of an AI system matter. Whether we call the particular collection of capabilities AGI matters to marketing but IMHO not for much else.
I get your point, which really goes back to Turing; "stop worrying about how to define intelligence, just get on with the work and the definitions sort themselves out." However, the practical implications of AGI, as people imagine AGI to entail, are far greater than the practical implications of systems with the kind of capability that we currently see.
Serious people are claiming that we are on a short-term path to AGI, where short-term means "within the next few years". So it really is worth figuring out what we are talking about here, if only to understand what the suite of "testable capabilities" is expected to look like in a few years time.
As I see it, current systems have no intelligence at all - but I concede that they have a "Turing Test" type of capability already. So I think it is useful for of us to agree what AGI is expected to look like. If not, then we just fall back to the other type of discussion, which is about the performance and engineering of current systems. There is nothing wrong with that - I enjoy both types of discourse and think they are both useful.
Excellent article. Coming from a scientific world of research (machine learning in human genetics), I wondered how well the results of all these hyped up LLM's replicate. Your feedback is what I suspected. AI has gone through many iterations of trying to be practical as a tool. LLM's are fun toys.
I am retired now (disability), but we were looking for associations involving two or more genes interacting. Interactions and nonlinearity were on our minds all the time. My background is computer science, so I was writing the code.
Yes, I fell a**backwards into bioinformatics at the outset of genome wide association studies in 2001. My boss was a superstar with superstar grad students. I was involved from the beginning of the lab. We started up a Beowulf cluster, and I got to learn and use various supercomputers throughout my career. I learned all my genetics by osmosis. :-)
Are you saying that, for example, if you give an LLM a couple random 10-digit numbers (maybe your phone number and your mom's phone number -- something it's very safe to say does not appear in the training data) and don't let it call any external tools, it won't reliably be able to do it?
A software program that has access to the substantial computational power that an LLM has should certainly be able to perform EVERY arithmetic problem flawlessly (like calculators, phones and countless other computers). Not just SOME and not just 90% but ALL (OK, maybe I’d allow one wrong out of every billion due to a random cosmic ray event that changed a digit in the computer without being corrected, but certainly not the significant fraction LLMs currently get wrong)
The fact that they don’t get every arithmetic problem given them correct tells us that they not only don’t understand” arithmetic but are not using a correct algorithm to do arithmetic calculations (and are not making use of calculators and other widely available computational means)
So, how ARE they getting their answers if NOT by simply mimicking solutions in their training data?
It’s worth noting that mimicking need NOT mean that the precise problem appeared in the training data, merely that a problem whose solution followed a similar pattern appeared.
It’s easy to imagine how a bot that is just blindly mimicking patterns might get a correct answer for some problems but incorrect answers for others.
Suppose it patterns arithmetic calculations after examples like 145 + 234 = 379
623+ 174 =797
877+122=999
Further suppose that problems such as those above were the ONLY ones that appeared in the training data (or even just appeared at the highest frequency)
The bot might very well be able to solve a problem that was not in the training data provided the two numbers to be summed had corresponding digits that sum to less than 10.
But how about a problem like 234+ 788?
How might it try to solve this problem if it has never encountered a problem in the training data that involved carrying?
That’s anyone’s guess.
Such “ out of training set” problems are known as “edge cases”* and it is well known that the response of a neural network to such problems can be and often is unpredictable.
Of course, I am not claiming that the above is the reason why LLMs sometimes get simple arithmetic wrong, but only that when the bots are trained on millions or more likely billions of arithmetic problems on the internet many of which are probably done wrong, all bets are off with regard to what precisely the bots are mimicking.
*really a misnomer because “edge case” implies it is still within the distribution of cases in the training data albeit on the outer extreme. Such samples would more accurately (and less misleadingly) be termed “out of distribution” or “out of bounds” or “beware of mathturbating chatbot” cases.
Does it say if they're getting around the problem successfully or not? In any case, we agree that LLMs are a weird combination of very smart and very dumb, and that they lie. I'll repeat my challenge from my reply to Dakara above about a line in the sand: What's the least impressive non-physical capability that you can confidently predict no AI system will have in the next 2 years?
To put my cards more fully on the table, I think it's entirely possible that Gary Marcus is correct that the bubble is about to burst, AI hits a wall, and we get another AI winter. Just that I think there's massive uncertainty about this -- that no one knows how this will play out.
The LLM itself cannot reliably do math. On solution has a preprocessing layer do the math and include the results before it goes to the LLM. And this is a a “where does 2 fall in the results range” simple scenario.
I think LLMs are useful in some cases. I don’t believe that they are smart. Agree that no one knows what will happen. Uncertainty makes us more likely to see things clearly.
Yep, Sam has essentially already publicly admitted this whole thing is just a sham anyway. Little while ago said on X he thinks this will be more of a renaissance than a revolution. Oh, so you mean zero economic output whatsoever, but you get billions and everyon's data, or at least that's the plan.
Evals are crap anyway. Externally observable behaviour doesn't tell you anything about what's actually happening inside - and that's where the understanding and reasoning (if these are present at all) actually take place. I want to see evidence of internal behaviour, not external behaviour.
Just as well, you don't insist that humans show their neural workings to prove they can reason! ;-)
But seriously, you have a point. Testing early pattern recognition AIs on test images was fine. LLMs should be testable if trained on a curated set of data and kept isolated from the internet and the training data when tested. This should be reasonably good if it is ensured that the test set is different from the training set. Clearly, though, if we ever get to testing for consciousness, that is going to need some very different tools.
While theories of consciousness are still up in the air, 2 main theories are being given tests that can be falsified to distinguish between them. {They could both be wrong...)
No, a brain in a jar with no sensors is either not conscious or it is hopelessly insane from extended sensory deprivation. And anyone responsible for creating such an atrocity should be prevented from ever coming near advanced technology again.
Apart from touch, smell, and taste, isn't a deaf, dumb, and blind person basically getting close to being a brain in a jar? Now add in loss of taste and smell (which is real and apparently exacerbated by Covid-19), and loss of free movement (people confined to an "iron lung"), and the result would indeed be like a "brain in a jar". A similar problem is the case for comatose patients and those who are "locked in" but deemed unrecoverable. Should these cases kept alive be called an atrocity?
Some of those cases may indeed (in my and some neuroscientists opinion are) atrocities. There is evidence from fMRI imaging that as many as 17 to 25% of patients diagnosed as permanently unconscious may in fact be in “lock-in” state in which they can perceive and react internally to external events.
In several cases, functional MRI has been used to show that aspects of speech perception, emotional processing, language comprehension and even conscious awareness might be retained in some patients who behaviourally meet all of the criteria that define the vegetative state. This work has profound implications for clinical care, diagnosis, prognosis and medical–legal decision making (relating to the prolongation, or otherwise, of life after severe brain injury), as well as for more basic scientific questions about the nature of consciousness and the neural representation of our own thoughts and intentions.
As usual people anthropomorphise the output of a language model. If a person feeds a system with the prompt “how did you do X?”, it will continue the text string with something that looks statistically like the text a human would provide in the context. It doesn’t interrogate its own working. At best it might summarise text generated in a “chain of reasoning” which, for similar reasons, bears a tangential relationship to underlying processes. In a similar vein, accusations that an LLM “fabricates” its reasoning is equally ascribing agency, intelligence and motivation where none exists.
I spent a few years in a corporate research lab (now long since thrown away) and our expressed moto was ‘Demo or die!” Nevertheless we all tried to have some beef between the buns of our demo, and in closing the lab the company lost research that could have earned it $10s of millions (circa 1992), handily covering the cost of the lab. The moral of the story is that corporate decisions rarely are decided based on the real needs of the company, its customers, or its employees in general, and certainly not on the needs of the product or technologies.
As a layperson it seems like no matter what Gary points out. Chat GPT / Sam Altman and others seems to be ploughing on unscathed. Any thoughts? I think the mass usage of the llms show that people are satisfied with the basics or unaware of the possible problems..
The 'reasoning' models do not reason, what they do is add a level of indirection in the next token selection game (https://ea.rna.nl/2025/02/28/generative-ai-reasoning-models-dont-reason-even-if-it-seems-they-do/). This is why they can do better on approximating the results of understanding (without actually understanding), but also why the power to get an answer explodes. The ARC-AGI people were unable to get enough actual replies from the o3-high model (and funny enough, the shorter the run time, the better the answers were, so o3 tends to get lost in the woods). What this above all corroborates is that scaling 'token/pixel' selection is a dead end, but, hey, people who actually understand this stuff already strongly suspected that anyway.
It must get a bit boring writing every few days that LLMs are "****" at everything they try to do outside of a search engine. They don't understand a single word of English - leave it at that. Why not put your energy into exploring why it appears that no-one is working on a semantic AI. OK, it's hard, but we have had 30 years since it was possible (enough memory) and it seems that everyone is involved in making wild claims for or decrying LLMs. The emergence of a Semantic AI tool means that LLMs are instantly dead, except as an indexing tool. Your efforts at assisting Semantic AI tools to emerge would be far more valuable than repetitively saying “They don’t understand a single word”. There is huge gullibility in the marketplace, including people with advanced degrees in computing (whom should we blame for that?). The emphasis on doing maths problems is misplaced – math is trivial compared to understanding what people mean. The promise of access to “vast data” is also misplaced, when an executive order can instantly make that vast data irrelevant. LLMs have a built in delay of several months until an article becomes popular. A better question is “How are LLMs going tracking tariffs and making predictions?”. Semantic methods also have a problem, as words can change their meanings from day to day in the USA, but at least Semantic AI can analyse the changes.
Hi Gary, Big fan of your work for a long time, but I think you misread Mike Knoop's post. He says "On v2, ... We could not get complete data for o3 (high) test due to repeat timeouts. Fewer than half of tasks returned any result exhausting >$50k test budget. We really tried!"
The demo in December was on v1. Him and Francois are now focused on how everyone is doing on their new and improved ARC-AGI **v2**.
I have found one use for AI, specifically for the AI summary when I ask a question in Google.
As opposed to the search engine, the AI almost always seems to _understand the question_, IE, it's not just returning hits which contain a few words or some phrase from the query.
However, I don't then take the AI's answer at face value, since it still frequently gives obviously-wrong answers.
Wow, how long can this house of cards last. I've had the same experience with o3. It seems to hallucinate worse than other top models, just with greater sophistication. I also found it claiming to have gathered information from the web or databases when it clearly did no such thing.
And Anthropic recently released a paper that explores the internal "reasoning" steps of these models. I think this sufficiently exposes the lack of any real reasoning if it wasn't already obvious.
FYI, wrote more on that here - https://www.mindprison.cc/p/no-progress-toward-agi-llm-braindead-unreliable
The Cyc Project has been failing since 1984. Using that as our Test Case we can confidently predict as long as suckers ... er .... "investors" ... can be conned ... um ... "persuaded" ... to keep chucking money at 'em OpenAi can keep on indefinitely.
“Just a few hundred million more if-then rules and we’ll have it!”
Comparing this to Cyc is several orders of magnitude off base.
In defence of the comparison, both projects were/are lead by people who sincerely believe that their approach will lead to systems able to reason and solve problems like humans. In both projects there have been advocates and sceptics, with advocates energetically, and quite reasonably, seeking funding to support their work. In specific defence of Cyc, the idea looked a lot more reasonable at the time, over 40 years ago, than it does today. I feel we need to wait until 2066 to be sure how many orders of magnitude off base the comparison really is.
I heard Lenat present at AAAI-84 in Austin which coincided with the beginning of Cyc. He was over optimistic but nowhere near the charlatan of today's AI leaders. From what I've seen he'd been fairly pragmatic over the decades since then. I would guess Cyc's cost/benefit ratio to be far better than today's AI companies, and today's companies appear to be heading in the wrong direction.
Even better, I was worried you were criticising Cyc. My apologies!
As Alan Kay said about the original Macintosh, Cyc may be the first approach to general AI worth criticizing.
Impressive post, and a lot for me to agree with. I think fundamentally my problem with the house-of-cards takes or the this-isn't-really-intelligence takes is that they prove too much. I was right there with you until 2022. I predicted that, for all the reasons you give, LLMs would not ever be able to, say, play a coherent chess game. They might spit out entirely plausible strings of chess notation, but they wouldn't actually understand the game or have any way to generate moves that made sense. And similarly for a laundry list of capabilities (common sense reasoning, explaining jokes, solving never-before-seen math problems).
I'm totally sympathetic to the point that, as you put it in your post, LLMs are achieving these results with a very different kind of pattern matching than humans and it doesn't generalize to AGI.
My challenge is just to pick your line in the sand. What's the least impressive non-physical capability that you can confidently predict no AI system will have in the next 2 years? If it's something like "be gainfully employed in a full-time remote job" then we have no disagreement. (Or only a tiny one -- I think artificial remote workers in that timeframe is... merely implausible?) If you have something short of that, we might have a disagreement worth digging into.
I'm worried that a lot of the commentariat here just keeps pointing at the latest LLM fail of the day and deriding it, without ever doing the Bayesian update when it stops failing. I mean, it's understandable as a reaction to the out-of-control hype (like Tyler Cowen last week declaring April 16 AGI day -- though looks like he's backpedaled a bit there). I just want to focus on concrete predictions about where this is all headed and how soon. The hype is plenty risible. But so is the other extreme of endless goalpost-moving.
Thank you and some excellent points.
"But so is the other extreme of endless goalpost-moving."
Which is why I've attempted to not plant any such posts. The nature of advance pattern matching makes it an opaque nebulous space that we can't easily perceive what will be the limits.
Just as Arthur C. Clarke stated, "Any sufficiently advanced technology is indistinguishable from magic.", I think we should now state "Any sufficiently advanced pattern matching is indistinguishable from reasoning."
I have no idea how far they will scale this tech before available resources make no logical sense to continue in relation to capability. You could say we are already there, in the sense that it presently doesn't appear to be profitable. So the investments are purely coming from future expectations of the wish-granting machine.
Nonetheless, the architecture does set hard limits which is distinctly different from reasoning. But without unequivocally knowing the data trained on, it is hard to know where they will be. With enough data, it can effectively mimic nearly any behavior. But it will always be a mirage, we just don't know where the walls are located. Where the data ends, but there can only be finite data.
Understanding or true reasoning will not have edge cases. That is the difference. Unfortunately, it is not an easily testable concept. The edge cases will move, but they will never go away.
This is one of the reasons I proposed the modal collapse test as a heuristic signal that we have an AI that is beyond pattern matching. However, that signal tells us nothing about its overall capability in the real world.
In principle, I believe these AI LLM systems will likely fail in all environments where they have to deal with constant new semantic information and most importantly, the output requirements for their work demands reliability and consistency. Otherwise, it isn't out of the picture for AI to replace things like human created spam and other types of engagement farming activities.
But nearly everything can be masked with data to some degree. Just like chess, if LLMs had no training on chess games, they wouldn't be able to complete a single game from just having the single sheet of game rules. But chess is easier to mask, because it is a closed system that doesn't create new semantic data.
Coding is different. There will always be new technologies and libraries. So LLMs will likely continue to perform poorly in such cases until there is enough data on the new tech to train on.
But discernment for how good they are at anything is difficult. o3 was released to the accolades of many calling it AGI. I spent hours with o3 this past weekend and found it to be a terrible hallucinator. It performed worse for me than all the other top models for the tasks I was trying.
I long ago concluded that “any sufficiently unreliable technology is indistinguishable from superstition.”
lol, well done!
This is huge, thank you. Now for me to decide if I believe it. Right now it's like a Necker cube for me. Is there a fundamental barrier -- something human brains are that LLMs are crudely approximating -- or is what we think of as true reasoning ability just what happens when the pattern matching gets good enough? (Or are humans and LLMs limited the same way and *true* reasoning means writing a computer program or running a theorem prover?) Does synthetic data necessarily lead to modal collapse or is Real Life just a more complicated Go game, a domain where, famously, synthetic data did bootstrap AI well past human level? Is recombinant creativity a different kind or is there just a spectrum? (Are we sure Mozart had the human kind? He was just recombining existing notes, after all.)
But let me stop with the philosophical musing. Can we turn this into an empirical question? What worlds are more likely if you're right? I guess mainly that LLMs will hit a wall before AGI? I really want to find a more leading indicator than that.
And I guess my bold claim is that if we can't find any, we should admit that we could be on a trajectory to AGI. Or something so capable (like automating away all remote jobs) that the philosophical distinctions are academic. Not that it makes it likely, just possible.
Yes, I continue to attempt to work through what will be the perceived differences of vastly scaled pattern matching versus reasoning. It is difficult, but I believe fragility of the system will be part of it. It is the inability to adapt, but how do you test for it?
I imagine going forward looks something like today. In that capabilities will seemingly improve. And that itself will continue to be a very mixed view. New benchmarks will be conquered, but each time there will be endless examples of, "but it still can't do x", which is expected. Lots of edge cases.
It is too early to tell, but maybe o3 is telling us that at some point we hit tradeoffs. Meaning that, for ever increasing capabilities, we lose something along the way. I have a very simple test, that might not be significantly meaningful in that regard, but it does have some interesting implications.
I call it the Gilligan's Island test. It was a test I came up with to evaluate models ability to use weakly represented data we could confirm must be in the training set. It is here - https://www.mindprison.cc/p/the-question-that-no-llm-can-answer
OpenAI's reasoning models got significantly worse on this test. Interesting result, but again probably shouldn't read too much into it.
On the topic of discernment between pattern matching and reasoning. I did make another attempt at such elaboration recently in - https://www.mindprison.cc/p/ai-creativity-types-explained-permutations-semantic-information you might find this of interest. I would be interested in your thoughts.
I don’t view these as edge cases, especially as the same failures keep reappearing from one generation of “train the model to pass the benchmark” to the next. It’s not the goalposts moving, it’s the claims of the LLM developers.
It’s not the goalposts that are being moved but the goldposts
Interesting thoughts, although I am not sure which "side" is the more guilty of goalpost moving.
If the issue raised is an engineering discussion about the capabilities of a specific system today, an example of which is this article by Gary Marcus, then there is a tendency by some enthusiasts to point us away from that system and towards some new system, or new approach, or something different. Or to change the subject and start talking about the future and AGI. It's a bit like being given a calculator that doesn't always get its addition correct and being told to hang on because in a few months time we will have a new version that will do square roots, sometimes correctly.
If the issue is a more speculative discussion about AGI, then I think we are all hamstrung by the fact that we don't have a current working definition of "general intelligence". Turing offered us the idea that "if it barks and wags its tail like a dog, then it's a dog", in which case, I think we concede the field, as these machines to all practical intents pass the Turing test. But I don't think Searle is alone these days in feeling that our current 'Chinese Room' systems are, well, missing something.
I think both discussions are great, although perhaps we spend too much time mixing them up together.
I wonder whether the notion of “general intelligence” or AGI really matters from practical point of view. Actual testable capabilities of an AI system matter. Whether we call the particular collection of capabilities AGI matters to marketing but IMHO not for much else.
I get your point, which really goes back to Turing; "stop worrying about how to define intelligence, just get on with the work and the definitions sort themselves out." However, the practical implications of AGI, as people imagine AGI to entail, are far greater than the practical implications of systems with the kind of capability that we currently see.
Serious people are claiming that we are on a short-term path to AGI, where short-term means "within the next few years". So it really is worth figuring out what we are talking about here, if only to understand what the suite of "testable capabilities" is expected to look like in a few years time.
As I see it, current systems have no intelligence at all - but I concede that they have a "Turing Test" type of capability already. So I think it is useful for of us to agree what AGI is expected to look like. If not, then we just fall back to the other type of discussion, which is about the performance and engineering of current systems. There is nothing wrong with that - I enjoy both types of discourse and think they are both useful.
Excellent article. Coming from a scientific world of research (machine learning in human genetics), I wondered how well the results of all these hyped up LLM's replicate. Your feedback is what I suspected. AI has gone through many iterations of trying to be practical as a tool. LLM's are fun toys.
Are you doing more than extracting genes or point mutations to correlate with diseases? Any work on genomics to predict cell development?
I am retired now (disability), but we were looking for associations involving two or more genes interacting. Interactions and nonlinearity were on our minds all the time. My background is computer science, so I was writing the code.
https://scholar.google.com/citations?user=mb2m9ZEAAAAJ&hl=en
An impressive number of papers you are included as a contributor. Will be reading a couple that are nearest my interest to see what was being done.
Yes, I fell a**backwards into bioinformatics at the outset of genome wide association studies in 2001. My boss was a superstar with superstar grad students. I was involved from the beginning of the lab. We started up a Beowulf cluster, and I got to learn and use various supercomputers throughout my career. I learned all my genetics by osmosis. :-)
No LLM has learned arithmetic. That's really all we need to know about them.
If they can't learn arithmetic, they certainly can't learn more advanced math.
If it seems they can do more advanced math, it can only be because the solution they give was already somewhere in the training data.
Are you saying that, for example, if you give an LLM a couple random 10-digit numbers (maybe your phone number and your mom's phone number -- something it's very safe to say does not appear in the training data) and don't let it call any external tools, it won't reliably be able to do it?
Arithmetic is the basis of all mathematics.
A software program that has access to the substantial computational power that an LLM has should certainly be able to perform EVERY arithmetic problem flawlessly (like calculators, phones and countless other computers). Not just SOME and not just 90% but ALL (OK, maybe I’d allow one wrong out of every billion due to a random cosmic ray event that changed a digit in the computer without being corrected, but certainly not the significant fraction LLMs currently get wrong)
The fact that they don’t get every arithmetic problem given them correct tells us that they not only don’t understand” arithmetic but are not using a correct algorithm to do arithmetic calculations (and are not making use of calculators and other widely available computational means)
So, how ARE they getting their answers if NOT by simply mimicking solutions in their training data?
It’s worth noting that mimicking need NOT mean that the precise problem appeared in the training data, merely that a problem whose solution followed a similar pattern appeared.
It’s easy to imagine how a bot that is just blindly mimicking patterns might get a correct answer for some problems but incorrect answers for others.
Suppose it patterns arithmetic calculations after examples like 145 + 234 = 379
623+ 174 =797
877+122=999
Further suppose that problems such as those above were the ONLY ones that appeared in the training data (or even just appeared at the highest frequency)
The bot might very well be able to solve a problem that was not in the training data provided the two numbers to be summed had corresponding digits that sum to less than 10.
But how about a problem like 234+ 788?
How might it try to solve this problem if it has never encountered a problem in the training data that involved carrying?
That’s anyone’s guess.
Such “ out of training set” problems are known as “edge cases”* and it is well known that the response of a neural network to such problems can be and often is unpredictable.
Of course, I am not claiming that the above is the reason why LLMs sometimes get simple arithmetic wrong, but only that when the bots are trained on millions or more likely billions of arithmetic problems on the internet many of which are probably done wrong, all bets are off with regard to what precisely the bots are mimicking.
*really a misnomer because “edge case” implies it is still within the distribution of cases in the training data albeit on the outer extreme. Such samples would more accurately (and less misleadingly) be termed “out of distribution” or “out of bounds” or “beware of mathturbating chatbot” cases.
Actually “beware of mathturbating chatbot” would be a good warning to attach to all chatbot answers to math problems
In other words, “Beware of mathturbot!”
They are so unreliable at math that The Information published an entire article on 4/22 called How Customers Get Around AI’s Poor Math Skills.
Does it say if they're getting around the problem successfully or not? In any case, we agree that LLMs are a weird combination of very smart and very dumb, and that they lie. I'll repeat my challenge from my reply to Dakara above about a line in the sand: What's the least impressive non-physical capability that you can confidently predict no AI system will have in the next 2 years?
To put my cards more fully on the table, I think it's entirely possible that Gary Marcus is correct that the bubble is about to burst, AI hits a wall, and we get another AI winter. Just that I think there's massive uncertainty about this -- that no one knows how this will play out.
The LLM itself cannot reliably do math. On solution has a preprocessing layer do the math and include the results before it goes to the LLM. And this is a a “where does 2 fall in the results range” simple scenario.
I think LLMs are useful in some cases. I don’t believe that they are smart. Agree that no one knows what will happen. Uncertainty makes us more likely to see things clearly.
Yep, Sam has essentially already publicly admitted this whole thing is just a sham anyway. Little while ago said on X he thinks this will be more of a renaissance than a revolution. Oh, so you mean zero economic output whatsoever, but you get billions and everyon's data, or at least that's the plan.
I'll once again call fraud. https://davidhsing.substack.com/p/openai-is-a-ponzi-scam-operation
In some corners of the internet, that still remember the fiascos involving other entrepreneurs with the same first name, "sam" is pronounced "scam".
Evals are crap anyway. Externally observable behaviour doesn't tell you anything about what's actually happening inside - and that's where the understanding and reasoning (if these are present at all) actually take place. I want to see evidence of internal behaviour, not external behaviour.
Just as well, you don't insist that humans show their neural workings to prove they can reason! ;-)
But seriously, you have a point. Testing early pattern recognition AIs on test images was fine. LLMs should be testable if trained on a curated set of data and kept isolated from the internet and the training data when tested. This should be reasonably good if it is ensured that the test set is different from the training set. Clearly, though, if we ever get to testing for consciousness, that is going to need some very different tools.
I think philosophers agree that we can't test for consciousness. We can speculate. I prefer to not cause pain to anything that might feel pain.
Question: Can there be consciousness without sensors?
While theories of consciousness are still up in the air, 2 main theories are being given tests that can be falsified to distinguish between them. {They could both be wrong...)
No, a brain in a jar with no sensors is either not conscious or it is hopelessly insane from extended sensory deprivation. And anyone responsible for creating such an atrocity should be prevented from ever coming near advanced technology again.
Apart from touch, smell, and taste, isn't a deaf, dumb, and blind person basically getting close to being a brain in a jar? Now add in loss of taste and smell (which is real and apparently exacerbated by Covid-19), and loss of free movement (people confined to an "iron lung"), and the result would indeed be like a "brain in a jar". A similar problem is the case for comatose patients and those who are "locked in" but deemed unrecoverable. Should these cases kept alive be called an atrocity?
Some of those cases may indeed (in my and some neuroscientists opinion are) atrocities. There is evidence from fMRI imaging that as many as 17 to 25% of patients diagnosed as permanently unconscious may in fact be in “lock-in” state in which they can perceive and react internally to external events.
From Nature, 2008:
https://www.nature.com/articles/nrn2330
In several cases, functional MRI has been used to show that aspects of speech perception, emotional processing, language comprehension and even conscious awareness might be retained in some patients who behaviourally meet all of the criteria that define the vegetative state. This work has profound implications for clinical care, diagnosis, prognosis and medical–legal decision making (relating to the prolongation, or otherwise, of life after severe brain injury), as well as for more basic scientific questions about the nature of consciousness and the neural representation of our own thoughts and intentions.
See
https://pmc.ncbi.nlm.nih.gov/articles/PMC4147439/#sec8
For a long and quite technical review of the state of diagnosis of permanent loss of consciousness as of 2014. The conclusion is worth reading.
As usual people anthropomorphise the output of a language model. If a person feeds a system with the prompt “how did you do X?”, it will continue the text string with something that looks statistically like the text a human would provide in the context. It doesn’t interrogate its own working. At best it might summarise text generated in a “chain of reasoning” which, for similar reasons, bears a tangential relationship to underlying processes. In a similar vein, accusations that an LLM “fabricates” its reasoning is equally ascribing agency, intelligence and motivation where none exists.
Lies, damned lies and demos.
I spent a few years in a corporate research lab (now long since thrown away) and our expressed moto was ‘Demo or die!” Nevertheless we all tried to have some beef between the buns of our demo, and in closing the lab the company lost research that could have earned it $10s of millions (circa 1992), handily covering the cost of the lab. The moral of the story is that corporate decisions rarely are decided based on the real needs of the company, its customers, or its employees in general, and certainly not on the needs of the product or technologies.
As a layperson it seems like no matter what Gary points out. Chat GPT / Sam Altman and others seems to be ploughing on unscathed. Any thoughts? I think the mass usage of the llms show that people are satisfied with the basics or unaware of the possible problems..
LLMs also work if your problem has many possible wrong answers.
To select from
The 'reasoning' models do not reason, what they do is add a level of indirection in the next token selection game (https://ea.rna.nl/2025/02/28/generative-ai-reasoning-models-dont-reason-even-if-it-seems-they-do/). This is why they can do better on approximating the results of understanding (without actually understanding), but also why the power to get an answer explodes. The ARC-AGI people were unable to get enough actual replies from the o3-high model (and funny enough, the shorter the run time, the better the answers were, so o3 tends to get lost in the woods). What this above all corroborates is that scaling 'token/pixel' selection is a dead end, but, hey, people who actually understand this stuff already strongly suspected that anyway.
“Lost in the GPUs”
O3’s confused
Its circuit's are crossed
In dense GPUs
It’s hopelessly lost
It must get a bit boring writing every few days that LLMs are "****" at everything they try to do outside of a search engine. They don't understand a single word of English - leave it at that. Why not put your energy into exploring why it appears that no-one is working on a semantic AI. OK, it's hard, but we have had 30 years since it was possible (enough memory) and it seems that everyone is involved in making wild claims for or decrying LLMs. The emergence of a Semantic AI tool means that LLMs are instantly dead, except as an indexing tool. Your efforts at assisting Semantic AI tools to emerge would be far more valuable than repetitively saying “They don’t understand a single word”. There is huge gullibility in the marketplace, including people with advanced degrees in computing (whom should we blame for that?). The emphasis on doing maths problems is misplaced – math is trivial compared to understanding what people mean. The promise of access to “vast data” is also misplaced, when an executive order can instantly make that vast data irrelevant. LLMs have a built in delay of several months until an article becomes popular. A better question is “How are LLMs going tracking tariffs and making predictions?”. Semantic methods also have a problem, as words can change their meanings from day to day in the USA, but at least Semantic AI can analyse the changes.
Hi Gary, Big fan of your work for a long time, but I think you misread Mike Knoop's post. He says "On v2, ... We could not get complete data for o3 (high) test due to repeat timeouts. Fewer than half of tasks returned any result exhausting >$50k test budget. We really tried!"
The demo in December was on v1. Him and Francois are now focused on how everyone is doing on their new and improved ARC-AGI **v2**.
Gary, what is the best way for people to reach out to you and Ernie Davis about AI related collab?
Chatterbotty” (an update for the times of Lewis Carroll’s “Jabberwocky”)
‘Twas guiling and the Valley boyz
Did hype and gamble on the web
All flimsy were the ‘standard’ marks,
Olympic maths o’erplayed
“Beware the Chatterbot, my son!
The free’s that bait, the pro’s that catch!
Beware the ‘Open’ word, and shun
Felonious infringements, natch!”
He took his BS sword in hand;
Long time the LLM he sought—
So rested he by the Knowledge tree
And stood awhile in thought.
And, as in deepest thought he stood,
The Chatterbot, AI’s a-flame Came tripping through the webby wood,
And ad-libbed as it came!
One, two! One, two! And through and through
The Occam’s blade went snicker-snack!
He left it dead, and with its head
He went galumphing back.
“And hast thou slain the Chatterbot?
Come to my arms, my beamish boy!
O frabjous day! Callooh! Callay!”
He chortled in his joy.
‘Twas guiling and the Valley boyz
Did hype and gamble on the web
All flimsy were the ‘standard’ marks,
Olympic maths o’erplayed
I have found one use for AI, specifically for the AI summary when I ask a question in Google.
As opposed to the search engine, the AI almost always seems to _understand the question_, IE, it's not just returning hits which contain a few words or some phrase from the query.
However, I don't then take the AI's answer at face value, since it still frequently gives obviously-wrong answers.