72 Comments
User's avatar
Gary Marcus's avatar

i would flip it around and say whenever AI fails it’s probably a pure LLM. all the big wins lately have in fact involved code interpreters, symbolic code (in addition to neural networks etc).

Ibon Urrutia's avatar

The joke that is becoming viral is exactly what some working class engineers are we saying to anybody that wants to listen every time some de says that he is 10x productive with code agents.

It does not matter that you go ahead 100 meters in 8 seconds, if after you have to go back 500 meters.

In 25 years I did not see any valid productivity metric for developers. Zero. All BS. Measuring software by quantity, supposed quality (infinite definitions here about what is supposed to be quality in software) or complexity (here the only metrics still valid are ciclomatic Complexity and dependencies graphs, because more ifs more hard to maintain, and more dependencies, harder to maintain).

There is not any metric that is long term and that is not bullshit. Time to market could be totally erased by time to go back from market with your reputation as brand damaged for ever. Number of features per month could be downgraded by number of users not using them and complaining about your bloated product.

Technical debt was called like that because the engineer proposing it wanted to explain to suits that any technical shortcut or less than optimal technical solution works like financial debt : you will have to pay it someday, and if you start collecting other debts to pay that debt... Well, you know how is finishing the people that need to use their credit card to pay their mortgage.

Software is always one step near a combinatorial explosion of states. Every day you have to "cut the grass" from it, to maintain the number of different states manageable. Because as soon as the explosion starts, states not contemplated arise always. And any new feature adds more probability to that to happen.

How to measure that? No idea. But people not understanding that, just prompting because they can finish their tasks faster and be at home with their kids, are not receiving the incentives that is going to produce more productive software. Not at all.

So yeah. Stop saying that LLMs might have a productive application in software when still nobody was able to probe it (and dev tech bros claiming it in Twitter is not proof of nothing)

Carlos M's avatar

The 10x developers have always existed and this has been well known for a long time. Some of them might be using Claude Code now, but these are the ones who might have been 10x coders in a previous generation with emacs macros.

Back when I started, the focus has always been to try to find and nurture these kinds of folks, as they'd be paid 3x as much for doing 10x the work and everyone is happy. Up to this day no one has successfully found the secret sauce to turning average devs to 10x ones, and not for a lack of trying.

On a personal note, I find debugging, in particular, to be difficult if not impossible to teach. Some people just have that knack for systematically hunting down problems that most people just don't have. It doesn't even seem to matter what methods or tools are used. Experience and training helps but only to get up to speed with an unfamiliar application/framework/platform. After that people seem to settle down at their "natural" bug fixing capacity. Good luck then with all the new AI generated code.

Martin Machacek's avatar

Debugging is mostly about having correct mental model of how the system works. The mental model is built by studying the purpose, design and eventually the very implementation of the system. Good mental model allows to build relevant theories explaining the observed incorrect behavior. Is is very useful to rank those theories by likelihood and complexity of verification and check them starting from the most likely and easiest to verify. And of course you need to adjust your mental model and theories based on new data you collect.

Gerben Wierda's avatar

True story: Job interview, once. First a test where among the questions I had to answer "what is *ls*" etc.. My answer on "what is *xdb*" was: "guessing, something with a database?". Their incredulous follow-up: "No that is the X-debugger. How on earth can you not know what xdb is?" My answer: "Because I hate debuggers with a passion because I feel like a failure as a software engineer if I have to use one. And thus I never use one.".

For every hour you spend in preventing poor code, you save 7 hours in debugging code. Still true (though it is possible to overshoot). Now: translate to this LLM-coding-era.

Bugs do happen (especially when LLMs produce code). Nobody is perfect. But truly understanding the code is — I personally think — still the single best prerequisite in finding them. LLMs with a scattershot approach can guess correctly at some too, so maybe my experience will change.

Sivan's avatar

This is so damn true (pardon my French) after many months of delicately and meticulous work with Claude Code it today introduced a regression that almost invalidated work of 6 months and many $’s in tokens. I now feel it was a huge mistake to let stochastic pattern prediction and matching machine write code, which requires, inherently infinite rigour and is anything but stochasticz

Ibon Urrutia's avatar

I'm just a working class engineer, and I don't think that I have the right to tell anyone how to do their job. I'm opinionated about how I do my job. For example, I do TDD because in my case TDD made me a better developer, it helped me to think into my tasks and design interfaces, APIs thinking in how they are used first. But there is people out there that don't need to write the test first to do the same. OK, no problem. No flame war from me.

I'm bored and tired of the "AI is here to stay, deal with it" of people imposing me their views in how to use LLMs for writing code. The last idea, you should NOT read the code, it is so so absolutely nonsense for me and my 25 years in the trenches that I can not believe how they are even defending it in public.

We do not have any evidence of using code agents massively in a corporate code base more than 6 months, except some that look evidently inflated (from some company starting by Spoti... That had a vast amount of open source already in place governing their sdlc) and guys are recommending to not look at the code? And? Start looking at it at 3am with an incident in production that your agent couldn't solve? "trust the process"?

What are we speaking about? Religion or engineering?

Suddenly 30 years lessons learned in the trenches are useless. I don't believe that. They have to probe it to me. And they are doing it really badly with their imposition of ideas using fear of being "behind" and not getting a job.

I don't like what they are doing to the profession. Not at all.

Sivan's avatar

Well said. Agree with every word.

Larry Jewett's avatar

“Mindless Coding”

Mindless coding

Mindless bots

Ill foreboding

Link the dots

Larry Jewett's avatar

“Botshit Productivity”

Manure ten-tastic

Is still manure

Despite gymnastic

It’s crap, for sure

Larry Jewett's avatar

From the beginning of chatbot coding, technical debt has been the AI-lephant in the room

Rick Talbot's avatar

Several years before covid I went to a conference by one of the big consultancies. They were so deep into blockchain talk that it was kind of ludicrous.

Oaktown's avatar

That explains a lot.

James's avatar

Accenture actually had some big partnership or JV thing with Facebook on Metaverse stuff a few years ago.

Rick Talbot's avatar

Oh, that's right, the consultancies were trumpeting about metaverse being a trillion dollar business, weren't they. So weird.

Romeo Lupașcu's avatar

Interesting indeed. Accenture and in general consultancy business are indeed in a perilous position though I suspect this is omly temporary.

Here is my take.

1. In the ideal AI case (AGI reliable), consultancy companies are eradicated because who will need them if the AI can do all they do for cheaper?

2. In real world AI (unreliable Ai) the AI we have now, the consultancy companies are expected to take a hit from the initial gullibility of their own customers that confuse this AI with "that" AI and drop or not renewing contracts. That's until they start to realize that this AI do need experts... that consultancy companies (if smart) won't fire to replace with ... AI.

So, watch their retention and hiring.

If Accenture fire experts then dump their stock cause they are idiots. But if they retain talent (humans) and improve them (even with AI) then buy their stock now when is low because its about to rebound in few months or years.

Just my gut talking...

JazzPaw's avatar

My guess is that any client of a a consultancy is probably unable to implement any new tech without help from the consultant. AI will be no different. I can’t imagine just dropping my consultant contract and depending on untrained employees to replace the work with ChatGPT. The consultants will become more efficient if they can actually find a use for AI and they will still get the work.

Romeo Lupașcu's avatar

You are thinking rational and correct, but... many management people these days are fooled by these models believing they can do more than they can do in reality. That injects irrationality in the system and at least for a while that can create a lot of trouble for consultants.

JazzPaw's avatar

Many in mgmt are watching the hype, and some of their eager employees are probably producing something with AI that looks promising. The AI companies are hyping their products with new announcements every week, so I’m not surprised to see mgmt lusting after lower costs.

The tokens are underpriced and the downstream problems have yet to arrive. Anecdotally, I’ve heard that many of these AI roll your own projects have been less than successful and have caused operations and maintenance problems.

I would not be surprised if there is ROI, if these tools are used successfully, but I’ll bet it will take the best employees to smoke out those cases.

I used to work in software as a mediocre developer of technical software. Many of the tools we gravitated to over time were touted as increasing productivity. Many of those tools worked eventually, but not always very well in the beginning. I just don’t see large scale replacement of humans. Software development went from 1’s and 0’s, to low level languages, to compiled code on punch cards, to terminals, to the PC and all of the integrated development environments, to the script languages and beyond. None of those massive increases in productivity killed jobs. We just expanded our goals to use up the new capability.

Profusion's avatar

If I was an Accenture client, I'd want to know what exactly I was paying for. Expertise or LLM prompts I could do myself. To me, their business case is exactly proportional to the amount of unique human expertise they bring to a problem.

Carlos M's avatar

I've said this before and I'll say it again. My two major points of frustration with AI coding hype are:

a) The most productive coders have already mostly automated their workflows anyway. Build tools, code generators, automated tests, and so on are a thing long before LLM coding agents have been around.

b) More importantly, the biggest productivity gains in software are had not with automation, but reuse. A computer is nothing if not a machine that replicates solved problems at scale.

The thing AI coding agents are good at - and Claude hasn't changed this no matter its advances - is giving you quick implementations of previously solved problems. Which is actually quite scary from a technical debt standpoint, as instead of a handful of proven solutions with known bugs, we now have countless variations of new attempts to solve previously solved problems, with the inevitable new and innovative ways for them to fail alongside.

Catherine Blanche King's avatar

"Funny you should mention it," from NATURE Cancer Briefing:

"China reels from research-integrity scandal: Several cancer researchers in China have been fired or disciplined after a vlogger exposed potential instances of data fabrication in papers published in Nature Cancer and other Nature-branded journals. In videos posted to Chinese social-media platform Bilibili, former PhD student Geng Hongwei, known online as ‘Student Geng’, showed that many digits in the figures of one paper were suspiciously identical, suggesting the data had been made up. Following an investigation into the co-authors by Nankai University, Chen Quan was removed from his position as dean of the College of Life Sciences, Zheng Hao was found to have committed academic misconduct and was fired from his role as a postdoctoral researcher and another scientist was cautioned. Nature Cancer editors are investigating the concerns. (Nature Briefing: Cancer is editorially independent from Nature’s research division.)"

Nature | 11 min read Reference: Nature Cancer paper (under investigation, from 2024)

Aaron Turner's avatar

Respectfully, Gary, I have to push back a little on "I really do think that Claude Code (which remember, is a special-purpose neurosymbolic system rather than a generic chatbot) will increase productivity for coders (if technical debt doesn’t swallow back those gains)." You quite regularly opine that Claude Code is a neurosymbolic system, and therefore (cognitively) head and shoulders above basic GenAI. But, to the best of my knowledge (and I admit I haven't properly read the recent paper where some researchers reverse engineered it), Claude Code is just Claude wrapped in a (fairly large) hand-coded (?) harness (which calls the underlying Claude LLM/chatbot and basically turns it into an "agent"). But, to me, the "symbolic" in neurosymbolic needs to refer to some kind of logic-based element, such as induction, deduction, or abduction, that manipulates fragments (such as terms, wffs, or sentences) of some formal logic (such as FOL or lambda calculus), and therefore performs some semantically significant (i.e. truly cognitive) reasoning -- but Claude Code's harness isn't that, it's just normal executable code. It might be fairly clever and might therefore significantly enhance basic Claude's programming utility, but, to me, it's not really "symbolic". Just sayin'.

Steve Robbins's avatar

Accenture was bad enough to deal with when brought into a big company with big systems to "help." I'd note again that system I worked with - the 5m line mainframe, 5m line Java code lending system serving hundreds of banks where absolute computational accuracy was an absolute requirement, where even the smallest change required massive testing, and latent, undiscovered errors subtly screwing up computations over time could cause nuclear disaster. There is ZERO chance Accenture would have been allowed to introduce AI slop, and any thought it could do so would have been an index of how little they understood the systems they were hired to work on.

Robert Keith's avatar

Tokenmaxxing. Or what we used to call busywork.

Larry Jewett's avatar

Above all else, AI is built on talkin’maxxing.

AI CEOs have raised it to an AIrt form

Robert Keith's avatar

Playing with their shiny new Rube Goldberg machine.

Larry Jewett's avatar

Or maybe it’s a “Gold Ruse-berg” machine?

Joy in HK fiFP's avatar

That seems to be what AI-ccenture did, to their current misfortune.

Patrick Logan's avatar

Accenture is so named because they used to be Arthur Andersen and got in big trouble being the "auditor" for Enron. I would take anything they ever say with more than a grain of salt.

Thomas Schmid's avatar

Absolutely! That type of leopard doesn't really change its spots.

Jonathan Grudin's avatar

A couple points. Without knowing details about Accenture's business, it could have been doing well as Levi Strauss did by selling (AI) tools to its customer-prospectors, which will dry up if the customer-prospectors don't succeed or you don't find new customers. Fortunately for Levi-Strauss, although not many prospectors found gold and new ones stopped coming, other people liked blue jeans. People seem unfamiliar with the details of the task-focused chatbot tsunami of 2016-2019 kicked off by Facebook's M (see Wikipedia for details). Facebook, Microsoft, and IBM all built task-focused chatbots and also built tools and platforms for developers to pay to use. None of their task-focused chatbots succeeded, but hundreds of thousands of developers used their platforms, so this was good business for Facebook, Microsoft, and IBM for a time. But when their customers failed, it ceased being a good business. Accenture (and perhaps IBM consultants and others) who are advising on use of genAI could be making money this way for now, but if we are right this will not be a successful path.

Second, a couple commenters are correct that Gary's attribution of successes to symbolic features is not thoughtful, and the statement that failure is due to being a "pure LLM" is wrong, because THERE ARE NO PURE LLMs OUT THERE. They have all had a wide variety of "harnesses," which was evident in their design and even more obvious in the quick adjustments made following embarrassing incidents reported in the early weeks, changes made before rebuilding the LLMs. You may recall that ChatGPT evolved while still using a model that had no information newer than when it was built many months before.

The sad effect of trying to argue that coding successes are due to symbolic elements is that it hides the real reason for coding successes. AI has for half a century been successful at tasks that required no knowledge of human, group, and organizational behaviors, which are wildly variable. AI could do well at checkers. Tower of Hanoi. Chess. Go. With ML and Deep Learning it could take on identifying and categorizing items in photographs.or drawings. Then speech recognitiin, protein-folding. Solving math test problems. And coding! A recent dramatic example is completing significant theoretical mathematics proofs that were eluding serious mathematicians, creating a crisis for some over the future of working in the field. An unfortunate consequence is that people may reason that "Coding is difficult. Theoretical mathematics is even more difficult. My organizational challenges are not rocket science so AI can surely help me" and be wrong, because organizational challenges generally involve a bunch of temperamental people. Good luck wih that. And symbolic AI failed for decades largely because it assumed people relied on cognition and reasoning. John McCarthy, who coined the term 'artificial intelligence' in the 1950s, many years later said exactly that, they underestimated the complexity of human behavior/intelligence.

Sure, LLMs make mistakes and confabulate, like people do. And they are very useful, as are people and some ML and DL systems. But they are limited, as are people. To expand their reach hugely they would need cognitive models as Gary says, but more than that, cognition alone won't do it. And they would need more detailed information about individual people than we are likely to be comfortable sharing. And all of that will be very expensive to collect and use. IMO we may drop back to the slowly rising and much less expensive Deep Learning trajectory. I can always be wrong, and I sometimes worry that Dr. Strangelove types in secret places with inexhaustible budgets may do unfortunate things.

DEEPAK CHOUDHARY's avatar

Thanks for information.

Jeff's avatar

If AI is reducing the need for billable consulting hours, then weakness at Accenture could be evidence that AI is working, not evidence that it is failing.

Jeff's avatar

That’s a genuinely difficult argument to dismiss because Marcus’s own caveat (“competition from AI might be eating into their business”) creates a contradiction. If AI is replacing work previously purchased from consultants, then AI is generating ROI somewhere.

I would tighten Claude’s response further and make it more pointed:

I don’t think Accenture proves what Gary thinks it proves.

Accenture didn’t miss earnings. EPS grew 9% YoY, beat expectations, and margins expanded. The stock sold off because bookings were soft and management trimmed growth guidance.

More importantly, Gary’s own caveat undermines his thesis.

He argues Accenture’s weakness proves AI isn’t delivering ROI, while simultaneously acknowledging AI could be taking business away from Accenture.

But if customers are using AI to do work they previously paid consultants to do, that’s evidence of AI ROI, not evidence against it.

Consulting revenue is not a proxy for AI value creation.

In fact, if generative AI reduces billable hours, automates routine work, or allows enterprises to self-serve tasks that once required outside consultants, consulting firms may be among the first businesses disrupted by successful AI adoption.

That’s a bear case for consulting. It’s not necessarily a bear case for AI.

The broader studies Gary cites mostly show that enterprise deployment is slower and harder than expected. That’s a valid criticism. But implementation delays are not the same thing as technological failure.

The strongest argument against AI today is that ROI is arriving more slowly than the hype suggested.

The weakest argument is that one government-exposed consulting firm’s softer bookings prove “the whole jig is up.” Hat trick Claude/ CHTGPT

And Then It Fell's avatar

Oh, for the love of little green apples . . . I will never understand people who post AI-generated text in comment sections. Either join the discussion yourself or STFU.

Jeff's avatar

I will never understand those that don’t proof their narrative prior to posting exposing all the obvious holes and errors. Human intelligence is so yesterday. :)

Larry Jewett's avatar

And then it fell

To borrow a bottism: “That [what you say] is a genuinely difficult argument to dismiss.”

JazzPaw's avatar

Accenture has been digesting a decline in US Federal spending on consultants like Accenture. They may also be dealing with client uncertainty that reduces budgets in the near term. Given the magnitude of the total drop in the Accenture stock price, it is clear that most of the damage is theoretical. The business is not off that much, and in fact it has not declined. The entire SaaS and consultant sectors are getting thrown away like trash, but the financial results don’t yet justify the pessimism. We will see how this plays out.

Jeff's avatar

Markets anticipate what is coming. Until proven otherwise, their stock is now marked to the narrative that AI is disrupting their business.

Peter Bona's avatar

"that Claude Code (which remember, is a special-purpose neurosymbolic system rather than a generic chatbot"

What makes you think that Claude Code is different?

Gary Marcus's avatar

i gave a link to my essay on that

Tom Gracey's avatar

The Anthropic codebase has been revealed to be a complete mess, probably because it "wrote itself" - and there's no doubt this was the same reason it leaked. i.e. Claude messed up and posted its own codebase. Have you ever known a major tech company to leak the full codebase of their main product? It's unheard of. But that's what you get when you trust a pattern matching algorithm to act as an “agent”.

As for "neurosymbolic" - if this means simply wrapping an LLM in with some simple conditionals and regexes, then this is hardly going to lead to something special, and definitely not “AGI”. Mammals for example (known for their actual intelligence) have far more intricate and synergistic brains than could ever be reproduced with a simple attention mechanism wrapped in some if/then statements. Their operation involves a complex interaction between a vast number of different neurotransmitters and synapses, and is still not well understood. It is really quite absurd - not to mention incredibly arrogant - to make out that wrapping a text generator in some if/then statements is at all comparable.

Real intelligence is going to need a view over reality, not just tokens - and In my opinion it’s also going to need an emotional framework, just like big brained animals. There is no emotion in text tokens.

Furthermore, why is it Anthropic’s product alone that gets the “neurosymbolic” label and the Gary Marcus seal of approval? *All* the commercial LLMs incorporate traditional conditional processing. Look at ChatGPT 5 - it routes through to different models depending on the type of input. Gemini is almost certainly doing the same. All of the models attempt to put “guardrails” around the input and/or output - how do you think they are doing that? Heck even the transformer python library itself is full of conditional processing. In fact I think it would be pretty difficult to put an LLM based system together of any kind *without* also including standard conditional processing. (Indeed we’ve created LLM-based systems a number of times ourselves, and they all included plenty of conditional processing. Should we be breaking out the champagne because they were “neurosymbolic” then?)

So what’s so special about Anthropic’s product, Gary? If you applaud it for being “neurosymbolic” surely you should applaud all of them? But coincidentally hyping Claude seems very much to be a theme in the media; perhaps they set aside some percentage of those VC billions to astroturf their “Claude is so amazing it’s dangerous” campaign? Are you a participant by any chance, Gary? It’s sad to see someone debase themselves like this, and is embarrassing to read. The money must be good, I guess.

Johndavis's avatar

He seems to think a NN model using external symbolic tools is a neurosymbolic system.

Gary Marcus's avatar

if you use symbolic tools, you are neurosymbolic yes. the whole claim of hinton etc was that you could do everything w feedforward nets and no symbolic tools and that just turned out to be wrong

NG's avatar

That essay only mentioned Claude Code using a bunch of If-Thens as a crutch. One could argue that this is just a case of tool use. Not the best example of a true neurosymbolic system IMO (admittedly the term is wide enough to include all kinds of uses).

I would imagine such a system to have symbolic/discrete rules encoded in the continuous weights in some way, not to simply have external manual code as fallback.

Gary Marcus's avatar

it uses over 50 symbolic tools

NG's avatar
Jun 18Edited

How do we know OpenAI and others don't? How do we know that's the differentiator for Claude?

Plus, on a technical level, I am willing to bet that whatever advantage is built into it, is not built into the frontend.

Larry Jewett's avatar

That’s closer to “Nero shambolic”, aka fiddling while Anthropic burns