Earlier this month I attended an AI conference called The Curve in Berkeley. A lot of people there were “AGI-pilled.”
For example, I participated in a role-playing exercise organized by Daniel Kokotajlo, a co-author of the AI 2027 report. That report argues that AI systems will soon achieve human-level intelligence. Then they’ll rapidly improve themselves, leading to superhuman AI capabilities and an extreme acceleration of scientific discovery and economic growth.
I also attended a talk by another AGI-pilled writer, Nate Soares, who believes that superintelligent AI will kill everyone.
At the opposite end of the spectrum are skeptics who believe AI is not just overhyped but practically useless. This perspective wasn’t as well represented at the conference, but I appeared on a panel with perennial AI skeptic Gary Marcus. He laid out his case in a New York Times op-ed a couple of weeks ago:
These systems have always been prone to hallucinations and errors. Those obstacles may be one reason generative AI hasn’t led to the skyrocketing profits and productivity that many in the tech industry predicted. A recent study run by MIT’s NANDA initiative found that 95% of companies that did AI pilot studies found little or no return on their investment. A recent financial analysis projects an estimated shortfall of $800 billion in revenue for AI companies by the end of 2030.
Marcus argued that companies should “stop focusing so heavily on these one-size-fits-all tools and instead concentrate on narrow, specialized AI tools engineered for particular problems”—tools like AlphaFold, the Google DeepMind model for predicting protein structures.
Another skeptic is Ed Zitron, who has built a following making the case that OpenAI will never turn a profit because LLMs can’t generate value commensurate with their high inference costs. Zitron expects OpenAI to collapse in the next few years, and suggests that this could set off a tech industry contagion analogous to the failure of Lehman Brothers in 2008.
My view is between these extremes. I think today’s AI has genuinely impressive capabilities that are likely to improve further in the coming months and years. I think the AI industry is likely to be profitable in the long run, and that OpenAI’s basic business model is perfectly reasonable.
But I don’t think we’re very close to human-level intelligence. And I don’t think AI is about to drive the kind of massive social and economic changes that AGI-pilled folks expect.
So I recently found myself nodding along as Andrej Karpathy was interviewed on Dwarkesh Patel’s podcast. Karpathy co-founded OpenAI in 2015 before joining Tesla to lead its self-driving team. Since departing Tesla in 2022, the 39-year-old has become something of an elder statesman in the AI industry. People credit him with coining the phrase vibe coding back in February.
A key theme throughout the interview was a palpable frustration with both extremes in the AI debate.
“When I go on my Twitter timeline, I see all this stuff that makes no sense to me,” Karpathy said. He believes that fundraising needs have pushed AI leaders to make unrealistic promises about the pace of progress. At the same time, he said he was “overall very bullish on technology.”
If you have time, I encourage you to listen to the full two-hour conversation; it was packed with deep insights about the state of AI and its likely impact on the broader economy. But for those who don’t have two hours, I’ll highlight the bits I found most interesting.
Then I’ll discuss that MIT study finding that 95% of enterprise AI projects fail. Skeptics like Marcus love to cite it as evidence that AI is useless. But the study’s actual findings were more interesting than that — and they don’t really support the views of either AI skeptics or the AGI-pilled.

“There’s so much work to be done”
AI leaders have described 2025 as the “year of agents.” But that doesn’t sound quite right to Karpathy.
“I was triggered by that because I feel like there’s some over-prediction going on in the industry,” Karpathy said. “In my mind this is more accurately described as the decade of agents.”
He acknowledged that agents like Claude Code and Codex are “extremely impressive.” However, he said, “there’s so much work to be done.”
Today’s agents were built using reinforcement learning, a technique that trains a model by judging whether it reached the right answer. This technique enabled the reasoning models that started to appear late last year. Karpathy argues that “reinforcement learning is terrible. It just so happens that everything that we had before is much worse.”
The problem, in Karpathy’s view, is that reinforcement learning tries to judge a long string of reasoning—potentially thousands of tokens—based on whether the model ultimately got the right answer.
“You’re sucking supervision through a straw,” Karpathy said. “It’s crazy. A human would never do this.”
Humans can analyze intermediate steps in our own reasoning processes. So if we make a mistake, we can often identify the specific point where we got off track. This allows us to learn with many fewer rounds of trial and error.
Karpathy believes that the AI industry will need to develop similar techniques if it wants to achieve human-level intelligence. It has been trying to do this for several years without much success.
“I expect there to be some kind of a major update to how we do algorithms for LLMs coming in that realm,” Karpathy said. “And then I think we need three or four or five more.”
The intelligence explosion started decades ago
AGI-pilled thinkers believe we’re about to enter a process of recursive self-improvement where today’s AI models are used to improve the next generation of models, leading to accelerating progress in AI capabilities.
Karpathy doesn’t deny that this is possible. Instead, he argues that something like this has been happening since the dawn of the computing industry.
“I see AI as fundamentally an extension of computing in some pretty fundamental way,” Karpathy told Patel. “I see a continuum of this kind of recursive self improvement or of speeding up programmers all the way from the beginning. I would say code editors, syntax highlighting, data type checking—all these tools that we’ve built for each other.”
“The human is progressively doing less and less of the low level stuff. For example, we’re not writing the assembly code because we have compilers. Compilers will take my high-level language in C and write the assembly code. So we’re abstracting ourselves very, very slowly.”
I made a similar point last year when I noted that Meta had made extensive use of Llama 2 to help generate and filter the data used to train its successor, Llama 3:
I count more than 30 times Meta used Llama-based models to either filter out low-quality training data or generate high-quality training data. Other leading AI labs have not been as transparent as Meta, but I’d be surprised if they weren’t using similar techniques.
So when it comes to filtering and augmenting training data, AI systems are already doing a lot of the work to build their successors. I expect this to become increasingly true over time. But it will be many years—if ever—before companies stop hiring human beings to oversee the process.
“We’re in an intelligence explosion already and have been for decades,” Karpathy told Patel.
Karpathy doesn’t limit this insight to programming.
“I looked at some of the other technologies that I thought were very transformative, like maybe computers or mobile phones or et cetera. You can’t find them in GDP,” he said. “We think of 2008, when the iPhone came out, as this major seismic change. It’s actually not. Everything is so spread out and so slowly diffuses that everything ends up being averaged up into the same exponential. And it’s the exact same thing with computers. You can’t find them in GDP because it’s such a slow progression.”
“And with AI, we’re going to see the exact same thing,” Karpathy argues. “It’s just more automation. It allows us to write different kinds of programs that we couldn’t write before. But AI is still fundamentally a program and it’s a new kind of computer and a new kind of computing system. But it has all these problems. It’s going to diffuse over time and it’s still going to add up to the same exponential.”
Despite these somewhat skeptical arguments, Karpathy insists that he’s not an AI pessimist.
“I think this will work. I think it’s tractable,” Karpathy said. “I think we’re going to work through all this stuff and I think there’s been a rapid amount of progress. For example, Claude Code or OpenAI Codex and stuff like that, they didn’t even exist a year ago.”
Karpathy doesn’t agree with those who think that the recent boom in data center investment is unsustainable.
“I don’t actually know that there’s overbuilding,” he told Patel. “I think that we’re going to be able to gobble up what is being built.”
The truth about that 95% failure statistic
Skeptics think Karpathy is wrong. They view the growth of data centers as an unsustainable bubble because AI models are not actually very useful.
Earlier I quoted AI skeptic Gary Marcus citing an MIT study that found that “95% of companies that did AI pilot studies found little or no return on their investment.” The July study really did include this finding. But in the context of the overall study, it doesn’t seem as damning. Here’s a key paragraph:
In interviews, enterprise users reported consistently positive experiences with consumer-grade tools like ChatGPT and Copilot. These systems were praised for flexibility, familiarity, and immediate utility. Yet the same users were overwhelmingly skeptical of custom or vendor-pitched AI tools, describing them as brittle, overengineered, or misaligned with actual workflows.
The study found that AI pilots based on “general-purpose” models like ChatGPT or Claude led to successful implementation 40% of the time. The much-quoted 95% failure rate was specifically for “custom enterprise AI tools.”
This result should be unsurprising to anyone who has studied the history of general-purpose technologies.
In the late 1980s, economists started to wonder why the proliferation of personal computers had not led to a jump in economic productivity. The economist Robert Solow famously quipped in 1987 that “you can see the computer age everywhere but in the productivity statistics.”
In 1990, the economist Paul David wrote a famous paper pointing out that it was common for it to take multiple decades for companies to take full advantage of new technologies. David’s go-to example was the electric motor, which was first commercialized in the 1880s but wasn’t widely used in factories until the 1920s.
A key reason was that factories needed to be redesigned to take full advantage of the new technology. Previously, all of the machines in a factory needed to be driven by a single giant shaft connected to a steam engine. Electric motors were small and flexible enough that there could be a motor at each workstation. This had numerous benefits: greater energy efficiency, more flexible factory layouts, and greater worker safety.
But while the advantages of electric motors were obvious in theory, it took years of trial and error to apply them across numerous industries.
“Implementation on a wide scale required working out the details in the context of many kinds of new industrial facilities, in many different locales, thereby building up a cadre of experienced factory architects and electrical engineers familiar with the new approach to manufacturing,” David wrote.
Writing in 1990, David suggested that something similar was happening in the computer industry. Intel had released its first general-purpose microchips in the early 1970s, yet it seemed the US had experienced few economic benefits from its technological leadership. David argued that the US economy was experiencing a “diffusion lag” analogous to the one that had occurred with electric motors 60 years earlier.
“The information structures of firms (i.e. the type of data they collect and generate, the way they distribute and process it for interpretation) may be seen as direct counterparts of the physical layouts and materials flow patterns of production and transportation systems,” he argued.
Companies needed to reorganize their internal processes to take full advantage of the potential of personal computers. This was a slow and error-prone process, but it seems to have happened eventually.
The US experienced a productivity boom from 1997 to 2004. It’s hard to say exactly how much this had to do with computers, but this was right around the time many companies fully integrated computers into workflows.
Businesses are now at the beginning of a similar learning process with respect to AI. It seems obvious that AI will eventually make many businesses more productive. But it’s going to take a lot of trial and error.
Just as electrifying a factory required more than just replacing a steam engine with an electric motor, adopting AI is going to require companies to change their internal processes. It probably won’t be possible to just replace an existing worker with an AI model.
And this seems to be exactly what the MIT researchers observed.
“Custom solutions stall due to integration complexity and lack of fit with existing workflows,” the MIT researchers wrote. In some cases, companies will be able to address that by designing new models that better fit existing workflows. In other cases, they will need to redesign their workflows to take advantage of the unique capabilities of AI models.
But either way, the process is likely to take years. And on an economy-wide basis, it could easily take decades, as knowledge diffuses from the most technologically sophisticated firms to the laggards.
In short, that 95% failure rate isn’t a sign that AI is useless. It’s a sign that companies are in the early phases of learning about a new technology.
This kind of thinking is anathema to people at both ends of the AI debate. The AGI-pilled believe it’s only a matter of time before AI systems become so smart that they can become “drop-in remote workers.” At the opposite extreme, skeptics are reluctant to admit that the technology is useful at all.
But I think the evidence so far supports a middle interpretation: this technology has immense potential. But like previous general-purpose technologies, it’s going to take years, if not decades, to fully unlock it.


Thanks for the coherent article. It is within my 70 years of life to see technical innovation lifecycle and maturation cycles are generational. Think about how long it too us to put wheels on suitcases, or move from mainframe computers to networked desktops. No biz plan stands contact with the consumer and the rate of change in technology. It is way too early to know the true value of Ai. It is evolving and it is a genie now out of the bottle. Remember the REDHAT business that grew from Shareware. There will be candy and blood before the answer is known.
I'm sure I've plugged this before, but the following essay by Cosma Shalizi expresses Karpathy's view in more generality: the singularity already happened, it was called the industrial revolution.
http://bactra.org/weblog/699.html