Dialogue with a16z partner: Large models do not have network effects and will ultimately only earn a "toll fee" like operators
Original Title: The Economics of AI Usage and What's Next For SaaS | Benedict Evans on a16z
Original Compilation: Yanlin Hang, Z Finance
In June 2026, Silicon Valley. Anthropic's annualized revenue has more than quintupled in the past 12 months, reaching $47 billion. The combined AI capital expenditure guidance from the four major tech giants this year exceeds $700 billion, nearly double the total investment in the global telecommunications industry.
As the entire industry convinces itself that "the risk of underinvestment is greater than that of overinvestment," Benedict Evans discusses the mobile data crisis of 2008 in the a16z podcast studio.
It was a chaotic time after the launch of the iPhone. AT&T introduced unlimited data plans, users went wild watching YouTube, and the network collapsed instantly, leading operators to spend hundreds of billions to expand capacity. Ultimately, all the cool applications were created by others, and operators only earned "pipeline fees."
Evans releases an annual talk titled "AI Eats the World," regarded in Silicon Valley as an important reference for observing technology cycles. Unlike most optimists, he tends to find bad news in history. In his view, the gap between today's $20/month ChatGPT subscription and the underlying token costs of over $10,000 stems from the same illusion that led to the $500 billion data bills of the past. Pricing is severely disconnected from costs, and everyone pretends not to see it.
"All bets are still open," he admits he cannot predict the outcome. But as model efficiency improves by 100 to 200 times each year, and nearly a trillion dollars of capital flows into this field, he believes at least one thing is certain: today's luxury of "ROI-based pricing" will not last long.
Here are Evans's six judgments on the core contradictions of AI economics:
Foundation models are not products; value will ultimately shift upstream. Model companies are likely to become mere water sellers, repeating the mistakes of chip makers, ISPs, and mobile operators. They build amazing infrastructure but fail to capture the most profit.
Programming is currently the only field that has truly found product-market fit (PMF). Agentic Coding has transitioned from "somewhat useful" to "transformative," but beyond that, most scenarios remain at the "open ChatGPT once a week to try it out" stage.
The pricing system is collapsing. As model efficiency improves by 100-200 times each year and CapEx floods in at the trillion-dollar scale, the luxury of ROI-based pricing today will not last. Tokens will eventually commoditize like mobile data, leading to price wars.
AI will not end SaaS, but it will redefine the boundaries of software. Should probabilistic LLMs be placed at the top or bottom of the tech stack? Enterprise software will enter a new round of chaotic games of "Excel vs. specialized software," leading to more software, more competition, and more uncertain profit margins.
History can explain but cannot predict. Analogies from the mobile internet, cloud computing, and PC eras are useful, but none can tell you whether OpenAI will become the next Windows or the next Netscape.
The real questions are leaving the tech circle. What does AI mean for law firms, investment banks, consulting companies, and Hollywood? These answers are not in San Francisco but in the hands of industry insiders who know "what junior employees are actually doing."
01 OpenAI and Anthropic's Strategic Divergence
Erik Torenberg: Benedict, welcome back to the a16z podcast. The last time you were here, we discussed the first version of your talk "AI Eats the World." It's been almost a year and a half since you finished writing it. You always start your talks with "What are the big questions?" But this time I want to ask first: What have we learned since you first gave that talk? Which predictions have come true? Let's review.
Benedict Evans: Let's talk about what has happened in the past year. I think we have a clearer view of the differentiation in product strategies and the competitive tensions—this competition is no longer just about "making the model bigger, faster, and investing more computing power."
OpenAI's strategy has undergone several shifts—from "betting on all directions at once yesterday" to "maybe we should double down on programming." Clearly, Agentic Coding has truly started to work. Thus, all the focus in the tech world is highly concentrated in this area, which has achieved absolute product-market fit, with strong customer demand almost overwhelming. Of course, this has also brought about supply shortages, and the contradictions surrounding capacity, pricing, supply-demand imbalances, and capital expenditure pricing are exactly what we see today. This is the point we are at now—once we thought this was interesting and exciting, but we weren't entirely sure what to do with it. Now it can indeed be used for programming, and whether it can be used in other fields is almost certainly yes, but programming is where it is truly effective right now.
Now the focus has narrowed significantly. Beyond that, data is continuously rising: models are getting larger, CapEx is growing, and usage is increasing, with people using it more and more. But most of the fundamental questions you raised two or three years ago still have no answers. For example, we don't know if there will be an absolute winner in the model space, whether they can capture value at the top of the value chain, where the boundaries of model capabilities lie, or when consumers will shift from weekly to daily usage with the current technology. So, all these questions remain unresolved.
Erik Torenberg: Speaking of programming, did we foresee that it would become the first truly explosive application scenario?
Benedict Evans: If you look at it from a deterministic perspective, you can reason this way: Who enjoys tinkering with these things the most? Software developers. And what do software developers most want to try to do with these things? Of course, it's software development itself. So from this very basic perspective, software development has the highest priority. I often liken this moment to the internet in 1997 or 1998, or the personal computer era in the late 70s and early 80s—everything was very exciting, but it wasn't clear what this was actually for because it hadn't truly matured. In the early days, the main thing people did with personal computers was to create more computers, and now the first thing people do with LLMs (and larger LLMs) is also to create more computing power. So this is not surprising.
But there was a noticeable shift at the beginning of this year: Agentic Coding transitioned from "somewhat useful" to "truly transformative." I'm not sure if anyone could accurately predict when this would happen or that it would be the first application to break out. Some people will claim they saw it coming, but I don't think anyone could predict with certainty when all this would happen, and that it would break through first in the form of programming.
Erik Torenberg: On an organizational level, what have we learned? What does it mean for junior engineers, senior engineers, and the job landscape and team organization?
Benedict Evans: I think we can't really say anything yet. Six months ago, this thing was not usable at all. Now everyone is frantically trying to figure out what it actually means. If you get too caught up in the noise and details, and take something someone said at an event as a sign that the sky is falling, you will get lost. It will take at least two or three years for all this to stabilize, not to mention that there are huge supply-demand contradictions just in pricing, leading to various surprises. So we really don't know what future teams will look like.
I think people are starting to ask some new questions, the most obvious one being: Are you still hiring junior employees? If so, what will they do? Why did you hire junior employees in the past? Were they hired to do what they could do themselves, or to do something else? If you've automated an entire category of work that was previously done by people, what happens? This question has become more real in the software development field because you are indeed automating a lot of tasks that were previously done by people. So these questions have shifted from theoretical to practical. But I don't think anyone can claim to know what the market structure will look like in the next three to five years, nor what the career paths for software engineers will become—if you think you know, then you are crazy.
Erik Torenberg: Let's talk about OpenAI. What has surprised you the most? How do you understand their strategic evolution and the future challenges they face?
Benedict Evans: This has always been a place full of dramatic conflict. Clearly, their CEO is on medical leave, which has changed the situation somewhat.
In the second half of last year and into the fourth quarter, the questions from the outside were: The model itself is pretty good, but what else? How do you get people to use these things for other purposes? It was almost like saying, "Go ask ChatGPT for 15 ideas on how to create value based on infrastructure, and then do them all," and OpenAI was doing just that. Meanwhile, Anthropic, with relatively weaker funding, said: We focus on coding. And then they really delivered on coding. Whether this was intentional or a lucky coincidence is up for others to judge. But clearly, this approach worked.
However, the problem remains: what is really working now is software development and some other specific scenarios. Many people are just excitedly testing the waters on the margins, using it in some scenarios. The differentiation within Silicon Valley is also quite evident—on one side are those who bought piles of Mac Studios and are running open-source models around the clock, while on the other side are the 40-50% who think this stuff is somewhat useful but only used it once last week. The question is, how do you bridge this gap? I don't think there is a simple answer. Software development has indeed crossed that gap, but many others are still scratching their heads, using it only superficially.
Additionally, many companies are using AI to automate specific backend processes—here, you are not letting users figure out what this new tool can do; instead, you are directly telling them: Here is a problem we can solve. When I talk to companies outside the tech industry in the U.S., and chat with consultants and investors, they are examining these point solutions one by one.
For example, I spoke with a commodities company a few days ago; they want to use LLMs to improve cash flow forecasting because they deal with many small producers and are uncertain about when they will receive payments, which is a low-margin business, so cash flow forecasting is a big issue for them. This is completely different from just going to ChatGPT or Claude and saying, "Help me write something."
Erik Torenberg: How do you see the situation compared to early adopters using it weekly or daily?
Benedict Evans: I think there are several different dimensions to answer this. First, we are always progressing on the shoulders of giants, and progress is accelerating. The mobile internet didn't need to wait for the internet to appear; it just needed the cellular data networks to be in place. The internet didn't need to wait for personal computers to become widespread, and personal computers didn't need to wait for consumer electronics and semiconductors to mature. So the adoption speed has been accelerating. When your boss Marc Andreessen was doing Netscape, there were only a few tens of millions of personal computers worldwide. You couldn't have 900 million weekly active users—because there simply weren't 900 million computers. So the acceleration has always been there.
The second point is: in the early stages of these transformations, no one can see how they will work, and in reality, nothing is usable. I am old enough to remember these things. I don't know how old you are, but people in their thirties may not remember that era—when you were halfway through working, everything on the screen suddenly froze, and you had to crawl under the desk to unplug the power and pray that you could at least recover part of what you had done in the past hour. That kind of thing doesn't happen anymore.
Back in the 80s, you could spend $300 on a sound card, and the computer might not even produce sound properly. That $300 spent would take a weekend to get it working. I remember those days of tinkering. The internet was the same. You had to first get a floppy disk with TCP/IP protocol, which was painfully slow, and the things you wanted to do didn't have ready-made tools. The mobile side was the same. We are currently at that stage. Of course, the key question is: which of these things will ultimately succeed? The same applies now: will browsers succeed? Will this thing stand firm? How will all this fit together? There is a gap between those extremely exciting things and the small group of people willing to invest energy to make them work, and what we need to do is turn it into something that can be done with one click.
02 Pricing Crisis and Lessons from Platform History
Benedict Evans: The third point is that the unit economics have become more visible. I observe that the pricing tightening we are talking about now is very similar to what happened in the mobile data space from 2009 to 2010. On one hand, people suddenly received data bills amounting to $5,000 or even $10,000; on the other hand, if you had an unlimited data plan—like the one AT&T launched exclusively with the iPhone—everyone bought the iPhone and started using it, and when the 3G network was turned on, everyone began watching YouTube, and the entire network collapsed because there simply wasn't enough capacity to support it. Interestingly, there are still people in the tech circle who don't understand that cellular networks have marginal costs. They had to expand capacity, and expanding capacity costs money. Operators had to frantically adjust their cost curves to align with the pricing system of the infrastructure, the underlying costs, and the perceived value by users—they roughly achieved this through tiered plans, reasonable use terms, throttling, and so on.
But the other side is exactly what we see now: you pay $20 a month but use tokens worth over $10,000; conversely, if you casually played around for two days and suddenly received a $10,000 bill, you would definitely exclaim, "What the hell is this?" You can now see the exact same news stories, which is precisely what happened from 2008 to 2010, and it is similar to what happened in the GPU space from 2001 to 2003.
But the more interesting part of this analogy or comparison is that since then, mobile data traffic has grown by about 1,500 to 2,000 times. The total revenue of the global mobile internet network is about $1 trillion, with annual CapEx around $200 billion, while ARPU has not grown for 20 years. All those cool new things were created by others. Operators once thought all the good things would be built by themselves—I once worked at a telecom company with a banking license because they thought they would create mobile banking themselves—which seems completely crazy now. But the key point is this: they built this amazing, global, extremely complex, and expensive infrastructure, usage continues to grow, changing everyone's lives, and we all pay for it—but they themselves haven't made much money from it because all the value has shifted upstream.
And this is the core problem facing LLMs: can the model itself do everything, or do you need to build 300 applications on top of it? Can you directly tell the model, "Help me file my taxes," or do you need tax software that internally uses 10 different AI methods to handle it? If the answer is the latter, then what is your positioning as the provider of the underlying foundational model?
Will it become a commodity infrastructure sold at marginal cost? Currently, this seems like a hard concept to accept. Because you can sell all the tokens you can create, you can price based on ROI. But in the coming years, with about $1 to $2 trillion in CapEx investment and model efficiency improving by 100 to 200 times each year, new models will emerge. Will models consume more or fewer tokens? We will reach a different equilibrium point. And when models perform similarly, do the same things, and use the same chips, what basis do model companies have for pricing power?
Looking back at history: chip companies did not capture value, ISPs did not capture value, and mobile operators did not capture value. Windows and iOS captured value, but what they were doing was completely different: they had various levers to move upstream and network effects, while models do not. So the question is: will model companies end up like the infrastructure layer, or will they capture value like the operating system layer and decide what others can do? Ironically, the story of Netscape illustrates this point—Marc Andreessen famously claimed he would turn Windows into a bunch of poorly debugged device drivers, and then Microsoft forced its way into the market. But it ultimately proved that the web browser itself was not key—the value was elsewhere. So these swirling big questions remain unresolved, ultimately returning to what I said earlier: there are some things you can know, but you have no idea how it will all end.
Erik Torenberg: Yes, it's still unclear whether this will resemble the internet (where most value and better profit margins occur at the application layer) or more like the cloud (where value seems to be at the hardware layer, and currently, Nvidia appears to be making the most profit). But will it always be this way? Or will it become more like the internet? How do we start predicting this answer?
Benedict Evans: I have two answers to this. There are many famous sayings about how history works—one of my favorites is: History teaches us only one thing: that something will always happen, and you can always explain afterward why it had to happen, but it wasn't obvious at the time. Especially, I remember about 15 years ago, many very smart tech people got their hands on the iPhone and Android and said: This is another round of competition between open and closed systems; we are going to take down the iPhone, but that obviously didn't happen. I can explain the reasons afterward, but all these analogies are useful, yet none are predictive. History is always obvious in hindsight.
Interestingly, I recently did a few podcasts and released this talk, and then there was a type of comment saying: Benedict, you haven't done your job; you should tell us what will happen in the future; your duty is to make predictions. But you seem to only say, "We don't know." There are two issues here. One is that in some places, I do say what I think won't work and what will work—for example, I believe foundational models are not products, chatbots are not products, and value will shift upstream. But on the other hand, at this stage of the cycle, there are too many variables, and you don't know which one will ultimately prevail. If you insist on saying, "I think it's that one," you might be right, but you must realize how much uncertainty there is and how many different paths could lead to it. That is the nature of this stage of the cycle: all bets are still open.
Once the S-curve starts to rise, it will narrow. There was a moment when Windows Phone seemed like it could succeed—of course, in hindsight, it seems unlikely. There was a moment when it was unclear how the mobile side would go, and then it became clear—this is what is happening, and then we turn to the next question.
One characteristic of technology is that when you truly understand something, know how it works and where it is going, that is precisely when you should turn to the next thing. You should always be looking for those quiet corners where we still don't know the answers. I haven't updated my Apple spreadsheet in five years—because we already know the outcome. I don't care what the next generation iPhone looks like; I'm not even concerned about their market share in China anymore. The outcome is set; it's time for the next question.
03 Foundation Models Are Not Products; Value is Upstream
Erik Torenberg: You just said you don't believe foundational models are products, and you think value will shift upstream. Please explain your reasoning and what that might look like.
Benedict Evans: I think we can discuss three or four building blocks.
The first is: currently, it is difficult to create a model that is fundamentally better than others and can maintain a differentiated advantage. It has no network effects, no levers you can pull, or strategic positions like Instagram, YouTube, or Google Search. LLMs do not have corresponding elements. Of course, different models have their strengths—maybe this one is a bit better than that one, or you prefer this one—but there is no fundamental difference between them; the only distinction is how much you are willing to pay.
The second issue: chatbots themselves are a strange, limited-functionality V1 UI. They are indeed very useful in certain scenarios, for certain people, and for certain tasks, but in most cases, you still need a whole bunch of other things. You need a toolchain, you need the right configuration, you need the appropriate data, and you need to set up various controls and user interfaces. Someone needs to have seriously thought about how this tool should operate because the people good at using tools to get work done and the people good at deciding what the tools should look like are usually two different groups.
People skilled in graphic publication design are not the ones who should be doing InDesign; that is a different set of skills. People skilled in financial consulting are not the right candidates to design financial robots—that requires different skills and different people. Now we are exploring that middle ground: Claude has this, Claude has that, and there are Skills, etc. To me, it feels a bit like, who will create the Skill? Another question is, this looks a bit like the templates you see when you select "New File" in Excel—they can only go so far, and at some point, you will break through the template. My presentation has a slide quoting something someone said to me on Twitter years ago: he said he was a consultant, and half his job was teaching others to use Excel as a database, while the other half was teaching others to use databases as Excel.
So there is a fuzzy, chaotic space here: do you need specialized software? Do you need horizontal software or vertical software? Or can you just get everything done in Excel? We have all seen examples of entire departments running on a 10MB Excel file—my own business also uses Numbers spreadsheets. But at some point, you will exceed that limit.
So can model companies do all this? Of course not, just as Microsoft or Apple cannot create every app on Windows and iOS. So do model companies have that kind of leverage? Are they Windows or iOS? Again, is there a network effect?
For example, if you are a law firm and you want to buy a piece of software—many enterprise software companies have been invested in by a16z—would that law firm or manufacturing company or bank ask, "Is this using Claude or OpenAI? Because we uniformly use Claude"? No, that is not how it works at all. In the cloud era, it wasn't like that either; you wouldn't say, "Our company uniformly uses AWS"; you wouldn't even know which cloud a certain SaaS product runs on. That is the key; it has been stripped away; it is not something you need to care about. So in that sense, foundational models are more like cloud service providers—they may have some competitive advantages, but they lack that kind of leverage, network effects, and control.
This also reminds me of another comparison: the semiconductor industry—each generation is getting more expensive, and there are fewer participants. Overall, models are essentially a different commodity; chatbots are not the correct UI or product, and model companies cannot create everything themselves. So they are the underlying infrastructure. Do they have pricing power?
In the future, there will probably be about 3 to 6 companies producing cutting-edge models, investing each year—no one knows for sure, but it is probably $200 billion to $2 trillion—plus a batch of edge models and open-source models. So ultimately, there will be five or six companies competing in this market to sell these products. Where does pricing discipline come from? Especially since some of those companies have completely different business models. For example, Google makes money from advertising, and their attitude towards pricing is different from OpenAI's.
I think the difficulty lies in the gap between the state we are currently in and the state we should ultimately reach—this is actually a topic discussed in freshman economics courses. We are currently in a period of extreme imbalance between supply and demand, pricing, CapEx, and capacity. But the demand for tokens is infinite, which does not mean you cannot reach a different price equilibrium—because mobile data has gone through this. Over the past 15 years, the demand for data bits has grown by 1,500 to 2,000 times, but you still see market price equilibrium between supply and demand, and there are still fierce price wars among operators in most parts of the world. Fundamentally, you are selling a commodity, and customers can switch suppliers at any time—developers will also switch back and forth.
Of course, I fully accept that this could be wrong. Maybe in the end, there are only two companies in the world that can produce LLMs, and they have pricing power. Or we enter a world where almost everything is integrated into the model itself, and the model has leverage upstream. But my core point is: just like the iOS vs. Android battle—you can certainly say it has gone this way the past three times, but you cannot prove it will be the same this time. But at least you should ask these questions and acknowledge that the current situation is temporary. We are in this state of extreme scarcity, and then there will be pricing systems, free markets, and about $1 trillion in CapEx flowing in. So these multiples will change.
04 Where is the Next Breakthrough After Programming?
Erik Torenberg: This leads perfectly to what you said earlier about "we already know what Apple looks like." So what is the next question you are most focused on? Or what should we be most focused on?
Benedict Evans: We have already discussed some issues, such as how far the capability stack of models can go, whether models can achieve differentiation, etc. Another obvious question is: at what point will we see more and more categories of use cases where models are good enough that we no longer need the most expensive, fastest, largest, and heaviest models in the cloud? We can use older models, open-source models, or models running on devices. This is precisely the story Apple will tell next—how much can be pushed to the device, where computing power is free (or at least free for you, with no marginal cost for developers).
Another classic question is: the problem itself has started to leave the tech domain. For example, if you look at a law firm, a consulting company, or an investment bank—basically all traditionally pyramid-structured professional service organizations—what happens if you can automate a large portion of the work done by those at the bottom of the pyramid? I can only say that if you haven't worked at a law firm or at Bain, BCG, or McKinsey, you probably won't be able to figure out what is going on because you may not even know what those junior employees are actually doing or what clients are really paying for. How will those roles be restructured? What does AI mean for finance? This includes both internal hiring structures and the types of products and profit margin structures you can create. What does it mean for the consulting industry? What does it mean for the Big Four, the Big Three, Accenture, large law firms, and advertising companies?
You might know some of these questions, but if you are not in the industry, you have no idea what the answers are. This reminds me of something I often say at a16z: "Content is not king." I also wrote that "Netflix is not a tech company"—what I mean is that Netflix's entire business is supported by the infrastructure built by the tech industry. But all the problems Netflix faces are Los Angeles problems (content issues): what shows to select, how many shows to produce, what types of shows? How much to spend on talent? Should they chase awards? Should they make movies? Should they buy sports rights? These are all Los Angeles problems, not San Francisco problems. San Francisco doesn't even know what the right questions are—they are media industry questions. All the truly important questions for Netflix have become media industry questions.
Similarly, whether Tesla is a car company or a tech company has always been a point of debate. What I want to say is: what does AI mean for the legal industry? This question is both a technical question and more of a lawyer's question—you need to understand how law firms actually operate and what clients are really buying. Similarly, what does generative video mean for Hollywood? Ben Affleck probably knows much more than I do—he founded a company and sold it for hundreds of millions. So this is the second type of question: the questions are leaving the realm of AI itself and becoming hybrid questions of half AI and half other fields.
The third level—perhaps I should have said this earlier—the fundamental difference between all this and past platform transformations is that during the 3G, iPhone, and internet eras, although you didn't know what would happen next, you knew the physical limitations. For example, in 1995, you knew telecom companies wouldn't install broadband for the entire world next week; you knew not everyone would buy a PC because a PC cost $3,000. So you knew where the basic impossible boundaries were.
But with generative AI, we don't know. It is possible that after we finish this episode, the phone will push a notification saying OpenAI's new model has been released at only 2% of the previous price—because of a new technological breakthrough. I don't think this is very likely, but we don't know the answers to these types of questions. How much will models grow? How much better will they get? How much faster will they become? How much cheaper will they be? In what areas? How will the features of models change? We don't know. This is different from all previous platform transformations—before, you knew where the basic limitations were. And this will give rise to a series of new questions.
In a sense, I mentioned earlier that the only area with product-market fit right now is programming. Other areas do not yet have an equivalent level of PMF. I can safely say that Anthropic's revenue has risen from a run rate of $9 billion last year to $47 billion now—all from software development. So what happens if someone in other fields creates something usable?
Erik Torenberg: For example, law firms, banks… If you had to guess, besides programming, what use cases might lead to daily active usage?
Benedict Evans: The talk I released a few weeks ago is roughly divided into three parts. The first part discusses capital, CapEx, infrastructure, and the differentiation of foundational models—this is what we just talked about. The second part is: how do you use these things to build software? What does this mean for the software industry? What will software look like? What changes will occur in profit margins and company structures?
The third part I call "change." I opened with a quote from Yogi Berra: "Predictions are hard, especially about the future." I think there is an interesting backtesting perspective: imagine asking questions about the internet in 1997; what would you get? What wouldn't you get? I think one way to look at it is: this is a form of automation that turns a category of things that people used to do but couldn't automate into something that can be automated. So what does that mean? I proposed three or four buttons you could press.
The first is price elasticity, which is what Gerber's PowerDNS is really doing: if the cost of doing things decreases, are you doing the same amount of work for less money, or are you doing more for the same amount of money? Or are you making more money because you are doing more? Is there something that was previously impossible that has now become cheap? Is there something that was very expensive, serving as a barrier to entry—like owning a printing press for a newspaper—that has now disappeared? Has anything been unlocked in the business model or competitive space because your costs have decreased?
The final question is: what things were previously completely impossible, entirely due to high costs that no one ever thought about, that have now become accessible? A common example I use is that the steam engine made trains possible; you could buy as many horses as you wanted, but you couldn't create a train that spanned east to west. A more modern example is YouTube or Spotify. Spotify says: look at the history of the music industry over the past 25 years; the first half was "you don't have to spend $1.50 to buy a CD just for that one song," and the second half is "for $15 a month, you can listen to all the music you can find," which was completely impossible before.
The problem with these types of predictions is that on one hand, you can be clever and say something obviously correct, but you actually don't know what it means in specific industries. For example, at the end of the 90s, we said the internet would destroy the value of physical distribution—this meant completely different things for newspapers and movie companies: newspapers were destroyed, while movie companies were hardly affected. So it still depends on the specific situation.
There is also a part where I think we can ask some more useful questions. One question I'm particularly curious about is: how will AI change advertising, e-commerce, branding, and our consumer behavior? Advertising is a trillion-dollar market, and retail is $25 trillion; this is a considerable market that can be reached. I have been thinking: Google, Meta, and Amazon actually don't really know what those products are. They know SKUs, they know what publishers entered in the metadata fields, they know "people who bought this also bought that," but they don't know why, nor do they know what those things actually are. That's why you get the joke: Amazon, I bought a toilet seat, but I'm not collecting toilets—because Amazon doesn't actually know what a toilet seat is, nor do they know that normal people wouldn't buy two. In reality, they should be able to know using frequency analysis, but they haven't done that. And with LLMs, in principle, you can know what those things are, why people buy them, and what else people might buy.
Of course, the word "know" is hard to define. But at least AI systems can provide a completely different dimension of statistical color—this is why Google and Facebook's advertising revenue and conversion rates have been skyrocketing every quarter—they have integrated AI into their advertising systems, recommendation engines, and predictive algorithms. You will see more things you like, and the ads you see are more likely to be things you want to buy. So their advertising revenue has seen a sudden acceleration.
Overall, if you look at how these systems operate now: they say, "people who bought that also bought this." And now you should be able to do this: here is a picture of a coat: what is this? Where can I buy it? Ten years ago, this was absolutely impossible; five years ago, it might not have been possible; now it should be possible. Then you can also say: Help me recommend 10 similar coats at different price points, tell me where I can buy them, and list the pros and cons of each; you should basically be able to get that too. Further: look at my Instagram, help me recommend a winter coat I should buy that will change my style but not too much. Three years ago, this was completely science fiction; now you would think it is indeed possible to create something that is almost usable.
And these changes—the things computers can know, what they can automate, what suggestions they can make—return to the fundamental question. Every time a new technology emerges, you first use the new technology to do old things: more spreadsheets, more PowerPoints, more emails, better emails. But the important thing is not to do old things better but to do new things that old technologies could not do at all. This is a very cliché observation, but we often forget it. So what are the things you can only do with this new thing, rather than just automating the old things?
The enterprise version might be: you recorded all your Zoom calls with clients, you have all the email flows in Salesforce, and you have all user behavior analysis data and metrics. So how should you adjust pricing to improve churn? This is what LLMs might be able to do, which is different from "doing sentiment analysis on call centers and telling me which customers are angry." You are experiencing multiple transformations at the abstract level of analytical capability. Of course, this will spawn new companies, destroy old companies, and create new businesses. But then again, we are now in 1997, and I want to predict Uber and Airbnb. If I could really predict that, then we would be living in a parallel universe. The hit rate for venture capital wouldn't be one in ten; it would be one in ten.
Erik Torenberg: Yes. One of the questions we are asking now is: what things were previously outrageously expensive that have now become possible? For example, rebuilding YouTube from scratch? Or rewriting the Linux kernel?
Benedict Evans: Interestingly, another observation is that new companies always say, "We want to redo the old things with new things. Of course, we want to redo Office with open source; we want to rebuild it on the web." What is the result? Look at Google Docs; its market share is about 20%—because that is not the key.
What is really interesting is to create something entirely new, to shift the level of abstraction, and to discover problems that didn't exist before. Sitting in a venture capital firm all day listening to pitch presentations, you will find that some things sound somewhat useful, while others seem unlikely to succeed. But some things feel like they fill a void in the universe—when others explain it to you, you immediately think: wow, why has no one done this before? Why has no one discovered that this problem exists? This is what makes watching startups so exciting. And this is precisely what people will use AI to do—people will suddenly discover a way, realize that a certain problem has always existed, and that no one, including those who have this problem, has realized that problem exists, and then they will create a tool to solve it.
This also brings me back to my earlier point: this is why I don't believe models can do everything. Think back to all the pitch presentations you have seen at a16z; how many of them were problems that people in the industry already knew existed? The answer is often no. In fact, no one in the industry thinks it is a problem; it usually takes two years to explain and convince them that the problem really exists before this new thing can help them solve it. That is the problem; you cannot expect an ordinary manager in a finance department to use this tool to solve a huge global industry problem—because no one even knows that industry problem exists, let alone come up with the right tool solution.
05 Will the SaaS Landscape in the AI Era Be More Decentralized or More Centralized?
Erik Torenberg: Does this mean that the SaaS environment in the AI era will be more decentralized than before? Fewer bundles, fewer giants like the Microsoft enterprise suite?
Benedict Evans: Back to the topic of SaaS. Let's set some building blocks first. Clearly, building software will become cheaper and faster. There will undoubtedly be a whole bunch of things that software can do that were previously completely impossible. Therefore, competition will be fiercer. Of course, this also comes with new profit margin structures. But as we discussed earlier, we don't really know what that profit margin structure will look like.
Will it lead to outcome-based pricing? It is very difficult to link every keystroke in enterprise software to the profit and loss statement; sometimes it can be done in Salesforce, but for the vast majority of software, it is hard to say, "The work I did today had this much impact on earnings per share, so we should pay this much for it." I don't think that makes sense, at least not in the long term. But how will the pricing structure evolve? There will definitely be more competition, and building software will be easier and faster.
I think there are two useful frameworks for thinking about this issue. The first is: look at today's enterprise software landscape; there are three main categories. The first category is large horizontal systems—SAP, Workday, CRM, human capital management software, payroll management software, etc. The second category is vertical software—a typical large American company probably has 300 to 400 SaaS apps, plus thousands of internally purchased or built apps running on Teams. The middle ground is a fuzzy impromptu space made up of Excel, email, and shared file systems, where things move back and forth between these three. In principle, every SaaS app is doing something you could also do in SAP or Excel, like managing campus recruitment in Workday.
But at some point, for example, I have talked to people about this: if you are PwC and you need to hire thousands of graduates every year to train as accountants, you might have a set of custom software, or you might hire Accenture to build a set, and you might even hate it. But if you are a company that only hires five graduates a year, you are just getting it done in email and shared Google Sheets because why would you buy a whole software suite? The middle ground can use Workday, Excel, or specialized apps to do it. Now you add ChatGPT into the mix: are you using LLMs to do this? Is there an LLM tool that allows you to accomplish something in Salesforce that you couldn't do before? Or in your vertical software that you couldn't do before? You are building a tool for yourself using LLMs—just like a department in a company operates on a 10MB Excel file built 15 years ago, and no one knows how it works, but everyone is still using it. So LLMs are entering this vast, fragmented, and complex landscape as another option for completing tasks.
I think another thinking framework is: should LLMs be at the top or bottom of the stack? On one hand, being at the bottom means it is a feature within Salesforce; you are in Salesforce, and the system looks at the history with that customer, the context of all other sales calls, business goals, and then helps you generate an email or suggests what to say when calling the customer. This is a controlled feature, with a toolchain, with guardrails, driven by that specific use case. On the other hand, it is the example I just mentioned: looking at Salesforce, Workday, all emails, and Google Analytics data, and then synthesizing an analysis that was previously impossible. So the dilemma is: where do you place probabilistic, potentially error-prone software? Where do you place deterministic system software? Where do you place the database? Where do you place LLMs? Is it at the top or the bottom? It could be both, depending on what you are specifically doing.
Ultimately, this is about what software means—more software, far more software. The very existence of all software companies is to solve the problems created by other software companies. This is the classic joke: the purpose of all security software is to solve the problems created by other security software. Clearly, the SaaS era has already given us an explosion of software by an order of magnitude or even two. This time, we should expect the same thing to happen.
As for the doomsday of SaaS, investors look at all these companies and say, we don't know which companies will be wiped out by all this. Some companies will definitely fail; a certain proportion of existing SaaS companies will be eliminated by this wave, but you don't know which ones, so you shouldn't directly devalue the entire industry by 50%. But you should definitely say: until I figure out what all this is about, I won't go all in on SaaS.
Erik Torenberg: You mentioned in your conversation with Ben Thompson that software is designed by someone sitting down and designing a workflow, then saying this is the correct way to do this from now on. But you also said that processes grow out of how businesses operate. Does this take time? Or do we need more experimentation and iteration—from those vertical AI startups—to find the right form of software for the future?
Benedict Evans: In a sense, there is an interesting overlap between what strategy consulting firms and software companies do: they both observe what is happening within a company and say, "This way is too bad; a better way can help you achieve your goals." Software companies encode this way into software, while strategy consulting firms encode it into workflows, job responsibilities, processes, training, and objectives, and they may also suggest that they buy a software suite to do this—or increasingly, directly help them build that software.
Another point to discuss is: how much of the work within an organization is implicit, undocumented, not in the training data, and not something anyone in the company can sit down and draw a flowchart to explain clearly. This constitutes a large part of the value of BCG and McKinsey. They have the authority to enter a company and talk to everyone, including those in different departments who are not allowed to talk to each other (and who won't be fired for it), to figure out "how this actually works," rather than "how it should work"; and why people are not executing the strategy because, in reality, their bonus targets depend on them not executing that strategy. Then they provide you with answers as an external team, and you can shift the responsibility onto them.
These are all issues in organizational management and personnel operations—how people operate, how they explain what they do—these are hard to write down and difficult to directly bake into a Skill and say, "Here you go, maybe make a PPT." So there is a bigger challenge: how to get people to use these technologies? How to get users to adopt new tools? How to help people adopt new tools and figure out what new things they can do with them, which has also happened in the cloud, web, mobile, internet, PC, and spreadsheet eras.
Erik Torenberg: Do you think there will be a common evolution between AI-native software and new interfaces? For example, new customer service AI platforms might not need as much user-facing UI, or system software might not have a frontend at all? Because its main users are the AI agents querying it directly?
Benedict Evans: These are all interesting ideas; I find it hard to have strong opinions because I haven't delved into the details of how enterprise infrastructure is built. What I am curious about is how new these questions are. I remember about 10 to 15 years ago, Chris Dixon said: APIs are the new frontier; software no longer needs software companies; you just need to open your API. So old things will always come back in new forms: now you don't need an API; you just need an MCP server, and the agent will connect directly. I don't know.
I think the biggest challenge with these things is that all decisions are essentially exception handling. The question is always: what can you not automate? What requires someone to make decisions, judgments, and have their opinions—because that thing may have never been written down, never happened, or looks different from before.
The distinction I use in my talk is between tasks and positions. The tasks used to complete a position may change, but the position itself may not change much, or what that position delivers to the client may not change much. Think about accountants 50 years ago and accountants today—the core things they do are almost none the same. But from the client's perspective, it is almost the same thing, just accomplished in completely different ways through a series of completely different tasks.
I think a deeper or more abstract way of thinking is: in what areas do you want the app to provide "the answer everyone does"? That is the answer everyone wants, the answer anyone can give, the answer any junior employee can provide, the answer anyone can give me. And in what areas do you not want that kind of answer? In what areas do you want an answer to a new question, a different answer, or a different idea? Because LLMs will be very good at anything you can describe how people do it, and what you want is just what an average person would do. In areas where they are not good, it is where you cannot explain why you are doing it that way, and what you are doing is different from what others do.
06 The Fate of Model Commoditization and Historical Lessons
Erik Torenberg: Many people, including Google's CEO, have said that the risk of underinvestment is greater than that of overinvestment. Is there a level of CapEx that would make this statement no longer valid? Are we approaching that point now?
Benedict Evans: First, there is a financial gravity issue: Microsoft's, Meta's, and Google's CapEx this year has reached about 50% of their revenues. The telecommunications industry is considered capital-intensive, but their CapEx is only 15% to 20% of revenue. The guidance from the four major companies this year is $700 billion. The total for the telecommunications industry is $300 billion, with mobile at $200 billion. Oil and gas, depending on how you calculate it, is probably between $700 billion and $1 trillion. So $700 billion a year is not an impossible huge number—this is the normal cost of large global infrastructure; it's just that there is indeed a lot of money.
Clearly, these companies cannot spend $1.5 trillion next year; if they do, they will have to borrow money, and it is impossible to maintain this level of spending in the long term. So at some point, growth will inevitably slow down because there is no more money. Of course, you can talk about ROI and investment returns. The capital markets are also willing to provide funding within a certain range. But just throwing out a number—you cannot spend $10 trillion a year on AI infrastructure—because there simply isn't $10 trillion to spend worldwide. So there is a kind of physical limit.
I am currently hesitant to say anything more specific than this. I almost have to return to what I initially said: we currently have a bunch of multiplicative questions—the demand far exceeds supply. But on the other hand, efficiency is also improving significantly. We don't know what the next model will look like. We don't know when edge computing and open-source models will join the fray. And you are always chasing the latest model. This is the overarching theme—one model only remains relevant for 3 to 6 months or 6 to 9 months, and it costs billions of dollars and how much infrastructure to create.
I believe this situation has not truly stabilized yet. Clearly, there are many very smart semiconductor analysts spending a lot of time trying to assign values to these numbers—this is a bit like valuing internet bandwidth at the end of the 90s—you don't even know what the rows in the spreadsheet are, let alone what they are worth. You can only say it cannot be infinite; there are physical limits.
Another way to answer this: if you are Google, Meta, Microsoft, Amazon, or Apple, this is a matter of survival to some extent. You have a kind of FOMO issue. On one hand, your current return on investment is very positive. On the other hand, you cannot let others run away in your absence, or the company will be finished. You don't want to be like Microsoft in the 2000s, IBM in the 90s, or Intel in the 2010s, constantly being beaten down by Apple. If this is the future of the computing ecosystem, you must participate. But at the same time, the CFO is sitting there saying: fine, but to what extent do we participate? Clearly, at some point, CapEx growth must slow down because you simply cannot get that much money.
Erik Torenberg: Will there be a reckoning moment regarding token waste? Is it possible that companies overuse AI, and when they do formal ROI studies, they will cut back?
Benedict Evans: Clearly, some people are chatting online using the most expensive models—just like in 2010 with mobile, you receive a $10,000 bill and then say, "Wait, I thought this was an unlimited data plan." So there will definitely be some silly and painful anecdotes. But I think the more interesting question is: as I have said several times, we are in a severely imbalanced Ponzi moment; pricing must realign with costs, and usage must align with pricing and ROI.
The difficulty is that at this early stage, it is hard to know what ROI is. It is a bit like the late 90s internet; you say, "Go become more efficient." If you look at the Deloitte and Federal Reserve surveys I quoted in my talk—you ask CFOs if they see returns; currently, most returns are those hard-to-quantify things: more accurate analytics, better customer support, higher productivity, you can do more slides faster, do analyses faster. It is very difficult to assign a financial value to it. It has financial value, but it is not the same as saying, "We created a new product with AI that brought in this much revenue or saved this much money." Establishing a new revenue line is far quicker than sending everyone to use AI to do Excel.
Another answer, of course, is consumer surplus—just like Excel back in the day. If doing a DCF valuation model took a week, you might only do one or two. If doing one only takes 10 seconds, you can do 50— but you can't charge extra for that. So part of the result is that these become competitive necessities—everyone must buy and use them. But the cost savings or productivity gains you get from them will be eaten up by competition. You can't charge more.
If you are McKinsey, Bain, or BCG, and an analysis that used to take a week now only takes a day, you might do five times the analysis volume but charge the same to clients. Your cost structure hasn't changed; this is what happened in investment banking and financial analysis: you did much more analysis with fewer people, charging clients the same fees.
Erik Torenberg: One of the core arguments in your paper is that models will ultimately become commodities. However, the fastest-growing and highest-funded area right now is precisely these foundational model companies. What advice do you have for them, whether overall or for a specific company?
Benedict Evans: I am not sure they will necessarily become commodities. My stance is more like: there is a chain of reasoning here; looking at it deterministically, these things seem likely to become commodities—so please explain to me why they won't. I only commit to this extent. As for raising so much money, I return to my earlier point about the mobile industry; while it is not predictive, it is a noteworthy observation: the mobile industry is very large, spent a lot of money, but profits are not high, and all the cool things were created by others.
Then you can look at capital return rates; the answer depends on whether you are in the U.S., Europe, India, or China. But at the same time, this is worth doing and has indeed brought returns for some, but ultimately mobile operators did not control the narrative; others captured greater value. What was Google's net profit last year? About $50 billion? What is the net profit of the entire telecommunications industry? I should subscribe to a Bloomberg terminal to answer that directly, but I can confidently say that the combined profits of Google, Meta, Amazon, Microsoft, and Apple exceed the entire telecommunications industry.
So this is a puzzle: you are pushing the frontier forward, you fall into a trap, and you must keep competing, or others will do it, and you will fall behind. One thing we haven't talked about at all is: are we building AGI? Are we going to create a "God box"? Some people already believe this, and while it is hard to analyze, maybe that is the case. So in any case, you will continue to build.
But the real question is: how do you create something that people want to use that is not software? Software is a good business, but is it the only business? If spending hundreds of billions makes the software industry more efficient, that is great; that is worth a trillion. Then what? How do you expand it to other parts of the economy? Expand it to everyone else? This is why you see these discussions—partnering with private equity, partnering with consulting firms—just as we discussed earlier, if you are running a physical company, figuring out what to do with these things is actually quite difficult. So you will seek out Bain, BCG, McKinsey, Infosys, Cognizant, IBM, Accenture, or private equity stakeholders. So on one hand, you are building larger and larger models, and you feel you must keep going. But on the other hand, what are people actually using it for? Why do most people open ChatGPT and still struggle to think of what to do today?
Erik Torenberg: Is there anything in your talk that you particularly hope the audience remembers?
Benedict Evans: I used an IBM ad last year and again this year, from the early 1950s, showing a large group of engineers holding slide rules, with the ad copy saying, "One IBM electronic calculator is equivalent to 150 professional engineers." How many of these slides have you seen at a16z? We remember that whenever these foundational technological transformations occur—every 10, 15, or 20 years—they bring about astonishing changes that completely alter everything and are entirely different from anything that has happened before. So AI is amazing, transformative, and completely different from anything that has happened before.
But mobile was also a huge transformation, the internet was, the PC was, and computers themselves were—all of these were very significant at the time, and it was also hard to predict what would happen next. So as a baseline scenario, we should assume: well, we are going through this again. This will produce a bunch of things that will ruin people's lives and cause a group of people to lose their jobs. There will be some things we are not too happy about, and there will be some things we all think are great. Twenty years from now, we will forget there was ever a world where computers couldn't do those things.
We have now been on this call for an hour; the computer hasn't crashed, and we are still transmitting high-definition video to each other. Of course, it works fine; in fact, I am doing this with my iPhone. My iPhone is wirelessly transmitting video to my Mac, like magic, and we no longer think it is remarkable. I think this is my description of what I believe the endpoint of all this will be: it will become magic; 20 years from now, we will say: of course, computers have always been able to do this.
Erik Torenberg: That is a great closing. This talk is called "AI Eats the World," available on Benedict Evans's website, and it is very insightful, with a lot of content we didn't get to discuss. Benedict, this conversation has been fantastic; thank you for being here.
Benedict Evans: Thank you; it was a pleasure to chat.












