Ben, I follow you and think you’re brilliant, but boy you can’t take feedback like ever.
Duplicate slide is there clearly to add the question, but it has a glaring typo—- “predications” instead of “predictions”.
A random internet stranger read to page 51, which is already a rare occurrence, and helped you find a typo that you can now edit.
But sure, the answer to that is “your comment is wrong”.
Ah, that’s different. Your question confused me because you said they were duplicate slides and they are not, and your word ‘predication’ looked like a typo by you.
The slides say "predication" and not "prediction." Just a typo :) I think the AI slop comment was excessive, and I liked your deck, but there is indeed a typo you could fix!
I didn’t make any comparison at all with the fibre bubble, for precisely that reason. The comparison is with mobile data, which was and is always behind capacity.
I think one of the things that the usage data shows us is that chatbots absolutely do not have infinite use cases - most users only use them a day or two a week or less.
That's fair, I may be conflating your takes on mobile data with others who've made the comparison to the telecom bubble, and if so, mea culpa!
But I also do disagree with the take that usage patterns indicate a fundamental shortage of use-cases. Yes, everyone reports WAU instead of DAU because WAU numbers look much more impressive, but I think the extreme shortage of compute plays a major role in this. I suspect all the AI labs are deliberately holding back from pushing AI adoption too much because of this. (Google execs have even made comments internally to this effect.) Note that even at such low frequency of usage all the model providers are desperately strapped for compute, which means there is insanely high demand from some quarters.
One way how capacity limitations could impact adoption is that the free-tier models are not as good as the frontier ones, so the free users come away less impressed with AI capabilities, leading to lower regular usage. This problem is larger than it appears, because it can take a long time to figure out how to get AI to work for your use-case, and people simply have not experimented nearly enough, partially due to first impressions. On the other hand, most companies seem to be OK with huge tokenmaxxing bills!
It seems to me the AI players are all playing a delicate balancing game across three fundamental dimensions: adoption, monetization, capacity. That is, they are simultaneously 1) pushing free / cheap AI usage as much as possible to hook users, capture market share and suss out new use-cases, while 2) carefully allocating token quotas for the most lucrative use-cases to satisfy investors, and 3) balancing available compute between those two competing priorities. I suspect as the compute bottleneck is alleviated and frontier models become more accessible cheaply, we'll see way higher DAU numbers.
I've made the semi comparison myself, but the amount of capital required to build a SOTA model today is clearly nowhere near enough to lead to a monopoly.
I'm aware that telecoms networks are standardised (I was once a telecoms analyst), but that isn't a precondition for a commodity.
Just like how starting a chip fab was relatively easy back in the 80s and 90s. There were dozens of chip fab companies in the 80s.
It turns out that fabs follow Rock's Law which is that the capital cost to build a new fab doubles every 4 years. This means it will quickly get rid of the less competitive players. This is not dissimilar to the LLM scaling laws where you need a magnitude more compute to get unlock a new tier of intelligence.
Today, Anthropic and OpenAI are clearly in the lead for models and then there is everyone else. Google is a close 3rd. No one else is challenging them anymore in SOTA models. Some models might beat them in one or two benchmarks but none can compete overall. I expect this gap to grow bigger as models cost more and more to train.
Evidence point to the same type of scaling law. Compute for a training run grows 4-5x every year.[0] I'm sure this will slow down but the premise remains that weaker competitors will not be able to maintain this pace. We already see labs like Cohere, Mistral, Inflection AI, Adept, Character.ai, and others bow out of the frontier race. I'm also skeptical that Meta, xAI can catch up. Even Google has trouble keeping up.
Even if this isn't true, comparing telecom bits to tokens is wrong. Bits are the same no matter what telecom transfers them. Tokens are not all the same. The quality varies.
We're already seeing a massive divide between frontier models and lesser models in growth rates. Anthropic is adding $10b - $15b every month in ARR. This figure likely dwarfs open source labs. This is all because its models are maybe 10-15% better.
The cost to inference a 1T param frontier model is the same as a 1T param open source model. Therefore, if the frontier model is even 10-15% better, it will gobble up the market over time.
Lastly, even though Claude Code and Codex are the biggest revenue drivers for Anthropic and OpenAI today, I don't believe this will be true 2-5 year from now. I believe selling their tokens via API will be their biggest. The sum of applications in the world will dwarf coding in market size. For example, biotech, finance, physics, engineering, robotics, sensor data, etc. This is why I think OpenAI and Anthropic are becoming more like iOS and Android than AT&T and Verizon. Applications will build on top of OpenAI and Anthropic just like iOS and Android.
How about the externalized intelligence around the model weights (skills, tools, harness, memory etc)? If the model weights are sufficiently intelligent, the focus might move to the external layers.
I agree with much of what you’ve written but think you are missing the correct alignment of the mobile data timeline — mobile data had standards because it was forced to. It was forced to early because it was not a fundamental innovation, telecom itself was the fundamental innovation, mobile was a constraint relaxation. Intelligence might be forced to have standards as well, we will see what form the regulations take when prices reflect costs and healthy margins and become existential threats for many businesses.
I agree with that as a premise, but again it seems to
me you are selectively jumping way into the end game. There were early networks that did not standardize, and these nonstandard networks had advantages, and some of those advantages were sacrificed in market-driven standardization.
Intelligence must have interfaces, and those can be standardized. Businesses will try to remain provider agnostic, which will also drive standardization via standard sales and marketing methods.
Separately, we are doing our best to standardize performances on benchmarks.
I don’t disagree that right now transport of standardized mobile data vs emulation of human intelligence is qualitatively different, but perhaps primarily because it is early in development, and our vantage point this time is relatively from within the network, instead of outside it.
You lay out some good arguments but I agree with both: the models relative to few years back really did become the commodity because today you could take the non-frontier model, maybe self-host it or pay the much less price per M tokens to get the performance of a ~2-year old frontier model. At the same time I do think that we are getting into the monopoly/duopoly/tripoly with the frontier models for all the reasons you already mentioned, and this scares me a little bit.
Lower intelligence LLMs can be a commodity, yes. But these won't make much money, if at all. At the end of the day, it costs the same to inference a 1T frontier model and a 1T free model.
OpenAI and Anthropic don't compete in the LLM commodity market. Hence, I had a problem with slide 22.
I doubt it. Compute costs tend to crater over time, and LLMs will almost certainly plateau. So the opposite is almost certainly the case: over time, it becomes cheaper to train.
There’s a bunch of fuzzy metrics here, which is one reason I turned it back into a monthly number.
The other issue (as you’ll see on the chart) is that Anthropic and openAI are recognising revenue in completely different ways.
You’ve missed the point completely - if the important experiences are things built on top of foundation models, where the model itself is just an API call, then you don’t need to have a foundation model for build them and the model is just commodity infra
Yes, but OpenAI has 900M+ user reach, plus staggering amounts of cash, plus early access + deep integration with the latest and greatest models. I hardly think that is tantamount to "just an API call".
Deep Research doesn’t give the numbers that are in statcounter and statista. It’s choosing the wrong sources, but it’s also failing to represent them accurately.
Wow, that's really surprising. My experience with much simpler RAG workflows is that once you stick a number in the context the LLMs can reliably parrot that number back out again later on.
Presumably Deep Research has a bunch of weird multi-LLM-agent things going on, maybe there's something about their architecture that makes it more likely for mistakes like that to creep in?
Have a look at the previous essay. I couldn't get ChatGPT 4o to give me a number in a PDF correctly even when I gave it the PDF, the page number, and the row and column.
ChatGPT treats a PDF upload as a data extraction problem, where it first pulls out all of the embedded textual content on the PDF and feeds that into the model.
This fails for PDFs that contain images of scanned documents, since ChatGPT isn't tapping its vision abilities to extract that information.
Claude (and Gemini) both apply their vision capabilities to PDF content, so they can "see" the data.
So my hunch is that ChatGPT couldn't extract useful information from the PDF you provided and instead fell back on whatever was in its training data, effectively hallucinating a response and pretending it came from the document.
That's a huge failure on OpenAI's behalf, but it's not illustrative of models being unable to interpret documents: it's illustrative of OpenAI's ChatGPT PDF feature being unable to extract non-textual image content (and then hallucinating on top of that inability).
Interesting, thanks.
I think the higher level problem is that 1: I have no way to know this failure mode when using the product and 2: I don't really know if I can rely on Claude to get this right every single time either, or what else it would fail at instead.
Yeah, completely understand that. I talked about this problem on stage as an illustration of how infuriatingly difficult these tools are to use because of the vast number of weird undocumented edge cases like this.
This is an unfortunate example though because it undermines one of the few ways in which I've grown to genuinely trust these models: I'm confident that if the model is top tier it will reliably answer questions about information I've directly fed into the context.
[... unless it's GPT-4o and the content was scanned images bundled in a PDF!]
It's also why I really care that I can control the context and see what's in it - systems that hide the context from me (most RAG systems, search assistants etc) leave me unable to confidently tell what's been fed in, which makes them even harder for me to trust.
1: It's the TITLE of a 100 slide presentation. It's not the only thing it said, and it's a way to think about what was happening.
2: Mobile replaced the PC as the main way people use the internet and do their day-to-day computing. The consumer Internet runs on smartphone apps, not PCs. In 2013 a lot of people didn't understand that that was happening, so it was worth saying.