Here’s a true story I find funny about scaling laws at Baidu.
From 2016-17 I did a projection using our scaling law equation with my coauthors about how many GPUs it would take to train an LLM with a step function in capability. Joel Hestness in particular did excellent experimental work to enable this.
I came out with a projection of about a $1 Billion GPU budget.
Baidu was in the middle of downsizing the US research center (SVAIL) in favor of AI in China and I was participating in the layoff of many of my colleagues while trying to keep the lights on long enough to finish our scaling law experiments, which I personally thought would change the world.
I actually wrote a report to Robin explaining the implications of scaling laws and asking for a $1 billion budget to train a Baidu LLM in 2016 and sat on it through 2017.
But I never sent it because I thought it would never have been supported in that environment. I sometimes wonder what Robin would have thought about it and how the world may have been different if Baidu had released ChatGPT.
We may be about to find out because the AI moat filled with simple algorithms and scale seems to be much more shallow than the processor and systems moat.
I have a huge amount of respect for Dario and Ilya for carrying on scaling laws at OpenAI or it may have never seen the light of day.
If there is one problem for the AI community to solve by 2030 I think it is the moat problem.
From 2016-17 I did a projection using our scaling law equation with my coauthors about how many GPUs it would take to train an LLM with a step function in capability. Joel Hestness in particular did excellent experimental work to enable this.
I came out with a projection of about a $1 Billion GPU budget.
Baidu was in the middle of downsizing the US research center (SVAIL) in favor of AI in China and I was participating in the layoff of many of my colleagues while trying to keep the lights on long enough to finish our scaling law experiments, which I personally thought would change the world.
I actually wrote a report to Robin explaining the implications of scaling laws and asking for a $1 billion budget to train a Baidu LLM in 2016 and sat on it through 2017.
But I never sent it because I thought it would never have been supported in that environment. I sometimes wonder what Robin would have thought about it and how the world may have been different if Baidu had released ChatGPT.
We may be about to find out because the AI moat filled with simple algorithms and scale seems to be much more shallow than the processor and systems moat.
I have a huge amount of respect for Dario and Ilya for carrying on scaling laws at OpenAI or it may have never seen the light of day.
If there is one problem for the AI community to solve by 2030 I think it is the moat problem.