i believe GP is referencing the Kaplan and Chinchilla scaling laws. we reference those in the podcast but iām not sure if some deeper insight is being hinted at here where different scaling laws apply for different domains/purposes
But these say exactly the opposite, the more tok/param the better. There is some optimum after which you need more training FLOPS to improve than if you add parameters but it is definitely not the other way around