Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Are you saying the less you train the model the better it is? I'm confused


i believe GP is referencing the Kaplan and Chinchilla scaling laws. we reference those in the podcast but i’m not sure if some deeper insight is being hinted at here where different scaling laws apply for different domains/purposes


But these say exactly the opposite, the more tok/param the better. There is some optimum after which you need more training FLOPS to improve than if you add parameters but it is definitely not the other way around




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: