Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Well, we do know why gradient descent works (for smooth data), at least for finding a local minimum, because finding the minimum is basically what it does by construction. Similarly, we certainly know how back-propagation works, because it's simple calculus, backwards application of the chain rule.

Perhaps what you're trying to say is we don't know why finding local minima of these problems is good at solving the problem?



Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: