Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Proof is not binary, it depends on the claim and the constraints you put around it, and the nature of the subject of your claim.

Most of the general LLM discourse in our industry is still closer to "proof of the pudding is in the eating" than to "double-blind studies on large cohorts, p<0.01, effect size is still so small that result is useless in practice"[0].

And we're not talking about curated demos either - most of the contested value can be proven for your own specific cases with little to no expenditure of money and time, at a PoC level (it gets more expensive once you try to operationalize it and find kinks that are hard to iron out).

And that is, the article claims (and I agree), the point of last 6-12 months of tokenmaxxing policies and top-down push - it's putting pressure on people to actually go and do those PoC-s for themselves, because just giving the opportunity and permission turned out insufficient for significant part of the workforce.

--

[0] - Ironically, I remember it was the opposite around the time GPT-4 came out. Back then people talked more about specific claims and demanded measured evidence, because it was hard to get the models to reliably do something interesting. But now that the models can handle bad prompting and can understand you even when you're drunk, suddenly people are denying the general capability of LLMs and asking for randomized control trials.

(For double irony, nowadays one can just ask an LLM for randomized trials; the current SOTA models will happily design you a bespoke eval pipeline if you ask them to.)



> And that is, the article claims (and I agree), the point of last 6-12 months of tokenmaxxing policies and top-down push - it's putting pressure on people to actually go and do those PoC-s for themselves, because just giving the opportunity and permission turned out insufficient for significant part of the workforce.

FWIW, I think most tokenmaxxing is, to riff of what you said earlier, turning a technique into a ritual and science into religion.

This isn't specific to AI, we've had it before with pretty much everything in software (and since well before software), from "object-oriented solves every problem" to "clean code [where every function is] two, or three, or four lines long", to reporting your daily kloc, to bounties for every bug reported and/or fixed.

Humans do what doomers are afraid AI will do: make a sounds-good utility function (tokens, lines of code, bugs, dead cobras) and get surprised when it is easily gamed for something far less helpful than the vision of whoever set the goal.




Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: