You are saying: it turns out that one-shot playable Minecraft clones are actually pretty simple. Maybe it seemed a hard problem to programmers, but why not just say that the verifiable-rewards training has shown that their skill is unusually simple?
Can't wait to see what unusually simple Erdos problem LLMs will expose next, hiding in plain sight for decades and seemingly intractable for professional mathematicians who weren't aware just how simple the problem was.
I am saying 3 years ago there wasn't a snowballs chance in hell I could one shot a playable Minecraft clone, and there has never been a snowballs chance in hell a human developer could do so in 45 minutes.
The difficulty of the task and human performance on that task hasn't changed. LLMs performance on that task has changed dramatically.
But nobody has nailed "this is a game anyone wants to play" or "professional game development life cycles are shorter" so it's not particularly impressive.
I'm still waiting for any evidence of LLM economic impact in the real world.
Yeah, well I doubt any economic impact will from LLMs will ever come. I doubt anyone is working on nailing those things as we speak. I doubt the results my team and I are getting with agentic engineering are real. I doubt anyone can take increasingly better tools and translate that to economic value.
Those demos are put out by non-engineer AI-bros who start from a blank slate and just test raw one-shot performance of the model.
In comparison I'm aware of how much additional performance you can extract out of models with a team of professional engineers that are balls-deep in harness engineering.