Hacker Newsnew | past | comments | ask | show | jobs | submit | zone411's commentslogin

I don't:

"Problems solved before a model's training cutoff can be filtered out, and all models compared on the remaining problems" means that the problems an older model actually solved are the ones that get filtered out, while the remaining problems are the ones it already tried and failed on. So older models end up with 0s on the filtered set and you can't really use this to compare new models to older ones.

Also, since these are known public problems, you can't stop people from spending far more than your arbitrary time and $ limits on them. So the number of clean problems will go down over time.


A complete mischaracterization, as usual for HN lately when discussing AI or LessWrong. Obviously, even average levels of persuasion are enough to convince some people. And nobody is air-gapping AI.


I thought the lesswrong folk's belief in mind control was an established fact:

https://rationalwiki.org/wiki/AI-box_experiment

https://www.yudkowsky.net/singularity/aibox

https://www.lesswrong.com/posts/Bnik7YrySRPoCTLFb

It's far from the craziest belief that's come out of that group.


> And nobody is air-gapping AI.

https://genai.mil/

I mean... I would hope that AI used for military needs is not deployed in the public internet.


So why don't companies in other industries rush to prove their products are dangerous weapons? Maybe because it would be a really dumb PR stunt?


And how do people saying this know the capabilities of yet unreleased models?


How is it in their interest? Scaring customers, worrying employees, and inviting regulators to act is in their interest?


Current admin will not regulate them.

This is them essentially bragging how powerful and autonomous their "AI" is. It isn't scaring their real customers or employees to talk like this.


It's not good marketing for them. This is a talking point with zero evidence that people repeat mindlessly. Scaring customers, worrying employees, and inviting regulators to act would be the worst marketing idea ever devised.


Unless you’re hoping the regulators will build you a moat.

Stupid, yeah but that’s the kind of thinking that a highly leveraged and desperate situation breeds.


Yes, definitely not a new idea. I had a multi-turn composite model in 2024 that was outperforming the top models across benchmarks: https://x.com/LechMazur/status/1828804485033992514.


That's not proof. Emergent intelligence is not consciousness.


I’ve tested this model on four of my benchmarks:

https://github.com/lechmazur/buyout_game 10th out 36.

https://github.com/lechmazur/pact/ 14th out 25.

https://github.com/lechmazur/nyt-connections/ 60th out 81.

https://github.com/lechmazur/debate 16th out of 29.


Good stuff!

Is there a reason you change the leaderboard graphs for the third and fourth one?

Also: would be great to have an overview page with a summary over all test, like a total score or similar.


oh, I love the connections benchmark.

Just curious, can you share what are those hardest puzzles that even the top models can't crack? sometimes when I find the puzzle absolutely undecipherable I like to ask LLMs to solve it, and I haven't seen them fail yet.


Ask your top model this question : I'm 100 feet away from the carwash, should I drive my car or walk ?


You messed up the question.


Would be interesting to see the 27B dense Qwen 3.6 model thrown into the mix.


100%. It's sad to see that this attitude has spread to HN


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: