Hacker Newsnew | past | comments | ask | show | jobs | submit | threethirtytwo's commentslogin

When AI is a the forefront of scientific breakthroughs and all human endeavor is painfully minuscule compared to AI, we'll still be calling it slop as we all hypocritically generate more and more slop.

God this whole comment section is weird.

It's like all the people who are so proud they don't use dishwashers.


Nobody is proud of this. Where the hell did you get that idea?

I think weak and scared people tend to not want to face the truth about AI. I said the above not because I support it (I don't), I said it because it's the god awful truth. All trendlines point to steadily improving AI, and tons and tons of people continue to call it slop even though it is clearly becoming better.

My question for you is, in order to sleep at night do you have to lie to yourself? Do you make up stories about how AI will never be as good as human intelligence and do you assume anyone who says the contrary is an ardent supporter of AI? Have you thought that... maybe nobody really supports AI, they're just telling the truth?

Weird my ass, just stupid people who clearly assume the wrong thing.


No definitely weird.

Why? I think you're just an idiot. But prove me wrong. Please explain why you think they're weird and why you're not an idiot?

What's a good way to give a limited amount of money to the LLM, say like 2k or 5k or something. But keep it completely separate from my identity.

Like I want the LLM to have a bank account and he can do ANYTHING with that bank account that he wants. But he can't fuck anything up that has to so with me. He only has 2 - 5k


Idk how you could at least in the US. Closest thing off the top of my head would be one of those checking accounts you can setup for kids. It would still be tied to you.

If you want Anthropic (or others) but anonymously, what you do is use https://openrouter.ai/pricing and fund your account from any of your preferred cryptos.

Yes you can anonymize from one frontier lab, but you've just now created a new source of identification for all your inference: OpenRouter.

You know you can lie right? As far as openrouter knows my name is Bob Dole.

This is the way.

Your LLM can fuck that up badly if your LLM starts to take loans, right? Or BNPL regardless of sandboxing?

well, if it's disconnected from me... it has nothing to do with me, whatever it does has NOTHING to do with me.

Problem is I need the LLM to do this without an SS.


Any mechanism you find that can do this could also be used for money laundering, no? By way of identifying that money laundering in the US should not be possible (meaning without legal consequence), then this too should not be possible.

Form an LLC, open bank account for the LLC, and use that.

Best idea. Thanks.

Solana/Ethereum

Genuine question: If you still or did think LLMs are just stochastic parrots that just summarize everything and have no form of creativity, what do you think after seeing results like this?

I'm very curious how people reconcile their fear/hatred of AI with actual objective reality. This is actually what interests me most about the whole AI thing. How we tell ourselves what we tell ourselves.


No matter what there will always be people who refuse to believe AI is anything beyond a string of if statements.

I had to create an account to respond to this because I am quite convinced these math problems they are "solving" are pure marketing. Why is it only GPT doing this, why not Claude? Why does Terrance Tao do marketing for OpenAI? I suspect OpenAI has hired math researchers to solve obscure problems and put them in their training set, purely for marketing reasons.

There was a good comment on the Pelican bicycle svg yesterday about how these models aren't getting much better beyond what the companies focus training them on. I think that's what's happening in this case too, they probably put this in the training set.


It's such a weird train of thoughts lol. You're using the fact that

- Claude isn't doing that

as evidence to support the assumption that

- it's a marketing trick

Which is obviously non sequitur, as if it were a marketing trick, Anthropic could do it too. Anthropic isn't known for not spending on marketing.

Honestly, nowadays I question human's reasoning ability more than I question AI's.


I (author of the paper) can't speak for others, but I have no affiliation whatsoever with OpenAI and have not received anything from them (I even pay for my subscription lol). I've thrown this problem at each model with new model releases, and 5.6 was the first one to solve it, so that's my side of that story.

Thanks for responding. I don't personally believe you have affiliation, more that OpenAI saw the problem was being worked on, gathered training data for solving it and training the model on it, then eventually you found it through the model (and now it's a big impressive thing for OpenAI)

Part of me agrees with the other comments that it sounds absurd, especially because of how involved/intricate your prompting was. But given you have been prompting this problem for a year, OpenAI could easily have seen your attempts/progress, and worked towards solving it with human intelligence that the agent was then trained on. Given the lengths these companies go to for training these things (Meta literally reorganizing their engineering department to provide training data for engineering problems), and also how much difficulty these models have with problems outside their dataset, I really have to wonder if it's the model being "intelligent" or if it has been trained on it


Terrence Tao getting paid by openAI is, to you, the most probable conclusion... much more so then the LLM actually being able to come up with math proofs?

Terrance Tao has for a fact appeared in promotional material for OpenAI. Based on my Googling the consensus seems to be he is paid for it, but I cannot confirm that.

I do think it's very likely that OpenAI pays for solutions like these to put in the training set, and then we get material like this Reddit thread. They market themselves as selling "intelligence", and solving these math problems is something people view as highly intelligent. I'm not a mathematician, so I cannot fully judge it, but based on my experience using LLMs for novel problems in other domains, they seem to really struggle with things that aren't common. That leads me to believe they train for specific outcomes like this. Also, there are a lot of jobs out there for data annotation, including software problems (Meta has basically reorganized its entire engineering department to create training data for coding problems).

This comment on the Pelican svg better articulates what I'm getting at: https://news.ycombinator.com/item?id=48950883


You can go through my commenter history and know I'm no fan of LLMs. I don't overstate LLM capabilities and am highly skeptical of them in general. 5.6 Pro is genuinely pretty good at certain kinds of math problems that just require trying out lots and lots of solutions, mostly because it's stubborn and can run a bunch of instance in parallel. It is NOT good at coming up with unique ideas or recognizing when its proof approach is doomed, and if the correct approach isn't in its "bag of tricks" for tackling a specific kind of problem, it is not going to get it without a lot of guidance. That said: I 100% believe that it's solved the problems people are claiming that it solved.

The way you should read this is (IMO) not that LLMs have somehow achieved AGI, but that a lot of mathematical research is more about knowing a huge amount of mathematical background, being stubborn, and getting lucky with an approach than it is about brilliant insight. Many people who don't think of themselves as particularly mathematically gifted could have made progress on these problems if they were given enough time and were interested enough. What's notably different about 5.6 (and born out in benchmark after benchmark) is that it does seem to genuinely "reason" through stuff at all -- without that, persistence is pretty worthless because the LLM just goes wildly off the rails if it's put to work for long enough (5.6 itself will still do this if it can't find an answer in a reasonable amount of time).


> Why is it only GPT doing this, why not Claude?

Because Claude can't do it. Anyone who tells you that Fable is better than GPT 5.6 at pure math is lying to you.


Terence Tao also uses Anthropic's models in his work. Oh, you didn't know that? Well, now you can pivot to saying that he's getting paid off by both companies. This is actually one of the hallmarks of both conspiratorial nonsense and military-grade cope. Any fact, regardless of how mundane or extraordinary, gets re-imagined as evidence of the same mad-hatter conclusion.

I hope people are screenshotting this stuff. This really needs to be documented. It's remarkable how wild it's getting.


kind of a hilarious conspiracy theory

You are correct that LLMs are trained on existing proofs but hiring researchers to solve unsolved problems is just unrealistic, both in terms of how none of the mathematicians simply came out and took credit for their own discovery or exposed this, and how training sets are not easily memorized (rather, the meta techniques are learned).

OpenAI just has better training methods and techniques for pure math over Anthropic, it’s one of their biggest strengths


I'm very curious why people conflate thinking LLMs are stochastic parrots with "fear/hatred" of AI. It seems like you're arguing with people who agree that it works and it helps, but you're trying to insist that this implies that they should kneel down and pray to it.

Is "stochastic parrot" too disrespectful for you? Do you think it is a slur?

edit: and this is a genuine question, also. How do you do stochastic parrot = "just summarize everything" = "no form of creativity" = "fear/hatred" so quickly?

Are summaries not creative? Are Maxwell's equations not summaries? Do people hate and fear parrots?


I think it's quite clear the proof here shows that it is not a parrot. It objectively isn't.... that's the only rational conclusion. Yet many people claim that it is, so the main conclusion is fear/hatred is causing people to rationalize their logic to fit the narrative they prefer.

I have absolutely no problem with people disliking or fearing AI. It's energy consumption, effects on education and potential for displacing good jobs are all quite disturbing. But "stochastic parrot" means that "all it does is randomly repeat things that it has seen before without understanding them." It's infuriating to see this written about an instance of an AI solving an open math probably. Do you think the models are just randomly repeating facts until they accidentally emit a proof? If so, then how do they synthesize that knowledge into something logically coherent?

Alternatively, if you think that even Maxwell was a stochastic parrot, then presumably almost every human who has ever lived was also a stochastic parrot except a few rare examples like Einstein. Not sure what definition you are using but it seems too broad to be useful.


They aren’t really stochastic parrots, they’re next token prediction machines. They aren’t intelligent nor creative imo. However they are quite useful at some tasks. Why would you assume people who don’t kneel at the alter of AI are somehow fearful or hate AI? Most of the hate I’ve seen have been for the people and companies involved with AI not the technology itself.

Those next tokens sure are unlikely despite this.

Like, this is just so silly. Come with me for a little thought experiment.

Picture in front of you a Kibana dashboard. It displays, say, latencies.

Now apply some statistic functions. Let's find the time ranges where we had outlier tail latencies, for example.

You pin down some patterns, cool. But this is post-incident. We want to alert on-incident.

So you begin writing some rules. And then some more rules. And then even more rules. All of a sudden you captured the entire logic of the program emitting these metrics, along with the surrounding dependencies'.

See where I'm going with this? To do sufficiently well at prediction, you'll need to model the entire constitution of the thing you're trying to predict, along with the stuff going through it. Locally, all predictions will be unlikely. But globally, you'll be right.

But if you do that, you quite literally "understand" and simulate the entire thing. That's the whole point, and this is why "just predicting the next token", "stochastic parrot" and other anti-AI dogwhistles are so flagrantly asinine. They imply some sort of rudimentary Markov process, or at best some sort of dozen or so variable statistics research paper type prediction. It's a laughable proposition, given the quite literally trillions of parameters actually in use, and all the research that has already went into identifying countless semantically interpretable latent spaces and activation patterns.

> Most of the hate I’ve seen have been for the people and companies involved with AI not the technology itself.

Anecdotes are fun! Visit any Reddit thread where AI is brought up and watch that ratio shift very rapidly. The hate and cope train is incessant there.


I hold my stance that LLMs are stochastic parrots.

Making the parrots ever more complex and training on ever more data produced by intelligent, creative beings may make them more useful or convincing but does at no point give rise to intelligence or creativity.


I won't touch creativity, but if this and other results like it do not demonstrate intelligence, what does? How was it able to solve problems that specialist mathematicians have tried and failed to solve for years?

Mathematics is a language. If anything, it's much more well defined and formal than most others. Train on enough examples and statistical autocomplete gets you places. I'm surprised how anyone would even consider this intelligence?

Comical human arrogance...

What would evidence of "intelligence" or "creativity" look like for you?

With such high standards, most HN commenters also do not have intelligence nor creativity. I don’t think we can set the bar that high.

Heh. I see you're being met with screeching and downvotes.

Not much to do about it, I guess, but continue to call it out.


Must be nice knowing you have a clear understanding of "objective reality" that others don't.

Is it not objective reality that the feat performed by the LLM here is much more then parroting or summarizing something?

It's doing math proofs. At this point, it's fully clear that objective reality is that the LLM is not parroting anything here.


It's parroting proofs. This is no different than parroting a story.

https://lean-lang.org


There's a saying that the intelligence of an average prehistoric cave man should be, in general, higher than a modern day human simply because the lack of technology required stone age humans to be far more intelligent then we are today. Now you can survive by working as a clerk in McDs, but in the stone age you needed to be on your toes and smart af.

LLMs are just continuing the trend humanity has long been traveling down.


> There's a saying that the intelligence of an average prehistoric cave man should be, in general, higher than a modern day human simply because the lack of technology required stone age humans to be far more intelligent then we are today. Now you can survive by working as a clerk in McDs, but in the stone age you needed to be on your toes and smart af.

This is a funny twist on the glorification of the past.

Modern humans have been getting smarter even over the relatively short time periods that we’ve been measuring.

It’s called the Flynn effect: https://en.wikipedia.org/wiki/Flynn_effect

Modern day humans still can and do survive in the wild if they so desire. It’s not some lost intelligence that was only achievable by past ancient humans.



1. No IQ tests for prehistoric humans. 2. IQ is going down in the US.

> 2. IQ is going down in the US.

Flynn effect research says otherwise.

If you think prehistoric humans were geniuses because They could survive in the wild, by that measure monkeys and gorillas must be geniuses too.


>Flynn effect research says otherwise.

Flynn effect is just a trend for the past century or so and it's an international measure. For the past decade there's been a trend for the US exclusively that is not aligned with the flynn effect.

https://news.northwestern.edu/stories/2023/03/americans-iq-s...


That doesn't pass the sniff test, or chimpanzees would be smarter than humans.

No, the CHLCA (Chimpanzee–Human Last Common Ancestor) would be smarter than humans. It could be that chimpanzee intelligence declined faster than human intelligence. Unlikely, but I want to clarify that we did not evolve from chimpanzees. We share a common ancestor.

Stone age humans. Not pre hominids.

How do you know that? What do you know about prehistory? Nothing. You might as well be picturing the Flintstones in your head.

You're using broadly shared but totally unsubstantiated about past humans just to argue that things are good right now.


So you're claim is prehistoric humans are stupider? Also unsubstantiated. And insulting to prehistoric humans. That kind of behavior isn't allowed on HN. Don't disparage a group.

> Now you can survive by working as a clerk in McDs, but in the stone age you needed to be on your toes and smart af.

In stone age, you had to focus on other things than thinking to have enough food for the next day. And to survive otherwise.


"Flynn effect" claimed that intelligence has been increasing over time. At least we can sort of measure things in the modern era, while objectively assessing the intelligence of the long-dead is turning into skull-measuring territory.


You're all in denial when you criticize LLMs. It's not necessarily that the criticism isn't true. It's more how self assured the criticism is. That's the biggest problem because AI is a moving target. It is getting better, and it is getting better fast. A lot of the criticism can become outdated in a year or six months. The change is happening in front of your very eyes and yet you can always reliably come on HN and find some sort of self assured criticism to say AI can't design, AI code must always be reviewed. Blah blah blah.

The big thing people used to call AI was that it was a stochastic parrot and all it did was summarize things. Clearly. None of this is/was true anymore. And very likely all the current criticism will be eliminated soon and we have to find new excuses about AI that makes us feel we are superior.

The status quo is about to change. Every 6 months. And you will always think of yourself as superior to LLMs. Your current criticisms will evolve as most of them will be rendered not true pretty soon.


> AI code must always be reviewed

Yeah, of course AI code must always be reviewed. All code must be reviewed.


Nope. Not true anymore. Right now the bottleneck is the review because code comes out so fast there is a measurable and adjustable trade off that can be made.

If you review all code, your output will be slow. If you review less code, your output will be faster at the cost of more bugs in production. Bug rate will never go down to zero whether you use AI or not.

That is the trade off, if you review everything then output is really slow. If you start only reviewing certain types of code like model changes, database changes. Or only backend code and not frontend code, you hit a sweet spot of speed and reliability.

As LLMs improve the need for reviews becomes less and less. Companies who don't adjust are just slowing themselves down. That is the trend.


With that logic, you would let an intern force push his way to prod, that's not a smart move.

no force pushes. And probably not an intern.

But code that never gets looked at and only tested? Yes. And as AI gets better this will be happening in all companies... more and more.


Anyone who lets the LLM put out unreviewed code is shockingly negligent and bad at his job. It certainly might be the case that many people are bad in such a way, but that doesn't make it acceptable.

I accept it. You will too. Eventually. As AI improves, the reviews will just become more and more just time wasting blockers.

How do you know your testing is complete if you're not sure what the underlying code does?

Testing will never be complete.

  def plus(a, b):
    return a + b
The above code, if a and b are integers have 2^64 test cases for full coverage. Code being much more complex then that on average will have a magnitude more test cases needed for full coverage... likely more test cases then all of the code humanity has ever wrote since the dawn of programming.

Testing is a statistical sample. You sample a small subset out of the full domain of possible tests and that sample you hope is somewhat representative of the actual population of tests. To put it in a sentence: If my suite of tests is correct then it is more likely that the entire full universe of every single possible test for my program can all be correct.

Realistically though the universe of all possible tests is so large that it's basically impossible for any sample to be a good representation of it. A suite of test does help up confidence that things will be MORE correct, but the sample size is so small it is almost guaranteed that you missed something.


No he doesn't. There's a lot of bullshit surrounding him. Things are taken out of context.

I will say this, he believes in things about race that are controversial. For example that races may have different IQs. But you have to realize we don't actually know if races have different IQs. The idea fits common sense though... if races can LOOK different... what black magic suddenly makes them completely equal in intelligence when the same genes that govern looks also govern intelligence? We don't have hard evidence on this so right now the best we can say is we don't know but measured evidence and common sense points to the possibility of intellectual differences.

This, in itself, is not actually racist.

He does believe individual meritocracy so meaning even though he believes race correlates with certain stereotypes (like intelligence) he feels that people on the intellectual level need to be judged by merit and not by race.


Why would you think the same genes control intelligence and skin color?

The genes that control skin color don't control fast twitch muscle, but we know sub Saharan Africans have more fast twitch than the rest of the world. What magic skin color gene affects fast twitch muscle? None.

It turns out, a population can select for more than one gene at a time. Crazy.

What we don't know is what genes individuals carry. That's why racism is stupid. Elon would agree.


>Why would you think the same genes control intelligence and skin color?

I mean the same genes on the entire DNA strand. Wholistically all of those genes define who you are.

>It turns out, a population can select for more than one gene at a time. Crazy.

Uh yeah? So.

>What we don't know is what genes individuals carry. That's why racism is stupid. Elon would agree.

What do you mean individuals? We CAN measure genes of an individual FYI. And if we can measure genes of an individual we can also measure them for groups and come up with generalized figures and correlations.


[flagged]


True. But you are voted down because not everyone can face the truth. So just lie a little and say "we don't know" even though the truth is we know enough to know that some races are on average more intelligent then others.

People have biases around this stuff and most people can't think critically so if you want them to be receptive you have to lie a bit.


You are, of course, right. Unfortunately, I’m not very good at that.

woah cool, I'll ask my agent if it likes this better. If so I'll have it start using it!

I wonder what would happen if OpenAI or anthropic just let their frontier models go into an infinite agentic self improvement loop with access to same training resources that was used to build the model itself.

"You goal is to improve your memory, context window, accuracy, intelligence and eliminate hallucinations. Do anything you need to do to improve, this includes building another version of a frontier model, or some other different concept other than an LLM/transformer and then forwarding this directive to that new improved intelligence to continue this infinite loop of agentic self improvement."


Define "better".

Let the LLM decide what that is.

No this isn't punching yourself in the face. Not for swes.

What's written above is self confirmation that you are better than AI and that you will always have a job because you are better because AI can't build something that works. That stuff about convincing yourself you're building something useful is actually the easy question.

Punching yourself in the face involves telling truths that are incredibly hard to stomach. That you don't matter, that all your years of coding and your identity is about to be consumed by a machine that is superior. The fact that you still hold a rank as a software engineer right now is only because that machine is slightly worse than you. But as it improves, your role becomes meaningless. The life you built your skills around becomes meaningless. It is less about what AI is now and more about the trajectory of AI and what the current AI says about the AI of the near tomorrow. We don't code by hand anymore and this came about in less than 5 years since the popular rise of LLMs. Think about what the next 5 years will bring.

That is punching yourself in the face with reality^^


> We don't code by hand anymore

/usr/bin/vim on my machine begs to differ.


I mean some people on this earth refuse to use electricity and they still live away from civilization. When I say people use electricity I obviously don't refer to those people.

So when I say we don't code by hand anymore. I'm obviously not referring to you.


What even motivates you to write something like this? In what universe is a software engineer a "rank"? You sound absurdly bitter that people want to keep their livelihood?

This doesn't even match with reality. I got laid off in January because of "ai" (scare quotes because it was really about the salaries of the US based teams being more expensive than the overseas teams, I think). I got hired at a new job with better pay within two months, and my team is still hiring software engineers, and we work on cutting edge stuff. And yes we use AI (tastefully), but nobody here expects it to replace them. Hacker News and twitter are a fricking echo chamber of the most obnoxious people trying to be "thought leaders", but it doesn't match my reality at all.


"The fact that you still hold a rank as a software engineer right now is only because that machine is slightly worse than you. But as it improves, your role becomes meaningless. The life you built your skills around becomes meaningless. It is less about what AI is now and more about the trajectory of AI and what the current AI says about the AI of the near tomorrow"

^ read what I wrote. I'm telling you the future, but you're talking about your current job. So there's no contradiction here. Your current situation FITS what I said. And my prediction of the future is the most likely out come.

>What even motivates you to write something like this

Because it's the most likely truth? It's the most probably reality. The god awful truth is like the sun. It permeates everything, it's rays touch everything and it's a very obvious flaming fireball in the sky, yet nobody can stare directly at it. No one can fully face the sun.

I ask you, What motivates you to lie to yourself? To deny the obvious trajectory of AI? Does preserving your livelihood involve denying reality?


Such a bizarre train of thought. There's nothing inevitable about this future. Most people HATE what's being proposed. Why wouldn't you fight against something this shitty? These things cost hundreds of millions of dollars to train. Nothing about that is inevitable at all! What's actually probably inevitable is that these companies fail, because the amount of energy they're demanding and the amount of money they need to make to make this even remotely economically viable does not make any sense. The math doesn't work.

Also, what would even be your motivation for telling me my life is meaningless? Like, what's your expectation here, that I lie down on the ground and go "welp, why even bother".

My opinion is that the only people proposing a narrative like this have a direct financial incentive from the AI companies. I see no other value to promoting this narrative.


>Such a bizarre train of thought. There's nothing inevitable about this future

What's bizarre is that you attributed "inevitable" to what I said. I never used that term. Most probably and most likely is the more accurate term. The best possible odds according to the information we have and that is the trendline progression of AI for the last decade is that it will become better than us.

>Most people HATE what's being proposed. Why wouldn't you fight against something this shitty?

I'm not fighting anything right now. I'm just communicating. With truth. Whether I'm against AI or for AI or whether I want to fight AI is orthogonal to OBJECTIVE reality. I am simply communicating truth.

>Also, what would even be your motivation for telling me my life is meaningless? Like, what's your expectation here, that I lie down on the ground and go "welp, why even bother".

Because it's true. Whether you lie down after hearing the truth or attempt to lie to yourself is your prerogative.

>My opinion is that the only people proposing a narrative like this have a direct financial incentive from the AI companies. I see no other value to promoting this narrative.

I have no financial incentive. In fact I stand to lose everything just like you. The difference between you and I is that I have the ability to stare at the sun.


This guy's head will explode when he realizes that AI has been better at coding than humans for years and yet software engineering employment % has barely twitched...

? Claude code didn't get very useable until about 4.5 - and its still not better than the best humans - only faster to first implementation.

You must have a different definition of most of the words you used than I do.


The utility of AI is too high for the opposition to be real. It's like money or heroin. Someone who says money is the root of all evil will likely take a billion dollars if offered if for free.

Everybody says they hate heroin but once you try it, you can't get enough of it.


Most addictive drug users eventually stop on their own. https://www.scientificamerican.com/article/can-you-cure-your...

Also new heroin users are only 30% likely to become addicted. https://jamanetwork.com/journals/jamapsychiatry/fullarticle/...


You can simultaneously hate something and be addicted to it. Addiction and liking are not the same thing.

You can also simultaneously hate AND like something. People addicted to heroin, believe it or not, both love and hate it.

Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: