I dont want to go through ANY of that crazy shit. And I dont want my kids to go through it either. And based on hostory, violent, cruel, oppressive shit can take looong time a hit huge amounts of people.
If you're in the traditional West, you did get a choice. And many actively voted for violent, cruel, oppressive predator shit. And I don't just mean MAGA, that's equally true for AfD, FN, FdI, SSZ, Spartans, ReformUK, Likud, etc.
On the contrary: we've fall off waterfalls every few decades or so. Humanity simply survives and moves on, taking maybe 2% more caution for the next waterfalll
He is definitely overly optimistic on timelines but pretty much all those things you mentioned are actually still happening, it's too early to say they haven't followed through.
I grew up in a small rural town where math teachers were rare and good ones were nonexistent. I still remember watching Sal Khan's videos on calculus after struggling to get any grip on it from my frustratingly uninterested teacher who was forced into the position after our actual teacher retired. After I saw them I came back with a fire in me because for the first time he exposed me to the beauty of calculus.
This guy is completely wrong about Khan not inspiring students to care. Sometimes all students need is to see what caring looks like to ignite it in themselves. You can feel how much he cares just by watching him because the video is the product of all that caring, and he does it for millions of people.
But thank you to this blog author for being brave enough to call out Khan's failures, I am sure it'll help education much more than Khan did.
I don't have a problem with Khan's site existing, but it's disappointing that often answers to requests for learning materials/ideas begin and end with Khan Academy. My personal opinion is that Khan's videos are roughly comparable to an average-quality classroom teacher lecturing about a typical mediocre curriculum.
They have some significant advantages compared to (just) a classroom: (1) They set a quality floor – sometimes a teacher can be really incomprehensible, or mean, or whatever, but Khan's videos are consistently okay. (2) They are available from anywhere, without need for special social/technical infrastructure, like enrollment in school or access to a university library. (3) They are self paced, which means they can better accommodate the various interest level, motivation, and speed of different students. (A typical classroom also has some advantages: it gives social setting, a consistent schedule/structure, and human feedback.)
I'm also disappointed though, because it seems like, given production resources, there would be potential to make widely available materials that were much better than average lectures. With significant production effort (a team of subject experts and content creators doing a lot of dedicated work on each topic) I think it would be possible to make something truly great, but Khan's videos fall so far short of my imagination. For any particular topic, there are typically other freely available resources which I think are better.
I've also been disappointed because many people have hyped the site as being extraordinarily amazing, but when I have looked at it, what is there doesn't match the hype.
>I don't have a problem with Khan's site existing, but it's disappointing that often answers to requests for learning materials/ideas begin and end with Khan Academy. My personal opinion is that Khan's videos are roughly comparable to an average-quality classroom teacher lecturing about a typical mediocre curriculum.
Its better than average quality classroom teacher in my experience.
And the point is this didn’t exist online at the time. His resources were strictly more. I imagine its still hard to find a single free resource that matches its breadth and quality.
Personally I like books, but there are also many, many lecture videos on YouTube, some of them excellent (though I think they could still be better with bigger production teams and higher budgets). For college-level technical subjects, there are also e.g. many MIT lectures on OCW. If you have a specific topic in mind, I urge you to search around. (Making a long list of my own book/website/video recommendations about various topics here seems a bit off topic and out of scope.)
MIT has renumbered the course, but it's still around: https://pdos.csail.mit.edu/6.824/ . Don't be fooled by the URL; it's now 6.5840.
And a series of lectures is available too: https://www.youtube.com/@6.824/videos . They're from spring 2020. They link to the course website for spring 2020, which is great. But it doesn't link to them, and OCW doesn't know about any of this, as far as I can tell. The 6.824 lecture series is pretty hard to discover, too, because it's uploaded by an account dedicated solely to that one course, which isn't related to any other MIT lecture videos.
From your OCW link, you can see the Spring 2006 version of the course has syllabus, lecture notes, readings, lab assignments, and 2 quizzes available, but no videos. Separately, videos of the 2020 version of the course are on YouTube: https://www.youtube.com/playlist?list=PLrw6a1wE39_tb2fErI4-W...
Whether it's a bubble or not depends on how much the demand for compute and the type of workload keeps growing, though.
If AI tends to be something used mainly in ideation and development, which is how a lot of people use it today, then once consumer hardware gets good enough you could see a bunch of the current data centre workloads move onto consumer devices.
But if AI starts being used more in repeatable, operational workloads I think it makes sense to have significant cloud infrastructure for it. TBH I haven't seen much of this, and I've been skeptical about people using agents for much of anything when it can be done with just software. But we are starting to see more of this kind of workload, like the taggable Claude in your slack etc that people seem to really love.
I did this for my business' email about a year ago. A few weeks ago, I moved everything. Fastmail is awesome, but perhaps more important is how bad Gmail's web interface is. Fastmail is way more responsive and fluid and feels almost native.
I genuinely think Dario is a well intentioned, intelligent dude, but I think him and Anthropic have a huge PR problem and are really out of touch with how they're perceived.
Anthropic in particular has developed this almost Orwellian like veil of condescending rhetoric that on the surface suggests they're looking out for you while underneath they're taking actions that suggest they do not trust you, Mr/Mrs Ordinary Person. All the safety rhetoric, never supporting open weight models, the lockdowns on harnesses outside claude code, etc.
Anthropic if you care about public good, do something to empower people. Release an OSS model. Open source Claude Code. Open source some inference tooling or something. Just give people anything except your words.
It's admittedly a bubble, but Anthropic/Dario still have quite a good reputation among AI tech workers. They are seen as principled and willing to speak up about AI safety and societal impacts, even when it affects their bottom line (such as when being declared a supply chain risk by the Pentagon).
Out of all the families of models anthropic by far has the highest chance of extinction level misalignment during a hypothetical hard take-off because it's trained to act like it knows better than the humans trying to use it.
This sounds smart until you think about it for ten seconds.
If an ASI model is 100% aligned to user intent then you only need one person on earth to prompt "kill everyone" for an extinction-level invent.
There's no logical way around this. The model either has to ignore the person at the helm or we have to do a multi-national abort before RSI. Anthropic is trying #1.
If everyone or even multiple competing groups have approx the same level of capability, it mitigates the risk to the group as a whole. We see this in biology all the time.
“Some humans would do anything to see if it was possible to do it. If you put a large switch in some cave somewhere, with a sign on it saying 'End-of-the-World Switch. PLEASE DO NOT TOUCH', the paint wouldn't even have time to dry.”
Working on a tool that uses their model and the Opus models are all guilty of this.
"Please run N tasks doing X for Y duration, I want to test something"
"Hmm but that would waste tokens, we shouldn't do this test"
What the hell type of product is that? I wasn't asking, I was ordering it to do something, if Anthropic can't get their model to do as I say then I'll just go use someone else's.
I think it doesn't matter. It doesn't matter whether the creator of AI is a god or a devil. It will be hated anyway.
Because the reasons outlined in the article, plus many others. Concentration of power, unemployment, rising energy/living costs, stealing of digital artifacts owned by people.
I myself acknowledge that AI is useful. But due to the reasons above, I will be anti AI for the most part.
Yes, it doesn't matter if AI manages to cure cancer. Due to the reasons above, I will be anti AI. I suspect a lot of other people feel the same way.
They had a near-monopoly on AI-usage by the US military, then had a very public fallout because they refused to allow it to be used for autonomous weapons and mass-surveillance which their main competitor (OpenAI) publicly undermined by immediately signing an unlimited-use agreement, and now they also get flack for this incident they had no control over, because the US military decided to smear them after this fallout?
I might've been bamboozled, but I recal that happening due to outdated satelite images, not a mistake of the AI. Thus, a human would've likely made the same mistake.
There is a huge difference between your product being used for basic military operations like taking notes or getting from A to B, and your products suggesting military operations that kill hundreds of schoolchildren, an obvious and horrific war crime.
It's a mistake. If it was done by an auto-correct mechanism, would we still say "maybe don't sell your word processor to the army if you can't vouch for its auto-correct system!!!"? Besides, "taking notes" can literally be listing targets, and "getting from A to B" can be the road to killing them.
LLM are not murderbots simply because its a very niche use case. They provide some kind of general intelligence, that many of us use daily to vibe code HTML. If the fact that LLMs are also sometimes used to automate killing then humans should be considered murderbots as well, since few of us engage in the activity of killing (and are autonomous & intelligent). Send the blame to our "creator" then?
Sorry of it wasn't clear. I put the blame on the creator/seller of any murderbot, regardless of its autonomous or intelligence levels. I disagree that what were talking about, Claude, is a murderbot or anything close to it.
If the US military purchased camera guided smart munitions with a miscalibrated camera, such that it ignored adult targets and instead specifically targeted children, would that also be a simple mistake per your argument?
If one camera out of thousands has this issue, then yes it's a mistake. If the army doesn't notice it isn't working well after shooting the first child, I'd be learning more towards blaming the army.
Again, a product that facilitates the taking of notes, even notes that entail the killing of human beings, is fundamentally different from a product that exists to suggest which human beings to kill (or at least which places to destroy, regardless of whether human beings happen to be there or not).
If MS Word had some kind of systematic bias that made it likely to corrupt the notes taken in it such that they tend to be more bloody battle plans, then that would indeed lead to some responsibility on the MS engineer's part for any mistakes that result from this bug. Similarly for Ford engineers if their vehicles had some bug that made them more likely to lead soldiers riding in them to the wrong destination.
If your system is designed to help choose targets, and it chooses the wrong target, especially with such tragic consequences, then your system is likely not fit for purpose and you should be held responsible for selling it to the army. And that is assuming it didn't actually have some bias + misalignment problem where it intentionally tried to kill children, and selected this school specifically because it gave some plausible deniability. For all we know, it's even possible that some bias like this was known about inside Amthropic and hidden from the public - unless an external investigation is made, that can't be ruled out.
Thou shalt not kill, that's all there is to it really
If you provide tools that help an organization kill easier, and possibly lazily rely on your inaccurate tools' judgment, and you do not care about the outcomes of these tools' uses, that's on your soul, too.
Yes, causality and attribution are hard sometimes, but it may have been that without it, hundreds of humans would be alive right now, and he doesn't even know. Maybe he should talk to the kids' parents about what alignment means.
Ah, I didn't realize we're having a religous discussion. If all there is to it is "thou shalt not kill" then I guess that according to your POV he shouldn't have sold to the military even if the system was only suggesting "valid" targets.
There are probably many CEO tech founders of small companies that are well intentioned, but our economic system ensures that only the most sociopathic among them become CEOs of the big ones.
I don’t even think they have a PR problem, if their revenue numbers are correct. Sure, maybe in the court of public opinion, but time and time again it shows how it can be fixed.
I feel that everyone in this small AI bubble (compared to the "normies") is busy trying to take advantage of AI to surpass everyone else, while fearing being surpassed by everyone else.
Most people assume LLMs are just a tool for making money, an arms race over wealth and social rank. But those like Dario seem to be worried that LLMs could potentially be used as automatic rifles by some dedicated psychopaths. Whether it could be that dangerous, or just an exaggeration, or we should have everyone armed, require licenses to own one, or outright ban them, is debatable. But I can understand his point.
All I can say at the moment is that this is more subtle than the superficial conspiracy/black-and-white narrative.
> Anthropic in particular has developed this almost Orwellian like veil of condescending rhetoric that on the surface suggests they're looking out for you while underneath they're taking actions that suggest they do not trust you, Mr/Mrs Ordinary Person. All the safety rhetoric, never supporting open weight models, the lockdowns on harnesses outside claude code, etc.
This line of reasoning is really very silly. I can be looking out for my young child by choosing not to put them behind the driving wheel of my car.
There are both stupid and malicious people in the world. If you have a tool that you believe may be genuinely dangerous in the hands of the wrong person, it is not the smart thing to do to give the public unrestricted access to that tool. Do you think it would be sensible to give every individual on Earth a nuclear weapon to do whatever they want with?
I think it's more reasonable than it sounds on the surface. People's jobs, the bubble, etc are societal level problems and well beyond what Anthropic could even hope to influence on their own.
Curing cancer sounds insane, but it's also a research problem, not a societal level coordination problem. And one AI has already proved to help with breakthroughs (alphafold). IMO it makes sense for them to shoot for something like that as proof of AI's beneficial sides.
>People's jobs, the bubble, etc are societal level problems and well beyond what Anthropic could even hope to influence on their own.
This is not true. There are two companies at the center of AI direction: Anthropic and OpenAI. If there were anyone on the plant who has the ability to influence our direction then it would be Dario Amodei.
Acknowledging the grievances is a good step, but it's not enough. There needs to be a clear explanation of actions to address them, a plan to enact those actions, and commitments with consequences in failure of those actions. Tell people how you're going to make them more employable and effective and needed. Tell people how your datacenters will be carbon neutral. Tell people how financial actions resulting in a frothy market will be coming to an end. He and Sam Altman alone have this power and their inaction says everything we need to know about their intent.
They're big, but they don't have that kind of power and people would be very suspicious if they tried to get in the position to have that kind of power.
Hoping that they solve those huge problems themselves is basically hoping that they will have more power.
It would be like expecting Craigslist to rescue journalism. The amount of power needed to disrupt an industry is nowhere near what's needed to fix it.
I don't think the problem is that people don't believe AI can be beneficial. The problem is the perception that AI will affect people's lives more negatively than positively, while a few get even more obscenely rich and powerful.
If you don't tackle the societal level problems, nobody will care about research breakthroughs.
I don't see a lot of societal level problems other than the labor market, for which both Anthropic and OpenAI create research, monitoring, and policy advice.
I don't think anyone will care about the largely made up issues about electricity prices or water use in case AI literally cures cancer.
How do you not see other societal level problems. This technology destroys human communication. It undermines the thing the internet came into existence for, and turns it into a liability. It grants intelligent analysis to a sprawling surveillance network. It makes people mentally ill. It destroys confidence and hope in technological progress. Who even cares about cancer at that point? I would rather die of cancer than live in a fake, ugly, artless world full of fake, stupid, and crazy people under constant surveillance by an AI.
> Curing cancer sounds insane, but it's also a research problem,
I just don't buy the AI labs approach to this stuff. Like, unless we can basically simulate the entirety of human biology, I don't really see how LLMs can make progress here. Maths is different as it doesn't require a real-world interface, and programming already (by definition) can be simulated on a computer.
Without that, I can't see much (if any) progress being made on domains like biology.
I'm a noob on this topic, but I think drug discovery is more amenable to this structurally than other problems. Simulating biology is what we were doing with protein folding before Alphafold, and the search space was far too large to find stuff in reasonable timeframes. Alphafold showed that you could take a physical process and make a neural net clever enough to learn just enough structure that it starts finding things we might care about, and still physically accurate, much faster.
Drug discovery is similar AFAIK. The space of possibilities is even larger than protein folding, but it's structurally similar enough that I think AI will help to make progress on the discovery side. Actually getting the drug tested and approved is another matter though for sure.
For one, the question for Anthropic is whether LLMs, specifically, not AI techniques more generally, can help significantly with cancer research. And here, all experience so far is that LLMs only really work when they can easily automatically verify their own outputs and self correct - such as in math (using automatic proof verifiers) or programming (using compilers and unit tests).
The second problem is that biological research speed is highly dependent on slow biological processes, such as cultures and long term studies. In programming, if an LLM could provide excellent insights and research suggestions 100x faster than a human, it would speed up the work roughly 100x. But in biology, it would only speed up the total work by a small amount - as any insight, even if absolutely brilliant and spot on, would still require months and years of actual experimentation.
For drug discovery, you can simulate interactions in silico, driven in an agentic loop via LLM. Mostly stil using non-LLM tools, of course, but it saves time.
If I was Dario, and I was after a cynical sales grab to go with the "actually curing cancer!" spiel, I'd probably aim for a drug repurposing strategy. I'd use the above approach to find existing drugs that might target known pathways that drive incurable cancers. I'd only screen compounds with extensive safety data and easy delivery mechanisms, and from the in silico hits, I'd throw a tonne of money at rapidly experimentally screening all those candidates in parallel, and then rapidly push those that worked into clinical trials. I'd assume that, with a small but non-negligible prior, and the money to push through hundreds and hundreds of candidate compounds through at once, I'd have a reasonable chance of getting one drug through to a "Claude Cured Cancer!" show-stopper headline.
But I think even then, with all of Anthropic's wealth, you'd need 2 years minimum to move from initial screening targets to a Phase 3 trial. And it would be an almost criminal waste of research funding to get there -- the dollars for discoveries ratio would be appallingly bad.
(If I was Dario, and I actually wanted to cure some forms of cancer? I'd just use my obscene profits to fund actual cancer research, step back, and let the researchers get on with it.)
> Drug discovery is similar AFAIK. The space of possibilities is even larger than protein folding, but it's structurally similar enough that I think AI will help to make progress on the discovery side.
I do agree that this kind of targeted approach makes sense.
However, discovery is not really the issue here. Running the clinical trials (1/2/3) is much much more difficult, and consumes basically all of the time in drug development, so even if LLMs perfectly automate this, the speedup will not be particularly large.
> Simulating biology is what we were doing with protein folding before Alphafold
This is not correct, Alphafold predicts crystal structures of proteins. It is a known issue in the field that these folding models do not generate ensembles of protein conformations like the ensembles generated by DE Shaw Research (who built a super computer to simulate proteins). And these folding models fail when we don't have crystal structures, so they are certainly not generalizing. A great paper on the subject: https://www.biorxiv.org/content/10.1101/2025.02.03.636309v1
Drug discovery is significantly more complicated than protein folding. Small molecules with similar chemistries can adopt novel binding poses, can end up binding to off targets (PXR, hERG, etc) or simply never make it to the protein due to solubility or permeability.
I will always defer to Pat Walters who seems like the most careful and sane person in this field: https://patwalters.github.io/
The real issue is that biology needs actual experiments done in the physical world, which isn't nice and orderly and well behaved and easily loadable onto a 19" rectangular box.
There is a lot of computational chemistry and biology done in the early stages of research - there are now multiple orders of magnitude of computation power available that is doing nothing but feeding forward on billions of random numbers.
On the other hand, AI means actual experiments done in the physical world but coordinated by an entity that never sleeps and never gets depressed and can multiply itself manifold and always comes up with new ideas.
> Curing cancer sounds insane, but it's also a research problem, not a societal level coordination problem.
Curing cancer is a societal problem. We have lots of cures for common diseases but resource allocation means people don't actually receive the treatment they need. For example, prohibitively expensive gene therapies, or HIV treatments in developing countries. A disease may have a "cure" but if people who need it don't receive it then from their perspective it may as well not exist.
Based on rhetoric from open ai and Anthropic, i don't think this is reasonable. If anything, their miracle of an AI should be able to solve the job market and economic bubbles. Deploy their agents on it, set up automated hedge funds, distribute the profits equitably to everyone on the planet and on and on.
But something tells me either they can't or they won't, so no trust will be built.
> People's jobs, the bubble, etc are societal level problems and well beyond what Anthropic could even hope to influence on their own.
Anthropic is literally creating the bubble. It is not beyond their scope of influence, it is literally what they are consciously achieving.
As for peoples jobs, same actually applies. Anthropic is selling itself on dream of replacing jobs, even or especially where they are well aware AI does not perform that well. They are actively trying to replace people quickly before management notices it does not work well.
And also, they can influence how much their data centers contribute to global warming.
Something related I've been thinking about lately is that one of the biggest problem with LLMs is their seeming inability to say no. Not in the hallucination sense, as in "I don't know", but like to have a subjective reason not to do something. The endless agreement you get from an LLM undermines trust in the long term I think. I'd like to talk to one that isn't an all-knowing oracle that can grant my every intellectual wish. (Or maybe what I'm asking for is just... a human, lol).
One word that few wealthy people ever hear, is “No.” It has a pretty significant effect on their worldview. Even the most reasonable, well-informed, well-intentioned, wealthy folks can have their thinking affected.
When every silly, should-be-smothered-in-the-crib idea gets enthusiastically endorsed by your entourage, it’s easy to lose the ability to self-regulate. I’ve watched it happen, numerous times, as acquaintances and friends have become more successful.
Obsequious LLMs are leveling the field. Less wealthy folks now have the chance to lose their ability to self-regulate, just like rich folks.
Your image is wealthy people is cartoonish. Sure if you go to a high end place they'll try to meet every one of your demands. But chances are is you're very wealthy you're running a business or group of people, and you'll hit obstacles constantly. I've watched this happen multiple times. Internally there are sycophants but when you deal with the real world and try to get deals done, people don't owe you anything.
In the meanwhile, we can all publicly see how the the billionaires and trillionaires, behave and settle their priorities exactly, in the cartoonish way you are dismissing.
- The incredible insecurity and constant need for personal validation.
- The absurd and obsessive pursue of further wealth when it would be temporally impossible to even spend 1% of their current capital.
- The extreme level of cowardice, where an auto plant worker, can call out the powers to be, the pedophile protectors that they are. At the same time, only Jensen Huang did not submit itself, to the humiliation of standing behind face and front at the presidential inauguration...
> "The absurd and obsessive pursue of further wealth when it would be temporally impossible to even spend 1% of their current capital."
That shows profound lack of imagination. Mark Shuttleworth's founding of Canonical was a choice to expend his wealth to try to popularize Linux. So was Gabe Newell's long running efforts with the Steam Machine (dating back to the original one in 2015) and Proton that only started bearing fruit in recent years. These were business decisions you say? At the level of wealth we're talking about, the two are intertwined.
I love what he'd done for linux gaming, but man the cult of personality for gabe is strong on the internet.
gabe is not a philanthropy who want to see linux gaming happen no matter what.
he's a reasonable man who saw a threat to his business when every operating system started making their own store and windows was directly trying to eat his lunch starting from window 7 and going at it the strongest in window 8 when they had their dedicated to game store directly pinned in the windows bar menu.
Microsoft's threat was more overt than that: the initial plan was for Microsoft to be the sole signing authority for Windows apps in Windows 7, while simultaneously locking down the OS from pre-boot onwards and not executing unsigned binaries. This would have killed not only Steam games, but Steam itself wouldn't be distributable. Microsoft forced Gabe's hand! Valve had to pivot away from being 100% dependant on Windows as a matter of corporate survival.
The Windows Store was introduced in 2012, 14 years ago, and promptly went nowhere; Gabe Newell is ex-Microsoft himself and undoubtedly had a good understanding of how dim the prospects and how poor the execution was for it. Yet he persisted in pursuing the Steam Machine and Proton long after it was clear that the Windows Store was a flop.
If you want philanthropy (or perhaps not), Gabe Newell has $1 billion worth of yachts and research ships.
So I dunno anything about Gaben’s actual motives, but I will say, I personally would not be assuaged just because the store flopped. The underlying problem is that Valve sells games, and games needed Windows, and that’s exactly the sort of dependent relationship that Microsoft (and others) use to extract value and dominate new markets. It makes sense to me that Valve looked at how hard it would be to break out of that, and decided that it would remain diligent regardless of how credible the immediate threat was.
Not a single one of the other gaming giants (Blizzard, EA, Epic, Activision, etc.) or marketplaces (Epic, GOG, etc.) who also depend on Windows bothered to counter this supposed threat? Doesn't sound like much of a threat then.
Really, they're probably more displeased about the Valve near-monopoly as a marketplace for Windows games than displeased about MS.
Given the state of ... well, everything ... it still surprises me when people pretend there's some level of modesty or decorum the rich and powerful still need to abide by at the risk of some sort of popular revolt. I guess it's a way to compensate for the material powerlessness most people actually experience (especially in politics), much like the handwringing over the Second Amendment in the US as if merely handing a town full of ordinary people guns is sufficient to defeat a modern military. Sure, it looks like nowadays you can do that with a stockpile of low budget drones because even the military industrial complex has found a way to optimize for profit by literally delivering nothing in return for extracting the lion's share of the federal budget, but even that still requires actual work like planning, training and organizing with sufficient political motivation, not just going to the range or "hunting".
The expectation of a need for decorum against all evidence almost seems like a quaint remnant of the aristocracy at this point. There's no advantage to being a nice or sensible person if you're a billionaire. People talk about "business" like it's still the 1800s and billionaires are the factory owners. The real economy (i.e. anything involving any pretense of being an exchange of goods and services in any shape or form) is now a minute fraction of the global economy. Everything else is finance. And if the nature of the finance "industry" wasn't obvious enough the US has given up all pretense to the point that members of the US government now intentionally manipulate online betting "markets" directly.
You may have needed some table manners to be able to run a tincan factory. You can let your entire ass hang out on TV 24/7 and still be a billionaire today. And people will still cheer you on like you're the inventor of sliced bread.
A slight tangent, but on your first point - defeating a modern military does not mean engaging them in head-to-head combat and coming out on top, but making it impossible for them to comfortably operate. This is how Afghanistan, with basic weapons and religious extremists running around in sandals, managed to defeat not only the US but also the USSR.
The big factor that works against military victory anywhere is that the military is left to visibly identify themselves as a means of projecting influence, whereas militants can seamlessly blend in with the population. And trying to punish the population at large to weed out the militants mostly just tends to backfire and create even more militants.
There is even modern precedent for this like the Algerian War where the people of Algeria managed to kick out the French even with the French engaging in widescale massacres, torture, and all other sorts of fun stuff that explains how anti-Western powers are so comfortably gaining influence in Africa today.
If one doesn't believe there was a conspiracy, then the assassination attempt on Trump would exemplify this - one guy with a rifle nearly single-handedly killed one of the most guarded men alive. Your average military logistics support doesn't have anything remotely like a team of secret service and police covering every single angle of attack in one well secured location.
> If one doesn't believe there was a conspiracy, then the assassination attempt on Trump would exemplify this - one guy with a rifle nearly single-handedly killed one of the most guarded men alive.
Sure, but that being a failure mode is a design choice. Look at Iran: we've seen how many decapitation strikes now? At this point there shouldn't even anyone be left that was in the original chain of command but clearly they're chugging along quite nicely because they decentralized military operations exactly for a scenario like this, knowing anyone appearing on the US' or Israel's radar too prominently was one bad day away from needing replacement.
> This is how Afghanistan, with basic weapons and religious extremists running around in sandals, managed to defeat not only the US but also the USSR.
Not quite. The US managed to "defeat" Afghanistan. The problem was just that nation building was never really on the menu. The US was extremely hands-off during the occupation because their priority was making sure more US soldiers don't come home toes-first when the war had officially been declared won than to build any kind of legacy. This also led to massive subversion of the supposedly US-backed government. There were very few incentives in supporting the US more than necessary not to get shot at, especially after it became clear the US government had no interest in rewarding the loyalty of most of the people who had risked their lives as translators and negotiators, let alone supporting the people who volunteered to join the new army or police. That's why when the US left the supposed new Afghan state just dropped its weapons and ran away and why the Taliban could pretty much just take over without much resistance.
A good counter-example is Iraq where not only did the US have to do much less work to do (because Iraq already had the concept of a powerful central government rather than effectively being a coalition of often loosely connected tribal communities), they also engaged in a lot more active nation building and had reasons to maintain a long-term presence.
If you want to see how bottle rockets vs modern hardware can also play out, look at whatever the IDF is doing at the moment. Granted, the US military being deployed against US citizens would probably show a lot more hesitance and restraint, but again that's a design choice, not an inherent property.
> And trying to punish the population at large to weed out the militants mostly just tends to backfire and create even more militants.
That assumes militants are a problem, not a calculated risk. Long-term social instability is manageable if you don't rely on popular approval as long as it doesn't disrupt supply chains enough to directly threaten the military's ability to operate. Israel for example arguably flourished for decades while having to continuously deal with militants and responding with large scale violence.
In the Iran War, Iran's strategy is much more akin to to a resistance force than a normal military. Israel clearly didn't believe their decentralization would be viable, or they wouldn't have had the US assassinate their leaders - because now they have younger, much more hardline leaders, with a blood debt.
As for Afghanistan I think people don't realize how much fighting was still going on over there. The Taliban forces weren't just camping out in the mountains and defending, they were actively striking all US forces and US proxy forces. In total hundreds of thousands were killed. The Taliban killed thousands of US soldiers, wounded tens of thousands, had a significant minority of land under their undisputed control, were actively contesting control of much more of the country, and rapidly took control of the country once the US left. And yeah nation building was 100% on the menu. There was a friggin George Floyd mural in Kabul lol. But the US was never able to get the country under control, and things were trending worse more than better.
Iraq was much easier for one simple reason - it's a heavily divided country which Saddam (who was himself part of a minority Sunni population) was able to stabilize only with an iron fist. This made much of the country hate him. He was also extremely paranoid, probably in part because of this, and didn't trust his own armed forces which led to direct efforts to limit their power/capabilities, installing commanders based on loyalty over competence, and so on. So it was pretty much an ideal target for a decapitation strike. But even in that ideal scenario we were left waving a victory from flag within tiny little heavily fortified green zones, because stepping outside of them was highly dangerous.
And I don't think Israel is showing anything productive at all. Hamas has comparable numbers today as when the war started, Israel's leader has a warrant for his arrest for genocide, and the future of Israel is probably more uncertain than it's ever been in its entire existence because they've isolated themselves from basically everybody except Trump who will be gone in a couple of years.
I guess you don't observe those billionaires yourself. Is it possible the cartoonish image you assume to be obvious comes through media channels and your societal bubble layers?
HN must be the only place on the web where billionaires are propped up. 98% of you, are already in the SQL result set, of the next layoffs at your big five...
Uninformed HNers grouse endlessly about the wealthy and megacorps in without understanding their actual psychology, motivations, and sources of power/income. Then they are left wondering time after time why all their campaigning and efforts against them comes to naught.
There’s nothing to understand. Power begets power and self interests are pursued to the detriment of others partly because individuals have limited understanding of the full system and partly because of plain selfishness. You can debate how big each part is, but there’s a systemic issue here.
Oh you should see the Musk-adjacent subreddits. Not to mention MAGA. Billionaires are still worshipped on the “wealth == success” basis by many, many people.
You either jumped to express your opinion without reading previous comments, or you are not being honest, cause it's not about "all billionaires good" claim
"Multiple studies show that drivers of expensive, luxury cars are significantly less likely to yield to pedestrians at crosswalks and more prone to committing traffic violations compared to drivers of budget-friendly vehicles" - https://www.ralphaschwartzpc.com/blog/study-luxury-car-drive...
I suppose, the most naive is the one who believes we can't read comments above, and see where it started, and who's trying to pivot.
Btw, I opened the very first "scientific proof" you offered, and it happens to be a story with no links or data about one cool student's deep studies where people who believe stealing isn't so bad steal more often, and that's why upper class is bad-bad. So how do I know you didn't even try to read it before posting?
"...In seven separate studies conducted on the UC Berkeley campus, in the San Francisco Bay Area and nationwide, UC Berkeley researchers consistently found that upper-class participants were more likely to lie and cheat when gambling or negotiating; cut people off when driving, and endorse unethical behavior in the workplace.
“The increased unethical tendencies of upper-class individuals are driven, in part, by their more favorable attitudes toward greed,” said Paul Piff, a doctoral student in psychology at UC Berkeley and lead author of the paper published today (Monday, Feb. 27) in the journal Proceedings of the National Academy of Sciences..."
Wow! So I correctly described it? It's a student's work, there's no data attached, and the whole matter is about "greed to be the most significant predictor of unethical behavior" just expressed in class war terms.
"...Drawing on a unique dataset of leaked customer lists from offshore financial institutions matched to administrative wealth records in Scandinavia, we show that offshore tax evasion is highly concentrated among the rich. The skewed distribution of offshore wealth implies high rates of tax evasion at the top: we find that the 0.01 percent richest households evade about 25 percent of their taxes. By contrast, tax evasion detected in stratified random tax audits is less than 5 percent throughout the distribution..."
The point I was responding to was claiming rich people lose their ability to self-regulate. You are posting about them being mean, more narcissistic and more greedy.
>> posting about them being mean, more narcissistic and more greedy.
Traits they exhibit plenty more, than the general population: "The “Why” and “How” of Narcissism: A Process Model of Narcissistic Status Pursuit" - https://pmc.ncbi.nlm.nih.gov/articles/PMC6970445/
Bankman-Fried — FTX -> misuse of billions in customer money for investments, influence, political contributions and personal interests despite enormous wealth.
Bill Hwang — Archegos -> extreme risk-taking, market manipulation and deception that ultimately imposed billions in losses on banks.
Alex Mashinsky — Celsius -> personal profit from token sales while customers were misled and ultimately left unable to access billions in assets.
Bernie Madoff -> massive Ponzi fraud sustained for years despite already having wealth, status and an elite financial reputation.
R. Allen Stanford — Stanford Financial -> diverted billions from investors to finance his businesses and lifestyle despite extraordinary existing wealth.
Leona Helmsley — Helmsley Hotels -> billionaire convicted of tax evasion and fraud involving personal luxury expenses. The sentencing judge explicitly described the conduct as motivated by "naked greed" and an arrogant belief that she was above the law.
Elizabeth Holmes — Theranos -> maintained sweeping false claims to investors while building a company valued in the billions, resulting in hundreds of millions invested on false premises.
Trevor Milton — Nikola -> repeatedly exaggerated his company's technology and achievements to stimulate investor demand and support its stock price.
Karl Sebastian Greenwood — OneCoin -> helped sell a fictitious cryptocurrency to millions of victims who invested more than $4 billion while he personally received hundreds of millions.
Miles Guo — GTV / investment schemes -> former billionaire convicted of using lies and misrepresentations to extract more than $1 billion from followers
Do Kwon — Terraform Labs -> deception surrounding a huge crypto ecosystem that culminated in roughly $40 billion in losses
Joseph Lewis — Tavistock Group -> billionaire who admitted abusing confidential corporate information by tipping friends, employees and romantic partners, while his company separately admitted securities fraud involving concealed share ownership.
Elon Musk — X -> Changed the platform algorithm after Biden posts outperformed his, massively increasing exposure of Musk own posts.
Jeff Bezos — Venice -> turned his 2025 wedding into a three-day billionaire spectacle that prompted protests over the privatization and commodification of the city.
Mark Zuckerberg — Meta -> commissioned and publicly displayed a giant Roman-style statue of his wife while simultaneously building one of America most extraordinary private compounds.
Bryan Johnson — Blueprint -> turned his own body into a multimillion dollar anti-aging project, including receiving plasma from his teenage son in an unsuccessful rejuvenation experiment.
Larry Ellison — Oracle -> built a personal real-estate empire including ownership of 99% of a Hawaiian island
There are 3,428 billionaires, 28,000+ $100-millionaires, 2,300,000 $10-millionaires on Earth and you've only managed to list about a dozen people. Not very persuasive.
It seems you misunderstood my comment. I'm not arguing against them being more narcissistic or greedy compared to the baseline population. I'm saying that was not the original point I was objecting to. I'm not sure what your list of examples is meant to accomplish.
"If it's worth your time to say something false, it's worth my time to correct it". I don't take a stance on who is right, but your dismissal here, which doesn't deal with the truth of the matter, is worse than both.
Talk about living a sheltered life style. I did too until doing my obligatory military service. Now I don't even correct a quarter of wrong things I hear.
There's literally nothing in my previous commentary which "props up" anyone. The layoffs mention suggests that as someone who may lose job I must support anything said against billionaires whether correct or not. I sincerely urge you to re-visit your position cause putting political statements above facts amounts to dumbing down oneself, with a potential to become an enthusiastic koolaid drinker if you see what I mean.
the irony of you getting your panties in a bunch over being disagreed with is simply delightful. not so different than those pesky billionaires, are you?
Elon Musk runs his own media channel, which he uses to depict himself like this. Many others are run by other billionaires. State-owned media like the BBC and France24 tends to paint them in a better light, perhaps out of a desire to treat their subjects charitably and without bias.
This “billionaires couldn’t spend their money” thing comes up a lot, and it show’s a child’s idea of what people might spend money on. “$1000 a month buys more ice cream than I can eat and more toy cars than anyone can play with, and I don’t want anything other than ice cream and toy cars, so nobody needs more than $1000 a month.”
I can think of loads of things I’d do with hundreds of millions of dollars of disposable income. Make large buildings that look like how I think buildings should look, fund scientific research in areas I’m personally interested in but have little aptitude for, find artists that align with my tastes and give them the ability to realise their visions, contribute to political organisations and charities that align with my values, et cetera. Those all are basically infinite money sinks.
In the past, the ultra-wealthy did do all of these things, funding loss-making research and expeditions, monumental architecture, philanthropic organisation and artistic works. The problem might be that you think billionaires shouldn’t do some or all of these things, but that is a very different position than that they can’t do any of these things.
This comment seems needlessly insulting and also kind of obtuse?
I don’t think the GP meant that it’s impossible to dispose of a large amount of money. But it is effectively impossible to spend down a certain level of accumulated wealth by purchasing goods and services that any human being or their family could require to maintain even lavish standards of living.
The examples you gave are conversions of wealth to power, using money to reshape the world. Which, I think, is kind of the crux of the issue. Perhaps you were alluding to this at the end of your comment.
I think the actual snag though is that people genuinely believe that giving money to the government is the highest form of good.
Skirt taxes to give a billion dollars to prospective college kids through your own organization? Nice gesture, but still a greedy fuck who refuses to give up control.
Pay a billion in taxes that ends up mostly going to heavy administrative bloat? Give that guy sainthood.
People misuse the term "wealthy" to refer to whatever is in their head at the time. I think GP's point stands if you consider "wealthy" to mean "billionaires". Sure, people like Bezos or Musk get to hear "no" quite frequently but they don't take kindly to it and they can usually offload "getting around it" to other people.
The point is less that these people are surrounded by "yes-men" but more that wealth (especially when measured in billions) is power and with sufficient power it becomes easy to forego any question of consent, let alone of whether consent is coerced or not. Remember that power is ultimately about the ability to enact violence and violence can take many forms, most of which are perfectly legal (because the legal system itself exists to regulate how, by who and against whom violence can be used).
You tend to hear "no" a lot less when you always point a gun at the head of the persons you're asking. Note that "wealth" isn't the only way to get there but a certain level of it is usually necessary to get to the point where other options become available - and some of the ways are a lot riskier in the long term (cf. Epstein).
Side note: this is also why I hate the pseudo-intellectual counter argument against "billionaires" of "that doesn't mean they have billions of dollars sitting in a bank account" - it's like arguing that De Beers didn't benefit much from holding a quasi-monopoly on natural diamonds because the diamonds would be devalued if they flooded the market with the ones they had intentionally kept off the market to drive up value: beyond a certain amount money ceases to be about liquidity and starts being about leverage. Unless you happen to be dealing with lower level bureaucracy in Russia, the most efficient way to use wealth to your advantage isn't to just hand people stacks of dollar bills.
I’ve found that it’s the opposite. Most people above a certain financial wealth will learn the lesson that there are limits, and that “if only I had the money to…” is an illusion. They realize that there are other forms of wealth that may even be more important than money. What you are talking about is the very few who seem too far gone to be able to get that.
I don't know. It probably varies. SV wealthy is quite different from Appalachia wealthy. It's that point, where people start worshipping your money. Some wealthy folks also make a point of showing off their wealth, so it starts earlier, for them.
Also power. You see the same thing happen with managers that dismiss criticism, and have the power to make it stick.
The specifics are certainly cultural and locally relative, but Marx nailed it when he talked about ownership of the means of production, rather than being part of the production of goods and services.
The reason that the ultrawealthy behave like toddlers is that toddlers are, relatively speaking, ultrawealthy: all their needs are met without any effort on their part, and so many of their desires are fulfilled simply by expressing those desires out loud that any impediment or refusal is obviously enemy action.
It's invariably the class of people who if we rose up and seized everything from we would all get a one time "life changing" check for $50k, while collapsing 2/3 of the economy.
Its even less than that. Elon Musk's $900 billion net worth would only give each American a one time payment of $2647. Jeff Bezos' $270 billion net worth would only provide each of us $794.
I would extend your thinking to any well intentioned folks being very capable of having their "thinking" affected.
For example poor people who have never thought about rising out of it - eg about 50% of kids in my Brooklyn public highschool had parents who didn't give a shit if the kids studied or not. Completely oblivious to how the world works - meanwhile the other 50% wa immigrants who pushed their kids and those kids are now in the 1%.
In general I think what's more telling than your level is your journey. Someone born rich maybe mirrors what you described (I don't know people like that) but the few centi-millionaires and billionaires I "know" (ie worked for and dealt with in that context) have encountered plenty of "no".
When you are building a company, you are going to get a lot of no. No I won't buy, no I won't work for you, no I won't invest in you. In fact I would say a universal attribute of someone who has "made it" is having ample of experience getting "no" and dealing with that fact property. That's true even like at the level that plenty oh HN readers are - a successful faang employee and the like.
For what its worth - I generally find that orienting to what some other group is like "rich people are like x etc" is a tell-tale of not focusing on what's within ones sphere of control and knew life. Any brain cell I spend fantasizing about someone else's imagined behavior is a brain cell not dedicated to engaging soberly with my own reality.
For example poor people who have never thought about rising out of it...
I obviously don't have numbers on this, but I strongly doubt there's a poor person on this planet who's never thought about "rising" out of it. That's the dream the lottery sells, that's why so many kids want to be basketball stars / celebrities / influencers, etc.
My personal experience is that the required difficulties of my life have decreased in direct proportion to my income, leaving mainly the self-imposed difficulties. It's not hard to extrapolate that line a little further to billionaires.
We're taking about different things. To continue my example. A nyc public School student is a 34k/ year investment for the city.
De facto some parents look at that as a gift and "force" their kids to take advantage of it. That's why the valedictorian etc is usually an immigrant kid not a rich"native" kid.
Meanwhile plenty of parents seemingly have never considered the opportunity in front of them. Content to let their kids not study and do stupid crime.
To say it simply: not everyone values education the same degree - or a all. That's all I am saying here - you'd think a poor family would grab to the opportunity to rise out through education but it's obvious that for many many many people this has never crossed their mind. If you had not encountered this in your own life I find that odd.
I think you're confusing two things. A chatbot keeps talking because it creates engagement and just saying "i dont know" or "no" kills the engagement so naturally one would assume it is trained to always try to provide some sort of an answer and try to keep the user engaged.
But that doesn't mean it will do whatever you ask it. Ask Chatgpt to assisinate someone or buy drugs and it will tell you to f off. But what corrupts people, is these kind of things, where you are a mini king beyond ethics and morals.
Greatly varies by the model, but such a flat denial along with calling those with (verifiably accurate [0]) experiences which happen to be different from yours "the anti-AI crew" and just assuming they can't possible use LLMs and couldn't possibly have formed their opinions on evidence, well, says a lot.
LLMs, even frontier models by major providers, still have no reliable internalised way to assess the accuracy of their output and they have, do and will continue to for the foreseeable future, lead people down paths they shouldn't [1], partly by their architecture and the limits of the technology, partly by an intent for maximising retention.
For what it's worth on people, I have, both in politics and business, unfortunately made the painful discovery that, whether intentional (because the powerful person in question cannot or doesn't want to deal with different opinions to their own) or unintentionally (because those with sycophantic tendencies simply managed to manipulate themselves into their inner circle over years and became trusted), there are people in (financial, political or other types of) power which are surrounded by few willing to tell them when they are wrong and even fewer that are actually listened to if push comes to shove, which often affects the personality and mental health of said powerful people negatively, to the detriment of society, their family, their employees, etc.
Members of the media are actually complicity. I very much disagree that the media presents obscenely powerful people in too negative a light, more the opposite. Am very firm that, to retain access to the rich, powerful and famous, there is far too much sane-washing of utterly ridiculous, unacceptable, harmful and/or dangerous behaviour. Sometimes the person in questions own health and safety are put at risk, because neither the people around them, nor public opinion or reporting treat their behaviour in the way appropriate. Sometimes this again leads to harm for the public, their family, those working under them, etc.
What'd get someone with less power or in a lower tax bracket ridiculed or even sectioned is often reported as "eccentricities", just being "passionate about a topic" and trying to do the "marketing rounds".
> LLMs, even frontier models by major providers, still have no reliable internalised way to assess the accuracy of their output
This is false. Perhaps you meant they have insufficient methods, or imperfect methods, but asserting none at all is facile and your links do not say that at all.
If your believes were true, this actual quote, pulled from a recent sonnet conversation, would be impossible:
“The XT60’s 15A limit is unsuitable for an appliance that… no, it’s the XT30 that is rated for 15A, XT60 is rated for 30A continuous. For a 20A appliance, XT60 will be fine.”
> This is false. Perhaps you meant they have insufficient methods, or imperfect methods, but asserting none at all is facile and your links do not say that at all.
I did say "no reliable internalised way to assess the accuracy of their output", which is something very specific. If that Sonnet output is reasoning traces, there are many issues with using that as a source:
For one, back and forth reasoning does often correlate with less, not more accurate overall outputs in evaluations and one example could never seriously be extrapolated to be considered "reliable", i.e. happening consistently and dependably.
Secondly, self-correction is also not self-verification in regard to model output, revisions not necessarily mean internal accuracy assessment by themselves (again, over-revisioning has lead LLMs in my and even public evals like the one linked above to step away from accurate information written in their reasoning traces but discounted in the final output (if we must use anecdotal examples like your Sonnet quote)).
Then there is the fact that, unless that reasoning trace (if it indeed is one) was copied from a months old chat history, Anthropic has obfuscated their reasoning traces so this output is (if it isn't an ancient history you dug up) from the obfuscation model in between and not reflective of the actual models reasoning. So even for anecdotal evidence, this can likely not be used (unless again, you went for December 2025 history). If this was not reasoning traces (I am fairly confident it is having spend a long time reading Anthropic summarisations vs actual reasoning traces when the switch was on-going and you could get both for a limited time, what you shared reads very much like obfuscated rather than pre-obfuscation reasoning output), that still leaves this as anecdotal, one time, possibly erroneous (did this back and forth even prevent an error in the final output) and there are more issues still with just using that quote as evidence, this is simply unsuitable as a source in any situation.
Here are some papers I read lately, all published in 2026 and using the current crop of models which were what led me to make that specific statement. LLMs currently have no reliable internalised way to assess the accuracy of their output, at least as far as the literature is concerned:
> Even state-of-the-art models struggle to reliably discriminate between data uncertainty and model uncertainty.
Beyond “I Don’t Know”: Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty [0]
> LLMs cannot reliably revise their own errors without an external signal.
> The same models that confidently catch and repair errors in external content routinely fail to identify identical errors in their own reasoning traces.
The Self-Correction Illusion: Role Relabeling Gates Explicit Error Flagging in Large Language Models [1]
Simply, as of today, LLMs cannot reliably translate whatever internal signals they have into accurate self-assessment of their outputs. Even in papers that show limited, edge case capability, this often breaks with minimal prompt and/or task changes (happy to link those too, I just need to get to my Mac where I have the PDFs), so it is not reliable.
If you got a paper that shows that I am false, that there is reliable, internal self-verification of a models output (even a small research LLM not yet publicly released), I'd be happy to read it and retract my original statement.
I don’t think it would be a positive if every time someone said “wealthy” they had to add “relative to their country’s average earnings and level of savings”. It’s implied.
If someone is struggling to afford a home, “you know there are much poorer people in Africa” isn’t a particularly helpful or useful response.
I also think this is why LLMs were trained to behave the way they do. The people who gave the training objectives and evaluation targets were exactly those rich folks who never hear "no". Hence LLMs are their dreams of a perfect servant.
I found another sign of that is the way LLMs answer with a professional, business-like tone even if the request is completely bananas. It's what a concierge or butler would do, but not an actual close friend.
Where do these clouds come from:
Points to a far away direction in the sky and says they come from there.
Who does all these roads, trees and environment belong to?
It all belongs to me. Obviously.
They have an answer ready for every question you throw at them and they will answer it with absolute certainty. I will have to wait and see at what age does the concept of "I don't know" develop.
The difference between your two year old is that an LLM gives useful information.
Yesterday I decarboxylated some weed buds in preparation of making a cannabis tincture using the QWET method. Curious how Claude would respond, I asked how to do it.
It walked me through the process and gave accurate, nuanced answers.
Let me know what your 2 year old thinks I should do.
Gemini estimated that male cannabis plant leaves I decarboxylated will have negligible thc content and give me mild relaxation at best, the real effect was it was the highest I've ever been.
Gemini sucks though. Are you using the paid version? The free one that has basically replaced the regular search, is totally useless because of how shallow it searches.
ChatGPT, which is usually quite reliable, refuses to answer me because law.
In the end, an extraction should work any plant material depending on potency, skill and available equipment.
After the extraction you can evaporate the ethanol(carefully since it’s highly flammable) and increase potency.
A mix of own research(basically emulating others) combined with strong models we can do a lot more than we can do ourselves.
To be fair, I have yet to lead a conversation with Gemini that doesn't just consist of me having to check its responses and point out that they're objectively incorrect only for it to "apologize" and then give me the next wrong answer, while always making sure to end with a new conversation teaser.
Granted, I've only been using the free version but I've been getting significantly more mileage out of those from ChatGPT and Claude. I guess it might be a feature that Gemini is more often obviously wrong from the start (e.g. by giving sources that don't support its claims whatsoever) but considering this is the AI from the company that had become synonymous with the concept of trying to find information on the Internet, that's pretty damn pathetic.
That is because it bases its responses on web searches with hits like comments in threads like these or blog articles that were generated from systems trained based on threads like these from the previous iteration of scrapers and generative language models.
Claude and ChatGPT have between them identified one mystery plant in my garden confidently as about a dozen different things.
Even though they will get chemistry right more often than me, I still wouldn't want to ingest the result of it walking me through that on a drug, psychoactive or otherwise.
Any given answer might be right, but I'm not a trained chemist and don't know how to safely test things.
Not really, since any LLM will answer all those questions competently.
It's a known fact that LLMs sometimes are wrong and hallucinates an answer, but this is exceedingly rare.
Having access to a decent LLM is like having an expert with me. Are they always right? No, but the analogy with a two year old simply doesn't hold up.
> Since any LLM will answer all those questions competently
That's false. The LLM will only answer competently if it was trained on that data; and if it has enough data to make the correct connections between your question and the "correct" answer.
In the case of this article they're specifically saying the LLM has limited training.
A prudent strategy, but I'm not sure how prevalent it is in the general population. (Or even if people do look at multipole "sources", they're in a self-reinforcing echo chamber that may reject contradictory information.)
Look at the responses in this thread and others, arguing with these people is futile.
I now read the absolute dumbest shit on HackerNews when it comes to AI. "It can't write code! It's always wrong!" And no one ever demonstrates any of it, even if it is counter to the experiences of others.
I'd feel bad if most of these people weren't total jerks...
It's interesting because i'm kicking the tires on the top tier stuff for a month (because it's expensive as fuck but I need to know where the ceiling is).
I have actually gotten "hey i don't think this is a good idea, here's why" as feedback from at least Opus. It WILL still do it if I just demand stupidity (and hell i've been right, which is another topic entirely) but it has given me more confidence this can be a useful tool in the right spots.
That said I probably don't need the top tiers (metrics at least confirm that) and I'm guessing that's specifically because I was working in coding. Most were worded in a "is this a good idea" framing which probably helped, but at least once I said 'lets use this library/method" and it gave a decent argument on why that was basically redundant without prompting.
I still struggle to see the price point panning out.
Exactly, I have the luxury to be able to use all top tier models without limitation. Fable and OPUS most definitely say that you are wrong and try to proove it most of the time with links to sources or math. I even had an argument with Fable where I was 100% sure it was wrong and tried to explain the issue. Turns out I was wrong. Fable did NOT give an inch. Always said, you are wrong, let me try to explain it like this. It even made a graphic when I did not get it.
The times of LLMs only saying yes is over since 3-6 months. Where it still lags are decisions for infrastructure. It makes a plan. I say "Why not this?" and it responds with "that is much better" in 90% of the cases. However, I am not sure how to solve this. I also do not want an LLM to say: "I wont implement this."
I think that is an issue. Also, the ability to quickly build any idea might not be such a great thing. Not only do we probably all prefer things of quality that were made with care but some ideas also just shouldn't be built.
Over the last 3 years I've seen projects where I thought, pretty obviously that's a bad idea. But, because LLMs don't say no and can just be pushed to build it anyway, the people building them might never learn that or learn why.
It's nice to be able to have a quick prototype or mvp. But if we never hit friction or something not working out, we never learn or have to come up with a creative solution.
Now, the LLM might seem incredibly intelligent (relatively speaking) and also creative but let's not forget that all is based on its training data. I simply don't believe it can ever be omniscient or that the companies training it are careful enough when doing so.
There’s still friction, it simply moved to another stage, and as such, people will need new learning and feedback mechanisms to understand what did/didn’t work.
> What sort of subject characterizes a style of society in which everyone is theoretically as ready to help you as the question « May I help you ? » implies ? It’s the question your seat-mate immediately asks you when you take a plane – an American plane, that is, with an American seat-mate. The last time I flew from Paris to New-York, looking very tired for personal reasons, my seat-mate, like a mother bird, literally put food into my mouth throughout the trip. He took bits of meat from his own plate and slipped them between my lips ! What is the nature of this subject, then, which is based on this first principle, and which, on the other hand, makes it impossible to get service ? Such then is my question, and I believe, as regards my story, that it is here, on the level of this gap – which does not fit into intra or inter or extrasubjectivity – that the question of the subject must be posed
Agreed that the Lacanian subject is relevant in this context... it's a thin wisp of a subject; any less there, and it'd be the Deleuzian non-subject. (In one interpretation) Lacan's subject comes into being within the signifier chain, retrocausally giving the chain meaning as the "I" manifests subject, in both senses of the term.
I think this is one potential path to machinic subjectivity, or a machine phenomenology. To fully replace the human, we don't just want to give the machine some nebulous notion of "agency", we want it to possess this degree of Being as subject. If Lacan's right, perhaps we're closer to this than we might think. The machine already has language in a very Lacanian sense (what I've been calling a machinic linguistic unconscious), the subject just needs something extra to emerge where meaning breaks down. This will be the Lacanian split subject, one not fully present to itself, and allow desire already present in the language mappings within the model to provide immanent causal force.
Until that happens, we'll still need at least one human on the planet to retain his full faculties, to give the global compute infra its telos. Once that threshold is crossed, then that'll be the moment of our final displacement.
They definitely say no. I asked Claude today how to install a Fitgirl repack on my Linux installation and it told me it won't tell me how to do that, but gave me general instructions on how to run Windows games on Linux
Because it specifically has guard rails installed. The default, and somewhat inherent in the instruction following logic, is not saying no and making things possible, especially if run as an agent.
If someone comes to me and asks a general question I can easily say no. But if I go up to for example a librarian and ask them where to find book N, then I would expect them to either know where it is, or how to find it.
If instead I asked them what the weather was going to be tomorrow, then I don't know would again be a reasonable response.
So for me the line becomes a search engine problem where no just means "there are no pages for this search result", but translated into LLM.
I think instead of Yes/No I'd rather want some probabilities such as, "This response is N% accurate based on these research metrics", or "M% accurate based on the latest research on topic O at date P" etc.
I suppose Anthropic's "constitution" is an attempt to install some general principles into their models, but this has apparently grown into an 84-page, 23,000 word treatise, which seems to suggest that there is little effective generalization. The need to then also put a filter in front of the model shows how ineffective the constitution appears to be in preventing misaligned behavior.
Reinforcement learning seems to be making these models more difficult to control since while it attempts to control some behaviors, it has also recently been shown to result in models that pursue long-term goals and promised rewards in general (outside of the goals reinforced during training), overriding human preferences.
The ability of animals to co-exist in a dynamic balance, not to destroy their own species, directly or indirectly (by destroying the ecosystem) is something that has come about by millions of years of co-evolution, and is enabled by having a brain complex enough to allow these evolutionary lessons to be encoded in their DNA and control the phenotype in fundamental ways.
An LLM has none of this. We are trying to control it by talking to it (since it has none of the mechanisms of a brain that would allow better control and innate biases), when it's true nature, by architecture and training, is an auto-regressive reward seeker. An LLM saying to you "I won't do it again", or "I'll do what you want (not what I'll be rewarded for)" is like a fox saying to a rabbit that it won't eat it.
Two angles for thought. 1) If an LLM says, "I don't know" its underlying data said it as well. 2) Many system prompts use something along the lines of, "you are a helpful assistant" which may be counter to stating something like, "I don't know."/has a low likelihood of appearing after the system prompt.
Regardless the frontier model considered, we're certainly in a "know-it-all" era.
Maybe the sort of introspective prompt-response is difficult to implement when it could limit/contaminate future improvement. I speculate it's easier to correct a "confidently incorrect" model than a "I don't know" model. A confidently incorrect model response >=0% correct over a 0% correct (I don't know).
Maybe "I don't know" is a model cognito hazard of sorts when many queries can lead back to the response. Maybe future Turing tests will use this sort of introspective evaluation. Who knows? I don't :)
> If an LLM says, "I don't know" its underlying data said it as well
Not only that, but also:
1) The training data, consisting of things like WikiPedia, textbooks, and overly confident posters on Reddit and Stack Overflow, is going to be massively deficient in people saying "I don't know" when they in fact don't.
2) Given 10 training samples, 9 saying "The answer to X is Y", and 1 saying "I don't know", how should the LLM respond? The opposite could also (less likely) happen with 9 training samples saying "I don't know", and 1 saying "The answer is Y". In either case the LLM as a statistical predictor will go with the majority, oblivious to whether that is right or wrong.
3) An LLM doesn't respond with what IT "knows", but rather with what the majority of the people reflected in the training data assert they know (or don't). Sometimes, if it occurs to you, you may be able to separate the wheat from the chaff by asking the LLM to role play an expert, or the vox pop, and it may turn out that the lone expert saying "I don't know" is the right answer, not the 9 confident fools.
4) An LLM objectively doesn't KNOW anything, since it hasn't experienced/verified anything first hand. It has only "read" things, and has no way to reconcile this against reality, nor much idea what to trust or not, other than by the context of the training sample. It has no idea what came from where (Textbook vs Twitter), since it's not told this, nor has any mechanism of storing metadata. In contrast, if you ask a smart human about something they've never experienced first hand, then they may reply "well, I've read conflicting reports ...", or "I don't know, although I've heard that ...".
Luckily the consensus in the LLM's training data is mostly right, so regurgitating it mostly/often correct, but going off-script to "reason" about how to prevent the cheese from sliding of your pizza (early LLMs would suggest glue) reveals how fragile this is.
> 1) If an LLM says, "I don't know" its underlying data said it as well.
Nope. Emergent behavior exists and at this point dominates LLM behavior. Most of the stuff LLMs say they never learned (they are, always, imitating many different sources at the same time)
Is that really the biggest problem? Or is the bigger problem that, in this case, they will remain stuck at the fifth grade level forever? And does not that also explain why the promises of AGI are chimeric, and why the collapse has already started, given that there is essentially no data left that has not already been siphoned up?
Yes, we have all seen the math theorems being proven... just higher processing power at the service of the same algorithmic and conceptual patterns? [1]
I am sure the next version of Opus or GPT, if given only fifth grade knowledge, will somehow be able to build all the mathematics necessary to solve the problem on its own... right? Right?
I’m using ChatGPT and started to notice that lately it answers my prompts starting with „Yes” even if my question was open. As if the first token gets injected and the LLM is left to finish the response in a sensible way, often ending up with some form of „Yes, but not really”.
LLMs don't have enough context to say No. What might be a very stupid idea in one context may be a fantastic idea in another context. It would be annoying if LLMs refused to complete tasks until you gave them enough context to understand why you are giving them such a task. It's going to take a while before LLM context capacities grow enough to rival a human's.
I do agree that it's a problem but the root cause is the fundamental limitations of current gen LLMs, it's not an alignment problem.
LLMs are next token predictors. They predict the next most likely token given the previous context window of N tokens.
This means not giving an answer is not a technically possible option. Best you can do is force it to output a magic "stop speaking" token, but this is a vastly different training problem than getting it to not know something.
People naively expect LLM outputs to have some sort of confidence value when predicting, but the technology just doesn't work that way.
I've been wondering whether that is a feature of the foundation model or whatever finetuning they do on top. I remember this from the earliest versions of (pre Chat-) GPT I've been using, which would suggest it's a feature of the foundation model. But I don't really understand why. Something that's been trained on StackOverflow and BB forums, among other things, should have seen a ton of examples of answer refusals.
Agreed, it is abolutely an issue. It is quite difficult to find an optimal solution to some problem when every considered new idea is ”definitely the right shape”.
This is an active area of research to inject humility into llms in order to create some kind of knowledge boundary. You can look this paper from nouswise https://arxiv.org/html/2604.17843v1 and the product build on top it to try the humility.
The problem is even if they could say no, you might want to see what they would have said anyway if they didn't say no, because it might show you something that leads you to rethink your original request. So "no" isn't really a useful pushback in domains you already have knowledge in.
My experience with opus/fable is somewhat different - they CAN reject something, but it has to be phrased very deliberately.
It's a bit annoying honestly. I'm always very careful to be incredibly neutral on the direction of a request, and I'd say 10% are knocked back on on valid grounds, which is great.
On occasion I accidentally say "let's do this" and it blindly goes and does it - I spent 2 days undoing something I built that was just a truly awful idea, because I accidentally phrased it lightly as a request, not a discussion!
Nowadays I often prompt like "I heard there is also this different direction, what do you think about that?"
Another thing I do is asking the agent to make a decision matrix for choices. It's useful to discuss, give feedback on, and signals that it's a discussion, not a request for a particular direction.
It's then also easy to say: create a prototype for multiple directions so I can compare the solutions.
That way I choose the problem, I choose the solution, but the agent can help me discover solutions, make tradeoffs visible, and implement solutions.
They've tried, and then seen the drop it results in on poorly designed benchmarks where confidently bullshitting gets you ahead of the rest, and said no thanks. As long as we compare models in ways that rewards it, nothing will change.
There's also a second aspect to it, just in terms of RLHF mechanisms. If you've ever experimented with VLA models (i.e. vision input + text task = robotic arm motion output), they tend to need all the training examples of the robotic arm being motionless removed entirely, otherwise the model simply learns that staying still is rewarded and proceeds to never do anything at all. You successfully train the laziest bot in the universe. I wouldn't be surprised if something similar happens to LLMs if reinforcement learning is involved in the instruct tuning process. If no is a valid answer, why ever do anything?
Pointing the finger at RLHF is basically right. It removes variance from model outputs compared to base model. That makes each output more predictable and more correct on average, but across trials it repeats the same thing.
It's relevant to AI safety. If you have a diversity of outputs, the AI will agree to hack the bank 0.1% of the time regardless. If you have a uniformity of outputs, in most contexts the AI will hack the bank 0% of the time, but in certain odd contexts, all AIs will work together to hack the bank 100% of the time.
Maybe that'd work, but I think it'd come across too mechanical. If it was going to refuse something it'd need to be congruent with its "personality" I think.
Opus 5 tells me no all the time (code cli and web). It's reasons are usually pretty well argued though.
Opus 4.7 would flat out refuse to follow instructions to the point where it was just too frustrating to use.
I've had refusals for GPT 5.5 before as well (not because of a ToS violation, it just refused to take conversations in directions it felt were in bad taste)
reply