Hacker Newsnew | past | comments | ask | show | jobs | submit | supern0va's commentslogin

Exactly. LLMs regress to the mean. If you want don't want average, you need to feed the context to push it away from the mean.

> For homes, it's a very different bar. You'd have a hard time convincing most American families to purchase anything with a >$1000 price tag.

There are plenty of white collar employees in cities paying $500+/month for someone to come clean their home. As one of those people, I could easily see paying low five figures for something like this if it actually worked well and could replace most household labor.


Most consumers don't think like that. Businesses do.

Part of the reason is upfront commitment is scary and risky for personal money, but not scary for a business.


Most consumers? Sure. Upper middle class consumers? They're the sort more likely to buy cars with cash if they're dissatisfied with the value of the financing offered.

As with most things, it'll cater to that demographic and then filter down.


>figure out what the percentage split between Opus/Fable/Sonnet is.

This may be misleading, since I suspect many are using a blend through sub-agents. I tend to bias for Fable to orchestrate and Opus for implementation via sub-agents.


You may be surprised to hear that the famously socialist paulg might also have a problem with this.

Can you provide a source where they said it was complete? I can't seem to find any evidence of this.

Hmm, you are right. The 11 days number came from the Bun blog, not the Anthropic blog. Anthropic has a history of making dishonest claims like this (for example their C compiler) so when I saw the news coverage I assumed this was just another one of those. My mistake.

Even then, can you point out where Jarred said this on the Bun blog? Because AFAICT, it seems pretty clear that he said that it took 11 days to get to all integration tests passing, but not to a full release candidate. In his blog post from a few weeks ago, he specifically calls outs:

>Bun v1.4 makes Bun faster, smaller, use less memory and gives the team incredibly powerful tools for systematically improving stability going forward: Rust's borrow checker, Miri (which runs for a growing chunk of code in CI), LeakSanitizer, and 24/7 coverage-guided fuzzing for parsers. There's still more to refactor, but things are off to a great start.

Perhaps you never perused beyond the HN headline?


And yet, that is in the conclusion, the section with the lowest chance of being read in any article. Much higher appears,

> I rewrote Bun in Rust using about 50 dynamic workflows in Claude Code run continuously over the course of 11 days.

Is this Schrodinger's rewrite?


Yep, I'm perfectly content to use the full lane if there's not another option. And that doesn't mean squeezing into the shoulder. No, that means riding left of center to prevent assholes from passing unsafely on my left.

The funniest thing is that tactical use of bike lanes, neighborhood greenways, etc, actually serves to funnel cyclists into specific protected routes such that the majority of other roads have fewer cyclists utilizing them. From my home in Seattle, I can ride several miles to work on roads that were already low traffic volume, but which are faster because cyclists get priority, and cars get diverted away. Before this existed, I just used the main roads that cars did, because it was faster.

Like, if you hate bicycles, you should love bicycle infrastructure. It's better for everyone.


You are assuming that with or without bike infrastructure, there is a fixed number of cyclists.

Many drivers either believe there are no cyclists; or that the number of cyclists should be reduced to zero, and they should drive instead, like normal people.

Also, everything is zero-sum, so any money spent on bike infrastructure is less money spent to widen the freeways or subsidize parking.


Can you expand on this? I can't imagine this is holding up an entire machine or materially impacting capacity. Isn't it common to have a tiered cache and to ultimately evict (or move to another tier) if there's an active inference request and no available capacity elsewhere?

Yes, however if a tool like this commoditizes a way to stay "higher" in the cache hierarchy to avoid getting dropped, then for the same total capacity, Anthropic have to become stricter about retention.

It's unfortunate that there isn't some better way to signal that there's a high chance that a turn is coming (ie, due to active sub-agents). If CC could send a ping saying "I don't need a turn, but please keep this cached at some tier because there will be a turn soon." that would probably help with cache efficiency. Instead, it seems like there's just the shotgun one-hour solution.

>I find it easy to envision a world, maybe 50 years from now, in which the very concept of "truth in advertising" is viewed as a lost, idyllic fantasy. Something people are nostalgic for, but feel powerless to regain.

Alternatively, perhaps we'll see models fine-tuned or steer-able towards accuracy that customers can themselves use to get a more honest view on what the product would look like in person.

The funny thing about these tools is that they can go either direction, but it sure seems like there's the potential for it empower individuals and shift the balance. Clothing sales shifting online has given sellers the advantage/ease to deceive without much customers can do other than hope the reviews aren't manipulated (they are) even before AI. Maybe this can turn things around as people start to shop with personal agents.


> Alternatively, perhaps we'll see models fine-tuned or steer-able towards accuracy that customers can themselves use to get a more honest view on what the product would look like in person.

we live in a world where the norm is to use filters on pictures of your own face to lie about how you look, imagine thinking than people actually care about truthness


>imagine thinking than people actually care about truthness

I think they actually care about "truthness" when it comes to how others will perceive them. You might filter your face for an online photo.

However, you're not going to want to look at yourself through your selfie camera with a filter that hides your imperfections when you're checking to see if something is stuck in your teeth on a date, don't you think?

There's nuance. People will fudge the truth to be perceived better, but don't want to be lied to when gathering data about how to be perceived better.


> However, you're not going to want to look at yourself through your selfie camera with a filter that hides your imperfections when you're checking to see if something is stuck in your teeth on a date, don't you think?

No, but you definitely use it in all the pictures you put on dating apps


Sure, but just because I may want to filter the photos I post, doesn't mean I want my mirror to show me a distorted view of reality. Isn't this use case more like the mirror, and less like the dating app?

That's a fun thought. A gradual division between "institutional models" trained to have capabilities that solve for the needs of corporations, vs. "personal models" trained to have capabilities that solve for the needs of individuals.

I wonder what kind of capabilities those might be?


People are scared about the personal impact from AI, then backfill in justifications without even realizing they're doing it.

If the equivalent numbers for electricity and water usage were being being used for streaming video, I seriously doubt people would be demanding no more Netflix data centers. The news story would immediately die.


Nobody wants their electric rates to go up, the local water utility to have to raise rates to build a bigger plant, all in exchange for also losing good white collar jobs. That’s currently what AI data centre builders are selling.


Not exactly. There are some data centers being built in places that don't have the power and water to support them, and obviously it's rational for the locals to oppose them.

But I live in a place where we have plenty of water and relatively cheap power (lots of renewables). There's not much risk to data center construction, but people are opposing it here, too. Because for most people, it's not actually about that.


An obvious question is if the cheap power is going to stay cheap after a large power-user comes in who has a proven track record of trying to make everything cheap for themselves with no regard for anyone else.

Or another question to ask is - how does this data centre benefit the people who live there? If it doesn't, there's no reason they should want one to be built. Rubbish tips are necessary. I still don't want one built next to my house and would fight such a thing tooth and nail.


Here's a better question: as we've had nearly as large of a data center build-out happen between 2005 and 2020 for non-AI purposes, with similarly high electricity and water demands...where has the concern been? Why is it only in the last 2-3 years that people are suddenly up in arms, as a very specific application is being deployed?


The amount of data-centre construction is far more than it was 2-3 years ago.

There seems to be an alarmingly high amount of questionable data centre construction going on, such as projects being built in places with no access to power with an assumption they can somehow force the utility to provide it later. These buildouts seem to be being done for financial reasons (they are not Meta, Amazon, Azure, etc. facilities) with the hope to lease them out or sell them half-completed in the future. People rightfully don't want that kind of thing in their back yard.

To give a feel for the scale involved, this one (the new Amazon east DC) in my podunk area of the state is 250 MW (the existing us-east-2 in Columbus is 200 MW, although I'm not clear if this will be a new region or is just an additional availability zone). But that's small potatoes compared to the speculative project in Piketon, which amongst more absurd things is planned to be:

- 10 gigawatts (equal to 50% of current power consumption statewide) - "Modular" nuclear reactors built on site - 35,000 construction workers needed to build it (in a county with a total population less than that) - $30-$40 billion for the data centre, plus another $33 billion to build the 9 gigawatt natural gas electric plant - Meta agreeing to build an additional 1.2 GW nuclear plant on site - OpenAI in negotiations to lease the facility

This is a really big project, of the scale of "nothing like this has ever been done before". Nobody has ever built that much power generation at a single site before, nor has a datacentre this large ever been constructed. There is a very real risk of the project getting halfway done and then being unable to be completed. The prospect of a state literally doubling its electric generation is a bit ambitious, too (doing such means basically a complete revamp of the power distribution grid, or else some very novel designs to only use the power locally). For example, the normal type of shutoffs data centres have to prevent eg an incoming have are unacceptable in this situation because the grid cannot cope with 20 GW of demand suddenly disappearing.


I recommend actually reading their recommendation, because they get into the weeds about precisely how the US and China could address this in a trustless/auditable way. The TL;DR is that basically all of the relevant compute can be tracked.

Edit: Also, definitely not a Chinese op. The authors are prominent Americans, and are the folks responsible for the AI 2027 forecast that has pretty accurately predicted the current state of affairs today: https://ai-2027.com/


Pretty accurately what? So now OpenBrain has an Agent-1 that makes their algorithmic progress 50% faster than other companies? If it was 50% more CVEs, that would be something, but I doubt any meaningful self improvement is achieved which the competitors are slower due that, which is the core of the prediction for 2026.


I wonder. There has been some headway in getting decentralized training runs to work.


Yeah, Cognition's work is interesting in that regard, but it still doesn't obviate the need for the chips--it just enables training on them when they're spread across multiple data centers.

The Plan A proposal estimates that the ownership of ~96% of AI relevant compute hardware can have its ownership traced, since the companies selling are very few.


I doubt we can even track all of the chip production capacity.


Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: