Unlike cloud infra in general which offers things like automatic backups, regional redundancy, and effectively unlimited scalability, it seems like the value proposition of cloud LLM gets ever shakier.
* Many businesses don't need frontier level intelligence anyway.
* It's completely stateless. If your local LLM machine catches fire? Nothing was lost. Buy another.
Centralized inference can easily increase batch size, leading to huge efficiency gains in the usual scenario where most users have just one or very few session. Using local resources efficiently requires some way to increase the batch size. I'm not sure if we are there yet.
I think we'd need something different to make that a reality today. It's likely that decentralized compute eventually wins out here as well. It usually does. But without architectural changes, it will take a couple of years before the current single-conversation flow is truly usable on consumer-grade hardware.
Jevons paradox: large purpose-fit data centers increase efficiency such that you can use AI in more places, and use more tokens for those tasks.
The future is not a single chat bot session of bs=1. The future is many agents performing many tasks in parallel for a single user. Large GPU clusters will always have the edge in efficiency.
Jevon’s paradox suggests that total datacenter resource consumption will increase. It doesn’t say that people will choose datacenters over their own personal hardware when the latter is sufficient.
Personal hardware is only sufficient today for some tasks. As data center power efficiency, and large sparse MoE task efficiency increase, personal computing will continue to lose out.
> The future is many agents performing many tasks in parallel for a single user. Large GPU clusters will always have the edge in efficiency.
Agentic AI has pluses and minuses for cloud efficiency. The plus is that usage could be very bursty, but the minus is that agents will more fully utilize a local system. The main disadvantage of local AI is that you would be paying a large amount for a system mostly doing nothing. If it's constantly working on different projects and integrating that data, you get use out of every penny that you spent. Every GPU you added would instantly make the thing smarter.
What's more, your local AI could offload an agent to the cloud if it needed to. It could do this rationally, based on your personal desire for privacy.
Also, an entire family can share one big box that can run a good model. Not an entire household, but an entire family no matter where they live in relation to the box. They could all access it over the internet, or even simply over the telephone. When it becomes another family member/servant, a price tag in the low $10s of thousands seems a lot more reasonable - and this is accessible now.
I could absolutely see local AI taking the place of voicemail/call screening completely. Call me and you get my AI, who will route the call to me if approved, or even choose to service the call itself. If a friend of mine calls who doesn't have $20K to spend on a great AI rig, they would certainly have permission to steal otherwise wasted cycles from mine. An answering machine isn't much different than an issue tracker, and some of those are 97% AIs having perfectly intelligible conversations with each other, and 3% humans being eagerly serviced by AI.
The old objection was "this is so hard to set up, nobody wants to run a server!" Now it can set itself up.
tldr; you don't have to rent from some data center. You can get high utilization out of a local box. This 1) puts a hard ceiling on what a data center can charge, and 2) they won't be able to compete on privacy, which locally can be complete.
Centralization without proper controls against monopolization becomes sloth and gluttony. If american labs were constrained like china, their models would benefit.
The abstract benefits are quickly outstripped. The same way adding more highway lanes never improves gridlock. Its inducement.
I mean, I might be missing something, but isn't part of the idea with cloud infra that you can scale down as well?
Large orgs with significant demand might go out and buy local LLM hardware, but most businesses probably don't want to bother dropping $2k on a box with a beefy GPU and would rather just pay the lowest subscription tier so their employees can occasionally make queries.
Plus, you know, the whole economies of scale thing. Local LLM has a lot of privacy and independence benefits, but I'm not really seeing the world where it becomes more energy- or cost-efficient to buy your own hardware (and use it 1% of the time) versus sharing a giant machine, or even the same machine, in a datacenter (where it has a much higher utilization factor).
I'm pretty sure a small-to-medium org has plenty of things that could be queued/scheduled to run when there is downtime. It requires some planning and thought though and I don't think most orgs are there yet.
Yeah, but again, why would you? When it comes to open weight models, there's a good amount of competition, so you already get a really good price, without any optimization or anything.
Seriously, you can do years of Deepseek inference for the hardware to run just 1 or 2 requests against a slower, dumbed down model on your own hardware.
It makes no sense to buy hardware right now, when the price is completely disconnected from any material reality. It's much better to use the cloud providers VC funding by using their cheap as F offering. Either AI becomes less useful, or hardware costs come down. Either way, you'll be in a better position in 3 years than you are today.
All white collar work will be done by AI soon, there won’t be scaling down, just scaling up. I do food trucks and festivals and I’ve got AI doing so much of my non-meatspace work now that I’m buying $10-$20 a day in tokens. At that rate, hardware starts to look cheap, and I have less of a use case for it than most white collar workers.
There. I have the business plan with two seats and I use them both and blow through it pretty fast. I think it’s because much of what I have it do involves using a browser. For instance I have it pull various permits from cities and there’s no API for that.
Computer use will blow through tokens because it's doing image capture for everything.
You may have better and more reproducible results using browser controls that aren't image based, or writing tools that completely sidestep browser use.
I do where possible. I build MCP servers into custom tools. But a lot of it is stuff like filling out permit applications on government websites, or reading festival websites and applying, etc.
I will say one thing that frustrates me is the opaqueness of the billing model. I basically just have to pay a random amount. I guessed that it navigating the web is relatively pricey from watching my usage as it does stuff. And the whole thing is worth 10x what I pay for it anyway, it just would be nice if the pricing were somehow more transparent, even in hindsight. It can explain to me what is thinking as it does stuff, it could also explain the tokens.
Yes. Bureaucracy for sure when I throw festivals. ChatGPT pulls permits for me. (I think that uses a lot of tokens because of the browser control?) Manages the admin side along with some tools I built in Lovable via mcp servers.
Marketing definitely. My food truck side is relatively high volume and it manages my kitchen and warehouse side, basically generating all of the instructions my employees follow, managing and updating my PoSes, creating signage assets for specials, etc.
Reels/posts production and managing ad spend.
Here’s a fun one. I switched payroll providers after several years and suddenly my unemployment insurance rate went from 0.8% to 12.75% which is borderline debilitating to me. I knew something was wrong but not what and I work a lot of hours and calling the state takes forever and is usually unhelpful.
ChatGPT figured out that it was a penalty rate and dug in for me. Turns out because I’m seasonal and have no payroll for one quarter of every year, Gusto did not file a quarterly wage report, so even though I owed nothing I was delinquent. Gotta love government, it’s the only place where you can be delinquent for $0.
ChatGPT filed the report and requested a retroactive re-rate, which they granted. I’m sure I would have figured this out eventually but it would have taken hours, or $5 in tokens.
Would you say ChatGPT does everything you need today or do you still see some gaps? i.e. things you think it should be able to do but currently doesn’t, or things that still take too much effort on your part to setup chatgpt to do it.
Also, I have certainly gone down some unproductive rabbit holes figuring out what to do with it. Time spent on it is an investment, just like automating anything. You spend hours upfront to save them on an ongoing basis.
But, I think I’m on the black on it already after just a few months of heavy use. For instance I just tell it to book my dumpsters, bathrooms, sanitation crew, and security for X event. It goes and pulls event details, looks through my email to see who I get those from, and emails them relevant details, with no prompting.
Another good example: we launched a really fancy hot cocoa concept last fall that was a hit and I wanted to try to go to all of the local pumpkin patches in October and Christmas tree farms after Thanksgiving to serve when they have big crowds.
I asked it to contact all of the ones in my area and it sent out 70 emails and I booked several spots. It made me a nice database so I can see who followed up and who I need to reach out to again, etc. and it just gets those from my inbox.
Hours of my time saved with simple prompts. So while some things take awhile to pay off, some are instant.
Huge gaps. I expect the tooling to get better for non-programmers. Codex and Cowork are great, but you still feel like you’re trying to hammer the square peg through the circle hole often when using it for non-programming tasks.
The AI is good enough to do a lot of tasks but the tooling just isn’t caught up to it yet.
I’d say it’s freed up ten hours a week of my time. And that’ll only improve.
Ah. Yeah. I can’t wait until the tooling for non-programming tasks catches up to the tooling for programming. I’m sure it will happen. Most of the software world started off aimed at developers (because they made it) and then found its way to us muggles.
$2k? LOL. That's not even the GPU budget, these days.
The economies are currently out of whack because of underproduction of components and memory, so LLM providers have a few years of runway to entrench. Plus, the whole capex Vs opex thing that helped AWS will help here too, for sure.
At some point, though, things will change. More production will come online, and providers will have to end the current speculative subsidizing and jack up prices.
It's a bit like the dot-com era: the initial rush to land-grab web portals and e-commerce sites eventually died, once enough skills and infrastructure came online, and the bubble burst.
- Hyperscaling “we are going to serve billions of people in our applications”, which is becoming increasing unlikely as regional tech companies become more dominant than than the global one (this one is as much about geopolitics as technology)
- Operations is hard, in which case non-frontier models should be increasingly capable. Devops for small-ish deployment is one of the few cases where it is hard to clam you need deep expertise and AI can’t do it. Previously, the claim is that you need people specialized in ops, which is expensive. Now…
My prediction is that not just cloud LLM, but cloud business general will have to change. Not yet in the next 5 years, but probably 8-20 years-ish
I don’t disagree but saying “cloud will change in the next 8-20y-ish” is a bit of a non-argument, you’re not really stating any thesis to speak of; Change is a given over that time frame.
I’m not so sure I agree, given the overwhelming concentration of capital and regulatory capture they have, I just don’t see them going anywhere; becoming more niche rather than even more of a standard is “going away” to a certain extent as far as I can see.
Wouldn't that cause an awful lot of confusion in practice? Instead of having one being a superset of the other, you now have two almost identical syntaxes, but probably not quite.
LLMs have enough problems distinguishing between major versions between packages, e.g. where all the imports were moved around and renamed.
are you referring to this section?
> Verification is different from generation: Models scoring 98 can solve problems but can't always explain why their approach works at the level a human mathematician would.
It doesn't say it can't explain why with any accuracy, it just says it can't *always* explain at the level of a mathematician, but most of us don't have such a mathematician at our beck and call to answer our questions anyways (thinking of the perspective of a self-learner outside of formal education)
It’s not just math. Anecdotally, LLMs struggle the same way with software engineering where the code they write is correct (compiles and passes tests), but reasoning is wrong often enough to eliminate most trust in these models’ ability to explain codebases or even features they themselves produce. It’s not about the _level_ of the supposed intelligence where a model struggles to summarize things succinctly or simply enough (responding to the “pocket mathematician” comment) or can’t grasp certain concepts at all (if so, how tf is it able to apply them?). It’s that by their design LLMs have no concept of truth and no concept of causality. They guess with every single inference and it’s very hard as a user to understand which guesses are more or less certain, since, you know, confidence ratings aren’t part of these models’ design either.
I get wrist pain as soon as I use anything taller than Apple’s flat Magic Keyboard and touchpad. I’ve used this combo for like a decade now and never experience fatigue. Once, I bought a fancy mechanical keyboard and I returned it within a few weeks as it immediately caused me wrist pain.
You can scoff at low travel and membranes all you want, but less movement = less fatigue, so I think they’re amazing.
Never thought about the climbing angle, but I’ll chime in and say that after I started climbing I’ve definitely not experienced more fatigue.
Everyone’s anatomy is different: wrist rests give me significant wrist pain and stiffness in minutes. I either need my wrists a touch lower than the keyboard (I’m okay resting on the desk if the keyboard is thin) or I must hold them elevated without resting them at all (if the keyboard thicker) which always makes me feel like an old-fashioned typist :)
That's interesting! I'm curious to know: is there a bend in your wrist while typing (either with or without the wrist rest)? Also, are your arms parallel to the floor? Also also, do you rest your elbows on anything?
The typical position seems to be that my elbows are just off the edge of the table/desk, my forearms just distal to my elbows rest on the edge of the table, and my wrists and hands are suspended slightly to be able to move freely to type. My forearms and wrists like to be straight, and I use keyboards without extending their rear legs so that they're as flat as possible. There's a slight angle between the table/floor and my forearms, sloping up from the edge of the table to a point just above the keyboard.
Thanks for replying! The ergonomics recommend not resting your forearms on the edge of the table, but as you say, everybody's anatomy is different; if it ain't broke, don't fix it.
In case anyone else comes into this comment section looking for ideas: the thing that helped me the most was adding soft pads to my chair armrests - I found that resting my elbows on the hard surface of the default armrests resulted in forearm pain.
I think shoulder height and arm length relative to desk height must play an important role in this.
I went the opposite direction- mechanical keyboard, large trackball- and my pain cleared up. The lack of feedback of the flat keyboard was too maddening to get accustomed to for me.
My carpal tunnel syndrome symptoms or RSI basically went away once I started using a macbook as my primary computer. It definitely doesn't hurt when the trackpad is right below my thumbs.
This reads like a nervous breakdown at the thought of having to buy a ticket from A to B. It’s really not that complicated in most countries with functional public transport.
> It’s really not that complicated in most countries with functional public transport.
They are not wrong, just look at Germany before the Deutschlandticket, but that one is useless for tourists so they still are stuck in the Holy Roman Empire of Public Transport Systems [1].
Working with ESL people makes it clear that grammar isn't so important. One of my friends would say, instead of "go shopping", would say "make some shoppings". Instead of "it's time to go", he'd say "time for go". All words that started with 's' were pronounced 'ehs', as in "gas ehstation".
The other people at work soon picked those lines up, and we used them regularly amongst ourselves. I still use them decades later!
This is grammar from their native language. They dont know the English grammar, so they just assume it's the same as the one they understand. You just understand them via the context.
Older Europeans used to say 'AND NOW WE MAKE PARTY!"
But your friend was applying consistent rules of grammar from his mother tongue. Having learned some vocabulary, he would mentally form the phrases in his head, and then literally translate out the words and tenses.
Grammar was so important to him that he was adhering to the grammatical rules he knew. It would be interesting to know how he would perform with writing or typing instead of contemporaneous conversation.
"go" in particular is an irregular verb, and also one of the most versatile words in the English language. Having studied how to conjugate "go" in other languages, I do not envy anyone trying to conjugate or just use it in an English sentence!
Tbh, that's how I learned English. Rudimentary grammar, knew a few words, and then just read with dictionary until I didn't need a dictionary. Also watched a lot of shows until I didn't need subtitles.
Languages are weird like that – if you use them enough, your brain just figures some things out. After all, natural languages are easy enough for toddlers to learn.
When the abstraction detracts more value than it provides.
I think this is an interesting direction to explore. Plotly is fine. So is D3 and plenty of other solutions, but neither represent the ultimate pinnacle of charting.
But your choice of funny variable names is not a craft, and you’re not a craftsman if you think it matters. What matters is extensibility, maintainability and value delivered.
I care how the food tastes, not that the chef has a really cool Japanese knife and is really fast at cutting onions.
I think you miss the point, but using your foodie reference: would you rather go to a couple of Michelin restaurants -- where each piece brings a reflection of the chef, the geographic area, and what quality ingredients were available on that day -- or would you rather only eat at McDonald's for an experience that is extremely consistent across days, seasons, and continents?
Now imagine a chef who has a nice little restaurant but is now being sent a couple of "quality assurance" guys from McDonald's who tell him that his choice of potato variety for chips does not exactly conform to the "standards" defined at the mothership.
A chef in a michelin restaurant, unless he/she is the sole employee (very rare) is likely in charge of a team of line cooks and while they may get some time and place to experiment you better believe it will not be when preparing the chefs carefully curated menu and recipes for paying customers who have been on a reservation waitlist for months. Bob decided to cook the broccoli with a blowtorch and substituted some ingredients in table 9s dinner. There is a fire in the kitchen and the NYT food critic is going into anaphylactic shock but Bob at least feels like an artist so we will let it slide.
The one day, your chef leaves, the food goes to shit and productivity/output/morale drops. When the new chef arrives, then everyone has to relearn everything.
I too want to be an artist, but I work in an economy that requires me to be an artist mostly in my off hours.
And honestly, this is such a weird hill to die on. Lint rules when automated are basically pure win in most scenarios, and for your example, just exclude the rule for this block.
We want maintainable code we can quickly onboard in, and be able to modify and extend. Your artistry is getting in the way of other people getting stuff done.
The hidden implication in your extension of the foodie analogy is that consistent food must have worse quality.
That's not true, and if anything in software having consistent software delivery makes it more likely to achieve Michelin-graded level of quality, not less.
This reminds me of the Go Grandmaster speaking out after losing to AlphaGo, that the model has no sense of "aesthetic play", as long as it would lead to a win within the rules.
Yes but it was so shocking because it was such an inhuman, "unaesthetic" play. It was considered to be "creative" in that no human would have thought of making the move, so it can't simply be copying human play.
Indeed. To get "aesthetic" play you would have to mimic real human play for an interesting reason: the space of possible Go configurations and games is so mind-bogglingly vast that all of human history has only ever seen a infinitesimal sample. So any sense of aesthetics is just a consequence of chance and memetics. AlphaZero is effectively exploring a new branch of Go history so its aesthetics are completely different. Although maybe that just shows that the idea of aesthetics here isn't very meaningful.
What properties does this have that somehow allows one to essentially predict the market or a proxy thereof, when seemingly nothing else can? Or did I get that part wrong?
It might take you from nothing to technically—a-market-fit, but from there to actually-a-market-fit?
My point is, there’s probably 1000s of companies doing the same things, following every playbook for success (do this, measure that, brand this, …), only for a single one of them to become a viral market hit, and the rest fade into obscurity.
* Many businesses don't need frontier level intelligence anyway.
* It's completely stateless. If your local LLM machine catches fire? Nothing was lost. Buy another.
reply