Hacker Newsnew | past | comments | ask | show | jobs | submit | anthonypasq's commentslogin

almost of their business is hosting Sol ultra fast or whatever for OpenAI to use internally

I think this is because the primary use case of gemini is google search overview and the gemini app. thats probably 98% of gemini tokens. they didnt predict how important agentic coding would become.

wow thats actually pretty sick, ty, I could use this in the app im building

the 3.5 pro pretrain was a complete disaster, they shelved it and are now working on gemini 4.

3.0 flash -> 3.8 flash is all post training which is pretty impressive.


Do labs come back from disasters like GDM’s 3.5 pretrain? I am thinking of Meta’s Llama 4. Meta is just now starting to be taken seriously again but they are definitely not at the frontier. And when I say “come back” I mean have an Opus 4.5 moment, which was really mind blowing for me at the time. Fable was a similar leap, just not as big.

Unless the company is going under, why not? Let's say Google releases Gemini Pro 4 tomorrow, and it's better than Fable and Sol; lots of people would switch over to it.

AI models are almost completely interchangeable, so the best/cheapest/fastest whatever will always have a market.


I agree we’d switch to it. I guess what I’m doubting is if a company can recover from that sort of stumble in the first place. And they might not want to either. They might think there’s more value somewhere else besides trying to get back to the absolute performance and capability frontier. Smaller models targeted to specific domains that large models would be too inefficient at no matter how large they get or how clever you are at distillation, for example.

OpenAI had such a disaster themselves before, GPT-4, so they replaced it with 4o.

gpt4.5 was also one such disaster for them iirc

it seems to me that OpenAI is the only actual lab that truly understands reasoning. they have the best reasoning efficiency, they get pretty uniform improvements with more reasoning compared to other labs. (theres been plenty of graphs where models do worse with more reasoning), and i suspect their models are a lot smaller than we think.

i think the next gen of openAI models are going to be quite insane tbh.


api pricing has 80+% margin.

Why would payments be low margin? You're making a small amount on each transaction, but to you thats essentially all profit no?

youve spent 100x the time arguing about the point as it would have taken you to read 2 sentences further.

for someone who goes on about how busy and valuable your time is, you are making some interesting decisions


I suppose it's a very human behavior to speak out against the communicated preferences of others.

Maybe I'm wasting your time now.


cursor still uses embeddings and theyve found it works better than just grep

https://cursor.com/blog/semsearch


they don't use it anymore which adds to my point that people tried it and largely gave up

source? nothing im seeing on the internet agrees with you

https://cursor.com/data-use

their data use policy from july 2026 explicitly mentions embeddings


Continued hardware improvements really make it hard for me to believe token prices will not continue to plummet.

This may just be a classic case of Jevons paradox: https://en.wikipedia.org/wiki/Jevons_paradox

In short, better hardware will drive down token cost in the near-term, but will drive up the demand for tokens as it gets cheap enough for other sectors to start to use it heavily.

It comes from steam engines where economists originally thought that coal demand would plummet with more efficient engines, but it actually just meant that we found more uses for steam engines.


This is exactly what I see happening now.

Codex keeps doing these usage resets. What do I do? Burn even more tokens than ever before. I know I'm not the only one.


Is this a normal thing now? I remember seeing this talked about as a surprising thing, but now I'm seeing posts about this as if it's normal.

(I switched to using local models as usage limits, api instability and the concept of paying per token stresses me out)


If we are applying Jevons paradox to this then the unit being consumed is not tokens but the inputs for token production - power, capex, something else. To draw an analogy to the steam engine, coal:electricity::mechanical-work:tokens. Jevons paradox does not talk about mechanical work becoming cheaper in the short term setting up a sort of rubber band of demand creating spiking prices for mechanical work. Compared to the renaissance, mechanical work was much cheaper throughout the industrial revolution and remains cheaper to this day. We can still definitely say that the easier it is to produce tokens, the cheaper they will be.

All Jevon’s paradox says is that as a resource becomes cheaper total consumption of that resource increases. It applies equally well to the inputs of token production as it does to the tokens themselves. The former would describe the effect the sellers into AI companies see (energy, GPU chips, RAM etc - if they lower their prices they’ll have more overall consumption) while the latter describes what the AI companies see with their customers (if they lower token prices consumers will use more tokens overall).

Jevons paradox states it might increase. It is not an ironclad law and there are many many cases where increasing efficiency wrt. a certain resource really will decrease the total consumption of that resource. Yes tokens are an input themselves, but this thread is discussing hardware that is more efficient at generating tokens. To increase the efficiency by which tokens are converted into some other product would require innovation in some other area - harnesses, the models themselves, better skill from the prompters, etc.

I think you're reducing a very complex thing (the global economy) into a very simplistic model (Jevons' paradox) and thinking both are the same thing. This has no predictive power or rigor. You're just wishing things would happen as they did before, without considering that conditions and situations change significantly, and instead of Jevon's paradox, we look back at today 50 years from now and talk about Jensen's paradox.

This doesn't mean the concept is BS, but one single concept cannot explain away everything in such a system.


That's when demand is higher than capacity. Now imagine places like Gigalab and Chinese labs are online and able to produce significant percentage of chips. That could cause real surge in prices.

the total cost spent on tokens may go up, but i just cant imagine per token costs going up

Depends on compute capacity. If we become supply constrained on tokens, then prices will necessarily go up.

no they dont because inference stacks are getting more efficient and models are getting more intelligent per parameter.

I would posit there’s no way in hell they’re getting sufficiently cheaper on a short enough time frame vs how much demand is sky rocketing. AI companies are seeing quarterly doubling of revenue if not more.

It is incredibly cheap now. What sectors are you thinking of?

[flagged]


There is something counter-intuitive about the idea that making an engine that accomplishes the same amount of work with half the fuel will result in MORE fuel usage overall. You might expect it to be the same, or decline slightly, but the paradoxical element is that overall consumption goes up.

And you can say of course, it's so obvious, how could a dumdum not see that! But then there are lots of examples of things where increased efficiency results in less usage overall, because demand is inelastic, etc. Jevon's paradox doesn't apply to everything.

I don't think we know yet what is going to happen as software development gets much cheaper. If in ten years we can produce software 1000x more cost effectively, will we need fewer software engineers, the same, or more? Guess we'll see!


> I don't think we know yet what is going to happen as software development gets much cheaper. If in ten years we can produce software 1000x more cost effectively, will we need fewer software engineers, the same, or more? Guess we'll see!

Adding onto it, I feel as if this relates to some points regarding predictions of future in general. It is easier for us to look from the future to the past and think that it must be very obvious (as you also mention) but its also very counter-intuitive at the same time and there are just so so much nuance about basically any situation within it that its hard to really capture it all, and even then, be prepared for surprises and counter-intuitiveness.

I really like the Peter Drucker quote about it.

“The only thing we know about the future is that it will surprise us.” — Peter Drucker

and, “The future is fundamentally different from the past.” — Frank Knight, Risk, Uncertainty and Profit (1921)


There is just so much downward pressure on token price, from every direction. We would need a completely new understanding of economics to explain why the price shouldn’t go down. Or market collusion/regulatory manipulation.

The demand for them is growing _per person_, not just across the wider economy, if tokens cost half as much but you want to use 3 times as much you're going to have to pay more.

Maybe 1000s of tokens per second unlocks realtime robotic decision making, and now every robot needs to continuously stream tokens to and from the cloud to operate. That could 1000x demand overnight, just to speculate :)

I would very much like it if anything that moves with appreciable mass is governed locally just in case the link drops and/or latency suddenly goes up. Motion is very unforgiving and accidents will happen if that's not taken into account.

Seems unsafe to make locomotive decisions remotely

I think you just found what we will see in the S-1 prospectus of OpenAI

Think about the agents buying computers for their agents. /s

The price has been going down for ages, its not clear what you are pointing at

Pointing at the nay sayers who say tokens are heavily subsidized and it’s all going to come crashing down soon, surely any moment now

I mean, it will obviously crash at some point. With so much pressure on token price to go down that means way less opportunity for margin for AI providers. OpenAI is in a pretty bad situation

What does this have to do with margins? It can remain the same once prices go down

At the price going down? And that it will continue to go down, even if the hardware improvements stop. Not sure what isn’t clear

Token prices coming down means nothing if the models keep wasting them

> continue to plummet.

Continue what? The cost per output token has kept going up for the past three years across the board, as thinking models keep leaning more on test-time scaling.

The quality of the said output tokens obviously increased, and arguably increased more than their price, but the price still went up. Or, on the flip side, the price of combined tokens went down (a bit, it did not "plummet" at all though) but so did the average token quality if you count thinking tokens.


With the corollary that old hardware valuations will plummet with them.

Although given we have marginal pricing we need to push through to those lower prices in the face of increasing demand, so timing of this is uncertain and the key to the AI financial markets


this is a story about a proprietary accelerator being built/designed by a token provider. and you think they're going to return the efficiency gains to the customer instead of capture the value for themselves? interesting take.

OpenAI just dropped the price of Luna by 80% and Sol by 20-30%

and amazon shipping used to be free without prime, and uber used to be cheaper than taxis, and airbnb used to be cheaper than hotels.

you really don't get it?


almost every pure tech commodity has gone down in price

- gpus

- retail computers

- laptops

- ~gpu~ appliances like washing machines

- cloud computing

i think you don't get how economy usually works in tech


I'm especially enjoying how RAM and SSDs are going down in price.

GPUs and memory have gone up in price. It's more expensive to buy a 1-2 year old video card than it was at launch, sometimes by a shockingly large factor. Laptop vendors have recently shipped flagship models with less memory than the previous model, because they can't match price expectations for a laptop.

I was checking laptops today for an upgrade from the model I bought back in 2019 and it's not gonna happen from how cheap they are.

i figured out why this comment is so confusing: this is actually a message from the past, around 2020. either that or simianwords is a time traveler that arrived today and hasn't read the news yet.

listing gpu's here is crazy considering the current prices

GPUs and laptops and memory and storage are all crazy expensive

We should be mindful of the context that many of these providers VERY likely have been selling their subscriptions at a substantial loss

So as much as i agree “more profits to stakeholders screw the customer”, i think its more of an emergency to get to profitability before the music stops.


> We should be mindful of the context that many of these providers VERY likely have been selling their subscriptions at a substantial loss.

what makes you think this?


Because everyone keeps saying this so it must be true. Real "it is known" kind of vibe with these statements.

He's a subscription truther. There's loads of them. OpenAI's profit increases with each subscription that is cancelled. Pretty soon they'll have more profit than God.

Yes, I can bet on this happening. If anything, this is a net gain for consumers as it is a competitive market.

go ahead and bet: alibaba is a publicly traded company

Hopefully this also means billionaires can stop trying to drop data centers into residential neighborhoods with zero noise control and polluting on-site generators, signing local politicians on with NDAs, calling for eminent domain to seize homes to build power lines to data centers, etc. etc. etc. Not to mention the water use controversy.

Token prices plummeting is probably a good thing, but not without the regulatory backstops that prevent these effectively industrial facilities from being operated with no regard for the externalities they impose on people who live near them.


>polluting on-site generators

how much pollution do you believe modern gas-turbine engines to produce?

>Not to mention the water use controversy.

what percentage of US water usage do you believe is by AI data centers?


Reducing everything to national aggregates provides no insight into the strong negative externalities, imposed on the immediate surrounding communities, of unregulated industrial facilities. That's literally the reason we have zoning laws in the first place,

Nah, Jevon’s Paradox says that cheaper tokens will mean increased overall energy consumption.

If we can’t even build data centers, the least disruptive industrial use possible, there’s no hope to reindustrialize the US or anywhere outside of China.


We already had plenty of data centers in the US before the AI boom that weren't severely harmful to their neighbors. Cutting red tape is not the same as eliminating meaningful regulation. There are plenty of old industrial sites that could be repurposed as data centers. It turns out it's cheaper to bribe some small town government to give you a tax cut and discounted electricity and water rate.

Counter point: Many AWS services barely decreased their prices (if at all) in the past decade despite advancement in hardware

Yeah but is it really even as good as Rubin? Seems just competitive.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: