LLM inference is a two-phase process. The first phase is prompt processing aka prefill. It's compute-heavy but requires relatively low memory bandwidth. The second phase is token generation aka decode, which doesn't require much in the way of FLOPs but wants as much memory bandwidth as possible.
This announcement is for a system to do the first phase on Helios and the second phase on WSE.
None of these are problems if you put data centers far away from population. No one does that right now because it's a lot more expensive to set up and maintain, because you have to set up all the infrastructure from scratch. You know what's 10000x more expensive than that? Setting up the infrastructure in space.
it just seems off-topic. we see the slot machine analogy a lot but i dont see how it has anything to do with this long, thoughtful essay about pricing power
Pricing in the article is the difference between a $.25 slot machine with a range up to the $10 slot machine.
Its always a probabilistic gamble that you get what you describe. And it's a pull of the lever each time. And you pay no matter what, for good or bad results.
reply