Stories about Tim Cook sitting there being surprised by sales are total marketing. They simply ran out of memory chips because of the global shortage, so they repacked supply delays into a nice story about insane hype among AI startups. And hit two birds with one stone by throwing shade at their competitors
Apple hasn't been selling just ram for a long time, they sell vram. Try getting 512 gb of HBM on current Nvidia cards - it's gonna cost way more than $ 24k. And here you get the same amount of memory for weights right in a quiet unit under your desk
Put together a similar build with a couple of rtx 6000 Ada cards and Apple's price tag suddenly looks pretty damn reasonable
HBM itself is very expensive but it’s not really fair to compare to LPDDR or GDDR
They’re very different things.
The more logical argument to me is that Apple uses its upgrade price points as more than just direct BOM and rather as a proxy for things that are amortized across all their sales like support/warranty/etc so higher SKUs subsidize the costs of the lower ones.
If a kaiju shows up in the prompt, the weights will immediately drift from game theory into fiction. And by the laws of the genre, the military is obligated to drop a nuke on it - just to make the monster even angrier so it goes and trashes Tokyo
Yeah that’s why they delegated code gen to deterministic tooling and saved the model for input fuzzing - let it fight the data instead of legacy syntax
For the article they just mocked it out for unit tests. But in reality you really can't - that’s a massive pain in the ass. Rewriting pure math is easy, but mocking CICS transaction isolation in java means spinning up these monstrous adapter frameworks that just tank performance
The whole point of Apple silicon isnt speed, it is capacity. You can grab a Mac Studio with 192 gb of unified memory and shove a model in there that would otherwise require building a rig with multiple rtx 4090 s. For local dev just being able to fit the weights in memory is often way more important than token generation speed
Obviously without a proper RAG pipeline or web search it is just going to hallucinate with maximum confidence. It would be more interesting to see the results with top tier models that have proper alignment to refuse to answer when token probability is low. As it is they just proved that people tend to trust well written text in a chat ui
reply