I wish this can run directly on my RTX 4090, seems like 30B is the sweet spot for dense model to run locally, sadly RTX 5090 is very expensive and I need a new PC and new power supply(and UPS) to run that, adding a second RTX 4090 is another option, but not sure if my PC can do that yet.
even a 3090 will give you the VRAM headroom. i run Q8 on an 3090/A6500 combo. well, Q8 of 3.6-27B. I'm building the Q8 GGUF for 3.8 now, assuming mine will finish before someone else's.
great for mobile apps, fine for desktops, basically unusable for browsers unless you do wasm, if web can be revamped/improved it can be the best option for cross platform GUI
I love this framework! It offers all the benefits of Dart and the Flutter widget paradigm on the web, while still providing access to the DOM. I'm using Jaspr for all of my marketing sites. It's beautiful, and I can ship in minutes.
true, switched from ollama to llama.cpp these days and it's good. wonder if this is also the best option for edge ai deployment(currently use it on desktop)
I don't know how people can say this with a straight face. Nvidia was selling desktop-grade ARM SOCs before Apple Silicon was ever announced, specifically for edge robotics, computer vision and ML.
The absolute fastest desktop Mac GPUs cannot beat an Nvidia laptop GPU in prefill or inference speeds. Apple Silicon is a non-entity for professional datacenter deployment and arguably unusable for frontier models at agentic context sizes. AMD is Nvidia's primary worry, and they're not doing much better in terms of GPGPU SOC compute.
>Nvidia was selling desktop-grade ARM SOCs before Apple Silicon was ever announced
You can believe all you want that the dinky little jetson boards were desktop grade when historically the ARM SoC portion of a jetson board couldn't even keep up with broadcom/rockchip SoCs. It's taken until recently for the actual arm compute portion of Nvidia SoC's to be worth a damn at all, and they still fall far behind Apple let alone the rest of the pack like Qualcomm/Samsung.
I don't have to believe. I've run KDE and GNOME on the Tegra boards, you get full-fat CUDA support without sacrificing Vulkan drivers. It's incredible.
You can believe all you want that good single-core performance will corner the edge compute market. It hasn't, Graviton has more buy-in than any Apple Silicon chip ever got.
Every time nvidia takes its fab time and uses it to build anything other than datacenter chips it is losing money due to the massive markups the datacenter products have. Expanding their consumer offering means the datacenter backlog is going down which is very bad for their margins. Consumers will never pay 10-100x what it costs to fab something like datacenter users will.
In many cases it isn't Nvidia paying for the fab time. For instance, the Nintendo Switch 2 is basically a pure-play design and support product for Nvidia, while Nintendo negotiates with Samsung for the actual SOC prices. Their IP philosophy is closer to AMD's than Apple's, Nvidia has long roped in 3rd party manufacturers to mark up, integrate and sell their hardware.
I would not say cheap per say, but in the context of "where is the business success going to be" that's one of the reasons why Mistral focuses on providing tuned model on premises to their customers.
reply