Hacker Newsnew | past | comments | ask | show | jobs | submit | somnial's submissionslogin
1.What happens when a GPU writes memory (doubleword.ai)
5 points by somnial 6 days ago | past | discuss
2.What happens when a GPU reads memory? (doubleword.ai)
15 points by somnial 24 days ago | past
3.The case for disaggregated LLM serving (doubleword.ai)
4 points by somnial 25 days ago | past
4.On-the-fly snapshot compression for elastic inference at scale (doubleword.ai)
6 points by somnial 26 days ago | past | 1 comment
5.NVLink, NVSwitch, and All That (doubleword.ai)
5 points by somnial 46 days ago | past | 2 comments
6.The Anatomy of an Instruction Pipeline Hazard (hiraditya.github.io)
9 points by somnial 54 days ago | past
7.Width vs. Depth: Speculating on the Margin (doubleword.ai)
17 points by somnial 66 days ago | past | 1 comment
8.Pushing memory bound CUDA kernels past the speed of light with data compression (fergusfinn.com)
2 points by somnial 3 months ago | past
9.Speculative KV coding: ~4× losslessly compressed KV cache using a small model (fergusfinn.com)
2 points by somnial 3 months ago | past
10.70x faster cold(ish) starts for SGLang (fergusfinn.com)
1 point by somnial 4 months ago | past
11.LLM powered data structures: A lock-free binary search tree (fergusfinn.com)
1 point by somnial 7 months ago | past
12.Parallel Primitives for Multi-Agent Workflows (fergusfinn.com)
1 point by somnial 8 months ago | past
13.Scheduling in LLM Inference (fergusfinn.com)
1 point by somnial 9 months ago | past
14.How fast can an LLM go? (fergusfinn.com)
2 points by somnial 10 months ago | past

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: