Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
|
somnial's submissions
login
1.
What happens when a GPU writes memory
(
doubleword.ai
)
5 points
by
somnial
6 days ago
|
past
|
discuss
2.
What happens when a GPU reads memory?
(
doubleword.ai
)
15 points
by
somnial
24 days ago
|
past
3.
The case for disaggregated LLM serving
(
doubleword.ai
)
4 points
by
somnial
25 days ago
|
past
4.
On-the-fly snapshot compression for elastic inference at scale
(
doubleword.ai
)
6 points
by
somnial
26 days ago
|
past
|
1 comment
5.
NVLink, NVSwitch, and All That
(
doubleword.ai
)
5 points
by
somnial
46 days ago
|
past
|
2 comments
6.
The Anatomy of an Instruction Pipeline Hazard
(
hiraditya.github.io
)
9 points
by
somnial
54 days ago
|
past
7.
Width vs. Depth: Speculating on the Margin
(
doubleword.ai
)
17 points
by
somnial
66 days ago
|
past
|
1 comment
8.
Pushing memory bound CUDA kernels past the speed of light with data compression
(
fergusfinn.com
)
2 points
by
somnial
3 months ago
|
past
9.
Speculative KV coding: ~4× losslessly compressed KV cache using a small model
(
fergusfinn.com
)
2 points
by
somnial
3 months ago
|
past
10.
70x faster cold(ish) starts for SGLang
(
fergusfinn.com
)
1 point
by
somnial
4 months ago
|
past
11.
LLM powered data structures: A lock-free binary search tree
(
fergusfinn.com
)
1 point
by
somnial
7 months ago
|
past
12.
Parallel Primitives for Multi-Agent Workflows
(
fergusfinn.com
)
1 point
by
somnial
8 months ago
|
past
13.
Scheduling in LLM Inference
(
fergusfinn.com
)
1 point
by
somnial
9 months ago
|
past
14.
How fast can an LLM go?
(
fergusfinn.com
)
2 points
by
somnial
10 months ago
|
past
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: