Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

I get 150t/s peak, 120t/s avg with Qwen3.6 27B Q4 with a 4090 on Linux. Now that MTP has landed into llama.cpp.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: