Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Are you using quantization? I’ve gotten very good results from the float16 13B vicuna model.


Did you use GPU inference? If so, how much memory is required?


Yeah I'm using GPU inference. The vicuna 13B model uses 26.3GB of VRAM with my setup. I'm running it split between two rtx4090s which gives me about 20 token/s.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: