256+: GLM-5.2 (swap with Kimi K3 when it goes open-weight on Monday)
<256: Actually, surprisingly, not a Chinese model but probably Laguna S2.1. The best Chinese model at this size is DeepSeek V4 Flash though
<96: Qwen 3.6 27B
<32: Still Qwen 3.6 27B (NVFP4)
<16: Oof, not sure. Nothing will feel great at this size TBQH without finetuning on a specific task. Pick your poison of tiny Qwen or tiny Gemma (although again Gemma is not Chinese)
And we can quickly see that the real problem is not the models, but HW to run them. You can build whole Enterprise on Gemma 4 31B full precision without significant problems. If you can afford not to lobotomise it by quantisation.
<256: Actually, surprisingly, not a Chinese model but probably Laguna S2.1. The best Chinese model at this size is DeepSeek V4 Flash though
<96: Qwen 3.6 27B
<32: Still Qwen 3.6 27B (NVFP4)
<16: Oof, not sure. Nothing will feel great at this size TBQH without finetuning on a specific task. Pick your poison of tiny Qwen or tiny Gemma (although again Gemma is not Chinese)