Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Yes, but the tokens are translated into bytes not characters. There are only 256 distinct bytes so GPT models can easily be trained to produce any character. Probably the problem will be how sensible or understandable the binary form of Chinese characters in Unicode are, but that will be a problem for the model, not the tokenizer


Ok, I understand. That helped, thanks.

And also suddenly the B in BPE makes a lot of sense.


The way I think about it: A token can be one to many bytes long, so they can be longer or shorter than a single character.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: