Character-level manipulation fails. I suspect this is because GPTs are trained on tokens rather than sequences of characters, means they're fundamentally dyslexic and can't really spell.
LLM can't do linguistic nuances like irony, humor and sarcasm. It also have difficulty with ambiguous statements and multiple negation. It may also have problem with technical jargon (though it work quite well with medical terms for some reason).
It keeps inventing features and command-line flags. A dozen times it presented me with code (Postgresql/PostGIS related) where it turns out a method cannot be called with those parameters.
https://imgur.com/if92ZIQ
https://imgur.com/sYXFPRu