My development platform of choice is Opus 4.8 --effort max. My 'in production' general model just switched from GPT-4o to GPT-5.2. I probably should test local inference for that part (some translation/knowledge extraction/summerization) next but it has not made it to the top of the priority list yet, even though 3 other small specialized task models in the app already run on prem.