Open weight models as a category might be a winner due to cost pressure, but I don't think Inkling is a top performer in that respect. Pricing on a token basis is 6-9x higher depending on the provider.
Can chime in on the support use case specifically: GPT OSS performs really well here and has been somewhat of a benchmark with our customers, limited testing [0] against Inkling reveals basically identical performance, but with a significant cost increase at scale.
I'd say that for real-world tasks that aren't coding most companies don't see value by being on the latest and greatest model.
Have been testing Luna against other small models for production customer support uses cases, finding negligible performance impact when comparing against the substantial cost increase. Claims re: reduced token use also don't seem reproducible. Wrote a quick blog with some sample findings (marketing content but the findings are real): https://valiopt.com/blog/gpt-5-6-customer-support-cost-perfo...
Can chime in on the support use case specifically: GPT OSS performs really well here and has been somewhat of a benchmark with our customers, limited testing [0] against Inkling reveals basically identical performance, but with a significant cost increase at scale.
I'd say that for real-world tasks that aren't coding most companies don't see value by being on the latest and greatest model.
[0] https://valiopt.com/blog/inkling-model-customer-support-revi...
reply