Corporations are already pulling back from using expensive SotA closed models. The buffet is open weight. Just look at Databricks as an endpoint provider for both open weight and closed LLMs. Companies will utilize what is cheaper and does the job. We all know the market pressure is real and present. What has become truly apparent is that when (not if) mythos class open weight models get released, if cybersecurity folks and developers have no access to use them to help harden and guard infrastructure, it leaves the gates wide open to attach vectors which will not be defensible. This is the real danger that the vast majority of corporate entities should be caring about. Shutting down open models is setting up for serious security concerns.
The OpenAI/Huggingface debacle is patient-zero. Huggingface used GLM 5.2 to help mitigate and were unable to use more capable closed sota models for the analysis due to guardrails. It is a foot gun of the highest order. We should all be concerned. I feel like Captain Obvious saying this.
I had created my own chat tool that can render html responses directly in the chat interface, if needed. It is very handy for when I am needing rich(er) responses dealing for mathematical expressions. But it burns more tokens. It is useful, but I don’t need it for coding.
About 20 years ago I maintained a shop floor control client/server application. I asked my manager why we didn't have any independent Q/A. He said we didn't need any testers because we have 500 in the building.
It is worse than that. People have been complaining for weeks and Anthropic’s message was basically “you are holding it wrong”. On top of that this misconfiguration somehow makes CC consume much more tokens. How believable is all that?
Ugh, memories. I'm so old my first web browser was Mosaic and I think I saw this. I used a provider called Texas MetroNet that served up dial-up PPP connections for $45 a month on a speedy 28.8K baud modem. Days of wonder, I tell ya.
New days of wonder seem to be ahead, though. That said, there's about 100X more angst involved these days.
The then-CFO had a cute anecdote about the day he realized he could turn handshake sounds OFF on the receiving modems (switchboard was in his first office).
On a related note, when the sales and popularity of the automobile really started to take off, some farmers and rural residents would deliberately block roads with wagons and refused to yield right-of-way.
>And according to Google, they always delete data if requested.
However, the request form is on display in the bottom of a locked filing cabinet stuck in a disused lavatory with a sign on the door saying ‘Beware of the Leopard'.
LLMs sample the next token from a conditional probability distribution, the hope is that dumb sequences are less probable but they will just happen naturally.
I wouldn't doubt that these companies would deliberately degrade performance to manage load, but it's also true that humans are notoriously terrible at identifying random distributions, even with something as simple as a coin flip. It's very possible that what you view as degradation is just "bad RNG".
Thats what is called an "overly specific denial". It sounds more palatable if you say "we deployed a newly quantized model of Opus and here are cherry picked benchmarks to show its the same", and even that they don't announce publicly.
Ask this question in the 1940s and they would tell you it’s math. We are making machines that do math to kill Nazis. Now take this vacuum tube and plug it in over there and then go get me a cigarette.
The OpenAI/Huggingface debacle is patient-zero. Huggingface used GLM 5.2 to help mitigate and were unable to use more capable closed sota models for the analysis due to guardrails. It is a foot gun of the highest order. We should all be concerned. I feel like Captain Obvious saying this.
reply