The link has 6 well-known benchmarks where this beats Fable (out of 14 I counted). If the numbers hold up scrutiny, this is scary good.
Forget about their pricing but the companies that do have means to host such models fully on-prem are also the same companies that are paying tens of millions of $ in inference cost every month, and are by extension the biggest customers of OAI and Anthropic
I don't want to cheer against my country, but we've given up on open source. The way Anthropic and OpenAI treat their customers as adversaries is embarrassing.
I will cheer for China, for Kimi, and for z.ai until we have something in the same category.
[1] I'd even be fine with open weights, fair source, or anything that let us have direct access to the weights. Even if that came with stipulations. Don't hide the weights from us.
I am with you in the spirit of openweights but I am trying to hard-avoid bringing countries into this. The narrative of US vs China only benefits those who want regulatory capture in the US since attacking China is politically much easier than attacking open-weights, so certain groups like to repeatedly call them 'Chinese models'.
I call them “Chinese models” without vitriol. I think they’re great. Although it’s good to remember that all models have ideological biases inherited from their creators.
It's much more a rallying cry for open weights funding than it is for regulatory capture.
The argument on our side wins - if America or the West don't do open source, China will. And that means -- with certainty -- that China wins the market.
Every politician and VC should hear that loud and clear.
This is weird and reactionary. Lots of organizations are continuing to refuse to use chinese models due to security and IP concerns. Anthropic/american models aren't going anywhere anytime soon.
I suppose this is like when Anthropic was using “prompt modification, steering vectors, or parameter-efficient fine-tuning” to poison the work of people working in the LLM field, including academic researchers.
When the model is open weights you can even pass every token (including the chain of thought) though a fourth-party lightweight model like gpt-oss-safeguard to check that it has not become adversarial.
I feel like that's a threat that isn't super difficult to block. Unplug it from the internet, require it to go through an API intermediary to access web pages.
It could, but exposing that would doom the company entirely, and AI doesn't generate code with near the quality needed to get a model to mass adoption, insert malicious underhanded code, ensure that consistently looks innocuous enough to never be noticed, and- most importantly- actually exfiltrate data without being noticed. Once it is noticed, it's game over across the board.
You don't even need that, all models are susceptible to prompt injection. You already need to take extreme security precautions and assume all models can essentially behave like attacker-controlled rootkits.
> Lots of organizations are continuing to refuse to use chinese models
Correction: Lots of organizations are refusing to use Anthropic Fable because they have forced opt-in data collection as part of their privacy policy, even for Enterprise.
Both things, and both reasons, can be true at the same time.
Not everyone's going to care about Anthropic requiring data collection (a similar debate plays out with regards to "pay or consent" on website tracking), just as not everyone cares about China with regards to security/IP issues (if they did, a lot more would be banned besides occasionally-Huawei).
> Lots of organizations are continuing to refuse to use chinese models due to security and IP concerns
These customers exist (e.g. US military) but there's not enough of them to justify a trillion dollar valuation.
Anthropic's valuation is predicated on growth. If they start going backwards and losing customers to open models, it hurts their ability to gather investment and with it the ability to train new models, leading to a death spiral.
The best they can hope for is that the US gives them state aid to compete with China, however their relationship with the current administration is not great.
The inverse is also true: many companies are refusing to use American models due to security and IP concerns. And it's more concerning here: American companies straight up say they will train on your IP; local Chinese models structurally can't. The security concern goes this way too: for American companies, you're relying on their own security infrastructure and essentially blind trust. For locally hosted Chinese models, you 100% control the security story.
The reality this demonstrates: most US companies don't give even 2 shits about their IP, and are fine willingly handing it to Anthropic et al. Those that do care largely must care contractually. For group 2, they're either using Chinese today or aren't using AI at all. Those are the only two valid options, there is no secret "use US but self host it" third option.
I would assume the opposite is true — with an open-weight Fable-class model, doesn't demand for GPUs go up? Plenty of companies can now look at what Anthropic is offering — high per token costs for a very intelligent model — and do the math, and at some point it makes sense to just rent the GPU yourself and run Kimi on it if you get similar intelligence without paying Anthropic's margins (albeit with high upfront capital cost).
This would drive down Anthropic's margins, but drive up demand for datacenter and GPU capacity. It's not that people would be using fewer GPUs, they'd just shift demand from high priced token vendors to direct GPU rental, which benefits datacenter companies while hurting Anthropic.
The cost to run Kimi is the cost of the GPUs (+ overhead of hiring humans for now to manage it). Kimi K3 does not change the demand curve for LLMs, it only changes the possible suppliers — and they're all competing for the same supply-constrained resource, which is GPUs in datacenters. Regardless of who is serving the model, or how, they're going to need to rent or buy GPUs in datacenters. Hence: this is great for datacenters.
Oracle is fine, it's just that they can't really expect political decisions that hindered it to accquire TikTok which will be slated to be the biggest customer if the deal went through.
Now they are betting with Project Stargate but it also seems to be crumbling down.
But don't forget that they literally hold the biggest databases, both in commercial and open source, that is, Oracle Database and MySQL. Plus Oracle Java they literally controls at least 30% of the internet's software infrastructure.
And also with a good team of attorneies enforcing the licenses, they can squeeze so much money at the cost of morality.
Also recently they downgraded the always free OCI ARM instance from 4C24G to 2C12G without telling anyone.
New enterprise java licenses are going to milk enterprise just like broadcom is doing. New license deals makes you pay for employee total number (including contractors) instead of for users of oracle java.
Even boring, slow moving companies are well on their way to eradicate the last traces of oracle java. We kicked the last of that 2 years ago, despite having >3000 different systems across the globe. (We started circa 5 years ago).
They're drowning in debt and risk is increasing. If these US models don't keep holding up their valuation will tank further and some will recall the loans or ask for different terms.
As much as I like GLM 5.2 it's clearly a step below Opus (or even Fable) for more complicated tasks. I would place it at Opus 4.6/4.7 level.
Having said that, the safety system on Fable makes it an extremely unattractive model. It feels that half of the time you're paying double for Opus level performance.
GLM has issues with tool calls and nested JSON and it wastes tokens pretty often. I see it being a bit above half the price of Opus in a bit more complex eval tasks. With some RL you could probably get the tool calls sorted and the price down.
If Chinese AI companies can train a model that's slightly worse than the frontier, then there's no reason why they can't train a model that is slightly better than the frontier.
Everybody can agree that K3 doesn't clearly surpass Fable. However, inevitably there will be a time in the future when a Chinese AI company releases a model that's better than any US model.
K3 isn't the knockout blow but it's the 2nd knockdown that makes everyone in the arena realize that the fighter is not winning the fight.
These "real world" examples are nothing like the way I use LLMs from within a harness. GPT 5.6 Sol and Fable are clearly more impressive, but how does this translate to interactive agent use, or use under an agent orchestration framework?
I think given how much benchmaxxing we're seeing - the anecdotal evidence of how competent this model is (and efficient) will depend on user's actual real-world use cases.
Given the pricing, it suggests that this model is much more efficient/competent than previous-gen OS/distilled models.
(As an aside, I don't know how it was professional of Arena to unmask an unreleased cloaked model on their platform. Also practically, upstream could have been A/B testing multiple variants under same endpoint, casting validity of such pre-announcement tests into question)
Distillation is not an attack. It simply a way to train a model. Not doing it when you are behind is akin to snatching defeat from the jaws of victory.
Thinking about, if I had a lab, I'd be trying to get training data from a combination of the open web, piracy and all of my revivals. I wonder how labs are doing that.
Ah yes, because if a person agrees with anything a chinese company does it must be because they love Xi. Get real.
Munching of the top student in a class is clearly prohibited. Distillation is not cheating, it's learning from your competitor. Akin to a company purchasing their competitor thingamajig to see if you can improve their own product.
If your entire competitive advantage is copying your rival's better product and not making any true innovation or improvement and only delivering a slightly worse product its little different than cheating, yes.
>Distillation is not cheating
Blatantly against ToS of any of the major labs hence their efforts to prevent it.
Not saying you are actually a Chinese astroturfer but this is essentially exactly what I would expect one to be saying.
These models are not copies, distillation is simply a part in the pipeline. suggesting that the Chinese models are full copies through distillation is simply wrong. Just because they are distilling does not mean they are not innovating. In fact there is pretty much a consensus in industry that in terms of model size/performance the chinese models are better. But of course we don't know the exact model sizes because OpenAI and Anthropic are not giving access to their models.
> Blatantly against ToS of any of the major labs hence their efforts to prevent it.
ToS are actively malicious and fortunately not worth the paper they are printed on. Violating a ToS is not illegal or "cheating".
> Not saying you are actually a Chinese astroturfer but this is essentially exactly what I would expect one to be saying.
The AI industry in china has been acting a lot more moral and forward looking than the, to be frank, mustache twirling evil US parties of Anthropic and OpenAI. Perhaps take that into account before randomly spewing accusations of astroturing.
It is an attack at a sufficient level of sophisticated analysis. If you destroy the game theoretic first mover advantage, then you destroy the economic incentive to improve things.
Given that model distillation has existed since the early days of the current AI boom, and no robust defense has been demonstrated, the available evidence does not support your theory.
Be that as it may, it would seem absurd if we start calling distillation out as antagonistic, but don't do the same for the SOTA models being trained on human-created data.
These things enormously benefit from economies of scale. I am fairly certain their margins might be low but they don't actually sell API at loss, however that doesn't mean your cost footprint would be anywhere as low.
No I think uv is to python what opam is to ocaml, it's mostly a package/dependency manager.
Superficially, both uv and dune are also project runners. But dune is mainly a build tool, most important things dune does such as pre-processing, linking, compiling etc., are not needed in python in the first place (at least talking about pure python). You can use uv to create tarball/wheel but it's more akin to simple bundling than building in the dune sense. Dune can also run tests, but in uv you would need to delegate to something like pytest etc.
It seems frontier, on the balance, would rather lose that segment of he market than lower the API price. They are getting the bag in the enterprise segment, those clients aren't ditching them for DeepSeek.
As for other segments, high API pricing gets people to switch to the subscriptions instead which is stickier than the API.
I've been hearing that Anthropic want all major AI providers to stop developing front tier models for a year for safety reasons. The real reason is they need time to get there models cheaper because of the DeepSeek threat or local llms or other even cheaper providers.
Yes it was good for its time, but 10 months old now which is a long time ago in this space. It was also a fine-tune (albeit a good one) of Qwen-2.5 72B.
I wish they did more smaller models. Kimi Linear doesn't really count, it was more of a proof of concept thing.
I think tree-sitter's relationship with JavaScript is entirely syntactic. You don't need any JS runtime installed to write grammars, because technically tree-sitter CLI already has a JS runtime included and using that it converts your grammar first to an intermediate JSON format, then it generates parser code in C. And then this C code gets compiled into a shared library, which is what editors like Emacs use, so to use tree-sitter modules you definitely don't need a JS runtime either.
Very impressive demo. From VM curation to vibe coding something running on port 8000 in Shelley just worked in minutes. I imagine quite a few technically impressive things happening under the hood, would be interested in reading more about those.
Small nit: I think you should make it more clear in the docs (if not in the landing page) that one can just use any key with the ssh command the very first time and it automatically gets registered. Also on the web UI one should have the ability to add the ssh keys. I logged into the web UI first, and was a bit confused.
I think the pricing is alright for the resource and remote development features, though might be a bit much if someone doesn't need higher level of resources for deploying something that's mostly already developed.
Anyway, this reminds me of a product called Okteto that had similar UX. They were focused on leveraging k8s for declarative deployment. But for some reason they suspended their managed cloud/SaaS offering for individual/non-enterprise clients, I wonder if it was because they couldn't make the pricing work. Hope that doesn't happen here.
That's the Kimi K2 Thinking, this post seems to be talking about original Kimi K2 Instruct though, I don't think INT4 QAT (quantization aware training) version was released for this.
I am going to try and stick with Prolog as much as I can this year. Plenty of problems involve a lot of parsing and searching, both could be expressed declaratively in Prolog and it just works (though you do have to keep the execution model in mind).
They just aren't in any hurry to forward those cost savings to you.
reply