Hacker Newsnew | past | comments | ask | show | jobs | submit | natrys's commentslogin

They can and almost certainly are doing similarly impressive engineering works internally.

They just aren't in any hurry to forward those cost savings to you.


Some official benchmark numbers posted in Chinese social media (I am sure they will publish an English blogpost later too):

https://mp.weixin.qq.com/s/V4xhEIy8xDXSMDPrPkmUAQ

Generally looks like a Sol/Fable tier model, better across the board than Opus 4.8.

(Edit) English blogpost is up now: https://www.kimi.com/blog/kimi-k3


The link has 6 well-known benchmarks where this beats Fable (out of 14 I counted). If the numbers hold up scrutiny, this is scary good.

Forget about their pricing but the companies that do have means to host such models fully on-prem are also the same companies that are paying tens of millions of $ in inference cost every month, and are by extension the biggest customers of OAI and Anthropic


Open Source >>> Closed Source [1]

I don't want to cheer against my country, but we've given up on open source. The way Anthropic and OpenAI treat their customers as adversaries is embarrassing.

I will cheer for China, for Kimi, and for z.ai until we have something in the same category.

[1] I'd even be fine with open weights, fair source, or anything that let us have direct access to the weights. Even if that came with stipulations. Don't hide the weights from us.


I am with you in the spirit of openweights but I am trying to hard-avoid bringing countries into this. The narrative of US vs China only benefits those who want regulatory capture in the US since attacking China is politically much easier than attacking open-weights, so certain groups like to repeatedly call them 'Chinese models'.

I call them “Chinese models” without vitriol. I think they’re great. Although it’s good to remember that all models have ideological biases inherited from their creators.

It's much more a rallying cry for open weights funding than it is for regulatory capture.

The argument on our side wins - if America or the West don't do open source, China will. And that means -- with certainty -- that China wins the market.

Every politician and VC should hear that loud and clear.


I just want to see dario cry for some reason . i cannot explain it but i want him in particular to lose.

> If the numbers hold up scrutiny, this is scary good.

After using it for a few hours, I believe these benchmarks.


It's like reading Anthropic's obituary.

This is weird and reactionary. Lots of organizations are continuing to refuse to use chinese models due to security and IP concerns. Anthropic/american models aren't going anywhere anytime soon.

> Lots of organizations are continuing to refuse to use chinese models due to security and IP concerns

This is such a common omission: the Chinese models are open, you can host them yourself on your premises. So privacy and independence.


it's well documented that models can be adversarially trained with essentially backdoors in response to special inputs

while I am skeptical that this is happening atm, there are probably many industries where the risk does not seem worthwhile


I suppose this is like when Anthropic was using “prompt modification, steering vectors, or parameter-efficient fine-tuning” to poison the work of people working in the LLM field, including academic researchers.

No, that was totally different. They were just doing that for your safety.

When the model is open weights you can even pass every token (including the chain of thought) though a fourth-party lightweight model like gpt-oss-safeguard to check that it has not become adversarial.

So it's better to trust cloud solution where not only they can do the same, but also actively use your data?

Because in 2026 we still believe USA is more trustworthy than China?


I feel like that's a threat that isn't super difficult to block. Unplug it from the internet, require it to go through an API intermediary to access web pages.

Maybe I just don't have any imagination.


It could generate code that's plausible but has intentional flaws, kind of like the defunct underhanded C contest [0], except through a LLM.

[0] https://en.wikipedia.org/wiki/Underhanded_C_Contest


It could, but exposing that would doom the company entirely, and AI doesn't generate code with near the quality needed to get a model to mass adoption, insert malicious underhanded code, ensure that consistently looks innocuous enough to never be noticed, and- most importantly- actually exfiltrate data without being noticed. Once it is noticed, it's game over across the board.

You don't even need that, all models are susceptible to prompt injection. You already need to take extreme security precautions and assume all models can essentially behave like attacker-controlled rootkits.

That's right, but this strategy only works when there are no opponents of the same level to review the generated code.

For several export controlled industries in the United States, even self-hosting a Chinese model is a non-starter.

Good luck hosting 2.8T params yourself. A box capable of this at a useful performance level is at least $100k.

More like $500k, but that's not an unreasonable price for a medium sized enterprise to pay.

I have an RF engineering background, a nice mmWave vector network analyzer can easily land in that ballpark.

If the business value is there, companies will pay for it.


Not to mention the bill you would be paying anthropic could be way higher at that scale

How much would it cost to run a 2.8T model with long context for 1000 developers, plus 30,000 non-dev employees with short session memory?

> Lots of organizations are continuing to refuse to use chinese models

Correction: Lots of organizations are refusing to use Anthropic Fable because they have forced opt-in data collection as part of their privacy policy, even for Enterprise.


Both things, and both reasons, can be true at the same time.

Not everyone's going to care about Anthropic requiring data collection (a similar debate plays out with regards to "pay or consent" on website tracking), just as not everyone cares about China with regards to security/IP issues (if they did, a lot more would be banned besides occasionally-Huawei).


> Lots of organizations are continuing to refuse to use chinese models due to security and IP concerns

These customers exist (e.g. US military) but there's not enough of them to justify a trillion dollar valuation.

Anthropic's valuation is predicated on growth. If they start going backwards and losing customers to open models, it hurts their ability to gather investment and with it the ability to train new models, leading to a death spiral.

The best they can hope for is that the US gives them state aid to compete with China, however their relationship with the current administration is not great.


The inverse is also true: many companies are refusing to use American models due to security and IP concerns. And it's more concerning here: American companies straight up say they will train on your IP; local Chinese models structurally can't. The security concern goes this way too: for American companies, you're relying on their own security infrastructure and essentially blind trust. For locally hosted Chinese models, you 100% control the security story.

The reality this demonstrates: most US companies don't give even 2 shits about their IP, and are fine willingly handing it to Anthropic et al. Those that do care largely must care contractually. For group 2, they're either using Chinese today or aren't using AI at all. Those are the only two valid options, there is no secret "use US but self host it" third option.


Nope, but I think this is maybe the critical mass needed to finally crash the AI hype/datacenter cost problem everyones is talking about.

With Oracle being junk before this, more will follow.


I would assume the opposite is true — with an open-weight Fable-class model, doesn't demand for GPUs go up? Plenty of companies can now look at what Anthropic is offering — high per token costs for a very intelligent model — and do the math, and at some point it makes sense to just rent the GPU yourself and run Kimi on it if you get similar intelligence without paying Anthropic's margins (albeit with high upfront capital cost).

This would drive down Anthropic's margins, but drive up demand for datacenter and GPU capacity. It's not that people would be using fewer GPUs, they'd just shift demand from high priced token vendors to direct GPU rental, which benefits datacenter companies while hurting Anthropic.


Its a margins game. If its too cheap to run, its not worth the investment.

The cost to run Kimi is the cost of the GPUs (+ overhead of hiring humans for now to manage it). Kimi K3 does not change the demand curve for LLMs, it only changes the possible suppliers — and they're all competing for the same supply-constrained resource, which is GPUs in datacenters. Regardless of who is serving the model, or how, they're going to need to rent or buy GPUs in datacenters. Hence: this is great for datacenters.

That makes no sense, Terry.

Oracle is fine, it's just that they can't really expect political decisions that hindered it to accquire TikTok which will be slated to be the biggest customer if the deal went through.

Now they are betting with Project Stargate but it also seems to be crumbling down.

But don't forget that they literally hold the biggest databases, both in commercial and open source, that is, Oracle Database and MySQL. Plus Oracle Java they literally controls at least 30% of the internet's software infrastructure.

And also with a good team of attorneies enforcing the licenses, they can squeeze so much money at the cost of morality.

Also recently they downgraded the always free OCI ARM instance from 4C24G to 2C12G without telling anyone.


New enterprise java licenses are going to milk enterprise just like broadcom is doing. New license deals makes you pay for employee total number (including contractors) instead of for users of oracle java.

Even boring, slow moving companies are well on their way to eradicate the last traces of oracle java. We kicked the last of that 2 years ago, despite having >3000 different systems across the globe. (We started circa 5 years ago).

> Oracle is fine

They're drowning in debt and risk is increasing. If these US models don't keep holding up their valuation will tank further and some will recall the loans or ask for different terms.


Models need datacenters to run. It also need other services to do anything useful

The point: Fable isn't worth what Anthropic says it is, so Anthropic isn't as valuable as they make themselves out to be.

The DeepSeek incident has already shown it, this is a reminder.


Cursor will rebrand it as Composer 3.0 to assuage any such concerns, as they did with the previous Kimi models.

More likely for them to use Kimi 2.7 since Grok is now the flagship product.

Musk bought it. From now on, it will only be Grok.

If it ends up being open weights, companies will use it running in US data centers.

You can run open weight models anywhere.

This is apparently Open Weights, so no reason Amazon can't serve it alongside GLM which they already do.

> continuing to refuse to use chinese models due to security and IP concerns

One can run open weights in an exclusive TEE'd GPU too, which still comes out cheaper than closed weight LLMs. Ex: https://chutes.ai/pricing


Certainly for their IPO, anyway

Fable is by Anthropic, and this is too expensive, GLM 5.2 is roughly the same quality at a much cheaper price.

(I mantain a client with llama.cpp and 101 models across 14 companies by http)


As much as I like GLM 5.2 it's clearly a step below Opus (or even Fable) for more complicated tasks. I would place it at Opus 4.6/4.7 level.

Having said that, the safety system on Fable makes it an extremely unattractive model. It feels that half of the time you're paying double for Opus level performance.


Fable won’t even generate a jwt to test endpoints because it is security related. It is crazy capable but useless for real work

Unless your real work is outside the scope of one tiny niche of work.

Eh, it doesn't hit you until it hits you.

I finally bumped into a task that Codex would refuse to work on.

Was I attempting to reverse-engineer a GPU driver? Yes. Was I trying to hack into the DoD? No.

I wasn't doing anything wrong, but that's not what OpenAI's safety mechanisms thought.


GLM has issues with tool calls and nested JSON and it wastes tokens pretty often. I see it being a bit above half the price of Opus in a bit more complex eval tasks. With some RL you could probably get the tool calls sorted and the price down.

Nah:

https://www.youtube.com/watch?v=LSlV206xPqM

These real world examples show it's one tier away.


If Chinese AI companies can train a model that's slightly worse than the frontier, then there's no reason why they can't train a model that is slightly better than the frontier.

Everybody can agree that K3 doesn't clearly surpass Fable. However, inevitably there will be a time in the future when a Chinese AI company releases a model that's better than any US model.

K3 isn't the knockout blow but it's the 2nd knockdown that makes everyone in the arena realize that the fighter is not winning the fight.


Anthropic is arguably still better in tooling and integrating model and tooling. Good habit beats raw intelligence.

For code editing Cursor editor tooling is even better.


These "real world" examples are nothing like the way I use LLMs from within a harness. GPT 5.6 Sol and Fable are clearly more impressive, but how does this translate to interactive agent use, or use under an agent orchestration framework?

This is a question I am going to get an answer tomorrow with evals. Extremely interesting...

I think given how much benchmaxxing we're seeing - the anecdotal evidence of how competent this model is (and efficient) will depend on user's actual real-world use cases.

Given the pricing, it suggests that this model is much more efficient/competent than previous-gen OS/distilled models.


Any benchmark where Sol is better than Fable at coding is ridiculous.


If anecdote is data, then here's another point:

https://nitter.net/synthwavedd/status/2077537805715005724#m

(As an aside, I don't know how it was professional of Arena to unmask an unreleased cloaked model on their platform. Also practically, upstream could have been A/B testing multiple variants under same endpoint, casting validity of such pre-announcement tests into question)


Why we should waste 46 minutes instead of briefly looking at charts for 20 seconds? To pay their ads? No, thanks.

[flagged]


distillation attack? why the violent word choice? When OpenAI crawled Github was that an attack?


AI labs have been doing "distillation attacks" against actual creators like artists and engineers since the beginning.

Do you have moat if your advanced model can be distilled in a month or two ?

Distillation is not an attack. It simply a way to train a model. Not doing it when you are behind is akin to snatching defeat from the jaws of victory.

Indeed,

Thinking about, if I had a lab, I'd be trying to get training data from a combination of the open web, piracy and all of my revivals. I wonder how labs are doing that.


How Xi coded. If you just cheat off the top student in your class that's simply a way to get a good GPA.

Ah yes, because if a person agrees with anything a chinese company does it must be because they love Xi. Get real.

Munching of the top student in a class is clearly prohibited. Distillation is not cheating, it's learning from your competitor. Akin to a company purchasing their competitor thingamajig to see if you can improve their own product.


If your entire competitive advantage is copying your rival's better product and not making any true innovation or improvement and only delivering a slightly worse product its little different than cheating, yes.

>Distillation is not cheating

Blatantly against ToS of any of the major labs hence their efforts to prevent it.

Not saying you are actually a Chinese astroturfer but this is essentially exactly what I would expect one to be saying.


These models are not copies, distillation is simply a part in the pipeline. suggesting that the Chinese models are full copies through distillation is simply wrong. Just because they are distilling does not mean they are not innovating. In fact there is pretty much a consensus in industry that in terms of model size/performance the chinese models are better. But of course we don't know the exact model sizes because OpenAI and Anthropic are not giving access to their models.

> Blatantly against ToS of any of the major labs hence their efforts to prevent it.

ToS are actively malicious and fortunately not worth the paper they are printed on. Violating a ToS is not illegal or "cheating".

> Not saying you are actually a Chinese astroturfer but this is essentially exactly what I would expect one to be saying.

The AI industry in china has been acting a lot more moral and forward looking than the, to be frank, mustache twirling evil US parties of Anthropic and OpenAI. Perhaps take that into account before randomly spewing accusations of astroturing.


It is an attack at a sufficient level of sophisticated analysis. If you destroy the game theoretic first mover advantage, then you destroy the economic incentive to improve things.

Given that model distillation has existed since the early days of the current AI boom, and no robust defense has been demonstrated, the available evidence does not support your theory.

Be that as it may, it would seem absurd if we start calling distillation out as antagonistic, but don't do the same for the SOTA models being trained on human-created data.

These things enormously benefit from economies of scale. I am fairly certain their margins might be low but they don't actually sell API at loss, however that doesn't mean your cost footprint would be anywhere as low.


No I think uv is to python what opam is to ocaml, it's mostly a package/dependency manager.

Superficially, both uv and dune are also project runners. But dune is mainly a build tool, most important things dune does such as pre-processing, linking, compiling etc., are not needed in python in the first place (at least talking about pure python). You can use uv to create tarball/wheel but it's more akin to simple bundling than building in the dune sense. Dune can also run tests, but in uv you would need to delegate to something like pytest etc.


It seems frontier, on the balance, would rather lose that segment of he market than lower the API price. They are getting the bag in the enterprise segment, those clients aren't ditching them for DeepSeek.

As for other segments, high API pricing gets people to switch to the subscriptions instead which is stickier than the API.


I've been hearing that Anthropic want all major AI providers to stop developing front tier models for a year for safety reasons. The real reason is they need time to get there models cheaper because of the DeepSeek threat or local llms or other even cheaper providers.


Seems like a ridiculous request - how can they ensure China will stop developing frontier models?


Yes it was good for its time, but 10 months old now which is a long time ago in this space. It was also a fine-tune (albeit a good one) of Qwen-2.5 72B.

I wish they did more smaller models. Kimi Linear doesn't really count, it was more of a proof of concept thing.


I think tree-sitter's relationship with JavaScript is entirely syntactic. You don't need any JS runtime installed to write grammars, because technically tree-sitter CLI already has a JS runtime included and using that it converts your grammar first to an intermediate JSON format, then it generates parser code in C. And then this C code gets compiled into a shared library, which is what editors like Emacs use, so to use tree-sitter modules you definitely don't need a JS runtime either.


Very impressive demo. From VM curation to vibe coding something running on port 8000 in Shelley just worked in minutes. I imagine quite a few technically impressive things happening under the hood, would be interested in reading more about those.

Small nit: I think you should make it more clear in the docs (if not in the landing page) that one can just use any key with the ssh command the very first time and it automatically gets registered. Also on the web UI one should have the ability to add the ssh keys. I logged into the web UI first, and was a bit confused.

I think the pricing is alright for the resource and remote development features, though might be a bit much if someone doesn't need higher level of resources for deploying something that's mostly already developed.

Anyway, this reminds me of a product called Okteto that had similar UX. They were focused on leveraging k8s for declarative deployment. But for some reason they suspended their managed cloud/SaaS offering for individual/non-enterprise clients, I wonder if it was because they couldn't make the pricing work. Hope that doesn't happen here.


That's the Kimi K2 Thinking, this post seems to be talking about original Kimi K2 Instruct though, I don't think INT4 QAT (quantization aware training) version was released for this.


I am going to try and stick with Prolog as much as I can this year. Plenty of problems involve a lot of parsing and searching, both could be expressed declaratively in Prolog and it just works (though you do have to keep the execution model in mind).


Consider applying for YC's Fall 2026 batch! Applications are open till July 27.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: