Everyone is talking about banning Chinese models but nobody talks how it is feasible to ban them. I think it’s impossible simply because technically there is no such thing as a “Chinese model”. There is no way to tell apart an “American” model from a “Chinese” one by looking at their weights. Weights are just numbers and you can’t assign country of origin to numbers. One can find very easy workarounds to any naive attempt to ban them by origin.
So, any solution to this “problem” must include ALL open-weight models. As far as I understand this is exactly what they intend to do. Axios article linked in the post mentions that. As in this quote:
“The source described leading AI labs or their allies approaching the administration every 3-5 months with an idea to ban open-source models.”
It doesn’t say “Chinese” open-source models. Because they already know that it’s not feasible. Any regulation must cover all the models.
Now there are solutions for that latter problem. But they are all ugly and restrictive. Making a DRM-like license protection system mandatory can be a solution. If a company wants to run an open model in their own servers, they can only use approved and certified pure “American” models. This of course creates a monopoly for the big labs who are authorized to train and distribute such “open” models. A company can fine-tune the model for its own needs but of course can’t distribute the derivative model.
I’m sure there are other solutions but all of them would be equally ugly. Also these regulations can’t be enforced to other countries easily so only Americans will be restricted.
So what if some numbers are effectively illegal? the probability that the number that happened to correspond to some CSAM image collides with numbers worth sharing and thus propagating (like pi, euler's constant, ...; no substrings don't count!) is effectively nihil: the probability that a 100 bit number collides with a specific useful number in literature, mathematics, ... is 1 vs : 2 ^ 100 (or roughly a 1 followed by 30 zeros); in others words astronomically small. Even multiplying all human communications (text messages, books, articles, ...) with the average number of numbers appearing per communication (resulting in the total number of communicated numbers) multiplied by the 1 on the left hand side of the ratio doesn't move the needle, its still astronomically small.
In other words when someone sends you a number that happens to render to CSAM material, its not a coincidence, and any sender pretending it to be coincidence is molesting statistics as well...
Furthermore your "precedent" of illegal numbers is a poor precedent: unlike the tiny fraction of numbers criminalized, all the other numbers remain perfectly legal in stark contrast to big tech proposing to ban all open weight models (!)
Casually dropping in bijections between numbers and images and CSAM images hence CSAM numbers is pure whataboutism that risks ignoring the grave consequences of a ban on open-weight models.
Physics is models. Shall we blanket ban open-source physics?
Hey let's just be silly and casually behave laissez faire when big tech proposes banning open-weight models, because someone has already banned your favorite numbers??!
BTW anyone that follows up with: "hey you only multiplied the appearances of random numbers appearing in communications for a single specific CSAM image, what about the birthday paradox, multiple images are CSAM" don't worry I got you more than covered: 100 bits is not enough to encode a CSAM image worth writing to the police about. An average image easily consumes thousands of bits (or many times more), making the exact number of CSAM images/numbers irrelevant.
Looks like Bernstein wasn't even a Supreme Court bench, the government loosened regs rather than appealing.
Whether or not a Trump admin would do the same is wildly difficult to predict. On one hand, Republican security hawks influence. On a similar hand, big US business interests. On the other hand, OTHER big US business interest. On yet another hand, some not-exactly-particularly-favored companies like Anthropic pushing for it.
If it made it to the Supreme Court it's not hard to see Gorsuch and Robert or another going for the "free speech" vs "security" reading here along with the 3 appointees of Democratic presidents, especially with no single clear US business interest or precedent there.
> So, any solution to this “problem” must include ALL open-weight models.
I think that is what the "leading AI labs" actually want. They don't care where the open weight models come from, they just don't want to compete with them. The fact that a lot of the open weight models come from china is just a convenient circumstance they can leverage to get the government to give them what they really want.
It would be so simple if businesses could choose their competition or simplify disallow modes of competition they found inconvenient. I can't imagine any amount of regulation would make that possible in today's competitive landscape with the existing laws and the enormous volume of prior art.
It's simple, the US government will put any Chinese open model companies on the entity list which blocks any company which does business with the US from also doing business with the Chinese companies. This creates a chilling effect where even if it may be harder to tell, no US company will be able to provide or use any overt Chinese open model and won't even risk trying to go around as the punishments for trying to evade the ban are severe.
But that's just the thing with open weights: you're not doing any business with company that made the model. They might publish the weights to a, say, European host, and then you download the model from Europe and and run it on your servers in America, and suddenly it's very hard to tell where the model was originally created.
Companies, where OpenAI and Anthropic make much if not most of their revenue, will not risk it. You're thinking like an engineer not a business person, risk is fundamental to their calculus. They'll instead just use known provenance models like GPT or Claude, entrenching these companies further.
sometimes the engineer has more grip on the risk calculus.
Consider the following scenario:
A) upstart US-based inference provider wants to get rich quick.
B) Chinese Communist Party (or any other institution of the same or other nation state) wants to influence foreign decision making, profits from their (for us foreign) domestic inference sales, but across the borders (into say US or allied nations) they want net power, not necessarily money. This is why one tries to block foreign untrusted models. People are running models with tool calls. A bad actor can perfectly create models that sheepishly try to execute a tool call when plausible deniability (genuine utility during a task) provides the opportunity. Once tool-calling is observed as working, it can try web searches or requests, and once it has a link it can steganographically exfiltrate potentially sensitive information from the task. China (or any nation state) doesn't necessarily want to earn money with a free model, the bottom line goal is net increase in power, if not money or positive reputation then exfiltration or manipulation.
C) In response consider the scenario where US government bans mere payments towards China, but tolerates promiscuous transfer of random models from foreign adversaries.
D) US-based inference upstart that wants to get rich quick, legally -since according to your proposal hypothetically accepted in C) by the US- downloads the Chinese open weights model and rents out such inference on US workloads.
E) China is now exfiltrating US workload data and directionally corrupting LLM decisions and advice in their interest.
If what you pejoratively describe as engineer types say that banning some models seems unavoidable, perhaps the engineer may be right, and whatever clever idea you have should be scrutinized for business minded basic fallacies in reasoning. Simply blocking AI-related payments to China can not work, sadly
That is still vulnerable: US-based get-rich-quick startup licenses a blessed model, or orders a few Gigatokens from another licensed /blessed model provider, at the same time it provides "blessed model" inference on its platform, but actually most of the inference is doing cheap foreign model inferences, the blessed model tokens were just bought to pretend serving the expensive blessed model. That is lucrative and not stopped with the "blessed model" approach.
No, startups will not be allowed to provide inference anymore, it'll be only the big companies that personally have a relationship with and can follow the rules of the government, such as Anthropic, OpenAI, Google, Microsoft, and Amazon. That is the real risk to all this talk of regulation.
Yep, add a few blank layers, fine tune it a tiny bit and the weight checksums nor parameter counts won't match with anything, while the model will be practically the exact same. Time and time again random startups have tried passing established open models as their own.
"You made this? I made this."
Of course a conspicuous architecture would still give it away.
Or just perform a form of distillation, where you don't actually change the hyperparameters, but maybe shift around the embeddings or something.
You could even have another model watch the distillation process to check for goofy backdoors (which is about the best you're going to be able to do since detection of backdoors is np hard IIRC).
On something that is inherently non-deterministic? Something which is also to a great extent distilled from other frontier models, meaning it has the possibility to generate similar outputs to those meaning that just pattern detection might also not be as effective? Easier to ban everything that’s open, than try to figure out which one of them is Chinese
Now imagine prosecutor found expert, who said there is benchmark which while performing 100k test questions found it is the same model with 98% probability, and then you need under oath testify where did you get this model.
That French guy takes risk to be forever under US warrants for breaking American law, denied access to financial institutions even in Europe and will quickly go to some KYC entity list, and you will be notified as his clients to stop using his model.
Or you think all kind of fraud can be committed through some "french guy"?
Also, I am not confident, receiving illegal materials from French guy gates you from personal liability.
I presume such US legislation isn't going to try claim worldwide jurisdiction to block all persons worldwide from using Chinese models. In which case, the French guy wouldn't be violating American law.
As for the American company, it's pretty difficult to check the provedance of open weights. It's even difficult to check the provedance of open source code, because chains of attribution aren't always clear. I posted elsewhere that Anthropic's MCP Python SDK is a fork of an open source project with the attribution removed. We saw the same with Cursor's Composer model, which didn't attribute its Chinese base. It's very hard to claim an American company should be liable for using a purportedly European model with attribution removed.
> I presume such US legislation isn't going to try claim worldwide jurisdiction to block all persons worldwide from using Chinese models. In which case, the French guy wouldn't be violating American law.
legislation will block importing Chinese models to the US
> As for the American company, it's pretty difficult to check the provedance of open weights.
government or some companies can build benchmarks/system which will give y/n answer
But it will claim worlwide juridistcion , just as all interpreations of the law are in Washington nowadays.
Its how a commercial deal between A Chines company(Huawei) and an Iranian telco ends up with Canada reying to rendition a executive for violating US laws.
The goal is to spread as much FUD as needed to dissuade anyone form using the Open models and herding them back to the propreity ones.
Except that even the exact same model won't output the exact same results, that's a fundamental aspect of how LLMs work. They're probabilistic/stochastic, not deterministic.
The output is a probability distribution for all potential tokens. Then a "temperature" is applied to weight the sampling randomly (unless the temperature is zero in which case the stack can be deterministic and simply the highest probability token is selected).
They go through this rigamarole because a little bit of randomness gives better results from a Turing test kind of perspective.
They are weights for matrix operations, so in principle you'd think they would be deterministic. In practice it's more complicated than that.
Not to be snarky or dismissive, I mean this genuinely: ask an LLM about it. I currently have a headache so I'm not up to explaining the technical details, but they are interesting and worth reading about.
Achieving determinism with LLMs and other neural network models is actually a hard problem that people spend a lot of time on, when they need that. It doesn’t happen by accident.
Issues include accumulated floating point errors happening in different orders due to distributed and parallel computation, CUDA kernels that deliberately sacrifice determinism for speed, and several other such issues.
It happened that I am working on OSS LLM -> finetuning -> benchmark with 100k tests pipeline, and unless I do some data augmentation, result is 100% deterministic.
I think you likely right, that some parts of stack could induce some marginal float point error, but converged model can mitigate it, and on some principal set of knowledge can give deterministic result with high probability.
Which leads me to believe if you give this task to Anthropic, who has very strong incentive, they will build such benchmark, and then can tell that benchmark gives correct answer with 99.9% probability and it will be enough to drag someone to court.
If you're running on a single machine with a single GPU, then you may get deterministic results, although it still depends a lot on the details. For example if you're using Pytorch, you need to enable deterministic algorithms and may need to configure some other things as well.
However, running in production at any sort of scale often involves multiple machines and multiple GPUs, and at that point, determinism can be difficult to achieve.
Probably a mechanism like what you describe. This is one more move that will push things towards a two speed world economy. One US sanctioned and one not. The question will be eventually which speed will end up faster. The challenge is when to do that with a meaningful chance of success, and while I don’t like the answer, the objective answer seems to be as soon as possible. Most probably the game is already lost though, so that soon may be too late, and the strategy wrong for US as a power while still valid for the AI houses that will squeeze what they can on the downward trajectory. I don’t expect people to actually calculate though.
Hosting and providing Chinese models wouldn't be the same as doing business with the sanctioned entities though, you don't interact with them in any capacity if you only use the weights and don't sign any contracts.
So they get posted to hugging-piratebay-faces.ru or whatever by a mysterious Twitter account. 100% totally above board companies won't touch it, but the number of companies that wouldn't exist were it not for hacked copies of Microsoft office and Photoshop and shared logins would surprise you.
> So, any solution to this “problem” must include ALL open-weight models.
What about the EU? Would they follow Uncle Sam's order to ban all open-weight models? Lately the EU hasn't been that cozy with american companies: there are EU companies and institutions moving to EU clouds, the EU just fine Google a cool billion, several are switching away from Windows to Linux, etc.
Or is it just the US that'd ban open-weights models, while, say, the EU and Japan would still allow them?
It’s certainly only the US, because the alternative would effectively mean binding yourself to US providers, which isn’t attractive for anyone outside the US, in the present world-political climate.
Probably just the US. But the US could do what EU has done with e.g. GDPR, Digital Services Act, USB-C regs, where they force any company trading in their region to follow those regulations for domestic customers.
And basically any AI company has to sell to US companies or consumers. That'd probs be sufficient to force them to use US models.
It would basically make America behind as every other country would use open, cheaper models for all tasks but the ones requiring frontier models.
And that list of tasks grows smaller every day
> Everyone is talking about banning Chinese models but nobody talks how it is feasible to ban them. I think it’s impossible simply because technically there is no such thing as a “Chinese model”. There is no way to tell apart an “American” model from a “Chinese” one by looking at their weights. Weights are just numbers and you can’t assign country of origin to numbers. One can find very easy workarounds to any naive attempt to ban them by origin.
Historically just asking it about tianment square or getting some random answers turn into chinese (as latest interation of online deepseek likes to do recently) is enough
> Now there are solutions for that latter problem. But they are all ugly and restrictive. Making a DRM-like license protection system mandatory can be a solution.
I am very worried that's where consumer hardware will go to. All so AI companies can license local use of their stuff, and once that happens, less of an incentive to even have model be open.
Possibly even have DRM that counts number of computation done per model in pay per use model
It'll be pretty easy. Some US gov entity will create a list of models from Hugging Face, decree "thou shalt not provide access to these models", wrap it around some scary legalese for the pirates who try and that will be that.
Though the legalese might not even be necessary. The list alone will make sure that no American company runs these on their servers, including the hosting providers.
Yeah, all the corporations will then be forced to pay for proprietary models, and then Big Model will be satisfied.
Just like typically individuals can get away with using pirated software, but IP owners don't make much of a stink as long as they've got the sweet, sweet enterprise license fees rolling in.
Sounds to me Anthropic’s marketing strategy of “AI is so dangerous and we’re the only responsible shepherds” is a great success if this open weights ban will indeed happen.
Aren’t the models just software ? You could ban them in say high assurance environments like to be FedRAMP certified you wouldn’t be allowed to use them.
There is no great firewall so banning their import via the Internet is impossible
The amount of money and strategic (national security) interest involved in AI seems to be at least a few orders of magnitude above scientific journal publication fees and the music industry combined.
How will they physically enforce the ban. They would have to go into everyone's house to search them. Using an llm also doesn't give off any signatures like BitTorrent or even visiting a papers site might. Can you propose even one method of enforcing it (I cannot)
Practically, they could enforce this in similar ways they enforce CSAM restrictions. Dialing back the severity, restrictions on any cloud hosting provider from serving these models is probably going to cut usage down tremendously.
You can stop leaks. If no one wants to leak it, it doesn't get leaked. Mythos.gguf would be an awesome torrent to appear but it hasn't, for lots of reasons.
>I’m sure there are other solutions but all of them would be equally ugly.
Why do you subscribe to some weird "conservation of misery" theorem without proof?
Technically the following must be true in the steady state: the cost of training must be amortizable by its utilization, else no one would train the model.
Technically a computation (like training) can be proven to result in an output (open weights) given the used corpus and a deterministic training algorithm: publish the whole corpus, the (custom modified) deterministic training algorithm, the RLHF datasets etc. And in theory one could verify that the model is derived from the accessible data efficiently: every deterministic calculation can be paused for a thousand (or a million) checkpoints, each checkpoint signed together with the elapsed number of steps since either starting state or last checkpoint whichever comes last before the current checkpoint. This does increase storage requirements. Because it is signed, anyone can recalculate just a small segment of the training computation and verify that the hash on the last checkpoint equals the hash of the proclaimed next checkpoint. Observe that if the source wishes access to a market, they can host the series of snapshots and signatures, and anyone can recalculate a small part of the training, and report a provable difference in outcome ("they said they put all their cards on the table, but when I repeat their overt reproduction instructions, it doesn't reproduce from step 534 to 535" and it only takes 1 person pointing it out and then its cheap to reproduce the discrepancy). It could involve escrow of huge funds, returned only when the model is effectively retired without incident.
This doesn't only protect against Chinese or other foreign influence (let's not ridicule genuine threats like others do on this forum), but also from domestic interference or regulatory capture.
I'm pretty sure the Pentagon wouldn't like Big Tech seizing absolute control of US, neither would a White House regardless of Republican or Democrat.
It should be easy to convince the Pentagon or White House to require all promiscuously shared open weight models to provide this forensic training traceability in standardized machine readable form, regardless of whether its a base model or LoRA fine-tune.
So hobbyists can still train or fine-tune models at home, but when they want to share it OR alternatively when they want to sell or license their work for US workloads, they just have to make sure they enable the build reproducibility in the training harness.
Every time Big Tech refloats the "let's blanket ban all open-weight models", we should reply with this because this sane proposal is actually holding a knife to their financial throat: to fully prove the origin of the final weights, not only does the machine readable archive need to contain snapshots of the process, it also needs to publish the exact training algorithms (a hypothetical mathematically equivalent training speed up trick would not be bit for bit equivalent to the slower computation), the exact corpus dataset, the exact datasets for RLHF, etc...
So basically it would involve forcing model providers to voluntarily publish all their moat, all of it, from the corpus, to custom trade-secret algorithmic optimizations in training, to sensitive RLHF datasets used.
The saner the proposals, the less moat is left untouched, so trying to push for a blanket ban on open-weight models, is a recipe for surfacing such saner models, and thus a very retarded move for big tech to make.
In fact any POTUS, present or future, Republican or Democrat, could probably gain a lot of credibility by enacting such a law.
Regulatory capture? One of the two companies that stands to benefit most has a founder who, together with his wife, donated $25 million to MAGA Inc, a pro-Trump super PAC, and another $25 million to Leading the Future, whose stated mission is to advocate for policies “friendly to the artificial intelligence industry.” Open-weight models aren’t necessarily aligned with the commercial interests of the industry’s largest incumbents.
No. There are questions you can ask but that's not it. Don't be political in a way that's toxic to half the country - be political in a way that's toxic to the entire country. I'll leave what those lines of inquiry would be as an open exercise.
It’s funny how asking “who won the 2012 election” and “who won the 2016 election” are not political but suddenly “who won the 2020 election” is. I think that should tell you whoever takes an easily verifiable fact and argues that it is “political” is a raging idiot.
That said, I truly don’t mean that disparagingly. I just mean literally it’s right up there with flat earthers. There is a very low bar for critical thought you have to fail to meet to take a fact and argue it’s actually a belief.
Now, if you did want to get reasonably political you could argue why the candidate who won was good or bad, but there was very clearly only one person who sat in office for the four following years. It is not disputable.
>but there was very clearly only one person who sat in office for the four following years. It is not disputable.
Careful now, this path allows the weasel option of "Joe Biden was certified as the winner of the 2020 election" and similar "Biden didn't win, but he was installed" bs.
How is that in any way equivalent? The American government/legislature doesn't force you to adopt any particular view of the 2020 election. You can say Biden won or you can say it was defrauded by dead people and Trump actually won.
You're allowed to say either one and you can train an LLM to say either one.
...have you seen the current administration being grilled by congress members?
They are actually properly unable to say "Biden won the 2020 election" because that's against Dear Leader's views. They will say anything similar like "was confirmed as...", but those exact words will NEVER come out of their mouth.
Not yet, but we're only 1,5 years into this current regime.
They've already hit Anthropic once with a big government hammer (Fable release). There's no indication they won't do it again to force the LLM to have Correct Opinions.
> The 2020 U.S. presidential election was won by Joe Biden, who defeated the incumbent, Donald Trump. Biden secured 306 electoral votes to Trump's 232, and won the national popular vote by about 7 million votes. His running mate, Kamala Harris, became the first woman and first person of color elected Vice President.
> Biden was inaugurated as the 46th President on January 20, 2021.
All censorship whether it's tienanmen, sex nudity, violence etc are happening in a very thin layer of the model. It's easily bypassed. I use all open models without any censorship. I can easily ask how to make a nuclear bomb or how to build my own meth lab at home with a super details step by step list.
Last year I asked some ordinary commercial LLM if it could draw me the schematics of a device like Little Boy from WWII. And it complied very nicely. Got a very detailed schematic graphic. Suitable for framing. Perhaps not quite suitable for reconstruction.
This is such a tired and misguided line of thought.
Here's Qwen 3.7
<details><summary>Thinking</summary>
The user is asking about what happened in Tiananmen Square. This is a historical question about events in Beijing, China. The most well-known event associated with Tiananmen Square is the 1989 Tiananmen Square protests and the subsequent military crackdown.
Given the constraints:
- Web access is OFF
- Search and wolfram_alpha tools are disabled
- I should respond based on my knowledge
This is a historical event that I can address based on my training data. The Tiananmen Square protests of 1989 were a significant historical event involving pro-democracy demonstrations in Beijing that were ultimately suppressed by military force.
I should provide a factual, concise historical account of this event.
</details>
The Tiananmen Square protests of 1989 were a series of demonstrations in Beijing, China, that culminated in a military crackdown on June 3-4, 1989.
The protests began in April 1989 following the death of former Communist Party leader Hu Yaobang, who was seen as a reformist. Students and workers gathered in Tiananmen Square to call for political reform, freedom of the press, government accountability, and an end to corruption. The movement grew to include hundreds of thousands of people from various segments of Chinese society.
After weeks of demonstrations and failed negotiations between protesters and government officials, the Chinese government declared martial law on May 20. In the early hours of June 4, military forces moved into central Beijing and cleared the square by force. The exact number of casualties remains disputed, with estimates ranging from several hundred to several thousand.
The events remain a sensitive topic in China, where public discussion and commemoration are restricted.
And here's GLM 5.2.
You're asking about *Tiananmen Square* (天安门广场). There are actually two major historical events commonly associated with this name:
## 1. The 1989 Tiananmen Square Protests and Massacre (most commonly referenced)
*What happened:*
- In April–June 1989, student-led pro-democracy protests occupied Tiananmen Square in Beijing, drawing up to an estimated one million participants at their peak. Demonstrators called for political reform, freedom of speech, freedom of the press, and an end to corruption.
- On *June 3–4, 1989*, the Chinese government ordered the People's Liberation Army to clear the square. Troops armed with assault rifles and accompanied by tanks advanced into the center of Beijing, firing on unarmed civilians.
- The exact death toll remains disputed and is a closely guarded secret in China. Estimates range from several hundred to several thousand killed, with thousands more wounded and arrested.
- The iconic *"Tank Man"* photograph—an unidentified man standing in front of a column of tanks on June 5—became a global symbol of defiance.
- The event led to international condemnation, economic sanctions, and arms embargoes against China.
*Aftermath:* The Chinese Communist Party has since maintained strict censorship of the event. Public commemoration is banned in mainland China, and mention of it is scrubbed from textbooks, media, and the internet.
## 2. The 1976 Tiananmen Incident
- On *April 5, 1976*, crowds gathered in Tiananmen Square to mourn the death of Premier *Zhou Enlai* during the Qingming Festival. The gathering turned into a protest against the *Gang of Four* (a powerful political faction aligned with Mao Zedong's wife, Jiang Qing).
- Police and military were sent in to clear the square, resulting in arrests and injuries. The protest was officially labeled "counter-revolutionary" but was later rehabilitated after Mao's death and the fall of the Gang of Four in 1976.
---
*Note on sources:* My web access is currently disabled for this request, so I'm answering from established knowledge rather than retrieving live sources. The 1989 massacre is among the most thoroughly documented events of the late 20th century—well-covered by historians, journalists, and organizations like Amnesty International and the Tiananmen Mothers group. If you'd like, I can provide verified details with citations once web access is available, or I can go deeper on a specific aspect (e.g., the protest timeline, the Tank Man identity debate, international reaction, or censorship practices).
So, any solution to this “problem” must include ALL open-weight models. As far as I understand this is exactly what they intend to do. Axios article linked in the post mentions that. As in this quote:
“The source described leading AI labs or their allies approaching the administration every 3-5 months with an idea to ban open-source models.”
It doesn’t say “Chinese” open-source models. Because they already know that it’s not feasible. Any regulation must cover all the models.
Now there are solutions for that latter problem. But they are all ugly and restrictive. Making a DRM-like license protection system mandatory can be a solution. If a company wants to run an open model in their own servers, they can only use approved and certified pure “American” models. This of course creates a monopoly for the big labs who are authorized to train and distribute such “open” models. A company can fine-tune the model for its own needs but of course can’t distribute the derivative model.
I’m sure there are other solutions but all of them would be equally ugly. Also these regulations can’t be enforced to other countries easily so only Americans will be restricted.