Whether or not you find Anthropic's behavior bad, theybhave been very loudly stating the foreign labs have been distilling their models for a while now. This seems like an obvious response to me that would be a mechanism to make that obvious.
From my understanding, distilling the model with another model is not illegal per se. Also, the output of the LLM is public domain by law, too.
So, why all this "effort" to protect the model? This is a free market, and moving fast and breaking things is the norm.
If they are so adamant on protecting their IP, maybe they can start by respecting others' IP, so we can start talking about ethics, equality and playing fair.
> distilling the model with another model is not illegal per se.
Just because it is legal, that doesn't mean Anthropic wouldn't reasonably want to prevent that from happening (which, from my understanding, isn't illegal either).
I love the asymmetry. When small fish tries to protect itself, big fish hits small fish with "It's not illegal" pole.
When small fish points out that what the big fish is crying about is "not illegal", big fish has the right to be above the law to prevent the problem themselves.
Having values requires equality. They have lost the right to cry foul when they trained their model with "but it's fair use" card. Life works by reaping what you sow. Now they are at the reaping stage.
"It's not illegal" is only an argument against lawsuits / law enforcement involvement. Those PoW anti-AI things people put on pages aren't illegal either.
No. From my interactions, I have understood that some people use the same argument to wash their consciences from any guilt. What they do is unethical, but not illegal, and they hide under the same argument to drown the ethical angle.
In other words, being honest to oneself is important.
Anti-scraping measures people utilize are neither unethical nor illegal. That’s the difference.
I’m still frequently shocked by the entitlement people feel to other people’s work/ideas/data/bandwidth/server load, to feed a multi-trillion dollar industry. I find the totally cynical “well when you’re making an omelet…” types to be a bit pathetic, but I understand their motivation— they’re simply greedy. But I just can’t understand the genuine indignation about people attempting to limit or stop ingestion of their own work, even if it’s just for the bandwidth costs. Go ingest your own shit.
It's good to agree that some don't have a conscience, and maintaining an appearance matters more. And appearances change based on what's legal or not.
They could detect the other AI labs and also silently burn the tokens at a faster rate providing fewer tokens for money, which does sound illegal to me.
The comments only further prove that without more regulation around this, big AI wouldn't have a "don't be evil" attitude going forward.
Anyone who called for regulations/guardrails of any kind were shouted down as Luddites who hate progress. We all knew this was going to be a mess but $$$ so screw it right?
Much as I hate to defend companies climbing to success and pulling up the ladder afterwards, this asymmetry you note is kind of the whole point a company would want to grow big. Growing an organization has some super-linear costs and generally sucks for most individuals living through it - including the management - but it's still considered worth it, precisely because big entities can do things small entities cannot, and escape the threats from smaller competitors.
It's so basic it's actually part of the reason we exist, and animals of various sizes exist, and generally why evolution didn't stop at single-cellular life.
> They have lost the right to cry foul when they trained their model with "but it's fair use" card. Life works by reaping what you sow. Now they are at the reaping stage.
Yup. Except what they're reaping is insane cashflow and ability to pull stunts like these. We can call out the hypocrisy until our throats run dry, and in ideal fantasy land this would've meant something, but here in the real world, they sow the seeds of success, and now are reaping the right to be hypocritical and continue to get away with it.
I've actually heard it quite a few times from different people who want to climb the greasy pole to get heard or resources. Idk it just seems rather soulless and slightly psycho to me. It also seems like that kind of system is rather broken and unstable, if the only way you get impact is to climb up the ladder and whatever that entails.
Change this from humans to companies and I still think it feels slightly wrong.
I disagree with your assessment that large organisations are beneficial.
We can see with our current crop of large organisations that they really struggle to create anything new; most of their new products or services were developed by a small organisation and then acquired. A lot of those products are then enshittified and badly managed because large organisation politics screws things up.
Large organisations are inefficient (everyone has stories of people in large organisations literally doing nothing all day). They are horrible to work for because of the politics. They mistreat their customers and their employees. Their executives tend to lose touch with reality, surround themselves with yes-folk and descend into authoritarian psychopathy.
My personal opinion is that we would be much, much, better off if we had fewer large organisations and more smaller organisations.
Needs a little more precision: not "those with an equity stake": those with a disproportionally large equity stake.
Otherwise it's just an opener for the old excuse of "they might be ruining your life, but it's all good, it's also your pension fund, little man, that's profiting from your life getting ruined, you should celebrate them!"
Large organizations are necessary if you want things like airplanes and rockets and computers and MRI machines to exist. And if you feel you benefit from those things yourself, then large organizations that create and operate[0] them are beneficial to you, too.
> A lot of those products are then enshittified and badly managed because large organisation politics screws things up.
That's not caused by org size. It's how modern economics work because of ad-backed business models and few other things (a tangent for another time). Importantly, small orgs and especially startups are very much complicit in this - the venture capital business model in software settled around a symbiosis, where startups create toys (er, MVPs) and growth-hack the shit out of them, in hopes of winning an acquisition or IPO lottery (aka. "exit"), where a big org buys the whole thing for ${a lot}, and enshittifies it further in an attempt of extracting a positive multiple of ${a lot} from the market. Both sides know what they're doing, exits are planned from day 1, and at no point in this process "creating useful products" is ever a driving goal.
Note this symbiosis: it's a recurring theme.
> Large organisations are inefficient
In some ways. Small organizations are inefficient in others. More at the end.
> (everyone has stories of people in large organisations literally doing nothing all day).
Some (not all) cases of this are about maintaining slack in the system, which is necessary for efficiency. A system at 100% capacity is extremely fragile to breaking completely due to tiny, random workload spikes. Breakage is inefficient. Some degree of idle capacity improves overall efficiency.
> They mistreat their customers and their employees. Their executives tend to lose touch with reality, surround themselves with yes-folk and descend into authoritarian psychopathy.
That description fits small business owners much better IMO. In our times, at least in non-failed western countries, there's a limit to how abusive or careless a large organization can be with their customers or employees - their very size makes them easy to target legally. It might be hard to get through their well-funded legal defense, unless the case is slam dunk, but that's still much better than the armies of small businesses flying completely under the radar, flagrantly violating basic health and safety regulations, or flat out lying to customers in their face, because they're not worth the effort of investigating.
(Of course I'm using a biased sample; I don't know many CEOs of big orgs.)
Symbiosis angle: for abusive practices they can't get away with on their own, big organizations are more than happy to outsource to small orgs and then look the other way.
--
Anyway, key point: *there is no categorical difference between "large organizations" and "small organizations". You need a certain amount of people and communication (and capital) to do high-complexity endeavors. The difference between a well-integrated big corporation, and a hundred of small businesses that kinda end up together delivering something big, is just that the latter is using the market as management layer.
And yes, you need big orgs to create things like commercial airplanes and MRIs, simply because the big org is a boundary layer, within which you have a non-market based incentive structure, and this lets you build things the free market just cannot reach on its own.
--
[0] - Airports and hospitals are themselves large organizations.
> That's not caused by org size. It's how modern economics work because of ad-backed business models and few other things (a tangent for another time
I disagree, it's not "modern" economics, it's one half of the Malthusian trap as it manifests in all economics; the other half is that profits tend to zero, both halves are a loss of systemic slack, to reuse the good point you make later.
> That description fits small business owners much better IMO. In our times, at least in non-failed western countries, there's a limit to how abusive or careless a large organization can be with their customers or employees - their very size makes them easy to target legally.
I think this is more like predator/prey size dynamics. One way to keep safe from predators is to be too big to hunt. The regulators and governments are more like predators than their peer-competitors are, cf. "too big to fail".
SpaceX built great rockets before it became large (though this is relative - a small rocket company is a large hairdresser, for example). There is a certain scale required for some types of business, agreed. But getting larger doesn't necessarily make them better.
> That description fits small business owners much better IMO. In our times, at least in non-failed western countries, there's a limit to how abusive or careless a large organization can be with their customers or employees - their very size makes them easy to target legally. It might be hard to get through their well-funded legal defense, unless the case is slam dunk, but that's still much better than the armies of small businesses flying completely under the radar, flagrantly violating basic health and safety regulations, or flat out lying to customers in their face, because they're not worth the effort of investigating.
Flat disagree with this. Small org CEOs are close to their customers and employees and if they behave like dicks then they get punished quickly. Obviously some still do, because people, but it's harder for a small company CEO to continue being a dick.
> Anyway, key point: *there is no categorical difference between "large organizations" and "small organizations". You need a certain amount of people and communication (and capital) to do high-complexity endeavors. The difference between a well-integrated big corporation, and a hundred of small businesses that kinda end up together delivering something big, is just that the latter is using the market as management layer.
There is a key step change when the first pure-management layer forms in an organisation. This is the management layer that only have other managers reporting to them, and only report to other managers. So no direct contact with front-line staff or shareholders. Personally, the presence of this layer is what classifies an organisation as "large". It's when the politics takes over from performance as the priority and the organisation starts to lose the connection between what the c-suite want and what the front-line actually do.
And all commercial airplanes, MRIs, anything, were built first by small organisations, and only later by large orgs. Large orgs just can't invent new things unless they form specialist small orgs to do it (skunkworks, or Palo Alto, or similar). Large orgs just don't work like that.
> Flat disagree with this. Small org CEOs are close to their customers and employees and if they behave like dicks then they get punished quickly. Obviously some still do, because people, but it's harder for a small company CEO to continue being a dick.
I'm thinking it might be both - depending on who the real customers are.
I've seen plenty of what I described in "boring" B2C like... grocery stores. But thinking about it, for a grocery store chain, customers are as much a commodity as the products they buy. Suppliers are where relationships (and power plays) matter.
(This might be fundamentally the same problem as the infamous case of "enterprise software procurement" - people using the software aren't the ones paying for it. For a grocery store chain, customers come and go all the time for many reasons, so it averages out anyway - but your suppliers and partners are what makes a difference in your bottom line.)
> And all commercial airplanes, MRIs, anything, were built first by small organisations, and only later by large orgs. Large orgs just can't invent new things unless they form specialist small orgs to do it (skunkworks, or Palo Alto, or similar). Large orgs just don't work like that.
Which is why I tried to point out the category error. "Large org with skunkworks" vs. "Bunch of smaller orgs forming an alliance and acquiring more smaller orgs to productionize a new technology" vs. "government megaproject" - they're all similar, arguably for a given invention they may very well be the same thing. Names and legal groupings are different, but the dynamics is (by anthropic principle) specific to what's needed for a given type of invention.
E.g. for stuff like airplanes or MRIs, you need individuals and small teams with lots of freedom (and a "hold my beer and watch this" culture often helps), but that gives you a prototype at best - scaling this so it works reliably, and then optimizing so it can be economical, both require throwing money at people doing boring work that mostly increments things on margin. And then the money has to come from someone, and someone must be willing to spend it to fund it all.
The actual org charts and legal charters don't matter - what matters is the incentives inside. I somewhat tentatively put forth a hypothesis: large orgs form to solve problems that the regular free market dynamics can't handle, by creating areas governed by different rule sets, within which that work can be done. Whether that's by fiat or corporate charter or a bunch of friends aligning their small businesses for the same goal, is window dressing.
yeah I see where you're going. But it doesn't address the politics/management problem - large orgs invariably generate internal politics as their incentives get misaligned with their objectives (the root cause of the 5-layer "pure manager" situation becomes a problem). Large orgs have to actually split off a smaller org (that doesn't have this problem) to actually do anything new.
The Innovator's Dilemma is a symptom of the same problem: a large org with an existing market cannot innovate because the incentives for management do not allow cannibalising the existing product sales to launch the new product.
At risk of losing the metaphor, they reaped stuff across all the lands, even ones that were not theirs, and it is questionbale that they even did most of the sowing in the first place
> It's so basic it's actually part of the reason we exist, and animals of various sizes exist, and generally why evolution didn't stop at single-cellular life.
It's also quite natural to want it to stop at individual human life instead of us getting absorbed by some next bigger thing.
Which I'm fairly sure is also the desire (as far as they can be said to have any) of these animals of various sizes you speak of.
> We can call out the hypocrisy until our throats run dry, and in ideal fantasy land this would've meant something, but here in the real world, they sow the seeds of success
Just because they pulled a mirage over people's eyes doesn't mean it suddenly became the "real" world.
No, the ordinary netizen, who runs their personal web servers, who are hit by crawlers, their content ripped from their hands and their servers chocked during the process.
What Anthropic is doing is illegal in many jurisdictions. I don't know about the legal situation for the Chinese domains they mark, but steganographic data extraction without user consent would definitely be illegal in the EU, for example.
AI probably should be. The bulk of its efficacy comes from the work of “everyone else” (in loose terms). AI also aims/hope/threatens to replace such a large number and range of jobs that it probabky should be a commons.
I would 100% support this as long as the same is true for human works. Every artist is trained on thousands of years of art tradition. Copyright wasn’t a thing for the first many millennia of artistic creation, and we still got Bach and Michelangelo.
In this case, the companies that make and provide AI models that are increasingly used to interact with me on critical things (banks, public sector services) then yes.
Abso-fucking-lutely they should be regulated like crazy.
In fact I'm really surprised by the amount of people that are not worried by how many parts of their lives are being handed over to be managed by a probabilistic system that is controlled by a private company with next to zero oversight.
There must be a greater liability than "oops, you're right to push back"
The code is not eligible for copyright. If they do not give you a copy of the source code, that does not matter. And if you don't know which parts were generated by LLM, you can't safely reuse the code.
> And if you don't know which parts were generated by LLM, you can't safely reuse the code.
I speculate this could be a real issue in future copyright infringement lawsuits.
The plaintiff bears the burden of proving that the code they claim is copyrighted by them actually is copyright. If it is known that large parts of it were generated by LLM, they’d need evidence to demonstrate sufficient human input to establish copyrightability. If they’ve kept highly detailed traces of the development process, that could be rather straightforward; if they haven’t, it could be really difficult.
Now, that’s true in the US, which never accepted mere “sweat of the brow” as a basis for copyright; the UK courts have, and most of the Anglosphere follows the UK on this more than the US.
The other factor: when dealing with an (almost) trillion dollar corporation, even if you’ll win the legal argument, they may bankrupt you with legal fees before the argument is ever properly heard.
But I suspect the precedents on this topic are going to be established by lawsuits involving far smaller actors.
(IANAL and I speculate only for myself, not any present, past or future employers.)
"The US Copyright Office and federal courts require human authorship for copyright protection; works created solely by AI are not eligible for registration under the current rules."
The Supreme Court declined to consider a challenge to this rule, and so for the moment at least, the rule remains in place.
This means that companies leaning heavily into their LLM use may very well find that they do not, under the law at least, actually own any of their code. As I've read elsewhere there's every possibility that AI code will be the asbestos of the Software Engineering world. Something we'll be trying to get rid of for decades, once everyone comes to their senses.
Or in other words: there's a big difference between public domain and copyleft and it looks like whoever came up with the asbestos analogy was underestimating that difference.
What they are trying to protect doesn't qualify as intellectual property. Only 4 categories of IP exist: (1) copyrights; (2) patents; (3) trade secrets; (4) trademarks.
The capabilities embedded in model outputs don't qualify. Machine-generated outputs are ineligible for copyright. They aren't covered by patents. They aren't trade secrets, because the model companies are selling them rather than keeping them secret. And of course, trademarks are conceptually inapplicable.
This leaves the model companies with contract law (ToS) which is pretty inept because it can't bind third parties. And technical measures, like the ones being discussed in the article. And, of course, politics.
Frankly, I think it's pretty ridiculous to even think that models can be protected from being learned from. I feel the Stanford Alpaca team demolished that idea 3 years ago.
The hypocrisy of the pro-AI mega corp arguments makes my head spin. For three years they’ve been using the example of a human reading books and then outputting creative works influenced by them as analogy for training AI on copyrighted works. Now suddenly we’re supposed to not draw the same parallel about a hypothetical person who learned from Claude and is now outputting creative work based on it.
The usage of the output is probably considered legal. The usage of the service for that purpose may not be, and using it at scale in a dishonest way is not, which is what China has been doing. Countless thousands of separate requests abusing the service (which is not a simple static HTML feed, but an AI service request) for every kind of query to soak up the results.
The post is about what's in the local code, but for a long time there has already been modifications made to the request outputs from the major cloud services as they work together to both curb adversarial distillation and to degrade the quality of training China can get from that distillation.
It's likely not to make the answer wrong or bad, but to make it so that any model trained on the output would not gain the benefit of the model's reasoning generalization skills as easily and also identifying markers that might even link back to request IDs.
The techniques talked about in this post are naive and simplistic, largely because they are released publicly.
It's not as much about protecting IP as much as it is about slowing China down or being able to track the effects of abuse. So many people are talking about greedy company this, greedy company that. The world is not made up of caricatured giant money pigs wearing suits with monocles and gold watches. That is a children's view of Marx's exaggeration on free markets. Bad, greedy people do exist, but if that is your only hammer for every nail then you have a problem.
The upper-bound for how good these models can be is so crazy that it is essentially dual-use military applicable to an extent most other technologies are not. It's not only cyber attacks or biological weapons. Most people are not even built to understand the possible threats.
Why does it matter if China gains those capabilities? I invite you to begin to learn about China's behavior around the world. The CCP is darkside material.
> Why does it matter if China gains those capabilities? I invite you to begin to learn about China's behavior around the world. The CCP is darkside material.
None of the superpowers in this world is innocent, and like MAD, more countries have the capability, the better.
I know some of the things CCP do/did. I know some of the things US does/did. I'm from neither, so I don't take sides.
AI's use has been confirmed, or more precisely boasted by two countries in two different wars, and China was not one of these countries.
We have seen the effects of "if they don't know them, they can't exploit them" mindset of NSA for years. Keeping information/technology private is neither beneficial, nor possible. It's only a temporary moat-ish gap. Not a definitive solution.
Certainly the world is full of actions and reactions, nothing is happening in a vacuum. You don't have to be from a country to take sides, but presumably you have some kind of moral compass, some kind of values around personal freedom or the worth of a human life.
There can be a very real cost, because one side comes from an ideology with a history that wants to conquer the entire Earth which caused World War 2 while the other side is trying to prune the planet like a bonsai to prevent it from descending into total chaos to preserve some sense of international order.
Europe was constantly at war, and we helped stabilize it. Middle East as been constantly at war, and if Iran can be sorted then it will be the closest to some sense of peace it's been in a long time.
We used to be in Japan, Philippines, Germany, Vietnam, South Korea, Iraq, Afghanistan and so on. How many are US territories? None. We aren't out there to conquer the globe and take land. We're usually fighting other people's wars for them, because they're up against better resourced opponents. Meanwhile China is over there building artificial islands, ramming other country's ships, creating ideological police stations in countries around the world to harass people and engaging in the most widespread international interference campaigns in human history.
They do not treat their people well and they do not have free speech. The internet is flooded with their propaganda now, because they have a human numbers advantage.
It's true that given time most advantages are temporary, but there's always that slim chance we could slow them down until the CCP collapses and they could become a more normal country.
You sound like you've swallowed pro-US-propaganda hook, line and sinker.
The reason the middle-east is at constant war is because colonialist machinations. Same goes for south-Saharan Africa. And the US is a big colonialist player, just ask Vietnam, South-America, Iran, Afghanistan, etc. They all have been attacked by the US because of US colonial interests. If anything, one could make the argument that the PRC is treading much more lightly than the US.
That said - I'm not defending the PRC by any way; it's a state-capitalist hell hole that's suppressing workers by denying them any ability to organize and whose political class is purely focused on furthering their own interests and that of the moneyed elite, the common person be damned.
I didn't make an exceptionalist argument, but if any country's behavior and values can be measured compared to others, you will always be able to make some kind of decision about where those fall in terms of goodness or badness. Do you not believe in good and bad?
I don't believe in US good, non-US bad. I also don't believe the same about my religion or my political party, for that matter.
How you measure depends on weights you assign (cultural system of values) and what information you use (media bias).
You can rank in the extremes (e.g. North Korea as worse than Belgium), since they come out that way by almost any set of information and values. Comparing the US to most other countries, there isn't a clear ordering. If you believe the things you wrote, I think the other comment summed it up well: "You sound like you've swallowed pro-US-propaganda hook, line and sinker."
Most countries have similar propaganda, by the way.
Most countries have some kind of story they tell about themselves. New Zealand doesn't have any illusions that it is a superpower. It doesn't have the resources, the talent pool or any of that to even begin to dream of it. That is natural. If New Zealand was powerful, then it would find itself in a position of greater responsibility.
You cannot compare countries that barely have the option of ambition with something like the US and even begin to imagine that it is meaningful. You have the UK, France, Germany, Japan, Russia, China and the US. You can rewind history to name other civilizations.
These kinds of countries are the only ones that matter, because they're the ones that have to answer about what people were thinking when they chose to make use of their power in a way that is relevant on a larger scale. They have to answer about what the reasons were when things went wrong and whether they agree things went wrong. If people are even allowed to talk about it.
It's very cheap to label anything as propaganda without taking the time to appreciate whether it has any merit in terms of the overall behavior of a country or its people. You can always find counter-examples, but how influential are they in the larger picture?
> These kinds of countries are the only ones that matter, because they're the ones that have to answer about what people were thinking when they chose to make use of their power in a way that is relevant on a larger scale.
This is an extreme claim, and incorrect. If you'd like to see counterexamples, you'll see many corrupt regimes in Africa, which did extreme harm to their own people. You'll see many regional powers.
One does not need to be a global superpower to be good or evil.
There's also nothing magical about Germany or Japan. Many countries had similar resources. Both chose to invest those resources into industrial militarism.
One can make the argument for a handful of countries which our outliers for land area or population, but in general, if any country chooses to invest in military and attack its neighbors, it has good odds of success.
> You have the UK, France, Germany, Japan, Russia, China and the US
Germany was an outlier, on the evil end, but otherwise, it's a selection of which facts one picks.
A comparison would require deciding which facts to compare on. For the US, the "evil" argument comes back to things like slavery and the genocide of the native peoples.
One can pick hundred of examples like Guantanamo Bay, fake vaccines in Pakistan, the Tuskegee Syphilis Study, the Tusla massacre, police violence, corrupt court, ...
The US does pretty nasty things, even if they don't always make US news or grade school textbooks.
> It's very cheap to label anything as propaganda without taking the time to appreciate whether it has any merit in terms of the overall behavior of a country or its people
This is an ad hominem, and a poorly placed one. You're discounting what people are telling you. A lot of the people here went through the US school system, learned "US rah rah rah" propaganda, and only deconditioned themselves as adults.
Many of us were where you are when we were younger.
It's largely correct, but you're missing the point. For many countries, it does not matter whether they are good or bad on a grander scale, because their power is so limited that they are largely confined. Countries that exist on the lower end of the intellectual development scale are expected to fall apart.
Not all countries are structured the same, are at the same stage of development, or resourced enough enough to arrive at a sense of responsibility.
That is why these countries are different. The US is the most different, because it is the only one of its kind. There is no other country you can compare the US to.
A lot of negative things about the US are being misrepresented or inflated by propaganda without sufficient context. It isn't all positive, but the US is a huge country and our history didn't begin 250 years ago.
> One can pick hundred of examples like Guantanamo Bay, fake vaccines in Pakistan, the Tuskegee Syphilis Study, the Tusla massacre, police violence, corrupt court, ...
I could come up with way more examples than that, but for each example you have to ask the questions. You don't simply label an event and point at a face to say "bad, they did it and this is who they are". Is it true? Why did they do it. What were they thinking. Was it an accident, intentional, what was the context? What else was going on in the world, and what was the world like then? What were they up against? Was the choice easy or hard? Was the badness mitigated, or was it unmitigated badness?
> You're discounting what people are telling you. A lot of the people here went through the US school system, learned "US rah rah rah" propaganda, and only deconditioned themselves as adults.
Your unbalanced schooling is not a compelling argument for why I am wrong as much as it is a piece of evidence for why you lost perspective when you got older.
The US doesn't have black sites anymore and when it did, the interrogation techniques were chosen to avoid physical harm. The results were bad, we didn't like it here in the US even if they were extreme measures for extreme times and so we shut it down. It had a high error rate and generally didn't reflect what we thought was right.
Meanwhile the CCP regularly abducts its own citizens and executes more people than the entire world combined.
> The usage of the output is probably considered legal. The usage of the service for that purpose may not be, and using it at scale in a dishonest way is not
This is literally what the "training AI on copyrighted works is just like a human learning/getting inspired" crowd has been arguing though.
Literally. People have been literally saying that it was wrong because they did this "learning" at scale in a dishonest way.
In some ways it's an offshoot of the honest benefit of search engines already crawling all this content. That has its own conflicts, like just how much of a page's content should you reproduce in the results before it's basically considered stealing their content without benefiting the site itself.
There is a balance to strike, both in search engine fair use cases and AI fair use cases. The major cloud LLMs do double as web search engines now, though they didn't originally. In many cases there's no reason left to click the links they sourced from.
That is a legitimate concern. At least within the US, I think there are nuances around fair use and contract law. A lot of companies are getting paid for having their content used in these models, but many websites had no particular rules you had to abide by and the content was simply public. I think if you're operating under an agreement, then even if there is fair use or public domain content being reproduced by the site you are still bound by that agreement.
Similar to old paintings digitized and hosted on some museum website. It's 300 years old, right? It should be public domain, yet the people who digitized it or provided a service to give you access have some say in how their reproduction can be used. These AI services are obviously very different, but there are laws that can govern how you are allowed to use a service if that service has laid out acceptable usage.
I'm not exactly comfortable with the mass scale that everything was soaked up to train these models even within the umbrella of search services, but I also admit that a lot of the usage was probably quite legal. The potential displacement caused by the resulting trained models on artists or writers is almost its own facet. In practice, whether they ONLY trained on strictly legally acquired fair use content with no errors and paid agreements to acquire even more content than they already do or not, there was enough legally accessible information for fair use that there was no escaping some kind of impact on artists, writers, etc.
With any luck, artforms and skills impacted by technology will adapt and continue to be valuable instead of complete displacement or the dilution of opportunity.
Well it was also problematic when the search engines started quoting the websites in such a way to disincentivize people from visiting the actual website.
> At least within the US, I think there are nuances around fair use and contract law.
The concept of "fair use" as it exists in the US-law system is completely dysfunctional (see e.g. nearly every educational music channel on YouTube), so utterly biased to favour large corporations, that there's very little room for whatever "nuances" you believe exist.
> Similar to old paintings digitized and hosted on some museum website. It's 300 years old, right? It should be public domain, yet the people who digitized it or provided a service to give you access have some say in how their reproduction can be used.
Yes 300 year old paintings are public domain. Indeed there are certain rules for the people/institutions who digitize them. It's not "they have some say", there's actually nothing mysterious about it and it is not similar to Anthropic's copyright heist at all because nearly all of the books they copied were not more than 100 years old.
> there are laws that can govern how you are allowed to use a service if that service has laid out acceptable usage
well where I live, there are laws about what a "service" can claim to "lay out as acceptable usage" instead of the other way around ...
> I also admit that a lot of the usage was probably quite legal
Let's disagree on that. I think it wasn't a lot and the vast majority was not legal. How do you think the LLMs "learned" to speak all these non-English languages? Unless your point is that it's probably quite legal to treat foreign IP like that. Which it may very well be in the US, especially if the corporation is large enough, but imvho it's still wrong.
> With any luck, artforms and skills impacted by technology will adapt and continue to be valuable instead of complete displacement or the dilution of opportunity.
And with any bad luck, these AI corporations will hold frontier models hostage for the rest of time.
oh no, the company that illegally used every possible media they could get their hands on is crying that some other company is doing something potentially shady but not illegal? And using that excuse to put in place hidden surveillance systems on their customers?
People keep throwing this idea around haphazardly, but U.S. courts have pretty consistently decided that training on copyrighted works falls under fair use. You may not like it, but that doesn't make it "illegal".
You have to admit that "downloading every book ever written for free from a repository of books that is itself illegal to compile and to run, in order to write a text generation tool" being legal is at least unintuitive, to put it mildly.
It wasnt, that's why they paid a >billion dollar settlement over it, and now license/purchase them. I don't know if the people distilling are licensing those books/etc today, though
Clearly paying that fine didn't do anything to stop Anthropic from doing it again.
Buying a book doesn't make it legal to publish lossy compressed copies of it.
Also, the vast majority of authors whose work was copied against their wishes didn't receive any of that fine.
It sounds like your argument is that they paid a fine for breaking the law, and therefore it is okay they reap the benefits of breaking the law and are allowed to continue to do so?
> The looser use of IP (eg, any characters/celebrities in AI video models) is increasingly mentioned as an advantage of overseas models.
UHmmm you remember when Sam Altman changed his profile pic to look like a Disney version of his own face? Yeah neither do I.
Clearly US AI models are playing loose with the use of overseas IP just as much, and even publicly flaunting it, as if US-based IP is more worthy of protection but Gibli can suck it.
The grandparent claim was that they were surprised downloading books was legal, I was saying that it's not, as they did need to pay.
Whether the law is enough is another question (some cases are still ongoing), and whether the courts are awarding it widely enough is another, but they are facing genuine legal backlash that international firms aren't right now and are more cautious. Several billion is a genuine cost that can move their prices higher in a time of strong competition (see also other announcements with media firms, it's not just books).
It's true that "in the style of" (eg. Ghibli) is not currently legally protected, only actual character IP or using the Ghibli name. That's not inconsistent with US IP treatment.
I don't think Anthropic argues that distillation violates copyright. AFAIK, their position is that it violates their terms and conditions for interacting with their servers.
I think Anthropic will argue whatever argument is likely to protect their interests. I don’t expect anything consistent or moral from them. My quibble is with all the Anthropic fanboys who repeat this crap.
I'd think the problem with fanboys is that they don't care about the truth of the underlying arguments. They just want to score points for their team. Do you have a different issue with them? If not, why not engage with the arguments yourself?
I know from reflecting on my own beliefs now compared with 10-15 years past that one's beliefs can change, and I don't want to be so cynical as to say that these fanboys don't actually believe what they are saying and only want to score points. I'm sure there's a great deal of commentary that is astroturf, but I think there are plenty of (hopefully young and naive) techno-optimists who sincerely think companies like Anthropic can do no wrong and only move humanity forward, or something like that.
In any case, online debate is not always about changing the mind of the single person you engaged with. To some degree, its performative debate so that other readers may be influenced by your ideas.
To be clear, I'm not saying that they don't believe what they say, just that they don't evaluate arguments critically. It's easy to agree with every argument that supports your existing position, but it's a recipe for disaster and they should be taken on a case-by-case basis.
>To some degree, its performative debate
No offense intended, and I'm certainly guilty of this myself at times, but this is a pretty gross way to talk to other people. It's certainly antisocial on the individual level and I think it's also pretty destructive on the community level- I look to Twitter as a case study here, which flipped from a left to a right-wing echo chamber without ever touching anything in between, which I blame on the design, algorithm, and culture being built around performance. Dunks are not truth-seeking behavior, but they perform very well.
For someone who finds this "not unintuitive" you sure are confused!
"Just like I can learn from a book" - ok. Are you allowed to go to libgen and download a book in order to learn from it, because learning is a fair use?
Maybe "U.S. courts have pretty consistently decided" used to mean something, but I don't think the opinion of US courts should be the standard for anything, anymore.
The courts have never said piracy, which is how the training sets were originally built, is legal. There are several court cases still ongoing over this.
There are plenty of good reasons to not use Anthropic's services. If you don't like their terms of service, do stop using them! I personally think Anthropic's increasingly successful attempts at regulatory capture are even more distasteful.
Oh Anthropic has shown their ugliness in more ways than one I agree. You have to have to done some pretty heinous shit for openAI to look good in comparison.
I miss the days when tech people were copyright skeptics. Remember when everyone was upset with Disney for our perpetual copyright regime and the destruction of public domain?
Now many tech people are copyright maximalists and 100% converted to the church of Disney. It’s depressing.
I don't think that's right. The problem is that Anthropic is hoarding it and that's hypocritical. If copyright doesn't count for Anthropic, they should publish Claude. If they wanna hide Claude behind copyrights and/or TOS, they don't get to screw with other people's copyrights and TOS and then profit from it.
To call that opinion "copyright maximalist 100% converted to the church of Disney" is, at the very least, hyperbole.
Except when it comes to image models. Imagine is extremely good and extremely cheap. I've been using it to generate book covers for ebooks (old novel short stories that never had a cover, for example) and it's phenomenal. Each cover is about 6 cents
I really doubt other labs are distilling Claude using the Claude Code CLI when they can way more easily use the API directly.
I also don’t get why the « protection » on ANTHROPIC_BASE_URL. If I change it to use a Chinese model, the Chinese model will not care at all about the modified prompt. On the contrary, if I’m distillating (which again, using CC CLI would be stupid), I’m not going to change ANTHROPIC_BASE_URL.
The obvious response is the realization that spending trillions on training LLMs is not a viable business model if they can be distilled for a much lower cost.
What does that have to do with CC? I'm not commenting on that being good/bad/legal/illegal, but CC is separate from the models. If they really are doing this maliciously it is because they are trying to ignore my 'CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1' flag (if that still means anything).
The article seems to state as much minus the obfuscation. However justified they are to respond, this can be a slippery slope. We're bound to hear more reports of hidden user data exfiltration.
The have been fucking distilling our websites and writing, even when behind TOS, aggressively bypassing protection mechanisms. They they can fuck right off