> but the act of training a neural network on datasets like this has already been decided to be legal
?
Has it?
Look, you can argue about whether it’s morally right or not to use models that are fine tuned explicitly to copy the style of someone else, trained on their art without their consent, to make a model that can generate images very similar to the training images.
You can argue about technically of that’s copying, or if lossy compression is copying.
…but legal and moral are different things, and right now, as far as I’m aware:
- it’s only legal because there are no laws specifically making it illegal currently.
- there are active (eg. Copilot) cases challenging this to set a precedent.
- it’s sufficiently ambiguous having a model that anyone can type “a naked picture of a 12 year old” in and get exactly that as output, that stability has nerfed the most recent mode release they’ve done.
- there is a reasonably obvious similarity to other fields where a thing is itself not illegal or bad, but it can enable people to do illegal or bad things, and therefore access, ownership and usage of said things (eg. Hand guns) is heavily legislated.
I think “this has already been decided to be legal” is a blatantly false assumption.
Yeah, we'll see how the courts come down on this one. But If we follow exactly the same process as you describe using a human instead of a computer it seems like that would be fine. Artists copy each other all the time.
Also, this looks to me like a situation where generating art is suddenly so cheap and easy it doesn't matter what the courts decide. People can and will ignore the law, because it is trivial to generate new pictures and it is cost ineffective to enforce any restrictions. We're going to see this tech take off. Who knew that art would be the next thing software took out?
This is basically my take - I'm mostly confused by the people up in arms about this (except that it likely makes their work more of a commodity, so I get that fear).
People look at art and make art in that style all the time, now a machine exists that can do that. Why is that unethical? Because they didn't consent for the machine to look at it/learn from it, but they did for humans? I don't think this argument will be able to hold the wave of change that's coming from this new capability. They'd be better off long term learning how to use it.
Nobody creates purely original things in a vacuum, machines won't either.
Automation is okay if it is about taxi drivers or warehouse workers. But when it targets artists some people turn Luddite.
My brother is an artist by the way who does a lot of work with AI, 3D printing and internet. Art will always survive and adapt.
If memory serves it took decades before photography was socially accepted among the art world.
The question now is what value can (human) artists bring besides merely producing images of a certain subject in a certain style. Software has clearly just solved that problem, although the buildup was the last 10-20 years (cnns, gans, style transfer, and now generative language-based models).
But when I think of the value and interestingness of art, there's a lot of intention and meaning in choosing what to draw, the form, etc.
Even from a first glance, the one by David is clearly superior. The scene is so much more striking and interesting to look at. You can see the emotions of the characters, and overall the composition underscores the significance of the philosopher's death, especially in the context of the enlightenment/romantic period. (This is my opinion as a pleb; not an art expert.) But what I do understand, as a computer person, is that these concepts are still beyond what a model can encode in an image as of today.
So although humans are no longer superior in the mechanics of producing images, I think in the higher-level/psychological aspects of "art", there's room for humans, at least for now.
Want to add, I understand the debate is also around the revenue loss in the commission/fanart/online art scene. But from having looked at lots of these over the years, I'd still argue the same thing: that IMO the really good series and artists are good because of their ideas and themes, and not their technique. But if a particular artist's revenue is 90% from drawing lewd fanart, unfortunately it seems like they'll have to adapt and compete, using the skills that humans are still dominant in.
And although I'm just an dumb anon on the internet spewing these ideas, I know I sound harsh, but I think I'm correct. Because the reality is that now, everybody's downloaded the SDv1.4 weights onto their hard drives, and the cat's out of the bag permanently.
I have never used any of those services before. I have never considered using one of the services before because it was always outside my price range.
Now that I have tried stable diffusion and have photo bashed and rendered some concept art for each of the main characters in my novel, I now want to commission an artist to create the 30 or so needed training images so I can ask stable diffusion to spit out my own character in various poses and expressions.
At minimum, that will require a human to render a model sheet of the character from front and back and side and 3/4 and above and below, as well as the emotions on the basic emotions wheel.
If I want the character to be able to wear different outfits, then I will also need to pay for renderings of that character wearing that clothing, all in service of trying to train stable diffusion to be able to remix that character into future images.
Let's also say that I do not have a killer graphics card to be able to train images into a model, luckily, for another $100 fee, the artist will use their existing graphics card to spit out an embedding or a hypernetwork or a VAE or whatever it is that you can use to add custom training to a model and send me that as well as the original set of input photos.
After all of that, I can generate the photos I want...but I will then have to slightly tweak each image so that it has human authorship, even if it's just removing noise and fixing the cursed loops that happen on limbs at times.
In short, I am considering something that is at least $200 for the crappiest cheapest artist out there, multiply that by my six or so main characters, and that is money that I am genuinely considering spending that I would not have even dreamed of entertaining for a moment.
The proof is in the pudding as to whether this thought process will be happening for other stable diffusion users who are able to get images they like, but do not have good rendering skills on their own, weather they too are willing to pay for this or not.
If so, there will instantly become a new type of artist job available, that of the AI art trainer artist.
At the very least, there will be AI art cleanup artists that remove the noise and so-called cursed elements of ai art when used in the concept art stage.
It’s really not that simple. If an artist picks 12 unusual colors to paint a sunset and you copy those exact same colors to also paint a sunset in his style then no that’s not ok.
People on HN have very strong options about this stuff without looking at any of the relevant case law.
I have no idea if that is legal or not, but it is an extremely common practice. For example, that is the scenario when someone on Deviant Art creates fan art of any cartoon.
And why should that be a problem? This is a King Canute and the tide scenario. We may as well bow to reality and admit that it is ok.
Now it’s so easy to share music, it’s meaningless to try to stop people doing it.
Apply 20 years of law cases and punishment for random people and now…
…everyone pays to stream their music.
Right? Wrong? Eh.
I’m just saying, you are kidding yourself if you think that the Powers That Be will just let people decide copyright isnt a thing anymore because of (insert reason here).
Once there is money involved, there will be court cases, and you know, I’ll be shocked if a combination of “needs bigger GPUs to run” and “legal issues” don’t cause these sorts of models to be locked away behind cloud APIs in the future.
It is what it is. Enjoy it while you can; some things are quite predictable, and:
“Law takes a while to catch up with new technology, but it eventually does, and when it does it favours the status quo”
The models are already locked away behind cloud APIs. Stable Diffusion wasn't supposed to happen; OpenAI thought that nobody else could afford to train a U-Net on CLIP at their scale and give it away for free.
I will point out that the usual copyright maximalists have been pretty silent on the issue of AI art. The biggest opposition to AI is coming from the Free Software community - i.e. the people who want to abolish artists' ownership over their work outright.
The Free Software movement originates in academia and has academic value - they don't care about receiving direct compensation, but attribution is critical, and attribution is what image generation models can't provide.
You don't need to pay to listen to music if you don't want to. The corpus of good music on YouTube for free is probably bigger than what you can listen to in a lifetime. If it isn't already it will be in time.
Napster's model won that war. Effortlessly. If you're paying for your music, you are paying on your terms based on the value you think is being provided to you. It isn't a legal framework making you do it.
I think your example doesn't quite work here. People pay for music now because handing Spotify or whoever ten dollars a month is a easier then torrenting, easier then managing directories of mp3 files and moving them to your phone, and has value adds in the form of discovery.
Spotify won because it's better than The Pirate Bay. The Playlist on Netflix explains why Spotify surpassed TBP and how the record labels still had to surrender even though they "beat" the pirates.
The difference between music and AI art is that music is still made by humans. The music industry currently has a monopoly on producing new music, so of course they have leverage to get paid for that new music.
My point is that Napster-to-SD is not a fair comparison. Napster didn't eliminate or replace human artists, while SD certainly did. Therefore, it's not fair to assume the government can regulate themselves out of this one like they did with Napster. Because even though Napster enabled widespread piracy, the human musicians still had leverage in that they were needed to create new music.
So my answer to the question posed by parent comment of "will SD play out the same way that Napster did?" is "no" because there are fundamentally different economics at play.
Obviously, if a good generative model comes out, the music industry will be in a similar boat to the art industry right now. Google was working on it in 2017 (https://magenta.tensorflow.org/performance-rnn) but I don't know if they've made any progress since.
> This is case with anything…anything not illegal is legal, at least in the US…
I think you're dancing around with words here. Let me be super specific:
There's no law specifically saying I can't use a jelly coated toasted to beat someone to death... but there are existing laws that cover 'beating someone to death'.
If you beat someone to death with a jelly coated toaster, you might argue for some obscure reason, your actions are not covered by the existing legal framework around beating people to death, but mostly likely you will not be protected by the claim you 'didn't know it was illegal' to do that.
...because the courts have not ruled that beating people to death with jelly coated toasters is legal.
ie.
a) There is no specific legal precedent or law around something but other laws related to it exist
and
b) Something has been determined to be legal by some precedent / law / whatever
Are not equivalent.
Regardless of what people want to believe, or have opinions one way or another, the assertion, made in the original article, that (b) was true, is not true.
Only (a) is true, and that is, legally, a much weaker statement.
Which makes it legal. When the courts interpret law it addresses a case and informs going forward. I don't believe, for instance, people can be retroactively tried for abortions during the period Roe V. Wade stood.
Pedandtly, this is not true. You might think something you’re doing is legal, but when someone takes you to court, it’s because they’re arguing that what you’re doing is already illegal. And if the court sided with them, then what you were doing was illegal the whole time. Courts don’t make laws, they interpret existing laws. Everyone can interpret laws in their own favor, but it doesn’t make things ‘legal’ in the definitive sense
Roe v. Wade is a ruling that people have a right to abortions. Removing it removes the right. It doesn’t retroactively remove the right.
A new law that is passed doesn’t have retroactive power
A court ruling that the thing you’ve been doing for a year has been breaking an existing law can absolutely punish you for it. But it’s unlikely to punish random individuals doing things on a non noteworthy scale when a clear understanding of the law in a new context has not been established.
I'm not sure this is true; it addresses a case, which means the interpretation already is ex-post-facto. I don't know the answer, but it wouldn't surprise me if, in states with anti-abortion laws standing for which the statute of limitations has not expired, abortion providers could be held legally liable.
Ex post facto laws are forbidden by the US Constitution, both at federal and state levels. [1]
At least, at this time, with current interpretation- "the Supreme Court has explained that people must have notice of the possible criminal penalties for their actions at the time they act" [2] (See also Weaver v Graham[3])
The laws are not ex post facto. They have been on the books, on some cases for over 100 years. States were not enforcing them, because they believed they were constitutionally prohibited from doing so.
I wouldn’t be so quick to say it’s legal because google image search is legal. Image search is legal because “thumbnails” and “image search” were decided by the court to be sufficiently transformative and in a different marketplace from “images.”
I could very easily see it being argued that “computer generated images” are in the exact same marketplace as “images” so already that case wouldn’t apply and new legal reasoning would be needed.
Legal departments are employed to tell you about legal risks, not to predict the future. Your corporate lawyer isn’t going to go “yeah that’s illegal but it’s totally awesome bro”. The execs might but not the lawyer.
In this case, the riskiest issue isn’t copyright, it’s that it can generate NSFW/CSAM, which governments and payment processors both get really upset about.
It’s not about predicting the future. The point of communicating legal risks is to make assessments about which risks might be worth taking because complying with every single law is impractical.
It’s part of the basic structures where the legal team has an advisory role.
It’s factual in much the same way the average person breaks multiple laws as written every day.
Much of this is simple ignorance and many thing aren’t particularly relevant. 14 states still had sodomy laws in 2003 when the Supreme Court reversed its stance and declared them unconstitutional. At this point there are hundreds of years of crap at the federal, state, and local level much of which changes based on where you happen to be.
What percentage of the US laws have you actually read?
I see two ways of interpreting "companies break the law constantly".
One, that there is always a company out there somewhere acting criminally. If that's the intended meaning, it is a factual statement, but doesn't not carry the original implication that companies do not try to avoid breaking the law. For instance, the fact that there is always a human out their somewhere committing crime doesn't not mean most humans do not actively avoid such.
Two, that any given company breaks the law often. This carries the implication that companies do not worry about breaking the law, but is also factually incorrect.
What I find weird is that humans do the same and no one mentions that; I have 2 professional (they have lived of it for decades) artist friends; one makes Giger art (paintbrush, exactly the same style but no copies, original but if you see it you are going to say Giger) and another one Vermeer, same thing. There is no discussion if that’s moral or legal, what’s the difference? The Giger one has bought a farm of GPUs a few months ago and is adding training to models for animations. He loves it.
Same with copilot; people copy shit from GitHub and SO all the time without mentioning copyrights and a lot of code people cough up is just ‘stolen’ from someone without remembering who/where it was; what’s the difference?
I'd say there's a good chance both of those artists are going to grow "beyond" the artist they are emulating (and by "beyond" here I don't mean better than but rather grow into something unique).
Lots of bands (nearly all?) learn their craft by playing covers. At some point the artists in the group start to find their voice.
In a lot of ways, whether it has or not is moot in the medium-term.
Court cases / legal clarification may outlaw Stable Diffusion or DALL-E for having been trained without the consent of the original artist on their copyrighted artworks. But if this technology proves valuable and viable, the likes of Disney, WPP, and Omnicom Group will pay a few dozen artists to create enough work to seed an engine that can generate 50 million wholly-owned lookalikes.
Copyright protection will ultimately shield the megacorps from competition by the common folk, not artists from competition by corporations.
> it’s only legal because there are no laws specifically making it illegal currently.
No, it's only legal because nobody has litigated a case all the way to the Supreme Court. We all thought that APIs weren't copyrightable, and then, suddenly they were (or worse, were "assumed to be" but not explicitly ruled upon).
Although not being explicitly illegal does make something legal, the phrasing "has been decided to be legal" suggests a debate, a conscious decision, and perhaps an overturned law. That is not the case as far as I know. It is legal (because that is the default), but no decision has been made.
AFAIK, from the discussion here, training the network has been established as fair use on the US. But that doesn't apply to use the network's results in any way.
As far as I know that hasn’t been established yet but is presumably what would be decided (especially in an academic environment). But yes the output is a totally different question.
> ... it’s only legal because there are no laws specifically making it illegal currently.
One might argue that any unlicensed use of copyrighted material is misuse (unless the courts have ruled it is fair use.)
Only the copyright holder may bring action against a potential misuse. Typically, individuals tend not to have deep enough pockets to hire the lawyers to bring the action. And during discovery it may be found that the model was trained on works-for-hire made by humans replicating the original style. Or appeals may continue on at more legal cost, without guarantee of recompense.
Most will decide it’s just not worth it. Until a company similar Getty buys up all the copyrighted material and, in addition to sensible suits, brings spurious legal action against individuals and turns the tables.
Yes. When artists say things like "this really feels like it should not be fair use" we are dismissed with "well it's legal!". Changing what "fair use" covers is a possibility.
Unfortunately art is a hard business to make a lot of money in and the vast majority of the people unhappy about this do not have the money to mount any kind of legal challenge to the corporations developing these databases, so we are basically fucked.
Which is in some ways nothing new, the average Internet user has no concept of creator's rights and will blissfully assume that if something's not behind a paywall, it's free for any use. It feels even shittier and worse when this is being done on an industrial scale by these AIs, though.
Even if artists could sue this is assuming there is someone making enough money to sue.
Open source models will likely be dominant (this was how Stable Diffusion got popular, it's not clear their new neutered 'safe' model will be as popular or a successful business).
This is the new reality even if the bigco gets sued out of business. Artists need to accept that reality eventually. Just like all "panics".
Fortunately AI generators are not a direct replacement for artists, I've seen it used as part of the design process, but that still takes talent. Maybe it will get better but it's not a one-stop shop for business use-cases. More likely it will be AI+artist, not AI replacing artist.
If an open-source model gets DMCA'd, then getting a copy becomes a hassle. This does not entirely stop things - people still have massive collections of Nintendo ROMs and downloaded films, but it's work to get it running in a way that, say the Stable Diffusion tool that Clip Studio Paint just announced is not.
A real open-sourced model trained on images that were explicitly licensed for such uses would be interesting. It would cost money but it would also not be hiding a ton of its actual cost by training it on images whose creator never imagined that "being dumped into a vast dataset" would be a thing that would happen, as well as images that were fine in their original context as "fair use" but got sucked in by the web crawlers feeding the dataset. You want to train an image-generating machine on my work? Ask my fucking permission and if I say "pay me" then you either negotiate a price we can both live with, or you do not put it in your dataset.
I do not accept your reality. Sue 'em all until you've gotta train your own damn datasets, or buy ones that have worked out the proper licensing. If Photoshop can be made to refuse to scan currency, then these things can be made to be fair to the artists whose shoulders they rest upon.
In the EU, training a model on copyrighted data does not infringe copyright. This is explicit, codified law as of the latest EU Copyright Directive[0]. In the US, AI researchers are more or less hoping that Authors Guild v. Google forms controlling precedent against lawsuits against the AI companies. While this is not done-and-dusted case law, I can see the logic. The model is capable of, and is intended to be capable of, creating novel art.
In neither case is it legal to use outputs of an AI that infringe an already-existing copyrighted work. This means anyone using it to generate production-ready artwork is exposing themselves to potential legal liability if their AI winds up regurgitating training set data. This is what I worry about way more than just "are the model weights infringing copyright".
As for morality:
- Absolutely none of the current generative art systems are trained from scratch on ethically-sourced datasets. The AI companies just assume that because they need ungodly amounts of training set data, that they are morally entitled to get it, because they spend more time staring into the eyes of a basilisk[1] than worrying about the world they already live in.
- Multiple companies getting into generative art have gone above and beyond in giving human artists the middle finger. DeviantArt decided to blow all their good-will from defending against NFT nonsense by making their AI training program opt-out, which is NOT HOW CONSENT WORKS. Mimic[2] and Dreambooth are basically artistic impersonation tools that have specific moral implications beyond the general practice of AI art.
Whether or not this becomes actual law is... well, I'll put it to you this way. Stability AI's Dance Diffusion is actually trained on an ethically-sourced, public-domain dataset. Why? Because they're afraid of being sued by the RIAA. Despite the name, copyright maximalists are less about maximizing the rights of artists and more about building moats around large publishers. And generative art does not threaten[3] those publishers, so it will not be banned.
[0] Yes, the same one that mandated upload filters on video sites.
[1] Roko's Basilisk posits the idea of a superintelligent AI - one that can eat everyone's brains and emulate them perfectly - constructing a perfect hell for people who didn't build it as a way to threaten those people in the past to build it.
Longtermists can burn in computer-generated hell. Time discounts exist for a reason.
> Section 23 (Adaptations and rearrangements): "(1) Adaptations or other rearrangements of a work, especially a melody, may only be published or used with the consent of the author. If the newly created work is at a sufficient distance from the work used, it does not constitute an adaptation or rearrangement within the meaning of sentence 1."
This amounts to the same as what TFA says about derivative works in general. If it's too close to the original, it may be infringing copyright. Gauging the boundary between "too close" and "not too close" is the bread and butter of copyright courts.
> Section 44b (Text and data mining): "(1) Text and data mining is the automated analysis of one or more digital or digitized works in order to obtain information, in particular about patterns, trends and correlations.
(2) Duplications of legally accessible works for text and data mining are permitted. The copies are to be deleted when they are no longer required for text and data mining.
(3) Uses according to paragraph 2 sentence 1 are only permitted if the right holder has not reserved them. A reservation of use for works accessible online is only effective if it is in machine-readable form."
Again, IANAL, but in my opinion this covers neural networks. "Obtaining information about patterns, trends and correlations" is as close to a literal description of the purpose and function of artificial neural networks as you're going to get in legalese. Paragraph 2 just means that you have to delete your data-mined stash if you ever deem your model complete (but it should not be too hard to argue that ML training is always an ongoing process and thus the mined data will always be required). Regarding Paragraph 3, if I'm not mistaken, the "machine-readable form" part of paragraph 3 is very very very specifically aimed at robots.txt. That is, if you allow your art to appear in search results (which most artists want), then it's also fair game for data mining.
Once again, IANAL, and this is only German law. But I think it's very easy to make a case here that this existing code of law covers AI training as well, as long as measures are taken to ensure that source images cannot be reproduced with any sort of fidelity (and obviously that's a big "if").
IAAL. Regardless of how the Copilot case develops at the trial court level, anyone who thinks the Roberts' Supreme Court is ultimately going to kneecap this technology needs to reboot.
YEah, I think people also need to give up on saying "well precedence says" with the current SCOTUS roster. They don't actually care about precedent when it doesn't jive with their agenda. It's hard to believe that the mightest lawyers in the land think precedent is unimportant but we've already seen it a few times and this SCOTUS is just getting started.
There is a lot of case law on what a derivative work is. A model that outputs an identically image to the training doesn’t mean the trained model is infringing or a derivative work.
?
Has it?
Look, you can argue about whether it’s morally right or not to use models that are fine tuned explicitly to copy the style of someone else, trained on their art without their consent, to make a model that can generate images very similar to the training images.
You can argue about technically of that’s copying, or if lossy compression is copying.
…but legal and moral are different things, and right now, as far as I’m aware:
- it’s only legal because there are no laws specifically making it illegal currently.
- there are active (eg. Copilot) cases challenging this to set a precedent.
- it’s sufficiently ambiguous having a model that anyone can type “a naked picture of a 12 year old” in and get exactly that as output, that stability has nerfed the most recent mode release they’ve done.
- there is a reasonably obvious similarity to other fields where a thing is itself not illegal or bad, but it can enable people to do illegal or bad things, and therefore access, ownership and usage of said things (eg. Hand guns) is heavily legislated.
I think “this has already been decided to be legal” is a blatantly false assumption.