> It's even being seen in image generation, where NatGeo cover images are reproduced in their entirety or where stock photo watermarks are emitted on finished images.
Can you cite sources? I’ve heard this claim repeatedly but have yet to see a good example.
The AI learns to generate watermark-like things, because those exist on a large fraction of its inputs.
It doesn't mean the rest of the output existed in the training set. It's perfectly capable of generating a completely novel picture, then slapping a watermark on it.
My point is that the watermark is deterministic. It is taken whole-cloth from an input, and reproduced as-is on an output. Thus, it is able to be attributed.
Stable diffusion is more like a paintshop artist who grabs bits and pieces of other art and melds them together, and less like a painter who creates from their imagination.
> It is taken whole-cloth from an input, and reproduced as-is on an output.
I respectfully disagree with this claim. It’s not as-is. It’s remarkably similar. That’s a big difference.
> Stable diffusion is more like a paintshop artist who grabs bits and pieces of other art and melds them together, and less like a painter who creates from their imagination.
I also disagree with this. The uncomfortable truth, imho, is that what Stable Diffusion does is FAR closer to what human artists do than we’d like to admit.
Human artists are perfectly capable of reproducing copyright infringing images. They just generally choose not to for legal/moral reasons.
> It is taken whole-cloth from an input, and reproduced as-is on an output
This is only possible for elements that are repeated hundreds or thousands of times. It's generally considered a failure of training data curation, if it's anything more than watermarks. The AI can learn how to reproduce the watermark, while being literally unable to reproduce the pictures it was on.
Can you cite sources? I’ve heard this claim repeatedly but have yet to see a good example.