Home AI & Web & Technology When AI art has no author: Study finds generated images often can’t be traced to training data – MIT News – Massachusetts Institute of Technology

When AI art has no author: Study finds generated images often can’t be traced to training data – MIT News – Massachusetts Institute of Technology

0
When AI art has no author: Study finds generated images often can’t be traced to training data – MIT News – Massachusetts Institute of Technology
At top left, a painted portrait of a man's face. More than 200 extremely similar images are shown in a grid next to the larger one.
MIT CSAIL researchers found that at large scales, you can often remove any single image from AI training data, every image by a given artist, or every photograph of a given person, and the generated output won’t change appreciably. At top left is an image generated by a model trained on public domain artwork created by 744 artists. The others are a sampling of images that would have been generated had any one of the 744 artists been omitted from the training set.
Credits: Collage courtesy of the researchers, showing images generated by AI.

MIT News | Massachusetts Institute of Technology – On Campus and Around the world Subscribe to MIT News newsletter Browse

When AI art has no author: Study finds generated images often can’t be traced to training data

A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.

Rachel Gordon|MIT CSAIL, Publication Date: August 18, 2026

Press Inquiries

When an artificial intelligence image generator produces a portrait, whose work went into it? The question sits at the center of lawsuits, licensing deals, and proposed regulations worldwide. Artists want credit. Companies want clarity. Policymakers want a way to assign responsibility.

New work from a team of researchers at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL) suggests that for models trained on large datasets, the question may often have no answer. It’s not that the tools for finding it are inadequate. The connection itself has disappeared.

The scientists identified a phenomenon they call attribution decay, where the more data a generative model is trained on, the less any individual training example matters to any particular output. It feels counterintuitive, but at sufficiently large scales, they find, you can often remove any single image from the training data, or every image by a given artist, or every photograph of a given person, and the generated sample doesn’t change.

And if removing something changes nothing, the researchers argue, it can’t be said to be responsible for anything. 

“If you take away a piece of data and the output of the model doesn’t change, then that piece of data didn’t affect the output,” says Zheng Dai SM ’21, PhD ’24, former MIT CSAIL researcher and lead author on the work. “So it doesn’t make much sense to attribute the output to that piece of data. And if you then do this one at a time for every other piece of data and find that the output doesn’t change for any of them either, then it doesn’t make much sense to attribute the output to any one of them.”

“All previous methods were approximate,” says MIT Professor David Gifford, who is an MIT CSAIL principal investigator. “They really could not absolutely show that deleting individual things did not change the output. This paper introduces the first method that is absolute. You’re actually deleting the inputs and deleting all influences of the inputs. This is the first exact method for doing large-scale deletion efficiently and showing that the results don’t change.”

Dai and Gifford’s project is described in an open-access paper published today in Nature Communications.

The retraining problem

Testing this idea directly meant answering a what-if question. What would this model have produced if it had never seen this particular image? Answering it honestly means retraining the model from scratch without that image, then doing it again for the next image, and the next. With millions of training examples, the math quickly becomes prohibitive, which is why prior work in the attribution field has relied on approximations that estimate a training example’s influence, rather than actually removing it.

Read more: When AI art has no author: Study finds generated images often can’t be traced to training data – MIT News – Massachusetts Institute of Technology

Continue/Read Original Article: When AI art has no author: Study finds generated images often can’t be traced to training data | MIT News | Massachusetts Institute of Technology


Discover more from DrWeb's Domain

Subscribe to get the latest posts sent to your email.

Leave Your Comments

This site uses Akismet to reduce spam. Learn how your comment data is processed.