
Credits: Collage courtesy of the researchers, showing images generated by AI.
MIT News | Massachusetts Institute of Technology – On Campus and Around the world Subscribe to MIT News newsletter Browse
When AI art has no author: Study finds generated images often can’t be traced to training data
A new method for surgically removing training examples from a model reveals that as datasets grow, the link between what a model learns and what it produces dissolves.
Rachel Gordon|MIT CSAIL, Publication Date: August 18, 2026
When an artificial intelligence image generator produces a portrait, whose work went into it? The question sits at the center of lawsuits, licensing deals, and proposed regulations worldwide. Artists want credit. Companies want clarity. Policymakers want a way to assign responsibility.
New work from a team of researchers at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL) suggests that for models trained on large datasets, the question may often have no answer. It’s not that the tools for finding it are inadequate. The connection itself has disappeared.
The scientists identified a phenomenon they call attribution decay, where the more data a generative model is trained on, the less any individual training example matters to any particular output. It feels counterintuitive, but at sufficiently large scales, they find, you can often remove any single image from the training data, or every image by a given artist, or every photograph of a given person, and the generated sample doesn’t change.
And if removing something changes nothing, the researchers argue, it can’t be said to be responsible for anything.
“If you take away a piece of data and the output of the model doesn’t change, then that piece of data didn’t affect the output,” says Zheng Dai SM ’21, PhD ’24, former MIT CSAIL researcher and lead author on the work. “So it doesn’t make much sense to attribute the output to that piece of data. And if you then do this one at a time for every other piece of data and find that the output doesn’t change for any of them either, then it doesn’t make much sense to attribute the output to any one of them.”
“All previous methods were approximate,” says MIT Professor David Gifford, who is an MIT CSAIL principal investigator. “They really could not absolutely show that deleting individual things did not change the output. This paper introduces the first method that is absolute. You’re actually deleting the inputs and deleting all influences of the inputs. This is the first exact method for doing large-scale deletion efficiently and showing that the results don’t change.”
Dai and Gifford’s project is described in an open-access paper published today in Nature Communications.
The retraining problem
Testing this idea directly meant answering a what-if question. What would this model have produced if it had never seen this particular image? Answering it honestly means retraining the model from scratch without that image, then doing it again for the next image, and the next. With millions of training examples, the math quickly becomes prohibitive, which is why prior work in the attribution field has relied on approximations that estimate a training example’s influence, rather than actually removing it.
Read more: When AI art has no author: Study finds generated images often can’t be traced to training data – MIT News – Massachusetts Institute of TechnologyContinue/Read Original Article: When AI art has no author: Study finds generated images often can’t be traced to training data | MIT News | Massachusetts Institute of Technology
Discover more from DrWeb's Domain
Subscribe to get the latest posts sent to your email.