5 Sources
[1]
When AI art has no author: Study finds generated images often can't be traced to training data
Collage courtesy of the researchers, showing images generated by AI. When an artificial intelligence image generator produces a portrait, whose work went into it? The question sits at the center of lawsuits, licensing deals, and proposed regulations worldwide. Artists want credit. Companies want
[2]
AI models get convenient amnesia about source material as they grow, MIT boffins find
The process of training an AI model becomes a paradox at scale - the more it remembers, the less it remembers about the source of its memories. MIT computer scientists went looking for a way to attribute AI model output to specific training data, in the hope that understanding could inform AI
[3]
Artists Want to Prove Their Work Was Stolen by AI. A New Study Says That's Impossible
One way to figure out how something works is to take it apart one piece at a time, and see how it changes along the way. That was the approach taken by two AI researchers who wanted to understand how image-generating AI models arrive at their final outputs: Do they reference a particular image, the
[4]
AI may be learning from billions of images without copying any one of them
MIT researchers have identified what they call "attribution decay" in large generative models. Here's an uncomfortable question for the generative AI era: if an AI creates an image, can anyone actually point to the specific images that influenced it? A new study from MIT's Computer Science and
[5]
Does generative AI actually copy artists? Researchers say it's up for debate
Researchers from MIT unpack the idea of 'attribution decay,' and what it might mean for artists looking to protect their work. Generative AI has long been accused of copying artists' work outright (see the numerous copyright lawsuits winding their way through court). But a new study out of MIT
Share
Copy Link
MIT researchers discovered that AI-generated images from large diffusion models like Midjourney and Stable Diffusion cannot be traced to specific training data due to attribution decay. The phenomenon means removing individual artworks or entire artist portfolios from training datasets produces identical outputs, raising questions about copyright lawsuits and intellectual property claims against AI companies.

MIT's Computer Science and Artificial Intelligence Laboratory (CSAIL) researchers Zheng Dai and David Gifford have identified a phenomenon called attribution decay that fundamentally challenges how we understand AI-generated images
1
. Published in Nature Communications, their study reveals that outputs of generative diffusion models become increasingly impossible to trace back to specific training data as model size increases2
. The finding complicates ongoing copyright lawsuits where artists attempt to prove their work was stolen by AI companies like Midjourney and Stability AI3
.The researchers tested this by removing specific images, entire artist portfolios, or photographs of particular individuals from training datasets. The result was striking: for sufficiently large models, these removals produced no measurable change in generated outputs. "If you take away a piece of data and the output of the model doesn't change, then that piece of data didn't affect the output," explains Zheng Dai
1
. This means you could remove the Mona Lisa or all of Leonardo Da Vinci's work from training data, yet diffusion models could still reproduce similar images or styles2
.The MIT CSAIL study introduced the first exact method for testing whether AI-generated content is causally linked to individual training images. Previous attribution methods relied on approximations that estimated influence rather than directly measuring it
1
. The team built a custom architecture called a diffusion ensemble, composed of multiple smaller components each trained on different data slices. This design enabled precise ablation—switching off components that saw specific images without retraining entire models5
.David Gifford emphasizes the breakthrough: "All previous methods were approximate. They really could not absolutely show that deleting individual things did not change the output. This paper introduces the first method that is absolute"
1
. The researchers trained 24 diffusion models on datasets ranging from 256 images to more than 160,000, using public collections including CIFAR-10, CelebA, MetFaces, and ArtBench. They measured differences through counterfactual analysis, creating what they term a "counterfactual universe" for each generated image1
.The study documented a consistent pattern: training data influence decreased along an inverse power law as dataset size increased. The team measured this through the counterfactual radius—the maximum difference between an original AI-generated image and alternate versions created after removing training data
1
. This held true whether differences were measured pixel-by-pixel or by semantic meaning, with statistical significance in both cases1
.The researchers stress-tested their findings by retraining 1,282 separate models using brute-force methods at smaller scales. Attribution decay appeared regardless of approach
1
. Larger datasets contain substantial visual redundancy, with many images sharing overlapping features. No single image becomes responsible for the DNA of generated outputs5
. The diffusion ensemble architecture demonstrated better data efficiency than conventional models, performing particularly well as training data increased1
.Related Stories
The findings carry significant weight for ongoing intellectual property theft cases against AI companies. In the 2023 lawsuit Andersen et al. v. Stability AI Ltd, plaintiffs argue that Midjourney scraped images associated with specific artists' names to enable mimicking their expressive content
2
. However, the MIT study suggests that even when AI-generated images closely resemble an artist's work, proving a direct causal connection to that artist's training data becomes impossible at scale3
.Cornell Law School professor James Grimmelmann notes that current copyright lawsuits haven't focused specifically on whether similar images can be elicited from models. "If attribution worked, it would reliably tell us whether similarities between a model's output and a copyright-protected work are due to copying or coincidence," he states
2
. The research suggests technologists and courts will need alternative methods for assessing copyright in AI-generated content beyond direct attribution2
.Gifford argues the findings raise questions about fair use and whether outputs qualify as copyrightable novel works: "If those outputs have nothing to do with any individual piece of training data, that raises questions about fair use, about whether the outputs are themselves copyrightable as novel works, and about how authors get compensated"
2
. The study also suggests a potential liability avoidance strategy—making models large enough that no output can be attributed to specific inputs2
.The research highlights fundamental differences between human and machine creativity. When humans create art, they typically remain conscious of external influences and may reference specific works directly. Diffusion models like Stable Diffusion operate differently, using the totality of training data to generate images through processes that remain mysterious even to researchers
3
. These black box AI systems don't store copies of training images but instead adjust millions of internal parameters based on patterns across entire datasets5
.The study doesn't mean training data becomes irrelevant. Models may learn composition, lighting, textures, and artistic styles from millions of images while leaving no single image with an obvious fingerprint on final results
4
. As datasets grow, AI-generated images may draw on patterns learned from enormous pools of material without having clear, identifiable source images4
. Watch for how courts adapt attribution standards as commercial diffusion models operate at scales many orders of magnitude larger than the test models used in this study3
.Summarized by
Navi
[1]
[2]
1
Science and Research

2
Policy and Regulation

3
Technology