π€ AI Summary
This study addresses whether generative AI systems, by virtue of memorizing training data, produce outputs that constitute legally actionable βcopiesβ of copyrighted works. Integrating insights from the memory mechanisms and probabilistic generation behaviors of large language models, the paper offers the first systematic interdisciplinary analysis arguing that copyright law should adopt a functional standard to determine whether an AI model contains a βcopy.β The research demonstrates that current legal frameworks typically recognize copying only when specific protected works can be readily extracted from the model, thereby exposing significant limitations in the applicability of existing doctrines in the AI era. Building on this finding, the work proposes targeted legal reforms to better align copyright enforcement with the technical realities of modern generative systems.
π Abstract
Recent work shows that it is possible to extract verbatim or near-verbatim text of some copyrighted works from some large language models (LLMs or models). That is evidence that the model weights encode the works in some form - that the model has "memorized" those works from its training data. But LLMs don't store information in the same format as familiar databases. Rather, their weights store statistical relationships between tokens that have been learned from the training data, and those relationships inform a generation process that is often probabilistic rather than deterministic. In the case of memorization, those relationships are strong enough that, in many circumstances, the model might generate a copyrighted work from its training data with some probability.
Copyright law has not previously had to decide whether storing information that might or might not produce output similar to a copyrighted work is itself a copy of the work. The answer to the question is important, because it may determine the legality of many LLMs. The statute and case law are largely unhelpful. We argue that copyright law will likely take a functional approach to the question, finding that LLMs contain a copy of a particular work only if it is straightforward to extract that work in outputs. That result is unsatisfying as a policy matter, and we suggest potential changes to the law, but it is the most likely outcome under current law.