Beyond Pairwise Comparisons: A Distributional Test of Distinctiveness for Machine-Generated Works in Intellectual Property Law

📅 2026-01-26
📈 Citations: 0
Influential: 0
📄 PDF

career value

193K/year
🤖 AI Summary
Current intellectual property law relies on pairwise comparisons to assess the novelty and originality of works, which is ill-suited for evaluating the fundamental differences between the output distributions of generative models and human-created content. This work proposes a two-sample test based on Maximum Mean Discrepancy (MMD) in semantic embedding space, enabling unsupervised, efficient, and distribution-level discrimination between human- and machine-generated content without requiring task-specific training or access to private data. Using only 5–10 images or 7–20 text samples, the method demonstrates consistent efficacy across three domains: handwritten digits, patent abstracts, and AI-generated art. Notably, it reliably detects the distinctiveness of generative model outputs in semantic space—even when human evaluators achieve only 58% accuracy—revealing these models to function essentially as stochastic interpolators rather than mere replicators of training data.

Technology Category

Application Category

📝 Abstract
Key doctrines, including novelty (patent), originality (copyright), and distinctiveness (trademark), turn on a shared empirical question: whether a body of work is meaningfully distinct from a relevant reference class. Yet analyses typically operationalize this set-level inquiry using item-level evidence: pairwise comparisons among exemplars. That unit-of-analysis mismatch may be manageable for finite corpora of human-created works, where it can be bridged by ad hoc aggregations. But it becomes acute for machine-generated works, where the object of evaluation is not a fixed set of works but a generative process with an effectively unbounded output space. We propose a distributional alternative: a two-sample test based on maximum mean discrepancy computed on semantic embeddings to determine if two creative processes-whether human or machine-produce statistically distinguishable output distributions. The test requires no task-specific training-obviating the need for discovery of proprietary training data to characterize the generative process-and is sample-efficient, often detecting differences with as few as 5-10 images and 7-20 texts. We validate the framework across three domains: handwritten digits (controlled images), patent abstracts (text), and AI-generated art (real-world images). We reveal a perceptual paradox: even when human evaluators distinguish AI outputs from human-created art with only about 58% accuracy, our method detects distributional distinctiveness. Our results present evidence contrary to the view that generative models act as mere regurgitators of training data. Rather than producing outputs statistically indistinguishable from a human baseline-as simple regurgitation would predict-they produce outputs that are semantically human-like yet stochastically distinct, suggesting their dominant function is as a semantic interpolator within a learned latent space.
Problem

Research questions and friction points this paper is trying to address.

machine-generated works
distinctiveness
distributional comparison
intellectual property law
generative models
Innovation

Methods, ideas, or system contributions that make the work stand out.

distributional test
maximum mean discrepancy
semantic embeddings
generative models
distinctiveness