How Many Samples Are Enough for Learning Across Domains?

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of theoretical guidance regarding per-domain sample sufficiency conditions in cross-domain learning. Overcoming the limitations of classical learning theory, this work establishes a theoretical criterion for per-domain sample requirements through generalization bound analysis and statistical learning theory derivations. The results demonstrate that the required sample size depends critically on the number of training domains, revealing an inverse-linear scaling law between domain count and sample complexity that elucidates the underlying principles of data sufficiency assumptions. These findings provide a rigorous theoretical foundation for evaluating existing datasets and constructing new ones, while offering deep insights into the intrinsic connection between in-domain learning and out-of-domain generalization.
📝 Abstract
Understanding the fundamental mechanisms of learning is essential for designing systems with strong generalization. Recent studies have shown that increasing the number of training domains, or enlarging the distribution shift among them, improves generalization when each domain contains sufficiently many data samples. However, the conditions under which the data samples can be considered sufficient remain unexplored. In this work, we fill this gap by establishing criteria for per-domain sample requirements based on the presented learning bounds. These criteria not only reveal an inverse linear scaling law between the number of training domains and the number of samples required per domain, but also explain the fundamental rationale behind the assumption of data sufficiency, thereby providing theoretical guidance for assessing the adequacy of existing datasets and constructing datasets. This differs from classical learning theory, as the number of samples required is highly dependent on the number of training domains. Additionally, we prove the close relationship between in-domain learning and out-of-domain generalization through the presented generalization bounds, and lastly discuss some key arguments.
Problem

Research questions and friction points this paper is trying to address.

domain generalization
sample complexity
cross-domain learning
per-domain sample requirement
generalization bounds
Innovation

Methods, ideas, or system contributions that make the work stand out.

Domain Generalization
Sample Complexity
Inverse Linear Scaling Law
Generalization Bounds
Out-of-Domain Generalization
🔎 Similar Papers
No similar papers found.