Low-Dimension-to-High-Dimension Generalization And Its Implications for Length Generalization

📅 2024-10-11
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates low-dimensional-to-high-dimensional (LDHD) generalization—a specialized out-of-distribution generalization problem where training data lies on a low-dimensional submanifold of the high-dimensional test space, equivalent in essence to length generalization. Theoretically, we prove that LDHD generalization is impossible without appropriate inductive bias; min-degree interpolation serves as a universal mechanism for SGD convergence; and chain-of-thought (CoT) reasoning improves length generalization via enhanced positional awareness. Building on these insights, we propose RPE-Square, a novel relative positional encoding that jointly models latent-space scalability and input-format robustness—marking the first unified approach to both properties. Grounded in Boolean function analysis, inductive bias modeling, and optimization dynamics characterization, empirical evaluation demonstrates that RPE-Square significantly outperforms standard RPE on long-sequence tasks. Our framework provides both a unifying theoretical foundation and a practical solution for LDHD and length generalization.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Search and Optimization: Learning to SearchNatural Language Processing: (Large) Language Models

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
Low-Dimension-to-High-Dimension (LDHD) generalization is a special case of Out-of-Distribution (OOD) generalization, where the training data are restricted to a low-dimensional subspace of the high-dimensional testing space. Assuming that each instance is generated from a latent variable and the dimension of the latent variable reflects the problem scale, the inherent scaling challenge in length generalization can be captured by the LDHD generalization in the latent space. We theoretically demonstrate that LDHD generalization is generally unattainable without exploiting prior knowledge to provide appropriate inductive bias. Specifically, we explore LDHD generalization in Boolean functions. We verify that different architectures trained with (S)GD converge to emph{min-degree interpolators w.r.t. different independent sets}. LDHD generalization is achievable if and only if the target function coincides with this inductive bias. Applying the insights from LDHD generalization to length generalization, we explain the effectiveness of CoT as changing the structure latent space to enable better LDHD generalization. We also propose a principle for position embedding design to handle both the inherent LDHD generalization and the nuisances such as the data format. Following the principle, we propose a novel position embedding called RPE-Square that remedies the RPE for dealing with the data format nuisance.
Problem

Research questions and friction points this paper is trying to address.

Explores LDHD generalization in Boolean functions
Analyzes role of inductive bias in LDHD generalization
Proposes position embedding design for LDHD generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

LDHD generalization captures scaling challenges
CoT changes latent space for better generalization
RPE-Square improves position embedding design
🔎 Similar Papers