Score
Design and analyze reductions and embeddings that map problem instances into finite-dimensional spaces while preserving or tightly bounding the ambient dimension (including finite-dimensional reductions and dimension-preserving embeddings). Construct efficiently buildable representations and proofs that these transforms avoid exponential dimension blowup and thereby transfer algorithmic or complexity properties, such as hardness, across dimensions.
This work addresses the topological mechanisms underlying dimensionality reduction and width in deep neural networks, specifically focusing on preserving connectivity and ensuring class separability when embedding compact topological spaces into Euclidean space. Method: We establish a rigorous mathematical framework grounded in topological degree theory, characterizing the intrinsic relationship between embedding connectivity and linear/nonlinear separability under dimension-reducing mappings. We quantitatively analyze how network width and input dimension compression affect classification separability and function approximation capacity. Contribution/Results: We propose the first unified theoretical perspective integrating topological degree, embedding connectivity, and deep network architecture design—providing principled, mathematically grounded guidelines for width selection and dimension compression. Our analysis significantly enhances interpretability of deep models in classification and approximation tasks and rigorously delineates their topological performance limits.
This paper investigates the impact of random dimensionality reduction on several maximization problems in Euclidean space—namely, maximum matching, maximum spanning tree, and maximum traveling salesman problem—as well as on dataset diversity measures. It introduces a novel analytical framework centered on the dataset’s *doubling dimension* λ_X, establishing that O(λ_X) random projection dimensions suffice to preserve optimal objective values within a (1±ε)-multiplicative factor—improving upon classical bounds dependent on the number of points |X|. The derived lower bound is shown to be tight. Theoretical analysis guarantees near-preserving solution quality, while empirical evaluation confirms high post-projection accuracy and substantial computational speedup. The core contribution is the first quantitative characterization linking dimensionality reduction efficacy directly to the intrinsic geometric complexity λ_X, yielding a finer-grained and more practically relevant theoretical foundation for optimization in high-dimensional spaces.
This work establishes a fundamental trade-off between embedding dimensionality and representational fidelity from an information-theoretic perspective. It rigorously proves that when the embedding dimension falls below a constant fraction \( cD \) of the true data dimension \( D \) (for some universal constant \( c < 1 \)), any such embedding necessarily violates at least half of all triplet constraints. Moreover, under the Unique Games Conjecture (UGC), even when the true dimension is one, no polynomial-time algorithm can achieve better than the trivial 50% accuracy in preserving triplet orderings. By integrating tools from information theory, contrastive learning theory, and computational complexity, the study reveals an inherent limitation in low-dimensional embeddings and establishes the first computational complexity lower bound for dimensionality compression in representation learning.
This paper investigates the theoretical limits of preserving neighborhood structure in low-dimensional visualizations of high-dimensional data. Addressing the fundamental question—“Can neighborhood relations of high-dimensional data be reliably preserved in constant-dimensional spaces (e.g., 2D/3D)?”—we introduce the *doubling dimension* as a geometric complexity measure for embedding difficulty. Leveraging graph embedding theory, metric space analysis, and planted cluster models, we systematically characterize visualization compressibility across graph classes. We prove: (i) almost all $n$-vertex graphs require $Omega(log n)$ dimensions to maintain neighborhood separability; (ii) sparse regular graphs still necessitate $Omega(log n / log log n)$ dimensions; and (iii) in normed spaces, nearly all graphs require $Theta(n)$ dimensions. This work provides the first information-theoretic and geometric characterization of intrinsic dimensional bottlenecks in common dimensionality reduction techniques (e.g., t-SNE, UMAP), establishing rigorous theoretical foundations for visualization design and interpretation.
In high-dimensional, small-sample regimes, conventional machine learning generalization bounds suffer from excessive looseness due to the “curse of dimensionality.” Method: This paper proposes an adaptive family of generalization bounds tailored to discretized Euclidean spaces. We first derive a non-asymptotic concentration inequality for finite metric spaces; introduce geometric representation dimension (m) as a pivotal parameter; and construct bounds with constant factor (c_m) scaling as (sqrt{m}). Tight analysis is achieved via metric embedding combined with discretization-based modeling. Contribution/Results: The proposed bounds yield significant tightening under practical sample sizes; retain the optimal (O(1/sqrt{N})) convergence rate; and break the exponential or polynomial dependence on ambient dimension inherent in classical bounds—achieving constant-factor improvement even in high-dimensional, low-precision settings.
This work addresses the challenge of efficiently preserving continuous curve distances—such as the Fréchet distance—under dimensionality reduction for high-dimensional polygonal curves. The authors propose a randomized projection method based on sparse oblivious subspace embeddings that simultaneously approximates multiple curve dissimilarity measures, including Fréchet, q-DTW, and Hausdorff distances, within a relative error of (1±ε) using a target dimension of O(ε⁻² log(nm)). By constructing a unified framework for generalized curve distance metrics, the approach extends dimensionality reduction theory to piecewise linear surfaces, substantially simplifying existing analyses and broadening applicability across diverse curve comparison tasks.
This work addresses the problem of constructing low-dimensional geometric embeddings for directed acyclic graphs (DAGs) with ancestor–descendant relationships, aiming to avoid embedding dimensions that scale explosively with the number of nodes or graph depth. By leveraging structural properties—such as treewidth and the number of cross edges—and insights from geometric embedding theory, the authors propose a compact representation whose dimension depends only on these structural parameters rather than the total node count. Key contributions include a proof that any directed tree admits an exact reachability-preserving embedding in three dimensions, and an upper bound of \( O(t \log n) \) dimensions for DAGs of treewidth \( t \), accompanied by a nearly matching lower bound that reveals fundamental limits on dimensionality. Experiments on real-world datasets demonstrate that the method substantially reduces embedding dimension while maintaining high recall, outperforming existing approaches with theoretical guarantees.
This study addresses the worst-case distortion incurred when embedding metric spaces into a fixed low-dimensional Euclidean space (\(d > 1\)) under adversarial online settings, with a focus on whether distortion necessarily grows exponentially. The work proposes deterministic online embedding strategies for solid graphs containing a \(K_5\) minor and tree-like metrics. It establishes, for the first time, that \(K_5\) metrics admit online embeddings into \(\mathbb{R}^2\) with only polynomial distortion, thereby refuting a conjectured exponential lower bound. For ultrametrics and other tree-like structures, the approach achieves \(n^{\Theta(1/d)}\) distortion in \(\mathbb{R}^d\), matching the offline optimum up to constant factors in the exponent and substantially narrowing the performance gap between online and offline settings. The methodology integrates metric embedding theory, graph structural analysis, and hierarchical well-separated tree (HST) techniques, enabling efficient derandomization of probabilistic embedding results into low-dimensional Euclidean space.
This study addresses the uneven distribution of instance-space complexity in higher-order atomic concept learning by introducing a locality-of-complexity perspective grounded in the geometric structures of hypercubes and hyperplanes. It reveals that logical complexity concentrates along the full diagonal, while complexity collapses on other hyperplanes due to constraint-induced simplifications. Through high-dimensional geometric modeling, analysis of logical equivalence classes, and construction of constrained hypothesis spaces—combined with canonical simple concepts, minimal orderings, and representative reduction mechanisms—the work systematically classifies the behavior of high-dimensional hyperplanes. The analysis fully resolves binary and ternary cases, characterizing properties of orthogonal families, partial diagonals, and full diagonals, and establishes an upper bound on the number of equivalence classes for non-full-diagonal hyperplanes that is independent of term depth.