Optimal VC Dimension of Contrastive Learning with Margin

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the suboptimal upper bound and the absence of a lower bound for the VC dimension in margin-based contrastive learning. Leveraging computational learning theory and the PAC framework, combined with geometric embeddings in Euclidean space and combinatorial analysis, we systematically derive tight bounds on its sample complexity. The core contribution is twofold: we establish, for the first time, an improved VC dimension upper bound of O(n/α²) in this setting by eliminating the redundant logarithmic factor, and we construct a matching lower bound of Ω(n/α²), thereby proving the optimality of these bounds. By providing optimal tight bounds on the VC dimension for arbitrary margin parameters, this work significantly enhances theoretical tightness and resolves a longstanding open problem in the field.
📝 Abstract
Contrastive learning is a successful paradigm for learning $d$-dimensional geometric representations from a collection of ``anchor--positive--negative'' triplets $(i,j^{+},k^{-})$, indicating that ``item $i$ is closer to $j$ than to $k$.'' Despite its success, understanding why contrastive learning leads to representations of high \textit{generalization} quality---beyond the often pessimistic predictions from PAC-learning---remains a central question. Recently, \citet*{alon2024optimal} proved that, for PAC-learning $d$-dimensional Euclidean representations of $n$-point datasets, $Θ(\min(nd, n^2))$ triplets are necessary and sufficient, while they posed as an open question whether their VC dimension bounds for the more realistic setting of \textit{contrastive learning with a margin} can be improved. For a margin parameter $α>0$, a triplet $(i,j^{+},k^{-})_α$ is satisfied by the embedding $φ:[n]\rightarrow \mathbb{R}^{d}$, if $\|φ(i)-φ(k)\|_2>(1+α)\cdot\|φ(i)-φ(j)\|_2$. In this work, we resolve their question by proving that the VC dimension of contrastive learning under any margin $α\in(0,1)$ is in fact $O(n/α^2)$, improving on the previous bound of $O(n\log(n)/α^2)$. We also establish that the bounds are optimal up to constant factors, by providing a matching lower bound of $Ω(\frac{n}{α^2})$ (the previously known lower bound was $Ω(\frac{n}α)$), for $α\geq \max(n^{-1/2},d^{-1/2})$.
Problem

Research questions and friction points this paper is trying to address.

Contrastive Learning
VC Dimension
Margin
PAC Learning
Generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Contrastive Learning
VC Dimension
Margin
Generalization
PAC Learning
🔎 Similar Papers
No similar papers found.