🤖 AI Summary
This study addresses the challenge that scientific foundation models, when applied to biological and physical systems, suffer from degraded geometric fidelity due to their discrete tokenized representations, which fail to preserve the intrinsic continuous geometric structure of such systems. By systematically comparing discrete-token and continuous-output heads under an identical encoder architecture, the work reveals the detrimental mechanism of the discrete bottleneck on geometric preservation. It introduces the novel concept of “geometric alignment tax” to explain how discretization induces geometric distortion and identifies three distinct representation failure modes. Through ablation studies on synthetic dynamical systems, rate–distortion theory, and MINE-based mutual information estimation, the authors demonstrate that while architectural performance gaps amount to only a 1.3× difference under continuous targets, they escalate dramatically to 3000× after discretization; moreover, no tested model simultaneously achieves low distortion, high mutual information, and global consistency.
📝 Abstract
Foundation models for biology and physics optimize predictive accuracy, but their internal representations systematically fail to preserve the continuous geometry of the systems they model. We identify the root cause: the Geometric Alignment Tax, an intrinsic cost of forcing continuous manifolds through discrete categorical bottlenecks. Controlled ablations on synthetic dynamical systems demonstrate that replacing cross-entropy with a continuous head on an identical encoder reduces geometric distortion by up to 8.5x, while learned codebooks exhibit a non-monotonic double bind where finer quantization worsens geometry despite improving reconstruction. Under continuous objectives, three architectures differ by 1.3x; under discrete tokenization, they diverge by 3,000x. Evaluating 14 biological foundation models with rate-distortion theory and MINE, we identify three failure regimes: Local-Global Decoupling, Representational Compression, and Geometric Vacuity. A controlled experiment confirms that Evo 2's reverse-complement robustness on real DNA reflects conserved sequence composition, not learned symmetry. No model achieves simultaneously low distortion, high mutual information, and global coherence.