🤖 AI Summary
This study addresses the limitation of multimodal models in symbolic regression that rely solely on global alignment, which hinders their ability to capture how substructures influence numerical behavior. To this end, we propose a compositional cross-modal alignment method that introduces, for the first time, a sub-expression-level alignment mechanism. By employing structural positional encoding to explicitly expose the local topology of expression trees and integrating multi-granularity contrastive learning objectives, our approach achieves fine-grained matching between symbolic representations and numerical behaviors, thereby overcoming the constraints of traditional global embeddings. Experimental results demonstrate that the proposed method significantly narrows the modality gap, accurately models the impact of local edits on functional behavior, and exhibits strong transferability across external symbolic corpora.
📝 Abstract
Mathematical expressions and the numerical behavior they produce are two views of the same underlying function, and connecting them is central to scientific discovery. Symbolic Regression (SR) relies on this connection directly: it searches for an expression that reproduces a given behavior. Recent multi-modal models learn this connection by embedding symbolic expressions and their numerical behavior in a shared representation space. We show that this embedding space is only globally aligned: complete expressions correspond to complete behaviors, but the contribution of individual parts of an expression is not represented. This granularity gap leaves the model unable to tell how a local edit to an expression changes its behavior, the central operation in SR. We introduce a compositional alignment method that closes this gap: a structural positional encoding exposes the substructure of an expression to the encoder, and a multi-granularity contrastive objective grounds each subexpression in the behavior it produces before propagating this grounding to the full expression. The resulting representations close much of the modality gap between symbolic and numerical embeddings, reliably distinguish the effects of local edits that the original alignment cannot, and transfer to external SR corpora.