🤖 AI Summary
In supervised learning, modular decomposition lacks identifiability when models exhibit equivalent local linear responses under input perturbations—termed the “mirage regime”—raising the question of whether internal module assignments are uniquely recoverable.
Method: We propose Modular Jets (MoJet), a differential-geometric framework that estimates first-order module-wise responses to infinitesimal input perturbations (empirical jets), integrating task manifold geometry and module-level representations to formulate a rank-based criterion distinguishing the mirage from the identifiable regime.
Contribution/Results: We theoretically prove that, in two-module linear regression, jets uniquely recover the ground-truth decomposition. Algorithmically, MoJet implements jet estimation and mirage diagnosis. Experiments demonstrate its effectiveness across linear and deep regression and classification pipelines. Crucially, MoJet enables the first *identifiability diagnosis*—inferring modular structure directly from input-output behavior—moving beyond conventional reliance on predictive risk alone.
📝 Abstract
Classical supervised learning evaluates models primarily via predictive risk on hold-out data. Such evaluations quantify how well a function behaves on a distribution, but they do not address whether the internal decomposition of a model is uniquely determined by the data and evaluation design. In this paper, we introduce emph{Modular Jets} for regression and classification pipelines. Given a task manifold (input space), a modular decomposition, and access to module-level representations, we estimate empirical jets, which are local linear response maps that describe how each module reacts to small structured perturbations of the input. We propose an empirical notion of emph{mirage} regimes, where multiple distinct modular decompositions induce indistinguishable jets and thus remain observationally equivalent, and contrast this with an emph{identifiable} regime, where the observed jets single out a decomposition up to natural symmetries. In the setting of two-module linear regression pipelines we prove a jet-identifiability theorem. Under mild rank assumptions and access to module-level jets, the internal factorisation is uniquely determined, whereas risk-only evaluation admits a large family of mirage decompositions that implement the same input-to-output map. We then present an algorithm (MoJet) for empirical jet estimation and mirage diagnostics, and illustrate the framework using linear and deep regression as well as pipeline classification.