Score
Designs and implements model architectures, prior distributions, regularizers, and loss terms that embed known physical symmetries, conservation laws, and other domain constraints. Builds measurable bias mechanisms and analyzes their effects on optimization, training dynamics, and generalization to ensure model behavior aligns with those physical constraints.
This work investigates the formation mechanism of implicit bias in machine learning optimization—specifically, how optimization algorithms prefer certain solutions among multiple feasible ones. By integrating the continuous symmetries of model parameterizations with the stochasticity inherent in optimization, the study provides the first unified explanation of implicit bias as a geometric correction induced by learning dynamics in the associated quotient space. Leveraging tools from differential geometry, stochastic differential equations, and Lie group theory, the authors develop a general framework that enables both forward prediction and inverse design of implicit biases. The approach accurately predicts and precisely controls diverse forms of implicit bias—such as sparsity and spectral properties—across various architectures, with numerical experiments showing excellent agreement with theoretical predictions.
This work addresses the challenge of enforcing physical symmetries—such as rotational equivariance—in machine learning models without imposing explicit architectural constraints. The authors propose a general, architecture-agnostic approach that introduces a novel metric to quantify the degree of symmetry learning, employs spectral analysis to diagnose failure modes, and leverages targeted data augmentation to guide unconstrained Transformers—including graph neural networks and PointNet-style architectures—to progressively approximate equivariance across layers during training. Experiments demonstrate that injecting only the minimal necessary inductive bias substantially enhances physical fidelity, numerical stability, and predictive accuracy, while preserving the model’s expressive capacity.
This paper addresses the unified modeling of symmetries in machine learning. It proposes a framework grounded in differential geometry and convex optimization to (1) enforce known symmetries, (2) automatically discover unknown symmetries in models or data, and (3) actively induce symmetry breaking via user-specified candidate groups. The core contribution is the first formulation of symmetry imposition and discovery as dual linear-algebraic tasks, leveraging the Lie derivative to characterize fiberwise linear Lie group actions on vector bundles, and employing nuclear-norm relaxation to construct convex regularization terms. The method is broadly applicable to neural networks, dynamical system discovery, basis-function regression, and neural operators. Empirically, it significantly improves generalization performance and parameter efficiency—particularly in low-data regimes—while preserving geometric structure and interpretability.
This work addresses the lack of a unified understanding of the design principles, applicability, and performance differences between Physics-Informed Neural Networks (PINNs) and Neural Operators (NOs), which hinders the development of reliable data-driven PDE solvers. It proposes the first unified analytical framework that systematically characterizes the design space of both approaches along three dimensions: learning objectives, mechanisms for embedding physical structure, and strategies for computational load distribution. By elucidating the intrinsic connections and fundamental distinctions between these methods, the study not only clarifies the positioning and performance origins of existing techniques but also provides theoretical guidance and novel pathways for designing efficient and robust PDE solvers that effectively integrate physical priors with data-driven learning.
This work investigates how task-relevant symmetries—exact or approximate equivariance—affect the generalization of deep learning models, particularly under symmetry mismatch between model and data. Method: We develop the first generalization bound that does not assume group structure, rigorously quantifying the interplay between model equivariance error and data equivariance error. Our approach integrates probabilistic generalization theory, function approximation theory, and symmetry metrics, accommodating non-group, non-exact, and non-global equivariance settings. Contributions/Results: We establish that precise modeling of task symmetries significantly improves generalization. We formally characterize the optimal error trade-off under approximate or local equivariance when model and data symmetries are misaligned. Furthermore, we derive an “error alignment” principle—a concrete, actionable theoretical guideline for designing robust equivariant models—thereby bridging abstract symmetry considerations with practical architectural design.
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.
This work investigates whether intrinsic symmetries in training data induce conserved quantities during gradient flow training of neural networks. By integrating tools from differential geometry and dynamical systems theory, the study establishes—for the first time—a systematic connection between data symmetries and conservation laws in training dynamics, employing tensorized networks (including linear, polynomial, and Lightning Attention architectures) as an analytical framework. The analysis demonstrates that, under general non-polynomial losses, data symmetries do not yield additional conserved quantities; however, when combined with data augmentation under mean squared error (MSE) loss, novel conserved quantities emerge. This finding uncovers a distinctive conservation mechanism specific to MSE loss and offers a new perspective for understanding the dynamics of neural network training.
This work addresses the challenge of generalizing partial differential equation (PDE) solvers to unseen geometric domains by proposing Geo-NeW, a novel method that jointly learns differential operators and compatible reduced finite element spaces within the framework of finite element exterior calculus to rigorously preserve physical conservation laws. Geo-NeW introduces geometry-aware neural Whitney forms that embed mesh geometric information into both Transformer encodings and basis function construction, thereby endowing neural PDE solvers with strong structure-preserving inductive biases. Furthermore, it devises a new constitutive model parameterization that guarantees the existence and uniqueness of solutions. Evaluated on multiple steady-state PDE benchmarks, the method achieves state-of-the-art performance and significantly outperforms conventional approaches on out-of-distribution geometries.