Score
Designs and implements self‑supervised pretraining pipelines that incorporate physics‑based objectives — such as masked physics prediction and equation‑consistency prediction — to learn representations from physical simulation or measurement data. Builds and evaluates self‑supervised learning objectives and pretraining strategies to improve downstream task accuracy, generalization, and recovery of localized physical structures.
Enhancing predictive and forecasting performance of physics-informed machine learning (PI-ML) for partial differential equation (PDE)-based modeling remains challenging due to heterogeneous physical knowledge integration strategies. Method: We systematically survey and unify over 120 physics-integrated ML methods, proposing a novel dual-path paradigm: “architecture embedding” (e.g., physics-constrained loss functions, structured neural operators, physics-guided data augmentation) and “data-as-knowledge” (e.g., multi-task learning, meta-learning, in-context learning, symbolic regression–assisted modeling), thereby decoupling physical knowledge injection mechanisms for the first time. Contribution/Results: We establish a theoretical framework covering seven transferable inductive biases; standardize interfaces across five mainstream open-source PI-ML libraries; release the first industry-oriented PI-ML tool landscape; and provide deployable practice guidelines for six domains—energy, climate science, fluid dynamics, materials science, biophysics, and geophysics.
Self-supervised representation learning in physical sciences—particularly in domains reliant on stochastic simulators (e.g., high-energy physics experiments)—faces challenges in generating physically consistent, diverse augmentations without compromising interpretability or fidelity. Method: We propose RS3L, a framework that performs *controlled interventions at intermediate layers* of the simulation pipeline and *re-executes downstream components*, yielding physically coherent multi-realization augmented samples. This “intermediate intervention + downstream resimulation” mechanism uniquely embeds domain-specific physical priors deeply into data augmentation, enabling interpretable, coverage-complete, simulation-driven contrastive learning. RS3L integrates causal augmentation with foundation model pretraining. Contribution/Results: RS3L significantly improves object discrimination accuracy and uncertainty quantification capability. We publicly release the RS3L benchmark dataset, establishing a new paradigm for simulation-driven scientific AI.
This study investigates the design of pretraining objectives for foundation models in simulation-based scientific domains such as high-energy physics, aiming to effectively support both classification and generation downstream tasks. Leveraging the JetClass dataset, the authors systematically evaluate three pretraining strategies—supervised classification, flow matching for generation, and self-supervised masked particle modeling—and fine-tune the resulting models on top-quark jet classification and JetNet conditional generation tasks. The findings reveal that classification and generation objectives are largely orthogonal, making it difficult for a single pretraining objective to excel at both. Supervised classification achieves optimal performance when labels are abundant, whereas masked particle modeling substantially improves results in low-label regimes. Flow matching benefits generation tasks only when included during pretraining. Built upon the OmniLearned framework, this work demonstrates that jointly optimizing multiple pretraining objectives is key to enabling effective cross-task transfer.
This paper addresses three physics-constrained regression problems in fluid mechanics: PIV velocity field super-resolution and data assimilation, data-driven turbulence modeling, and system identification for digital twin predictive control. Methodologically, it proposes a physics-informed regression framework that incorporates conservation laws—such as the Navier–Stokes equations—as soft constraints into supervised learning objectives; gradient-based optimization is enabled via automatic differentiation, and differentiable physics-informed models are implemented in Python. Key contributions include: (1) a unified approach to modeling under multiscale dynamics, limited data, and high noise; (2) substantially improved model generalizability and physical consistency; and (3) publicly available, reproducible educational case studies and code, demonstrating the efficacy and pedagogical versatility of physics-informed learning in scientific discovery and engineering closed-loop control.
Addressing the scarcity of large-scale labeled data and the limited generalization capability of self-supervised learning (SSL) in materials property prediction, this paper proposes a proxy-label-based supervised graph neural network pretraining framework. The method follows a two-stage paradigm: proxy-label supervised pretraining followed by task-adaptive fine-tuning. Key contributions include: (1) the first introduction of class-level categorical information as proxy labels for supervised pretraining in materials science, significantly enhancing downstream multi-task generalization; and (2) a graph-structure-preserving noise augmentation strategy that injects perturbations while maintaining physical plausibility and topological consistency. Evaluated on six diverse materials property prediction tasks, the proposed approach reduces mean absolute error (MAE) by 2.0%–6.67% over state-of-the-art SSL baselines, establishing new performance benchmarks.
This study addresses critical challenges in applying deep learning to scientific computing—namely, poor interpretability, heavy data dependency, and insufficient physical consistency—within physics-based simulation scenarios. We propose a physics-driven AI modeling framework integrating physics-informed loss functions, differentiable simulators, diffusion-based generative models, physics-guided reinforcement learning, and custom neural architectures, implemented via an interactive Jupyter-based experimental platform. Crucially, we pioneer the systematic embedding of physical priors across the entire deep learning pipeline—model formulation, training, and inference—enabling high-fidelity, data-efficient, and verifiable scientific modeling. The resulting methodology is modular, reusable, and immediately deployable, significantly enhancing model generalizability and interpretability. This work establishes a novel paradigm and technical foundation for next-generation scientific foundation models.
本文探讨了在机器人学习中嵌入物理先验知识的方法,以解决数据有限、复杂交互和可靠操作需求的问题,通过整合物理法则来提高学习算法的泛化能力、可解释性和样本效率。
This work addresses the limitations of traditional spatiotemporal physical system modeling, which relies on pixel-level next-frame prediction and suffers from error accumulation and high training costs, thereby hindering effective support for downstream scientific tasks such as physical parameter estimation. The study proposes evaluating general self-supervised learning methods based on their ability to yield physically meaningful representations, using downstream task performance—particularly physical parameter estimation—as a benchmark. By comparing objective functions rooted in pixel-space prediction versus those operating in latent spaces (e.g., Joint Embedding Predictive Architecture, JEPA), the authors demonstrate that latent-space modeling substantially enhances both the physical interpretability of learned representations and downstream task performance. Empirical results reveal that certain general-purpose self-supervised approaches outperform specialized physics-based models, highlighting their significant potential for scientific representation learning.
Standard data augmentation in self-supervised learning overlooks the symmetries and acquisition constraints dictated by the physical measurement process in scientific imaging, often leading to distorted representations. This work proposes a physics-aligned self-supervised learning framework that, for the first time, systematically incorporates measurement operator constraints into augmentation design. The approach establishes a modality-agnostic, label-light pipeline for selecting augmentations, thereby transforming augmentation into a controllable inductive bias. Evaluated across five prominent self-supervised methods—DINOv2, SimCLR, MAE, VICRegL, and I-JEPA—and applied to real-space electron microscopy and 4D-STEM datasets, the method significantly improves downstream task performance, reduces geodesic error, enhances robustness to detector gain variations and resolution loss, and reshapes the geometry of learned representations.
本文提出了一种基于能量搬运距离(EMD)的数据驱动方法,用于大型强子对撞机(LHC)自监督预训练中的事件相似性配对,避免了手工设计的数据增强问题。
This work addresses the limited generalizability of conventional neural surrogate models in computational fluid dynamics (CFD), which rely on explicit boundary conditions and struggle with scenarios involving altered boundaries or local geometric modifications. The authors reformulate steady-state CFD inference as a context-driven image inpainting task and introduce, for the first time, a reusable velocity field prior learned through self-supervision. Their approach employs a local neighborhood tokenizer to compress high-resolution velocity fields into compact spatially implicit tokens, trained via a masked autoencoder combined with latent-space flow matching. Evaluated on intracranial aneurysm hemodynamics, the method accurately reconstructs full flow fields from only sparse boundary context, significantly outperforming supervised surrogates under varying boundary conditions and distribution shifts, while enabling efficient local geometry editing and context reuse.