in-span learning

Design and analyze methods that update a model’s internal representation or parameters online using only the model’s own predictions or outputs, constraining updates to lie within the model’s existing representational span. These methods include reweighting and realigning basis modes or subspaces so the model better absorbs corrective signals and improves future self-predictions without expanding its basis.

in-spanlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Conditional updates of neural network weights for increased out of training performance

Dec 03, 2025
JS
J. Saynisch-Wagner
🏛️ GFZ Helmholtz Centre for Geosciences

Neural networks often suffer significant performance degradation under distributional shifts—such as out-of-distribution generalization, spatiotemporal extrapolation, and cross-domain transfer—due to mismatches between training and deployment data distributions. To address this, we propose a conditional dynamic weight update framework. Its core innovation is a weight anomaly regression mechanism: sensitive weight change patterns induced by distribution shifts are identified via subset retraining; an interpretable regression predictor is then constructed to map input features to weight increments; finally, model parameters are conditionally extrapolated. The method integrates weight difference extraction, regression modeling, and extrapolation techniques, and is empirically validated on multi-source climate observation datasets. Across temporal, spatial, and cross-domain extrapolation tasks, it substantially improves prediction accuracy and robustness on out-of-distribution data, while preserving interpretability and practical applicability.

Address pattern and regime shifts in training versus application dataEnable temporal, spatial, and cross-domain extrapolation of neural networksEnhance neural network performance on out-of-distribution data

Improving Unlearning with Model Updates Probably Aligned with Gradients

Nov 04, 2025
VD
Virgile Dine
🏛️ Centre Inria de l'Université de Rennes | AMIAD

This paper addresses machine unlearning—the efficient removal of a model’s dependence on specific training samples while preserving performance on the remaining data. We propose a constraint-optimization-based feasible update framework. Our core innovation introduces a parameter masking mechanism to select an updateable subspace, jointly incorporating gradient noise modeling and directional constraints on parameter updates to yield locally feasible solutions satisfying both unlearning objectives and utility preservation. The method operates as a plug-and-play module, enhancing the robustness and accuracy of diverse first-order approximate unlearning algorithms. Experiments on image classification tasks demonstrate that our approach significantly improves unlearning accuracy (average gain of 12.3%) while incurring negligible utility loss—less than 0.5% drop in test accuracy on retained data—validating its effectiveness and practicality.

Designing feasible parameter updates preserving model utilityFormulating machine unlearning as constrained optimization problemProviding statistical guarantees for gradient-based unlearning methods

This work addresses the degradation in accuracy of conventional reduced-order models when online dynamics deviate from the training distribution, a limitation stemming from their reliance on external information to update the reduced subspace. The authors propose an intrinsic span-learning mechanism that, for the first time, reveals endogenous signals embedded within the model’s own trajectory, which can be leveraged for adaptation. By employing incremental singular value decomposition with forgetting, the method dynamically reweights and aligns the subspace basis, recasting basis reconstruction as a dynamic preconditioner from the perspective of dynamical systems. This enables in-context learning without external supervision. The approach demonstrates significantly enhanced adaptability and predictive accuracy in out-of-distribution scenarios across three benchmark problems: three-dimensional helical flow, the viscous Burgers equation, and Fisher–KPP dynamics.

dynamical systemsin-span learningmodel drift

Maintaining Structural Integrity in Parameter Spaces for Parameter Efficient Fine-tuning

May 23, 2024
CS
Chongjie Si
🏛️ Shanghai Jiao Tong University | Shanghai AI Laboratory

To address structural distortion and topological inconsistency in high-dimensional parameter spaces (e.g., 4D tensors) induced by low-rank approximation in parameter-efficient fine-tuning, this paper proposes a structure-preserving low-rank core space modeling method. Unlike conventional low-rank adapters (e.g., LoRA), which are restricted to linear weight matrices, our approach explicitly models and preserves the intrinsic topological structure of the original high-dimensional parameter space—achieving compact and accurate reconstruction of N-dimensional parameter updates via high-order tensor decomposition. Evaluated across CV, NLP, and multimodal benchmarks, the method yields an average accuracy improvement of 1.8% under identical parameter budgets, while reducing structural distortion by 37%, significantly outperforming existing baselines.

Enable parameter-efficient fine-tuning across diverse dimensional spacesModel changes via low-rank core space with consistent topologyPreserve structural integrity in high-dimensional parameter spaces

In parametric dynamical systems, the Proper Orthogonal Decomposition (POD) basis drifts with parameters, degrading the accuracy of reduced-order models (ROMs). Method: This paper proposes the Projected Gaussian Process (pGP) framework—the first to formulate subspace adaptation as a statistical learning task mapping parameter space to the Grassmann manifold. It employs a two-stage geometric mapping: Euclidean space → horizontal space → Grassmann manifold, integrating POD, exponential/logarithmic maps, horizontal-space projection, and Gaussian process regression to enable uncertainty-aware POD subspace prediction while preserving manifold structure. Contribution/Results: Numerical experiments demonstrate that pGP significantly improves ROM accuracy and robustness in both parametric extrapolation and interpolation scenarios, and provides interpretable, calibrated confidence quantification—establishing a new paradigm for parameter-sensitive model reduction.

Adapting POD basis for parametric Reduced-Order ModelsMapping parameters to Grassmann manifold subspacesPredicting optimal subspaces using Gaussian Process regression

Latest Papers

What's happening recently
View more

This work addresses the feedback loops that arise after model deployment due to performativity—wherein the model’s predictions influence the data distribution—particularly under strong interventions where the convergence behavior of retraining remains poorly understood. The paper introduces the “stable signal principle,” positing that the prediction target contains an intrinsic component independent of the model (e.g., inherent item quality), and leverages this insight to analyze the dynamics of regularized repeated risk minimization. Theoretically, it establishes that as long as a non-zero stable signal exists, retraining converges geometrically to its direction, even when model-induced effects dominate. This reveals a novel role for regularization in mitigating performative feedback and extends the framework to nonlinear, heterogeneous, and time-varying settings—including language models—thereby explaining the observed stability of training on generated data.

feedback loopfixed pointperformativity

This study addresses the performance degradation caused by model collapse during the self-training of generative models. To mitigate this issue, we propose a geometric correction method based on singular value reweighting of the Jacobian matrix. This approach explicitly reinforces negative guidance signals for the first time by amplifying mode-seeking behavior and distortion characteristics, thereby overcoming the limitations of conventional direct fine-tuning. When integrated with algorithms such as Neon and SIMS, the proposed technique significantly enhances both self-training performance and iterative stability across various one-step generative models, particularly in data-scarce scenarios.

Generative ModelModel CollapseNegative Guidance

This work addresses the challenge of online learning and real-time adaptation in existing deep neural state-space models by proposing a unified framework that integrates recursive identification with batch online learning, tailored to subspace-structured encoders. It introduces, for the first time, a provably convergent recursive online learning mechanism specifically designed for encoder-based neural state-space models, accompanied by rigorous theoretical convergence guarantees. By jointly leveraging subspace encoding, recursive system identification, and an efficient batch update strategy, the proposed approach enables simultaneous real-time optimization of both latent states and model parameters directly from input-output data. Simulation results demonstrate that the method achieves high modeling accuracy while significantly enhancing online computational efficiency and dynamic adaptability.

neural state-space modelsonline learningrecursive identification

This work addresses the computational bottleneck in traditional approaches for large-scale parametric dynamical systems, which require repeated and expensive eigenvalue solves to construct modal bases, thereby hindering efficient design optimization. To overcome this limitation, the authors propose a coupled architecture that integrates a rank-reduced autoencoder (RRAE) with a deep neural network. The RRAE leverages truncated singular value decomposition to define a unified low-dimensional parameter space, which serves as input to the neural network for jointly reconstructing all modal bases. This framework enables a nonlinear, physics-consistent parameterization of the modal basis while entirely eliminating the need for repeated eigenvalue computations. Validated on both 1D and 2D dynamical systems, the method demonstrates high accuracy and computational efficiency, effectively capturing dominant physical features and mitigating overfitting.

computational costdynamical systemseigenvalue problem

This work addresses the lack of reliable and reproducible tools for managing weights in large-scale deep learning models, a gap often filled by fragile ad-hoc scripts. The authors propose BrainSurgery, the first framework to introduce declarative programming into model editing, enabling users to specify tensor surgery operations via YAML configuration files. By leveraging regular expressions for precise parameter targeting, BrainSurgery supports structured transformations such as reshaping, low-rank decomposition, and precision conversion. The system incorporates built-in assertions to guarantee correctness and has been validated on tasks including model upgrading and LoRA extraction. Experimental results demonstrate that BrainSurgery significantly enhances the reliability, reproducibility, and efficiency of complex weight manipulation workflows.

checkpoint managementmodel editingreproducibility

Hot Scholars

DB

David Bryant

University of Otago
MathematicsStatistics and Evolutionary Genetics
ZQ

Ziyue Qiao

Assistant Professor, Great Bay University
Data MiningGraph Machine LearningKnowledge GraphAI for Science
SY

Sangpil Youm

Ph.D Student, University of Florida
Natural Language ProcessingArtificial IntelligenceNetwork Science
JL

Jie Liu

Huazhong university of Science and Technology
Fault diagnosispredictive maintenance