Score
Designs, implements, and evaluates methods that identify and modify or remove subspaces, components, or directions in model representations or parameters that encode a targeted concept or capability, using interventions such as subspace projection, component attenuation, principal-component collapse, manifold- or scale-constrained edits, or targeted fine‑tuning. The work demonstrates reduced recoverability of the target concept while preserving non-target capabilities and minimizing overall performance degradation.
This work addresses the limited deployability of large language models in high-stakes settings, which stems from insufficient understanding of how internal capabilities are encoded within their parameter space. Existing approaches struggle to precisely characterize the distribution of such capabilities across model parameters. To overcome this, the authors propose SCALPEL, a framework that challenges conventional modular assumptions by revealing that model capabilities can be modeled as low-rank subspaces spanning multiple layers and modules. Leveraging LoRA-based low-rank parameter editing, SCALPEL enables selective ablation of specific capabilities—such as distinguishing correct from incorrect answers—while preserving general language modeling performance. Experiments on the BLiMP benchmark demonstrate the method’s efficacy, offering a fine-grained tool for interpretability through targeted intervention and analysis.
This paper addresses the challenge of jointly compressing high-dimensional input and output spaces in goal-oriented applications—such as sensor placement and sensitivity analysis—where conventional dimensionality reduction methods treat inputs and outputs independently. We propose an input–output co-dimensional reduction framework that jointly optimizes coupled input and output subspaces. Crucially, it reformulates the NP-hard combinatorial selection problem into a differentiable optimization over diagonal entries of a diagnostic matrix, obviating costly evaluations of the objective function. By integrating gradient-based upper-bound optimization, expected information gain, Sobol’ sensitivity indices, and spectral analysis of the diagnostic matrix, our method achieves substantial computational efficiency gains. Experiments demonstrate its effectiveness and scalability in sensor layout optimization and parameter importance ranking. The approach establishes a novel paradigm for high-dimensional, goal-oriented experimental design and global sensitivity analysis.
Neural network representations often entangle sensitive attributes with task-relevant information, compromising fairness and interpretability. To address this, we propose a linear concept removal method that completely eliminates linearly predictable sensitive attributes while strictly preserving the covariance between representations and primary task labels—thereby ensuring lossless retention of task-relevant second-order statistical signal. Our key contribution is the construction of a unique optimal oblique projection operator, rigorously justified through covariance-preserving optimization and theoretical analysis of concept predictability elimination. Experiments on benchmarks including Bias in Bios and Winobias demonstrate over 95% removal rate of sensitive attributes, with less than 0.8% degradation in main-task accuracy—significantly outperforming existing post-processing debiasing methods.
This work addresses the performance degradation commonly observed in multi-task model merging due to interference among tasks. The authors propose an efficient, training-free merging method that first constructs an intrinsic feature subspace dominated by task-specific parameter updates and projects individual task models onto this low-rank subspace for fusion. To further enhance knowledge retention and suppress redundancy, a multi-level polarized scaling mechanism is introduced to amplify critical parameters while attenuating less informative ones. By integrating principal component analysis, low-rank decomposition, and parameter projection, the approach substantially mitigates task interference across diverse task sets and model scales, preserving essential functionalities and achieving state-of-the-art performance in multi-task model merging.
To address the “curse of dimensionality” arising from high-dimensional design spaces in functional surface design, this study systematically reviews and innovates dimensionality reduction (DR) methods for shape optimization. Adopting the scoping review methodology—novel in engineering design—we establish a unified taxonomy encompassing linear techniques (PCA, kernel PCA), nonlinear approaches (t-SNE, autoencoders, VAEs), and physics-informed methods (physics-constrained embedding networks, surrogate-model-coupled frameworks). We propose a novel physics-informed embedding framework that significantly enhances interpretability and engineering applicability of DR results. Experimental evaluation demonstrates that optimization efficiency improves by 3–10× after DR, design space explorability increases, and physical consistency of optimal solutions is markedly improved—validating both computational efficacy and domain fidelity.
Existing additive-update-based concept erasure methods struggle to precisely remove specific concepts while preserving the generative capabilities of diffusion models. This work proposes Orthogonal Concept Erasure (OCE), the first approach to decouple concept semantics from generative capacity. OCE employs layer-wise closed-form orthogonal transformations to multiplicatively manipulate neuron directions, enabling precise erasure while preserving both magnitude and angular structure. Furthermore, it introduces subspace-level objectives and structured operations to support efficient multi-concept erasure. Experiments demonstrate that OCE significantly outperforms existing methods in both single- and multi-concept tasks, erasing up to 100 concepts in as little as 4.3 seconds while better maintaining the generation quality of non-target content.
Full fine-tuning of large language models often degrades their pre-existing capabilities, and current approaches relying on proxy metrics in weight space struggle to precisely preserve functionally relevant directions. This work proposes FORA, a function-space preservation mechanism that shifts capability protection from weight space to function space. FORA constructs a right projection matrix from the principal components of input activation covariances and combines it with the left projection derived from weight SVD, thereby structurally disentangling update paths for retaining original capabilities and learning new tasks. By imposing orthogonality constraints within capability-induced activation subspaces, FORA significantly outperforms weight-space projection and standard regularization strategies. Evaluated on Qwen3-1.7B, it effectively balances retention of translation and mathematical reasoning abilities with new task acquisition, exhibiting only minor trade-offs in mathematical retention scenarios.
Existing linear probes struggle to uncover the internal encoding structure of geometric information in self-supervised vision Transformers (ViTs). This work proposes a controlled subspace intervention framework that leverages singular value decomposition (SVD) on converged linear probe weights to isolate a low-rank subspace carrying explicit geometric signals. For the first time, subspace analysis reveals distinct differences in geometric representation between DINOv2 and MAE, demonstrating that geometric information is highly compressible, peaks in accuracy at intermediate network layers, and exhibits pronounced low-rank characteristics. These findings provide both theoretical grounding and practical design guidance for lightweight decoders and efficient feature selection strategies in self-supervised vision models.
This work addresses the challenge of precisely controlling specific behaviors—such as refusal or sycophancy—in large language models, where targeted interventions often produce unintended side effects. The authors propose a low-rank subspace diagnostic framework that reveals, for the first time, that distinct behaviors share internal representations in activation space. Through geometric analysis of decision subspaces and the mean squared cosine of principal angles, they demonstrate that intervention effects propagate asymmetrically, depending on the degree of subspace overlap and the angular proximity to the decision subspace. Experiments across multiple instruction-tuned models (7B–70B) show that behaviors exhibiting high representational overlap and closer alignment with the decision subspace are more susceptible to intervention, thereby explaining the fundamental difficulty in achieving independent behavioral control.