Score
Designs and implements layer-wise activation analyses and neural representation probes to measure and visualize how a model’s internal features change as a result of fine-tuning. Builds experiments, similarity metrics, diagnostic classifiers, and visualizations that compare activations across layers, units, or directions to quantify which representations are preserved, adapted, or newly formed after fine-tuning.
This paper systematically investigates the quantification of neural network model similarity, addressing both representational similarity (intermediate-layer activations) and functional similarity (output behavior). It unifies and comparatively analyzes mainstream metrics—such as CKA, SVCCA, PWCCA, linear probes, and top-k output agreement—across these two complementary paradigms. A structured taxonomy is introduced, accompanied by theoretical analysis of each metric’s mathematical properties, interrelationships, and applicability boundaries. Empirical evaluation assesses their explanatory power and limitations in downstream tasks including model compression, ensemble learning, and robustness analysis. The core contribution is a cross-paradigm benchmark enabling rigorous metric comparison, revealing how metric selection critically influences downstream conclusions. The work further identifies open challenges and proposes principled evaluation criteria, thereby advancing methodological foundations for model behavior interpretation and trustworthy AI. (149 words)
研究通过分析注意力模式和层激活来探索微调如何改变大型语言模型的内部表示,并检查这些变化与任务相关组件的关系。
This study investigates the mechanistic degradation and reversibility of language models under toxic data fine-tuning. Toxic fine-tuning induces model corruption, yet its underlying neural mechanisms and potential for recovery remain poorly understood. Method: Leveraging causal tracing and circuit localization—key techniques from mechanistic interpretability—alongside task-specific fine-tuning and clean-data reverse retraining, we conduct controlled ablation and reconstruction experiments. Results: We establish, for the first time, that corruption exhibits *circuit-level specificity*: only critical computational pathways are selectively impaired, while peripheral circuits remain intact. Crucially, we demonstrate *neuroplastic-like recoverability*: clean-data retraining reconstructs original functional mechanisms with >89% restoration fidelity; this recovery generalizes across fine-tuning epochs. Contribution: Our work identifies precise circuit-level localization principles governing corruption and empirically validates the reversibility of mechanistic damage—providing both theoretical foundations and actionable strategies for robust alignment and trustworthy fine-tuning.
This work addresses the problem of interpretable concept probing—i.e., detecting human-defined semantic concepts—in neural network representations. Conventional approaches rely on heuristic, manual layer selection, resulting in unstable and poorly generalizable interpretations. To overcome this limitation, we propose an automatic layer selection method grounded in two complementary representational properties: *informativeness* (quantifying a layer’s discriminative power for the target concept) and *regularity* (measuring structural consistency of concept-related representations). We formulate a joint evaluation framework that jointly optimizes these metrics to identify the optimal probing layer. Extensive experiments across diverse architectures (e.g., ResNet, ViT) and benchmarks (ImageNet, CUB) demonstrate that our method significantly improves probing accuracy, cross-model robustness, interpretability, and reproducibility. By replacing ad hoc layer selection with a principled, representation-aware criterion, this work establishes a new paradigm for trustworthy AI layer localization.
This study addresses the challenges of scarce labeled data and limited model generalizability in early Alzheimer’s disease (AD) detection by systematically evaluating and fine-tuning large language models—including BERT, T5, and Llama-1B—using a novel multi-loss supervised fine-tuning strategy. The approach is trained and validated across three heterogeneous clinical corpora: Pitt, CCC, and ADRC. Through linear probing and cross-corpus transfer analyses, the work demonstrates that fine-tuning substantially enhances the models’ ability to encode AD-related linguistic signals. Notably, decoder-only architectures such as Llama-1B exhibit competitive or even superior performance compared to encoder-decoder models on this task. The method achieves new state-of-the-art results on both the Pitt and CCC datasets and shows strong performance on ADRC.
High barriers to adopting pre-trained models and a lack of empirical guidance for strategy selection hinder practical deployment in few-shot image classification and object detection. Method: We systematically compare linear probing versus fine-tuning across ResNet, MobileNet, and EfficientNet, and propose an end-to-end TensorFlow framework integrating multi-scale feature-space visualization (PCA, t-SNE, UMAP) to unify analysis of representation evolution. Contribution/Results: Linear probing significantly outperforms fine-tuning under extreme data scarcity (≤100 samples per class) while accelerating training by 3–5×. The framework enables high-accuracy, rapid deployment (<1 hour for fine-tuning) on standard benchmarks (ImageNet-1K, CIFAR-100), balancing beginner-friendly usability with expert-level extensibility. It bridges the gap between theoretical representation analysis and real-world engineering practice.
This work addresses the lack of an end-to-end theoretical understanding of how pretraining initialization influences feature reuse and learning during fine-tuning. By constructing an analytical framework for pretraining–fine-tuning dynamics in diagonal linear networks, the authors derive exact expressions for generalization error as a function of initialization scale and task statistics, thereby revealing— for the first time—how initialization shapes the inductive bias of fine-tuning. The analysis identifies four distinct fine-tuning mechanisms governed primarily by the scale of initialization and demonstrates that shallow, small-scale initializations confer an advantage in subset-feature tasks. These theoretical predictions are validated through experiments on CIFAR-100 with nonlinear networks, confirming that the distribution of initialization scales significantly modulates fine-tuning generalization performance.
This work investigates the degradation of intermediate-layer representations in Vision Transformers (ViTs) under out-of-distribution (OOD) conditions and identifies optimal probing locations. Through large-scale linear probing experiments, the authors systematically evaluate the representational capacity of different layers and modules in pretrained ViTs across varying degrees of distribution shift. They find that performance deterioration in deeper layers is primarily attributable to distributional shifts. Notably, at the module granularity, the internal activations of the feedforward network and the normalized outputs of multi-head attention emerge as optimal probing points under strong and weak distribution shifts, respectively—challenging the conventional practice of probing only block outputs. Extensive validation on multiple image classification benchmarks demonstrates that this probing strategy significantly improves downstream OOD performance.
Neural network decision interpretability is critically needed in safety-critical applications, yet existing methods lack rigorous validation under industrial conditions. Method: This paper systematically evaluates Feature-Guided Analysis (FGA) for industrial applicability, conducting empirical assessments on MNIST and the Label-Specific Classification (LSC) benchmark. Using neuron activation monitoring and rule extraction, we analyze FGA’s robustness across diverse neural architectures, training strategies, and feature selection schemes. Contribution/Results: We find that model architecture significantly affects recall but has limited impact on precision. FGA achieves superior precision (+3.2% over state-of-the-art) and cross-dataset stability on both benchmarks. Crucially, it demonstrates consistent, trustworthy explanatory capability under heterogeneous industrial conditions—marking the first such validation. This work provides essential empirical evidence bridging the gap between FGA’s theoretical promise and real-world deployment.
Although supervised fine-tuning (SFT) exerts only a subtle effect on the cosine similarity of hidden activations in large language models, their internal representations may nonetheless undergo substantial changes. This work proposes a high-resolution mechanistic analysis framework based on pretrained sparse autoencoders (SAEs), integrating representational geometry with layer-wise feature tracking. For the first time, it reveals systematic semantic feature shifts induced by SFT within sparse latent spaces and identifies layer-update patterns uniquely associated with safety alignment. The method precisely localizes key semantic features whose distributions are altered by SFT. All code and analyses are publicly released.
This study addresses the challenge that intrinsic structures within neural network representations, independent of task outputs, remain difficult to observe directly. To this end, this work proposes a loss-invariant passive probing paradigm that employs fixed, untrained random projections as passive probes. Combined with S² parameterization analysis and ensemble learning techniques, this approach enables real-time monitoring of geometric evolution and feature accessibility in neural representations without interfering with training. The proposed method reveals distinct evolutionary patterns of information accessibility across regression and classification tasks, effectively disentangling the coupled effects of representational geometry and task difficulty on accessibility. Furthermore, it demonstrates that these passive probes exhibit superior structural independence compared to conventional linear probes.