diagnose representation shifts after fine-tuning

Designs and implements layer-wise activation analyses and neural representation probes to measure and visualize how a model’s internal features change as a result of fine-tuning. Builds experiments, similarity metrics, diagnostic classifiers, and visualizations that compare activations across layers, units, or directions to quantify which representations are preserved, adapted, or newly formed after fine-tuning.

diagnoserepresentationshiftsafter

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.04
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Neuroplasticity and Corruption in Model Mechanisms: A Case Study Of Indirect Object Identification

Feb 27, 2025
VK
Vishnu Kabir Chhabra
🏛️ The Ohio State University

This study investigates the mechanistic degradation and reversibility of language models under toxic data fine-tuning. Toxic fine-tuning induces model corruption, yet its underlying neural mechanisms and potential for recovery remain poorly understood. Method: Leveraging causal tracing and circuit localization—key techniques from mechanistic interpretability—alongside task-specific fine-tuning and clean-data reverse retraining, we conduct controlled ablation and reconstruction experiments. Results: We establish, for the first time, that corruption exhibits *circuit-level specificity*: only critical computational pathways are selectively impaired, while peripheral circuits remain intact. Crucially, we demonstrate *neuroplastic-like recoverability*: clean-data retraining reconstructs original functional mechanisms with >89% restoration fidelity; this recovery generalizes across fine-tuning epochs. Contribution: Our work identifies precise circuit-level localization principles governing corruption and empirically validates the reversibility of mechanistic damage—providing both theoretical foundations and actionable strategies for robust alignment and trustworthy fine-tuning.

Identification of primary corruption mechanisms during toxic fine-tuning.Impact of fine-tuning on poisoned data and mechanism changes.Neuroplasticity behaviors in models retrained on clean datasets.

Concept Probing: Where to Find Human-Defined Concepts (Extended Version)

Jul 24, 2025
MD
Manuel de Sousa Ribeiro
🏛️ NOVA LINCS | NOVA School of Science and Technology | NOVA University Lisbon

This work addresses the problem of interpretable concept probing—i.e., detecting human-defined semantic concepts—in neural network representations. Conventional approaches rely on heuristic, manual layer selection, resulting in unstable and poorly generalizable interpretations. To overcome this limitation, we propose an automatic layer selection method grounded in two complementary representational properties: *informativeness* (quantifying a layer’s discriminative power for the target concept) and *regularity* (measuring structural consistency of concept-related representations). We formulate a joint evaluation framework that jointly optimizes these metrics to identify the optimal probing layer. Extensive experiments across diverse architectures (e.g., ResNet, ViT) and benchmarks (ImageNet, CUB) demonstrate that our method significantly improves probing accuracy, cross-model robustness, interpretability, and reproducibility. By replacing ad hoc layer selection with a principled, representation-aware criterion, this work establishes a new paradigm for trustworthy AI layer localization.

Evaluate layer representations for informativeness and regularityIdentify optimal neural network layer for concept probingValidate method across diverse models and datasets

This study addresses the challenges of scarce labeled data and limited model generalizability in early Alzheimer’s disease (AD) detection by systematically evaluating and fine-tuning large language models—including BERT, T5, and Llama-1B—using a novel multi-loss supervised fine-tuning strategy. The approach is trained and validated across three heterogeneous clinical corpora: Pitt, CCC, and ADRC. Through linear probing and cross-corpus transfer analyses, the work demonstrates that fine-tuning substantially enhances the models’ ability to encode AD-related linguistic signals. Notably, decoder-only architectures such as Llama-1B exhibit competitive or even superior performance compared to encoder-decoder models on this task. The method achieves new state-of-the-art results on both the Pitt and CCC datasets and shows strong performance on ADRC.

Alzheimer's diseaseearly detectionlarge language models

Deep Learning and Machine Learning, Advancing Big Data Analytics and Management: Tensorflow Pretrained Models

Sep 20, 2024
KC
Keyu Chen
🏛️ Georgia Institute of Technology | Indiana University | Kyoto University | AppCubic | Rutgers University | Purdue University | University of Wisconsin-Madison | National Taiwan Normal University

High barriers to adopting pre-trained models and a lack of empirical guidance for strategy selection hinder practical deployment in few-shot image classification and object detection. Method: We systematically compare linear probing versus fine-tuning across ResNet, MobileNet, and EfficientNet, and propose an end-to-end TensorFlow framework integrating multi-scale feature-space visualization (PCA, t-SNE, UMAP) to unify analysis of representation evolution. Contribution/Results: Linear probing significantly outperforms fine-tuning under extreme data scarcity (≤100 samples per class) while accelerating training by 3–5×. The framework enables high-accuracy, rapid deployment (<1 hour for fine-tuning) on standard benchmarks (ImageNet-1K, CIFAR-100), balancing beginner-friendly usability with expert-level extensibility. It bridges the gap between theoretical representation analysis and real-world engineering practice.

Comparing linear probing versus fine-tuning approaches in transfer learningExploring TensorFlow pre-trained models for image classification tasksProviding practical guidance and code examples for deep learning implementation

Latest Papers

What's happening recently
View more

This work addresses the lack of an end-to-end theoretical understanding of how pretraining initialization influences feature reuse and learning during fine-tuning. By constructing an analytical framework for pretraining–fine-tuning dynamics in diagonal linear networks, the authors derive exact expressions for generalization error as a function of initialization scale and task statistics, thereby revealing— for the first time—how initialization shapes the inductive bias of fine-tuning. The analysis identifies four distinct fine-tuning mechanisms governed primarily by the scale of initialization and demonstrates that shallow, small-scale initializations confer an advantage in subset-feature tasks. These theoretical predictions are validated through experiments on CIFAR-100 with nonlinear networks, confirming that the distribution of initialization scales significantly modulates fine-tuning generalization performance.

feature learningfine-tuninginductive bias

This work investigates the degradation of intermediate-layer representations in Vision Transformers (ViTs) under out-of-distribution (OOD) conditions and identifies optimal probing locations. Through large-scale linear probing experiments, the authors systematically evaluate the representational capacity of different layers and modules in pretrained ViTs across varying degrees of distribution shift. They find that performance deterioration in deeper layers is primarily attributable to distributional shifts. Notably, at the module granularity, the internal activations of the feedforward network and the normalized outputs of multi-head attention emerge as optimal probing points under strong and weak distribution shifts, respectively—challenging the conventional practice of probing only block outputs. Extensive validation on multiple image classification benchmarks demonstrates that this probing strategy significantly improves downstream OOD performance.

distribution shiftintermediate layersout-of-distribution

Feature-Guided Analysis of Neural Networks: A Replication Study

Oct 28, 2025
FF
Federico Formica
🏛️ McMaster University | University of Bergamo

Neural network decision interpretability is critically needed in safety-critical applications, yet existing methods lack rigorous validation under industrial conditions. Method: This paper systematically evaluates Feature-Guided Analysis (FGA) for industrial applicability, conducting empirical assessments on MNIST and the Label-Specific Classification (LSC) benchmark. Using neuron activation monitoring and rule extraction, we analyze FGA’s robustness across diverse neural architectures, training strategies, and feature selection schemes. Contribution/Results: We find that model architecture significantly affects recall but has limited impact on precision. FGA achieves superior precision (+3.2% over state-of-the-art) and cross-dataset stability on both benchmarks. Crucially, it demonstrates consistent, trustworthy explanatory capability under heterogeneous industrial conditions—marking the first such validation. This work provides essential empirical evidence bridging the gap between FGA’s theoretical promise and real-world deployment.

Analyzing how neural network architecture and feature selection affect FGA performanceAssessing applicability of Feature-Guided Analysis on MNIST and LSC benchmark datasetsEvaluating effectiveness of FGA in computing rules explaining neural network behavior

Although supervised fine-tuning (SFT) exerts only a subtle effect on the cosine similarity of hidden activations in large language models, their internal representations may nonetheless undergo substantial changes. This work proposes a high-resolution mechanistic analysis framework based on pretrained sparse autoencoders (SAEs), integrating representational geometry with layer-wise feature tracking. For the first time, it reveals systematic semantic feature shifts induced by SFT within sparse latent spaces and identifies layer-update patterns uniquely associated with safety alignment. The method precisely localizes key semantic features whose distributions are altered by SFT. All code and analyses are publicly released.

Activation GeometryRepresentational DivergenceSafety Alignment

This study addresses the challenge that intrinsic structures within neural network representations, independent of task outputs, remain difficult to observe directly. To this end, this work proposes a loss-invariant passive probing paradigm that employs fixed, untrained random projections as passive probes. Combined with S² parameterization analysis and ensemble learning techniques, this approach enables real-time monitoring of geometric evolution and feature accessibility in neural representations without interfering with training. The proposed method reveals distinct evolutionary patterns of information accessibility across regression and classification tasks, effectively disentangling the coupled effects of representational geometry and task difficulty on accessibility. Furthermore, it demonstrates that these passive probes exhibit superior structural independence compared to conventional linear probes.

learned representationsloss-invariant projectionspassive probes

Hot Scholars

FS

Fei Shen

National University of Singapore
Controllable GenerationMultimodal Safety
JR

Jean-Rémi King

Meta
neuroscienceartificial intelligencehuman cognitiondecoding
SL

Sharon Li

University of Wisconsin-Madison
Machine learningReliable AI
HS

Helmut Schmid

Associate Professor, Ludwig-Maximilians-Universität München
Natural Language Processing