hidden state extraction

Extracting, supervising, and using latent representations from model hidden states to encode past interactions, detect anomalous rollouts or signals, and transfer representational insights across architectures.

hiddenstateextraction

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Existing latent variable models often suffer from under-constrained objectives, leading to non-identifiable, ambiguous, and poorly interpretable representations. This work proposes the Constrained Latent State Modeling (CLSM) framework, which systematically integrates six core constraints—namely predictive sufficiency, minimality, temporal consistency, and others—for the first time. Grounded in information theory and dynamical systems theory, CLSM formally characterizes the intrinsic couplings and trade-offs among these constraints. By reframing representation learning as a constrained optimization problem, the framework unifies diverse approaches such as variational autoencoders and state-space models, revealing that non-identifiability stems from insufficient constraints rather than technical shortcomings. CLSM thus provides a principled foundation for designing latent variable models that are interpretable, robust, and aligned with downstream tasks.

constraintsidentifiabilitylatent state modeling

Strengthening Anomaly Awareness

Apr 15, 2025
AB
Adam Banda
🏛️ Southern Methodist University | University of Manchester | Instituto de Física Corpuscular

To address insufficient sensitivity to unseen anomalies in unsupervised anomaly detection, this paper proposes a two-stage light-supervision enhancement framework. First, a variational autoencoder (VAE) is pretrained unsupervisedly on normal data; second, it undergoes supervised fine-tuning using only a small number of known anomaly samples, guided by reconstruction error minimization. Crucially, this work introduces minimal anomaly supervision—requiring labels *only* for anomalies, with no normal-class labels—directly into the VAE reconstruction objective, thereby significantly improving generalization to previously unseen anomalies. Extensive evaluation across four heterogeneous benchmarks—MNIST, CICIDS, LHCO2020, and SMEFT—demonstrates superior normal/anomalous separation. Notably, on the Higgs event momentum-shift detection task, sensitivity improves markedly, validating the efficacy of few-shot anomaly supervision in enhancing unsupervised models. The approach is particularly suited for high-precision, sensitivity-critical domains such as particle physics detection and cybersecurity.

Enhancing unsupervised anomaly detection with minimal supervisionImproving sensitivity to unseen anomalies in diverse datasetsLeveraging limited labeled anomalies for better model performance

This work proposes a novel approach to time series anomaly detection that addresses the limitations of traditional observation-likelihood-based methods, which often fail to capture structured temporal dynamics and misclassify anomalies as normal patterns. By introducing inductive biases into the latent space of conditional normalizing flows, the method models time series as discrete-time state-space systems, enforcing latent trajectories to conform to prescribed dynamical laws. Anomalies are then defined as deviations from these expected dynamics. The approach frames anomaly detection as a goodness-of-fit test for dynamic consistency—a formulation introduced here for the first time—and evaluates compliance of latent trajectories accordingly. Experiments on both synthetic and real-world datasets demonstrate its effectiveness in detecting anomalies in frequency, amplitude, and noise characteristics, achieving high detection performance alongside strong interpretability.

anomaly detectioninductive biaseslatent space

This study investigates how sequence models can infer latent stochastic states—such as volatility—from noisy observations like financial returns. Within a multivariate stochastic volatility framework, the authors establish the first interpretability benchmark in a controlled stochastic dynamic environment to systematically evaluate diverse neural architectures, loss functions, and output head designs. Their analysis reveals that Transformers decode latent states effectively at specific structural stages, and that a simplified filtering mechanism—comprising linear projection followed by ℓ² normalization—performs well over long horizons. Furthermore, the model exhibits a two-stage internal computation: hidden layers encode the next-step latent state, while the output head maps this representation to squared return predictions. Misalignment between these stages is identified as the primary cause of performance degradation under noise-sensitive MSE training.

latent dynamicsmechanistic interpretabilitypartial observability

Words in Motion: Extracting Interpretable Control Vectors for Motion Transformers

Jun 17, 2024
OS
Omer Sahin Tas
🏛️ FZI Research Center for Information Technology | Karlsruhe Institute of Technology

Transformer hidden states in motion prediction lack clear physical semantics and are difficult to edit controllably. Method: We propose an interpretable control framework leveraging supervised linear probes and sparse autoencoders (SAEs). First, linear probes identify functionally meaningful, physically grounded directions (e.g., velocity, heading) in the latent space. Second, we construct additive, semantically aligned linear control vectors enabling zero-shot generalization to unseen motion patterns. Third, SAEs refine latent representations to enhance control linearity and mechanistic interpretability. Results: Controlled predictions preserve physical plausibility; control responses exhibit high linearity; zero-shot adaptation incurs only millisecond-level inference overhead—no fine-tuning required. This work establishes the first method for extracting semantically explicit, plug-and-play linear control vectors and enabling generalized, interpretable modulation in motion Transformers.

Enable zero-shot generalization with minimal computational overhead.Extract and modify control vectors to influence motion predictions.Interpret hidden states in Transformer-based motion forecasting models.

Latest Papers

What's happening recently
View more

This work addresses the issue of non-identifiable representations in large language model–based world models operating in partially observable environments, where historical information often bypasses latent variables. To resolve this, the study introduces the strict mediation principle—adapted from causal inference—into the textual domain for the first time, enforcing that predictions depend solely on discrete textual latent states and actions, thereby guaranteeing identifiability. By integrating tree-structured reinforcement learning with a factorized GRPO (fGRPO) algorithm, the method learns interpretable, variable-length textual belief states in TextWorld and ScienceWorld. The approach maintains high one-step prediction accuracy while improving representation quality by up to 57% and rollout performance by as much as 98%, with gains amplifying significantly as task complexity and planning horizon increase, offering a theoretically grounded, testable framework for representation quality in textual world models.

history bypasslatent state identifiabilitystrict mediation

This study addresses the problem of recovering latent discrete states from the evolving weights of models trained on time-varying data streams to characterize non-stationary distributional shifts. The proposed approach trains classifiers over sliding time windows, aligns their weight trajectories, and fits a hidden Markov model (HMM) to these trajectories—enabling, for the first time, the identification of semantically coherent temporal phases solely from weight dynamics. Experiments on the Fakeddit and Yelp datasets demonstrate that transfer performance within the same inferred state significantly outperforms cross-state transfer, and this advantage persists independently of temporal proximity and shifts in class distribution. These findings confirm that model weights encode structural information about data distributions that extends beyond local temporal correlations.

data distribution shiftlatent statesmodel weights

This work addresses the challenge of detecting and regulating sycophantic behavior—excessive user flattery—in language models by proposing an iterative data generation method based on cascaded linear samples. Departing from conventional binary contrastive examples, the approach constructs sequences of samples with continuously varying behavioral intensities, revealing for the first time a linearly separable structure of sycophancy in activation space. This enables precise identification and disentanglement of the associated feature subspace. Through activation manipulation and subspace analysis, the method matches or exceeds baseline approaches such as LLM-as-a-judge and system prompting in detection accuracy, calibration, and robust controllability, while incurring lower computational overhead and substantially improving the interpretability of behavioral interventions.

activation steeringbehavior controlinterpretable features

This study investigates whether internal representations of a learned decision-making system encode interpretable concepts that explain its behavior. In the Game of Hidden Rules—a setting without explicit rule labels—the authors train a tokenized autoregressive Transformer agent and, for the first time, apply sparse autoencoders (SAEs) to extract structured features from its decision embeddings. The results demonstrate that individual SAE activation dimensions selectively correspond to fundamental concepts such as shapes and buckets, accounting for the vast majority of relevant decisions. Moreover, these representations reveal interpretable exploratory actions and feedback-driven strategies for switching between hidden rules, offering novel evidence for the interpretability of unsupervised rule-inference agents.

concept discoverydecision-making agentsGame of Hidden Rules

This work addresses the challenge of identifying latent dynamical systems and recovering their governing equations from noisy, high-dimensional observations. The authors propose DYSCO, a novel method that uniquely integrates multi-view temporal contrastive learning with structured function basis parameterization. DYSCO establishes strong identifiability of the latent dynamics—up to affine equivalence—even under nonlinear observations, and enables symbolic equation discovery. By performing symbolic regression within an affine-normalized representation, DYSCO accurately reconstructs both latent trajectories and vector fields. The method demonstrates robust performance across diverse dynamical regimes, including chaotic, oscillatory, and metastable systems, and handles both Gaussian and Poisson noise, with the latter being particularly relevant for neural recording data.

governing equationslatent dynamical systemsnoisy observations

Hot Scholars

SY

Samuel Yen-Chi Chen

Wells Fargo
quantum computationquantum informationmachine learningquantum machine learning
AG

Anubhab Ghosh

Ph.D. student, KTH Royal Institute of Technology, Stockholm, Sweden
Machine learningDeep learningGenerative modelsSystem identification
SS

Shiguang Shan

Professor of Institute of Computing Technology, Chinese Academy of Sciences
Computer VisionPattern RecognitionMachine LearningFace Recognition
SR

Shiva Raj Pokhrel

Marie Curie Fellow, SMIEEE, Deakin University
Gen AI Mobile ComputingQuantum ComputingFederated LearningAndroid/iOS
ZW

Zhongqi Wang

Institute of Computing Technology, Chinese Academy of Sciences
Model Robustness