learn robust representations

Design and implement representation-learning methods and training objectives that produce latent features which are stable and predictive across changing environments. This work includes algorithms and losses that infer or enforce shared latent variables, balance per-environment objectives, and disentangle prevalence or distributional shifts from underlying mechanisms so the learned features generalize to new environments.

learnrobustrepresentations

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses a critical limitation of traditional causal invariant representation learning, which fails when environmental variables directly influence the prediction target due to its reliance on the “no direct environmental effect” assumption. To overcome this constraint, the authors propose a novel paradigm that explicitly models distributional shifts across environments and marginalizes out environmental variables during representation learning. Grounded in a generalized random intercept model, the method provides theoretically analyzable guarantees on generalization performance. Extensive experiments across diverse and challenging multi-environment settings demonstrate that the approach significantly outperforms existing invariant learning methods, exhibiting superior robustness and effectiveness in out-of-distribution generalization.

causal invariancedistribution shiftenvironmental variation

This work investigates how predictive representation learning objectives often discard exogenous features that, while unpredictable, are relevant for control due to their inherent bias toward predictability. Through a 2×2 controlled experimental design that independently manipulates feature controllability and control relevance, the study systematically evaluates six representation learning objectives—including JEPA, action-conditional JEPA, inverse dynamics models, and their reward- or controllability-augmented variants—on their ability to preserve such features. The analysis reveals, for the first time, the precise failure mechanism by which predictive objectives neglect control-relevant information. To address this, the authors propose Reward-Anchored JEPA, which leverages as little as 2% reward-labeled data. Empirical results demonstrate that all purely predictive, reward-free objectives fail to retain the feature (performing near chance), whereas the proposed method robustly recovers its representation across multiple environments and latent space dimensions.

control-relevanceexogenous featuresJEPA

This work addresses the degradation in generalization performance in multi-environment prediction tasks caused by distributional shifts in latent variables. The authors propose a Bayesian modeling–based approach for learning environment-robust representations, grounded in the assumption that while the data-generating mechanism remains invariant across environments, the distributions of latent variables may vary. Their method introduces a variational objective augmented with a cross-environment balancing term, employs empirical Bayes to automatically determine the prior, and leverages amortized inference for efficient computation. Evaluated on diverse real-world tasks—including astronomical source identification, microbiome-based disease detection, and sepsis prediction in intensive care units—the proposed approach consistently outperforms existing models, demonstrating superior cross-environment predictive accuracy and transferability.

distribution shiftempirical Bayesenvironment-robust representation

Existing latent variable models often suffer from under-constrained objectives, leading to non-identifiable, ambiguous, and poorly interpretable representations. This work proposes the Constrained Latent State Modeling (CLSM) framework, which systematically integrates six core constraints—namely predictive sufficiency, minimality, temporal consistency, and others—for the first time. Grounded in information theory and dynamical systems theory, CLSM formally characterizes the intrinsic couplings and trade-offs among these constraints. By reframing representation learning as a constrained optimization problem, the framework unifies diverse approaches such as variational autoencoders and state-space models, revealing that non-identifiability stems from insufficient constraints rather than technical shortcomings. CLSM thus provides a principled foundation for designing latent variable models that are interpretable, robust, and aligned with downstream tasks.

constraintsidentifiabilitylatent state modeling

Training objective drives the consistency of representational similarity across datasets

Nov 08, 2024
LC
Laure Ciernik
🏛️ Technische Universität Berlin | Aignostics | Anthropic | Google DeepMind

This work investigates whether cross-dataset consistency in model representation similarity stems from intrinsic model properties or is confounded by biases inherent in common benchmark datasets. To address this, we conduct systematic representation comparison experiments across multimodal (image, image-text) and multitask (self-supervised, classification, image-text contrastive) models, using Centered Kernel Alignment (CKA) and linearly weighted similarity analysis on diverse domain-shifted datasets. Results demonstrate that training objective is the dominant factor governing cross-dataset representation similarity stability—significantly outweighing influences of data modality and network architecture. We propose the first evaluation framework explicitly designed for cross-dataset representational consistency. Furthermore, we reveal that self-supervised vision models exhibit the strongest generalization of representation similarity across datasets, and that the correlation between representation similarity and task performance is maximized on single-domain benchmarks.

Analyzes link between model representations and task behaviorExamines impact of objective function on similarity consistencyMeasures how representational similarity varies across datasets

Latest Papers

What's happening recently
View more

This work addresses a critical gap in representation learning: while much research focuses on refining existing representations, little attention has been paid to the mechanisms that trigger the emergence of representations at new levels of abstraction. The paper proposes that when current representations fail to account for observed organizational structures or dynamic patterns, the system actively initiates the generation of novel representations, thereby enabling a recursive, self-bootstrapping process of representational evolution. For the first time, the authors formalize “explanatory insufficiency” as a positive driving force behind this evolution and introduce a five-stage framework grounded in cognitive and systems theory to model representational emergence. This framework applies broadly to systems such as world models and foundation models, offering a new paradigm for AI design—one that endows systems with the capacity to recognize the limits of their own representations, thereby enabling more autonomous representation learning and scientific discovery.

anomaly detectionexplanatory insufficiencyrepresentation learning

This work addresses a critical limitation in self-supervised dynamic representation learning, where existing contrastive predictive objectives often misinterpret slowly varying noise within trajectories as genuine dynamical signals, leading to noise-dominated representations and degraded downstream performance. The authors identify this issue as stemming from an inherent inductive bias flaw in standard contrastive objectives and propose a general corrective principle: sampling negative examples from within the same trajectory to eliminate predictive shortcuts introduced by slow-varying noise, thereby compelling the encoder to focus on the true dynamical variables governing system evolution. Experiments based on frameworks such as JEPA and DySIB on synthetic moving-point and rigid-pendulum video datasets demonstrate that the proposed approach effectively disentangles slow noise from authentic dynamics, yields representations whose quality improves with trajectory length, and significantly enhances downstream task performance under strong noise conditions.

contrastive learningdynamicsinductive bias

Traditional performance-based analyses struggle to uncover the intrinsic organizational principles and evolutionary trajectories of adaptive biological systems. To address this limitation, this work proposes a guided five-tier progressive representation framework that systematically integrates observable performance, dynamic organization, latent structure, longitudinal feasibility, and internal predictive approximation. By innovatively incorporating the notion of “bootstrapping” at both methodological and epistemological levels, the framework synergistically combines latent space representation learning, longitudinal dynamic modeling, and multi-level abstract reasoning. Its efficacy is demonstrated through an illustrative case study on gait occlusion. This study formalizes a principled pathway from superficial performance metrics to the system’s intrinsic feasibility, offering a generalizable representation learning paradigm for adaptive biological systems.

adaptive biological systemslatent-space representationlongitudinal viability

This study addresses the challenge of ensuring physical quantity recovery and command response consistency in latent world models for vehicle control. We propose a decoupled evaluation framework for action-conditioned latent predictors based on the temporal Joint Embedding Predictive Architecture (JEPA). Leveraging IPG CarMaker simulation data and a physics readout mechanism, this framework independently assesses state retention, geometric organization, prediction accuracy, and local responsiveness using untrained encoder baselines, thereby precisely localizing error sources within either the representation or the predictor. Our analysis reveals an inherent trade-off between retention and responsiveness, demonstrating that prediction error alone is insufficient for evaluating control suitability. Ultimately, this work establishes a rigorous foundation for closed-loop control assessment in latent space.

controllatent world modelslocal response

研究通过提出预测一致性假设,利用DINO世界模型实验,探索了不同模型如何趋向共享的潜在结构,解决了世界模型学习表征本质理解不足的问题。

Latent StructurePlatonic Representation HypothesisPredictive Consistency

Hot Scholars

ZZ

Zhuocheng Zhang

Institute of Computing Technology, Chinese Academy of Science
Natural Language Processing
YZ

Yifan Zhu

Beijing University of Posts and Telecommunications
PEFT of LLMsGraph RAGGraph mining
ZL

Ziwei Liu

Associate Professor, Nanyang Technological University
Computer VisionMachine LearningComputer Graphics
JL

Jeng-Lin Li

AI Center Lead of Inventec; Adjunct Assistant Professor of NTHU-BAI
machine learningreliable AImultimodalhealth analytics
ZG

Zhen Gao

Beijing Institute of Technology
Generative AI6GMIMO communicationsIoT edge computing