Score
Designs and implements models and training objectives that mask parts of latent feature maps, aggregate (pool) remaining latent vectors into summary targets, and predict the masked pooled latent representations. This includes building backbone-agnostic encoders/decoders, loss functions and data pipelines to support pooled masked-latent prediction across varying spatial or temporal scales and long sequence durations.
Existing research on masked image modeling (MIM) for self-supervised visual representation learning lacks a unified formalization of pretraining paradigms and standardized, comparable evaluation. Method: We formally categorize MIM into two principal paradigms—reconstruction-based and contrastive-based—and construct an interpretable, hierarchical taxonomy via expert curation and agglomerative clustering. We conduct systematic, unified benchmarking of over 20 state-of-the-art models on ImageNet and other standard datasets. Contribution/Results: Our analysis identifies critical open challenges—including cross-paradigm integration, long-tailed masking strategies, and compute-accuracy trade-offs. To foster reproducibility and standardization, we publicly release a structured literature repository and an extensible evaluation framework on GitHub. This work establishes foundational infrastructure for rigorous, comparable advancement in MIM research.
This work investigates the root cause of poor “out-of-the-box” performance of masked image modeling (MIM) representations, identifying that the [cls] token in standard Vision Transformers (ViTs) fails to effectively aggregate semantic information due to uniform attention distribution. To address this, we propose Selective Aggregation: instead of relying on a single [cls] token, our method dynamically selects the most discriminative patch tokens based on token-level semantic importance and performs lightweight, learnable aggregation. Crucially, the approach introduces no additional parameters, requires no extra training data, and operates without fine-tuning. On ImageNet-1K linear probing, it achieves an 8.2% relative improvement over baseline MIM representations. This significantly enhances both the generalization capability and plug-and-play usability of self-supervised visual representations, establishing a new paradigm for downstream adaptation of MIM features.
Existing MAE-based skeleton action recognition methods predominantly reconstruct raw joint coordinates, resulting in weak semantic representation and computational redundancy. To address this, we propose the Generalized Feature Prediction (GFP) framework, which replaces low-level coordinate reconstruction with high-level semantic feature prediction to enhance both representational capacity and efficiency. GFP introduces a lightweight dynamic target generation network that constructs multi-level supervision signals in real time, coupled with a constraint-optimization mechanism that ensures feature diversity and prevents representation collapse during end-to-end training. Built upon a spatiotemporal hierarchical masking autoencoding paradigm, GFP eliminates the need for offline precomputation. Evaluated on NTU-60, NTU-120, and PKU-MMD benchmarks, GFP achieves state-of-the-art accuracy while accelerating training by 6.2×, significantly improving downstream task performance and computational efficiency.
This work addresses the limitations of conventional pooling strategies—such as [CLS] or mean pooling—which often bias information toward the beginning of variable-length sequences or dilute locally salient features, thereby struggling to balance short- and long-range contextual modeling. To overcome this, we propose Landmark (LMK) pooling, a novel approach that partitions the input sequence into chunks and inserts learnable landmark tokens between them. Global representations are then derived by mean-pooling the landmark embeddings, effectively integrating local saliency with global context. Experiments on Transformer encoders demonstrate that LMK matches state-of-the-art performance on short-context retrieval tasks while significantly outperforming existing methods in long-context scenarios, confirming its effectiveness, balance, and scalability for dense embedding pooling.
The impact of pooling operations on representational capacity and task performance in Transformer models has long been overlooked. Method: We establish, for the first time, theoretical expressivity bounds for pooling methods and propose a unified analytical framework that characterizes how distinct pooling strategies—e.g., [CLS], mean, and attention-weighted pooling—affect input discriminability, contextual modeling capability, and optimization dynamics. Our analysis spans three modalities—NLP, computer vision, and time series—and encompasses multiple attention variants across diverse downstream tasks. Contribution/Results: Empirical evaluation reveals that pooling choice significantly influences accuracy, gradient sensitivity, and convergence stability. Crucially, we identify task-agnostic, high-performing pooling patterns that generalize consistently across modalities and tasks. This work provides both theoretical foundations and practical guidelines for task-aware pooling design in Transformer architectures.
Conventional Transformer embedding pooling methods (e.g., Avg, Max, CLS token) suffer severe performance degradation under varying signal-to-noise ratio (SNR), limiting robustness in noisy real-world settings. Method: This paper proposes an adaptive attention pooling framework grounded in vector quantization (VQ) theory. Unlike static pooling strategies, it formulates embedding aggregation as an optimal VQ problem for signal reconstruction, derives the first theoretical bound on its reconstruction error, and proves that adaptive attention mechanisms can asymptotically approach this theoretical optimum. Contribution/Results: Evaluated on a synthetically generated SNR-controllable dataset and cross-domain benchmarks—including relational reasoning, multi-agent reinforcement learning, and visual recognition—the method substantially mitigates signal distortion under low-SNR conditions. It improves model robustness by 23–41% across multiple benchmarks and reduces performance variance by over 50%, effectively overcoming the SNR sensitivity inherent in traditional pooling schemes.
This work addresses the high storage and memory costs of late interaction models, which generate numerous token-level vectors per document. To mitigate this, the authors propose a lightweight, pooling-aware fine-tuning approach that incorporates a compression objective during training, enabling flexible compression of multi-vector representations at inference time. By integrating k-means pooling with multi-factor training, the method demonstrates strong transferability across pooling strategies and datasets, and allows a single model to support multiple compression ratios. On the BEIR SciFact benchmark, the model maintains or even improves retrieval accuracy compared to an uncompressed baseline, despite achieving a compression rate of up to 83% (pooling factors 1–6).
Masked image generation models suffer from inefficiency due to multi-step bidirectional attention and loss of continuous semantic information in discrete sampling, while existing acceleration methods introduce significant approximation errors at high speedup ratios. This work proposes MIGM-Shortcut, which, for the first time, formulates feature evolution as a controlled dynamical system. It employs a lightweight network to learn an average velocity field derived from the fusion of historical features and already sampled tokens, enabling efficient prediction of future features. By transcending the representational limitations of conventional caching-based approximations, the method achieves over 4× acceleration on mainstream architectures such as Lumina-DiMOO while preserving generation quality, substantially advancing the efficiency–quality Pareto frontier.
Existing cross-layer encoders struggle to capture high-level semantics that span multiple layers, as their latent variables are often confined to a single or few layers, leading to representations biased toward superficial patterns. This work proposes fmxcoders, which construct shared cross-layer bases via low-rank tensor decomposition and incorporate stochastic layer masking regularization to enforce coordinated activation of latent variables across layers, thereby recovering genuinely cross-layer semantic features. By combining factorized parameterization with layer-axis denoising regularization, the method substantially enhances functional consistency and semantic interpretability. Evaluated on models ranging from GPT2-Small to Gemma2-2B, fmxcoders achieve average probe F1 score improvements of 10–30 points, reduce reconstruction MSE by 25%–50%, double functional consistency, and increase the number of semantically coherent latent variables by 3–13 times.
This work addresses the limited predictability of latent spaces in existing world models, which stems from the decoupling of representation learning and dynamics prediction. We propose an end-to-end joint training framework that integrates vision foundation models with flow matching generative models to synergistically optimize the latent encoder and the generative dynamics model, thereby directly shaping representations amenable to temporal prediction. Furthermore, a collapse-prevention mechanism is introduced to eliminate the reliance on two-stage training pipelines. The proposed approach significantly enhances long-horizon temporal coherence, consistently outperforming existing baselines across multi-task and high-resolution scenarios.
This study addresses the prohibitive computational overhead of generative modeling with representation autoencoders caused by dense token grids. To this end, we propose PoolDINO, a framework that introduces a learnable affine pooling operator to merge adjacent tokens, achieving efficient compression by exploiting local feature correlations. Furthermore, PoolDINO jointly trains an RGB decoder with an internally guided diffusion model, eliminating the need for an independent feature autoencoder while preserving the standard two-stage pipeline and significantly simplifying the architecture. Experiments on ImageNet demonstrate that PoolDINO achieves 4× token compression without compromising generation quality, yielding a 3.7× to 9.0× improvement in sampling throughput. These results indicate that the proposed method effectively balances generative efficiency with visual fidelity.