pooled masked latent prediction

Designs and implements models and training objectives that mask parts of latent feature maps, aggregate (pool) remaining latent vectors into summary targets, and predict the masked pooled latent representations. This includes building backbone-agnostic encoders/decoders, loss functions and data pipelines to support pooled masked-latent prediction across varying spatial or temporal scales and long sequence durations.

pooledmaskedlatentprediction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.52
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Beyond [cls]: Exploring the true potential of Masked Image Modeling representations

Dec 04, 2024
MP
Marcin Przewiȩźlikowski
🏛️ Jagiellonian University | Brown University | AGH University of Science and Technology

This work investigates the root cause of poor “out-of-the-box” performance of masked image modeling (MIM) representations, identifying that the [cls] token in standard Vision Transformers (ViTs) fails to effectively aggregate semantic information due to uniform attention distribution. To address this, we propose Selective Aggregation: instead of relying on a single [cls] token, our method dynamically selects the most discriminative patch tokens based on token-level semantic importance and performs lightweight, learnable aggregation. Crucially, the approach introduces no additional parameters, requires no extra training data, and operates without fine-tuning. On ImageNet-1K linear probing, it achieves an 8.2% relative improvement over baseline MIM representations. This significantly enhances both the generalization capability and plug-and-play usability of self-supervised visual representations, establishing a new paradigm for downstream adaptation of MIM features.

Analyzing ineffective attention aggregation in MIM representationsInvestigating poor out-of-the-box performance of Masked Image ModelingProposing Selective Aggregation to enhance MIM feature utilization

Towards Efficient General Feature Prediction in Masked Skeleton Modeling

Sep 03, 2025
SS
Shengkai Sun
🏛️ Hefei University of Technology | Jilin University | Zhejiang Gongshang University | University of Science and Technology of China

Existing MAE-based skeleton action recognition methods predominantly reconstruct raw joint coordinates, resulting in weak semantic representation and computational redundancy. To address this, we propose the Generalized Feature Prediction (GFP) framework, which replaces low-level coordinate reconstruction with high-level semantic feature prediction to enhance both representational capacity and efficiency. GFP introduces a lightweight dynamic target generation network that constructs multi-level supervision signals in real time, coupled with a constraint-optimization mechanism that ensures feature diversity and prevents representation collapse during end-to-end training. Built upon a spatiotemporal hierarchical masking autoencoding paradigm, GFP eliminates the need for offline precomputation. Evaluated on NTU-60, NTU-120, and PKU-MMD benchmarks, GFP achieves state-of-the-art accuracy while accelerating training by 6.2×, significantly improving downstream task performance and computational efficiency.

Ensuring feature diversity while preventing model collapsePredicting high-level features in masked skeleton modelingReplacing low-level reconstruction with semantic representation

This work addresses the limitations of conventional pooling strategies—such as [CLS] or mean pooling—which often bias information toward the beginning of variable-length sequences or dilute locally salient features, thereby struggling to balance short- and long-range contextual modeling. To overcome this, we propose Landmark (LMK) pooling, a novel approach that partitions the input sequence into chunks and inserts learnable landmark tokens between them. Global representations are then derived by mean-pooling the landmark embeddings, effectively integrating local saliency with global context. Experiments on Transformer encoders demonstrate that LMK matches state-of-the-art performance on short-context retrieval tasks while significantly outperforming existing methods in long-context scenarios, confirming its effectiveness, balance, and scalability for dense embedding pooling.

dense embeddingslong-contextpooling

Pool Me Wisely: On the Effect of Pooling in Transformer-Based Models

Oct 02, 2025
SE
Sofiane Ennadir
🏛️ King AI Labs | Microsoft Gaming | NXAI GmbH | Kreditz AB | Amazon

The impact of pooling operations on representational capacity and task performance in Transformer models has long been overlooked. Method: We establish, for the first time, theoretical expressivity bounds for pooling methods and propose a unified analytical framework that characterizes how distinct pooling strategies—e.g., [CLS], mean, and attention-weighted pooling—affect input discriminability, contextual modeling capability, and optimization dynamics. Our analysis spans three modalities—NLP, computer vision, and time series—and encompasses multiple attention variants across diverse downstream tasks. Contribution/Results: Empirical evaluation reveals that pooling choice significantly influences accuracy, gradient sensitivity, and convergence stability. Crucially, we identify task-agnostic, high-performing pooling patterns that generalize consistently across modalities and tasks. This work provides both theoretical foundations and practical guidelines for task-aware pooling design in Transformer architectures.

Analyzing pooling's impact on Transformer model expressivity and capacityEvaluating pooling strategies across vision, NLP, and time-series tasksProviding theoretical and empirical guidance for pooling mechanism selection

Robust Noise Attenuation via Adaptive Pooling of Transformer Outputs

Jun 10, 2025
GB
Greyson Brothers
🏛️ Johns Hopkins University | Applied Physics Laboratory

Conventional Transformer embedding pooling methods (e.g., Avg, Max, CLS token) suffer severe performance degradation under varying signal-to-noise ratio (SNR), limiting robustness in noisy real-world settings. Method: This paper proposes an adaptive attention pooling framework grounded in vector quantization (VQ) theory. Unlike static pooling strategies, it formulates embedding aggregation as an optimal VQ problem for signal reconstruction, derives the first theoretical bound on its reconstruction error, and proves that adaptive attention mechanisms can asymptotically approach this theoretical optimum. Contribution/Results: Evaluated on a synthetically generated SNR-controllable dataset and cross-domain benchmarks—including relational reasoning, multi-agent reinforcement learning, and visual recognition—the method substantially mitigates signal distortion under low-SNR conditions. It improves model robustness by 23–41% across multiple benchmarks and reduces performance variance by over 50%, effectively overcoming the SNR sensitivity inherent in traditional pooling schemes.

Design pooling methods for transformer outputs to handle noiseDevelop adaptive pooling to minimize signal loss in noisy inputsImprove robustness of transformer models in varying signal-to-noise conditions

Latest Papers

What's happening recently
View more

This work addresses the high storage and memory costs of late interaction models, which generate numerous token-level vectors per document. To mitigate this, the authors propose a lightweight, pooling-aware fine-tuning approach that incorporates a compression objective during training, enabling flexible compression of multi-vector representations at inference time. By integrating k-means pooling with multi-factor training, the method demonstrates strong transferability across pooling strategies and datasets, and allows a single model to support multiple compression ratios. On the BEIR SciFact benchmark, the model maintains or even improves retrieval accuracy compared to an uncompressed baseline, despite achieving a compression rate of up to 83% (pooling factors 1–6).

late interaction modelsmulti-vector compressionretrieval accuracy

Masked image generation models suffer from inefficiency due to multi-step bidirectional attention and loss of continuous semantic information in discrete sampling, while existing acceleration methods introduce significant approximation errors at high speedup ratios. This work proposes MIGM-Shortcut, which, for the first time, formulates feature evolution as a controlled dynamical system. It employs a lightweight network to learn an average velocity field derived from the fusion of historical features and already sampled tokens, enabling efficient prediction of future features. By transcending the representational limitations of conventional caching-based approximations, the method achieves over 4× acceleration on mainstream architectures such as Lumina-DiMOO while preserving generation quality, substantially advancing the efficiency–quality Pareto frontier.

AccelerationFeature RedundancyImage Synthesis

Existing cross-layer encoders struggle to capture high-level semantics that span multiple layers, as their latent variables are often confined to a single or few layers, leading to representations biased toward superficial patterns. This work proposes fmxcoders, which construct shared cross-layer bases via low-rank tensor decomposition and incorporate stochastic layer masking regularization to enforce coordinated activation of latent variables across layers, thereby recovering genuinely cross-layer semantic features. By combining factorized parameterization with layer-axis denoising regularization, the method substantially enhances functional consistency and semantic interpretability. Evaluated on models ranging from GPT2-Small to Gemma2-2B, fmxcoders achieve average probe F1 score improvements of 10–30 points, reduce reconstruction MSE by 25%–50%, double functional consistency, and increase the number of semantically coherent latent variables by 3–13 times.

cross-layer featuresfunctional coherencelatent interpretability

This work addresses the limited predictability of latent spaces in existing world models, which stems from the decoupling of representation learning and dynamics prediction. We propose an end-to-end joint training framework that integrates vision foundation models with flow matching generative models to synergistically optimize the latent encoder and the generative dynamics model, thereby directly shaping representations amenable to temporal prediction. Furthermore, a collapse-prevention mechanism is introduced to eliminate the reliance on two-stage training pipelines. The proposed approach significantly enhances long-horizon temporal coherence, consistently outperforming existing baselines across multi-task and high-resolution scenarios.

Future Scene PredictionLatent RepresentationsTemporal Predictability

This study addresses the prohibitive computational overhead of generative modeling with representation autoencoders caused by dense token grids. To this end, we propose PoolDINO, a framework that introduces a learnable affine pooling operator to merge adjacent tokens, achieving efficient compression by exploiting local feature correlations. Furthermore, PoolDINO jointly trains an RGB decoder with an internally guided diffusion model, eliminating the need for an independent feature autoencoder while preserving the standard two-stage pipeline and significantly simplifying the architecture. Experiments on ImageNet demonstrate that PoolDINO achieves 4× token compression without compromising generation quality, yielding a 3.7× to 9.0× improvement in sampling throughput. These results indicate that the proposed method effectively balances generative efficiency with visual fidelity.

Computational EfficiencyDiffusion ModelsImage Generation