appearance-conditioned cross-attention

Designs and implements cross-attention modules that inject, modulate, and control appearance- and view-conditioned signals into feature representations and decoders. This includes building interleaved or cross-layer attention schedules, global pre-modulation and gating schemes, view-selective heads or GNN-based cross-attentive encoders, and interfaces for early-fusion or pre-modulated control to govern rendering or reconstruction of visual appearance across viewpoints and lighting variations.

appearance-conditionedcross-attention

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.44
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work investigates how to explicitly control and quantify the information flow within the attention mechanism of Vision Transformers to uncover the evolutionary process from independent local patch processing to the emergence of global representations. The authors introduce a variational information bottleneck along all pathways where attention writes into the residual stream, without altering the model architecture. This approach models the transmitted information as an explicit, tunable variable, enabling precise constraint and intervention on internal communication. Experiments on ImageNet-100 demonstrate that the method effectively modulates the degree of attention coordination and reveals a clear relationship between classification performance and information routing strategies, offering a novel perspective for understanding the internal mechanisms of Vision Transformers.

attention mechanisminformation bottleneckinformation flow

Rethinking Patch Dependence for Masked Autoencoders

Jan 25, 2024
LF
Letian Fu
🏛️ UC Berkeley | UCSF

This work investigates the roles of masked-patch self-attention and masked-to-visible cross-attention in the MAE decoder for representation learning, revealing that image reconstruction primarily relies on global semantic representations extracted by the encoder—not on intra-masked-patch interactions within the decoder. Motivated by this finding, we propose CrossMAE: a streamlined framework that retains only the cross-attention mechanism while entirely removing self-attention among masked tokens in the decoder. This design is the first to empirically demonstrate that MAE’s effectiveness stems from the encoder’s strong global modeling capacity, challenging the prevailing assumption that decoder-side modeling of dependencies among masked patches is essential. Evaluated across ViT-S to ViT-H architectures, CrossMAE matches or surpasses standard MAE in performance while reducing GPU memory consumption by 37% and FLOPs by 42%. Code and pretrained models are publicly available.

Challenges mask token interaction necessity in masked pretrainingExamines inter-patch dependencies in MAE decoders for representation learningProposes CrossMAE using only cross-attention to reduce computation

This work addresses the prevalent hallucination issues in current large vision-language models when handling multi-image tasks, which stem primarily from the locality of attention mechanisms and insufficient cross-image modeling capabilities. To mitigate this, the authors propose a co-optimization framework that integrates architectural and training-level innovations. At the architecture level, an optional image-token interaction attention mechanism is introduced to enable fine-grained cross-image alignment. At the training level, a cross-image contrastive preference learning strategy is designed to reinforce the model’s reliance on authentic visual evidence. This approach represents the first effort to jointly suppress multi-image hallucinations through coordinated structural and objective design, achieving significant performance gains across diverse multi-image tasks while maintaining or slightly improving single-image task performance, thereby demonstrating strong generalization capability.

attention mechanismcross-image modelinghallucination mitigation

Controllable Coupled Image Generation via Diffusion Models

Jun 07, 2025
CY
Chenfei Yuan
🏛️ UC Berkeley

This work addresses the problem of controllable multi-image co-generation, where the goal is to generate multiple images sharing a consistent background while enabling the central object to vary flexibly according to distinct text prompts. The proposed method is a text-driven diffusion model that explicitly decouples background and foreground representations. Its key contributions are: (1) the first introduction of a time-varying weight decoupling mechanism within the cross-attention layers of diffusion models, enabling explicit separation of background and foreground features; and (2) a multi-objective sampling optimization framework that jointly enhances background coupling, text–image alignment, and visual fidelity. Extensive experiments demonstrate that the method significantly outperforms existing approaches in background consistency, text fidelity, and image quality. By enabling precise, prompt-conditioned foreground manipulation without compromising background coherence, it establishes a novel paradigm for controllable image generation.

Control background coupling in multi-image generationDisentangle background and object via cross-attentionOptimize weight parameters for alignment and quality

Information Bottleneck Approach to Spatial Attention Learning

Aug 01, 2021
QL
Qiuxia Lai
🏛️ The Chinese University of Hong Kong | University of Electronic Science and Technology of China

Deep neural networks lack biologically plausible selective attention mechanisms, limiting both efficiency and accuracy in image recognition. To address this, we propose a spatial attention module grounded in information bottleneck theory. Our method explicitly optimizes mutual information: it minimizes the mutual information between the attention representation and the input to suppress redundancy, while maximizing the mutual information between the attention representation and task labels to enhance discriminability. Crucially, we introduce learnable anchors to quantize continuous attention scores—a novel design that strengthens information constraints and improves interpretability of attention maps. By integrating variational attention modeling with deep network embedding, our approach achieves significant performance gains across image classification, fine-grained recognition, and cross-domain classification tasks. The resulting attention maps exhibit high discriminability, strong background suppression, and enhanced interpretability.

Deep Neural NetworksImage RecognitionSelective Attention

Latest Papers

What's happening recently
View more

This work investigates how vision-language models (VLMs) attend to relevant image regions when generating captions, revealing a class of sparse attention heads—termed “gaze heads”—whose attention distributions closely align with the spatial locations of the described content. Through multi-scale analysis of VLM attention mechanisms, the study demonstrates that intervening on fewer than 9% of these gaze heads can precisely steer caption generation toward specified image regions without fine-tuning. By integrating attention relevance analysis with targeted masking, the method achieves 83.1% accuracy in region-directed captioning on both comics and COCO benchmarks and enables dynamic switching of focus during generation. This mechanism proves robust across model scales from 2B to 32B parameters and multiple VLM architectures.

attention mechanismgaze headsimage description

This study addresses the lack of mechanistic understanding regarding object hallucination in VQ-tokenized vision-language models by attributing hallucinations to specific architectural circuits rather than generic calibration biases. Through activation patching, a three-gate diagnostic framework, and univariate architectural substitution experiments, we identify early attention routing circuits across architectures as the locus of hallucination, confirming vector quantization as its source and proposing targeted interventions. Experiments demonstrate that ablating only the L0 layer reduces CHAIRi, an open-ended generation hallucination metric, by 31%, significantly outperforming decoding-time mitigation methods such as VCD and DoLA. This work establishes a novel mechanism-driven perspective for hallucination research.

cross-architecture circuitobject hallucinationvector quantization

This work addresses functional mismatch and redundancy in the attention mechanisms of current large vision-language models, which fail to efficiently exploit visual context. By establishing a unified framework grounded in information theory and information geometry, the study quantifies the geometric structure and entropy characteristics of residual updates, revealing a functional decoupling between attention mechanisms and feed-forward networks (FFNs) in subspace operations. For the first time from an information-geometric perspective, it clarifies their distinct intrinsic roles and demonstrates that attention can be replaced by predefined weights—such as those derived from Gaussian noise—without performance degradation. Empirical results show that this simplified model matches or even surpasses the original architecture across multiple benchmarks, challenging the prevailing design paradigm reliant on dynamic attention and confirming its substantial redundancy.

Attention MechanismLarge Vision-Language ModelsModel Redundancy

This study addresses the limitation of existing video encoding models that overlook the spatiotemporal structure of neural responses, hindering the prediction of dynamic visual processing mechanisms in the brain under naturalistic video conditions. We propose the first video-oriented joint spatiotemporal cross-attention framework, which leverages features from the self-supervised model V-JEPA-2 for spatiotemporally joint routing to construct an interpretable fMRI encoding model simulating the cortex’s adaptive weighting of dynamic information. This model significantly improves brain response prediction accuracy in higher-order visual areas and generates interpretable attention maps capable of tracking moving objects. Furthermore, it reveals dynamically shifting attention allocation patterns within motion perception and semantic selection networks as objects move, thereby elucidating the computational mechanisms underlying dynamic visual perception.

cross-attentionfMRI encoding modelshigher visual cortex

Hot Scholars

XH

Xiaobin Hu

Tencent Youtu Lab;Technische Universität München (TUM)
Deep learningComputer visionVLMAgents
SL

Songhua Liu

Shanghai Jiao Tong University
Computer VisionMachine Learning
PW

Peter Wonka

King Abdullah University of Science and Technology (KAUST)
Deep LearningComputer VisionComputer GraphicsMachine Learning
ST

Sergey Tulyakov

Director of Research, Snap Inc.
computer visionmachine learning
JT

Jin Tang

Anhui University
Computer visionintelligent video analysis