resolution-mismatch adapters for images/video (pixel-to-patch alignment)

Design and implement adapter modules that reconcile spatial resolution differences between pixel-grid inputs and patch-based model inputs for images and video by mapping pixels to patch-aligned embeddings; this includes learnable resampling/interpolation layers, convolutional or attention-based pixel→patch mapping, and adjustments to positional embeddings so tokens correctly correspond to image regions. For video, build variants that also preserve temporal consistency across frames (handling frame-rate or resolution changes) via motion-aware resampling, temporal alignment layers, or cross-frame attention to maintain spatio-temporal correspondence.

resolution-mismatchadaptersforimages

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.19
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Ultra High-Resolution Image Inpainting with Patch-Based Content Consistency Adapter

Oct 15, 2025
JZ
Jianhui Zhang
🏛️ University of Electronic Science and Technology of China | Megvii Technology | Dzine AI | SeeKoo

To address global structural distortion, local texture degradation, and poor text-prompt alignment in ultra-high-resolution (4K+) image inpainting, this paper proposes Patch-Adapter—a dual-stage adapter framework that operates without modifying pre-trained diffusion models. The method decouples global semantic coherence from local detail fidelity: Stage I employs dual-context adapters to learn coherence from downsampled features; Stage II introduces reference-image patch attention to enable adaptive, full-resolution patch-level feature fusion. Evaluated on OpenImages and Photo-Concept-Bucket, Patch-Adapter achieves state-of-the-art performance, significantly suppressing large-area inpainting artifacts while improving perceptual quality and text–image alignment accuracy.

Achieving 4K+ resolution in text-guided image inpaintingMaintaining content consistency with increasing texture complexityPreserving local detail fidelity while scaling diffusion models

Efficient Neural Video Representation via Structure-Preseving Patch Decoding

Jun 15, 2025
TH
Taiga Hayami
🏛️ Waseda University

Traditional uniform patching in neural video implicit representations causes discontinuities at patch boundaries and global structural distortions. To address this, we propose the Structure-Preserving Patch (SPP) decoding framework. Its core innovation is a PixelUnshuffle-inspired spatial rearrangement that reorganizes video frames into structured patch sequences, enabling patch-wise implicit modeling under global consistency constraints. This mechanism facilitates end-to-end differentiable video reconstruction. Evaluated on standard benchmarks, SPP achieves PSNR gains of 1.2–2.8 dB over prior methods, outperforms existing implicit neural representation (INR) approaches in compression efficiency, and significantly enhances boundary continuity and motion coherence. To our knowledge, SPP is the first method to realize structure-aware, patch-level neural video representation—effectively bridging local patch modeling with global structural integrity.

Address discontinuities in patch-wise video decodingEnhance spatial coherence in neural video representationImprove reconstruction quality and compression performance

Efficient Temporal Consistency in Diffusion-Based Video Editing with Adaptor Modules: A Theoretical Framework

Apr 22, 2025
XS
Xinyuan Song
🏛️ Emory University | University of Minnesota | Peking University | University of Electronic Science and Technology of China | University of Warwick | Washington University | Tsinghua University | Cornell University | Santa Clara University | Stanford University | Hong Kong University of Science and Technology

Temporal inconsistency across frames remains a critical challenge in diffusion-based video editing. Method: This paper introduces the first theoretical adapter framework tailored for the DDIM sampler, featuring a lightweight adapter module that jointly learns shared and frame-specific prompts, coupled with a differentiable temporal consistency loss. We formally prove the Lipschitz continuity of this loss’s gradient and derive stability bounds for DDIM inversion alongside monotonic convergence guarantees for gradient descent. Contribution/Results: Our analysis bridges a fundamental gap—prior adapter methods for generative video lack rigorous temporal consistency guarantees. Experiments demonstrate that the proposed framework significantly improves inter-frame coherence under low-overhead editing, maintains controllable inversion error, and achieves both computational efficiency and reliability.

Differentiable temporal consistency under bounded feature normsMaintaining temporal coherence in diffusion-based video editingStability analysis of adapter modules in DDIM inversion

RAW-Adapter: Adapting Pre-trained Visual Model to Camera RAW Images and A Benchmark

Mar 21, 2025
ZC
Ziteng Cui
🏛️ The University of Tokyo | Nanyang Technological University | RIKEN AIP

RAW image visual understanding is constrained by the sRGB pretraining paradigm, limiting performance on native RAW data. Method: This paper proposes an end-to-end RAW adaptation framework featuring a learnable ISP input adapter and model-level adapters operating in synergy to decouple physics-based ISP modeling from semantic understanding. It introduces RAW-native data augmentation and cross-domain robust training, built upon a lightweight adapter architecture, differentiable ISP modeling, and multi-granularity feature alignment. Contribution/Results: We establish RAW-Bench—the first benchmark covering 17 realistic RAW degradations—and demonstrate state-of-the-art performance with 62% fewer parameters and 2.3× faster inference. The method exhibits strong generalization and robustness under challenging conditions including low-light, rain, fog, and motion blur.

Adapting pre-trained visual models to RAW images.Benchmarking RAW-based corruptions for model evaluation.Integrating ISP modules with downstream vision tasks.

VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing

Jun 18, 2024
JG
Jing Gu
🏛️ University of California, Santa Cruz | Snap Research | KAUST | University of Texas at Dallas

Existing video editing methods struggle to simultaneously preserve global temporal coherence and local spatial fidelity in long videos, leading to spatiotemporal inconsistencies. This paper proposes a unified spatiotemporal adaptive editing framework that enables semantically and geometrically consistent editing of minute-long videos for the first time. Our approach comprises three core innovations: (1) a test-time editing adaptation mechanism that dynamically fine-tunes image editing models to video-specific dynamics; (2) recursive spatiotemporal attention propagation, which extracts and propagates attention maps across frames via keyframes to ensure temporal consistency; and (3) mask-guided latent variable refinement, enhancing local controllability and precision. Experiments demonstrate substantial improvements in editing fidelity, spatiotemporal coherence, and inference speed—achieving end-to-end second-level processing while maintaining full-video consistency.

Achieves precise long video editing in minutesEnsures local consistency in video frame editingMaintains global consistency across video sequences

Latest Papers

What's happening recently
View more

This work addresses the challenge of parameter-efficient fine-tuning (PEFT) for video tasks under low-resource conditions, where the optimal allocation of temporal contextual information remains unclear. The study systematically evaluates PEFT methods applied to both image- and video-pretrained models across appearance-, motion-, and spatially dense tasks. Leveraging probing techniques, it analyzes how temporal modeling is distributed across the backbone network, adapter modules, and probes. The findings reveal—for the first time—the varying necessity of temporal context across different video tasks and its substantial impact on adaptation performance. Based on these insights, the paper establishes practical guidelines for designing efficient adaptation strategies tailored to low-data video scenarios.

low-resourcemodel adaptationparameter-efficient fine-tuning

This work addresses the challenge of transferring image pre-trained models to video tasks, where maintaining temporal consistency across frames and semantic discriminability across videos is difficult to achieve simultaneously. To this end, the authors propose the Co-Settle framework, which introduces a lightweight projection layer atop a frozen image encoder and jointly optimizes the representation space through a temporal cycle-consistency loss and a semantic separability constraint. Co-Settle is the first method to explicitly model and theoretically analyze the trade-off between these two objectives. Remarkably, it achieves significant performance gains across multiple video understanding benchmarks after only five rounds of self-supervised training. Consistent improvements are observed across eight widely used image pre-trained models, demonstrating the framework’s efficiency and broad applicability.

image-to-video transferrepresentation learningself-supervised learning

This work explores how to achieve video frame interpolation using only an image foundation model with spatial editing capabilities, without introducing explicit temporal modeling or motion estimation modules. By applying parameter-efficient fine-tuning via Low-Rank Adaptation (LoRA) to the pre-trained Qwen-Image-Edit model, the method activates its latent temporal reasoning ability with merely 64–256 training samples. This study is the first to demonstrate that static image editing models inherently possess transferable temporal understanding, enabling cross-modal generalization from spatial editing to video interpolation. Notably, this approach achieves data-efficient video synthesis without any architectural modifications, offering a novel paradigm particularly suitable for resource-constrained scenarios.

Few-Shot LearningFoundation ModelsImage Editing