Score
Design and implement adapter modules that reconcile spatial resolution differences between pixel-grid inputs and patch-based model inputs for images and video by mapping pixels to patch-aligned embeddings; this includes learnable resampling/interpolation layers, convolutional or attention-based pixel→patch mapping, and adjustments to positional embeddings so tokens correctly correspond to image regions. For video, build variants that also preserve temporal consistency across frames (handling frame-rate or resolution changes) via motion-aware resampling, temporal alignment layers, or cross-frame attention to maintain spatio-temporal correspondence.
To address global structural distortion, local texture degradation, and poor text-prompt alignment in ultra-high-resolution (4K+) image inpainting, this paper proposes Patch-Adapter—a dual-stage adapter framework that operates without modifying pre-trained diffusion models. The method decouples global semantic coherence from local detail fidelity: Stage I employs dual-context adapters to learn coherence from downsampled features; Stage II introduces reference-image patch attention to enable adaptive, full-resolution patch-level feature fusion. Evaluated on OpenImages and Photo-Concept-Bucket, Patch-Adapter achieves state-of-the-art performance, significantly suppressing large-area inpainting artifacts while improving perceptual quality and text–image alignment accuracy.
Traditional uniform patching in neural video implicit representations causes discontinuities at patch boundaries and global structural distortions. To address this, we propose the Structure-Preserving Patch (SPP) decoding framework. Its core innovation is a PixelUnshuffle-inspired spatial rearrangement that reorganizes video frames into structured patch sequences, enabling patch-wise implicit modeling under global consistency constraints. This mechanism facilitates end-to-end differentiable video reconstruction. Evaluated on standard benchmarks, SPP achieves PSNR gains of 1.2–2.8 dB over prior methods, outperforms existing implicit neural representation (INR) approaches in compression efficiency, and significantly enhances boundary continuity and motion coherence. To our knowledge, SPP is the first method to realize structure-aware, patch-level neural video representation—effectively bridging local patch modeling with global structural integrity.
Temporal inconsistency across frames remains a critical challenge in diffusion-based video editing. Method: This paper introduces the first theoretical adapter framework tailored for the DDIM sampler, featuring a lightweight adapter module that jointly learns shared and frame-specific prompts, coupled with a differentiable temporal consistency loss. We formally prove the Lipschitz continuity of this loss’s gradient and derive stability bounds for DDIM inversion alongside monotonic convergence guarantees for gradient descent. Contribution/Results: Our analysis bridges a fundamental gap—prior adapter methods for generative video lack rigorous temporal consistency guarantees. Experiments demonstrate that the proposed framework significantly improves inter-frame coherence under low-overhead editing, maintains controllable inversion error, and achieves both computational efficiency and reliability.
RAW image visual understanding is constrained by the sRGB pretraining paradigm, limiting performance on native RAW data. Method: This paper proposes an end-to-end RAW adaptation framework featuring a learnable ISP input adapter and model-level adapters operating in synergy to decouple physics-based ISP modeling from semantic understanding. It introduces RAW-native data augmentation and cross-domain robust training, built upon a lightweight adapter architecture, differentiable ISP modeling, and multi-granularity feature alignment. Contribution/Results: We establish RAW-Bench—the first benchmark covering 17 realistic RAW degradations—and demonstrate state-of-the-art performance with 62% fewer parameters and 2.3× faster inference. The method exhibits strong generalization and robustness under challenging conditions including low-light, rain, fog, and motion blur.
Existing video editing methods struggle to simultaneously preserve global temporal coherence and local spatial fidelity in long videos, leading to spatiotemporal inconsistencies. This paper proposes a unified spatiotemporal adaptive editing framework that enables semantically and geometrically consistent editing of minute-long videos for the first time. Our approach comprises three core innovations: (1) a test-time editing adaptation mechanism that dynamically fine-tunes image editing models to video-specific dynamics; (2) recursive spatiotemporal attention propagation, which extracts and propagates attention maps across frames via keyframes to ensure temporal consistency; and (3) mask-guided latent variable refinement, enhancing local controllability and precision. Experiments demonstrate substantial improvements in editing fidelity, spatiotemporal coherence, and inference speed—achieving end-to-end second-level processing while maintaining full-video consistency.
This work addresses the challenge of parameter-efficient fine-tuning (PEFT) for video tasks under low-resource conditions, where the optimal allocation of temporal contextual information remains unclear. The study systematically evaluates PEFT methods applied to both image- and video-pretrained models across appearance-, motion-, and spatially dense tasks. Leveraging probing techniques, it analyzes how temporal modeling is distributed across the backbone network, adapter modules, and probes. The findings reveal—for the first time—the varying necessity of temporal context across different video tasks and its substantial impact on adaptation performance. Based on these insights, the paper establishes practical guidelines for designing efficient adaptation strategies tailored to low-data video scenarios.
本文提出了一种分辨率灵活的解码框架,用于解决混合神经视频表示在高分辨率视频中因上采样因子大而不均匀导致的问题。
This work addresses the challenge of transferring image pre-trained models to video tasks, where maintaining temporal consistency across frames and semantic discriminability across videos is difficult to achieve simultaneously. To this end, the authors propose the Co-Settle framework, which introduces a lightweight projection layer atop a frozen image encoder and jointly optimizes the representation space through a temporal cycle-consistency loss and a semantic separability constraint. Co-Settle is the first method to explicitly model and theoretically analyze the trade-off between these two objectives. Remarkably, it achieves significant performance gains across multiple video understanding benchmarks after only five rounds of self-supervised training. Consistent improvements are observed across eight widely used image pre-trained models, demonstrating the framework’s efficiency and broad applicability.
This work explores how to achieve video frame interpolation using only an image foundation model with spatial editing capabilities, without introducing explicit temporal modeling or motion estimation modules. By applying parameter-efficient fine-tuning via Low-Rank Adaptation (LoRA) to the pre-trained Qwen-Image-Edit model, the method activates its latent temporal reasoning ability with merely 64–256 training samples. This study is the first to demonstrate that static image editing models inherently possess transferable temporal understanding, enabling cross-modal generalization from spatial editing to video interpolation. Notably, this approach achieves data-efficient video synthesis without any architectural modifications, offering a novel paradigm particularly suitable for resource-constrained scenarios.
研究通过分析V-JEPA 2和VideoMAE-v2模型,探讨了视频基础模型中时空表示的编码内容、出现位置及几何组织方式,并使用轻量级探针来发现三种时间属性。