PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenges of contextual redundancy caused by disentangled visual representations and pixel-space modeling difficulties for videos in unified multimodal models. To this end, we propose an encoder-free, pixel-level spatiotemporal unified architecture. Specifically, this method maps images and videos into patches and spacetime tubes, respectively, which are processed by a shared backbone network, achieving pixel-space alignment via a single-layer linear projection. Furthermore, a hybrid Transformer architecture is designed to integrate autoregressive text prediction with flow matching-based video generation, thereby unifying understanding and generation capabilities within a single framework. Experimental results demonstrate that the proposed model achieves highly competitive performance across both image and video understanding and generation tasks.
πŸ“ Abstract
Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language pretraining pipelines. Recent advances in pixel-space modeling offer an encoder-free alternative, but extending this paradigm from images to videos is non-trivial: video understanding and generation adopt different temporal representations, leaving the design of a unified visual interface an open question. We present PixelUMM, an encoder-free model for unified image and video understanding and generation directly in pixel space. PixelUMM represents images as spatial patches and videos as spatiotemporal tubelets, connecting raw pixels to a shared multimodal backbone through single-layer linear projections. Its Mixture-of-Transformers architecture combines shared attention with task-specific parameters and extends clean-pixel prediction to video generation, jointly supporting autoregressive text prediction and pixel-space flow matching. Experiments show that PixelUMM achieves competitive performance across image and video understanding and generation tasks. We further conduct empirical studies of key design choices, including decoder design and spatial-temporal patch size, providing insights for future pixel-space unified multimodal models.
Problem

Research questions and friction points this paper is trying to address.

Unified Multimodal Models
encoder-free
pixel-space modeling
video understanding and generation
unified visual interface
Innovation

Methods, ideas, or system contributions that make the work stand out.

Encoder-Free
Unified Multimodal Model
Pixel Space
Mixture-of-Transformers
Flow Matching
πŸ”Ž Similar Papers
2024-02-20International Conference on Machine LearningCitations: 30