Score
Designs and implements decoder-only autoregressive models that represent, condition on, and generate multiple modalities by treating all inputs and outputs as interleaved token sequences and performing next-token prediction across modalities. This work includes training a single unified next-token predictor on interleaved multimodal corpora, removing modality-specific adapters and heads, and enabling any-to-any cross-modal generation and encoding.
This paper addresses the challenges in developing general-purpose multimodal large language models (MLLMs) capable of cross-modal coordination across six generative modalities: text, image, music, video, human motion, and 3D objects. To this end, it proposes a novel unified architecture integrating Transformer-based and diffusion-based paradigms, augmented with self-supervised learning (SSL), mixture-of-experts (MoE), reinforcement learning from human feedback (RLHF), and chain-of-thought (CoT) reasoning. The work introduces the first taxonomy covering all six modalities and identifies shared enabling mechanisms for cross-modal transfer. It further argues that structured reasoning and modular decoupling are critical to improving interpretability and generalization. The resulting comprehensive MLLM technology landscape clarifies common bottlenecks and transferable methodologies, providing both theoretical foundations and practical guidelines for building universal, adaptive, and interpretable multimodal systems.
Existing multimodal models struggle to jointly support understanding, generation, and editing, while suffering from low efficiency in high-resolution processing and excessive autoregressive decoding steps. To address these limitations, this paper proposes OneCAT—a unified decoder-only multimodal model. Its key contributions are: (1) eliminating vision Transformers and visual tokenizers by directly modeling raw pixel sequences; (2) incorporating modality-specific Mixture-of-Experts (MoE) layers and multi-scale visual autoregressive modeling to enable dynamic-resolution input handling; and (3) unifying understanding, generation, and editing into a single end-to-end training framework via one shared autoregressive objective. Experiments demonstrate that OneCAT consistently outperforms leading open-source multimodal models across all three task categories, achieves significantly fewer decoding steps, and sets new state-of-the-art performance.
This paper addresses the challenge of fusing interleaved text and time-series modalities for financial forecasting. Methodologically, it proposes a unified multimodal neural architecture featuring modality-specific expert networks to independently model news semantics and stock price dynamics, augmented by a saliency-guided token-level cross-modal alignment mechanism implemented via cross-attention for fine-grained semantic correspondence; an integrated interpretability module further enables decision attribution. The key contribution lies in the first integration of saliency-driven token-level alignment with a mixture-of-experts design, jointly optimizing temporal modeling accuracy, linguistic understanding depth, and model transparency. The approach achieves state-of-the-art performance across multiple large-scale financial forecasting benchmarks, and investment backtesting confirms statistically significant improvements in risk-adjusted returns.
This work addresses the challenge of unifying training and inference in general-purpose multimodal large language models, which typically rely on multiple expert decoders. The authors propose a single autoregressive Transformer decoder capable of arbitrary-to-arbitrary generation across text, images, and streaming speech—without requiring additional expert modules. To mitigate modality imbalance, they introduce task-aware loss reweighting; a lightweight token-level image-perception alignment loss is employed to enhance visual fidelity, and a finite-state decoding mechanism ensures generation stability. The method achieves high-quality generation across all three modalities, with speech synthesis operating at real-time performance (RTF = 0.88), significantly advancing both the efficiency and effectiveness of unified multimodal generation.
This work addresses the limitation of existing unified multimodal models, which rely on separate visual tokenizers that decouple the representation spaces for understanding and generation. To bridge this gap, the authors propose UniAR, a framework that, for the first time, enables end-to-end autoregressive modeling of both tasks within a shared context using a single discrete visual tokenizer. Key innovations include a multi-level feature-fused visual encoder, lookup-free bit quantization, a parallel bit prediction mechanism that substantially compresses visual sequence length, and a diffusion-based visual decoder. Combined with large-scale pretraining, supervised fine-tuning, and reinforcement learning, UniAR achieves state-of-the-art performance in image generation and editing while maintaining competitive results on multimodal understanding benchmarks.
This work proposes NextFlow, a decoder-only unified autoregressive Transformer that addresses the limitations of existing autoregressive multimodal models in image generation speed and cross-modal alignment. Trained on 6 trillion interleaved discrete text-image tokens, NextFlow achieves native multimodal understanding and generation. Its key innovations include replacing raster scanning with a “next-scale prediction” strategy to dramatically accelerate high-resolution image synthesis, alongside a multi-scale training stabilization approach and a prefix-tuning-based reinforcement learning mechanism. The model generates 1024×1024 images in under five seconds—significantly faster than comparable autoregressive models—while attaining state-of-the-art visual quality among unified architectures and matching the performance of specialized diffusion models.
This work addresses the absence of a unified theoretical framework for autoregressive decoding strategies in speech processing, which has led to ambiguous definitions, inconsistent taxonomies, and difficulties in fair comparison. The paper introduces, for the first time, a general formal framework that precisely specifies inclusion criteria for autoregressive search and systematically categorizes and describes decoding strategies employed in neural speech generation models. By clarifying conceptual boundaries, the framework enhances comparability and evaluation consistency across strategies, streamlines the design of decoding-centric benchmarking protocols, and enables ablation studies focused specifically on search mechanisms. Consequently, it facilitates standardized analysis of inference-stage behavior in speech generation models.
This work proposes a unified, decoder-only architecture for arbitrary-to-arbitrary multimodal modeling that treats all modalities symmetrically, eschewing modality-specific components such as dedicated heads, loss functions, or task pipelines. By employing a single tokenization strategy and end-to-end training, the model fully leverages powerful pretrained decoders without requiring modality-tailored structures or training procedures. This approach achieves, for the first time, fully symmetric arbitrary-to-arbitrary generation, enabling cross-modal chained generation and self-verification within a single framework. Evaluated across multiple benchmarks, the model matches or surpasses specialized and multitask baselines out of the box, demonstrating exceptional zero-shot and general-purpose performance without task-specific adaptation.
This work proposes a unified autoregressive multimodal large model framework designed to jointly advance image understanding, generation, and editing while aligning with human preferences. Leveraging a discrete semantic visual tokenizer, the approach converts visual content into a unified discrete representation, enabling multi-task modeling under a single next-token prediction paradigm. The framework integrates multi-objective supervised training with reinforcement learning–driven preference optimization, substantially enhancing both generation quality and instruction-following capabilities. Experimental results demonstrate state-of-the-art performance on text-to-image generation and instruction-guided editing tasks—evidenced by improvements in WISE from 0.50 to 0.56 and GEdit-Bench-EN G_O from 5.75 to 6.68—and reveal notable cross-task synergistic gains.
This work addresses the limitations of prevailing multimodal systems, which are predominantly language-centric and treat non-linguistic modalities as secondary, resulting in architectural fragmentation and insufficient modality fusion. To overcome these issues, the authors propose the Discrete Native Autoregressive (DiNA) framework, which unifies text, vision, and audio into a shared discrete space and enables native multimodal modeling through a single autoregressive objective. Key innovations include the first discrete native Vision Transformer (dNaViT) supporting arbitrary input resolutions—surpassing performance bottlenecks in discrete visual understanding—and a unified multimodal tokenizer coupled with a native multimodal Transformer architecture that seamlessly integrates both understanding and generation tasks. Experiments demonstrate DiNA’s strong performance across multiple multimodal benchmarks, achieving unified capabilities in image understanding, generation, and dialogue, with models and tokenizers released to foster community advancement.
This work proposes Wallaroo, a unified multimodal model grounded in standard autoregressive next-token prediction, which for the first time enables simultaneous support for multimodal understanding, image generation, and image editing within a purely autoregressive framework. By decoupling the visual encoding pathway and employing a four-stage training strategy, Wallaroo effectively handles multi-resolution image inputs and outputs while maintaining compatibility with both Chinese and English languages. The model achieves competitive or state-of-the-art performance across multiple benchmarks compared to existing unified approaches, demonstrating the strong potential and viability of the next-token prediction paradigm for unified multimodal modeling.