Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing

📅 2026-07-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high computational cost of training and inference in high-resolution image generation and editing, which hinders efficient user interaction. The authors propose a compact 4B-parameter generative model stack comprising a lightweight, high-fidelity Mage-VAE and a native-resolution multimodal diffusion Transformer, jointly designed for efficient text-to-image synthesis and instruction-driven editing. Key innovations include a one-stage VAE with anchor-latent regularization, rectified flow matching, adversarial-aware guidance for few-step distillation, and CUDA kernel fusion, collectively reducing computational overhead. The resulting Base, RL-aligned, and Turbo model variants achieve strong performance on standard benchmarks: Mage-Flow-Turbo generates 1024² images in just 0.59 seconds and completes edits in 1.02 seconds on a single A100 GPU, with low memory consumption, making it suitable for interactive high-resolution applications.
📝 Abstract
Large-scale visual generators are increasingly capable but costly to train, fine-tune, and deploy. We introduce Mage-Flow, a compact 4B-scale generative stack for efficient text-to-image generation and instruction-based image editing. The stack is built from two co-designed components: Mage-VAE, a lightweight high-fidelity latent tokenizer, and a Native-Resolution Multimodal Diffusion Transformer trained with rectified flow matching. Mage-VAE uses one-step diffusion-style encoding and decoding with anchor-latent regularization, preserving the reconstruction quality of strong public VAEs while reducing tokenization cost by more than an order of magnitude. Together with native-resolution packing and stack-level CUDA kernel fusion, the stack supports flexible-resolution training and improves end-to-end training throughput by about $2.5\times$. Built on this foundation, we develop a complete model family with Base, RL-aligned, and Turbo variants for both generation and editing. Diffusion-NFT improves prompt following, text rendering, aesthetic quality, and editing fidelity, while few-step distillation with adversarial perceptual guidance produces 4-step Turbo models for low-latency inference. Despite its compact scale, Mage-Flow and Mage-Flow-Edit achieves competitive performance across standard generation and editing benchmarks. More importantly, the Turbo variants make high-resolution generation and editing practical for interactive use: at $1024^2$ resolution on a single NVIDIA A100 GPU, Mage-Flow-Turbo generates an image in 0.59s, and Mage-Flow-Edit-Turbo edits an image in 1.02s, while maintaining a small memory footprint. These results show that careful tokenizer--backbone--system co-design can deliver strong high-resolution generation and editing within an efficient 4B model family.
Problem

Research questions and friction points this paper is trying to address.

efficient image generation
high-resolution editing
compact generative model
low-latency inference
visual foundation model
Innovation

Methods, ideas, or system contributions that make the work stand out.

Native-Resolution Generation
Lightweight VAE
Rectified Flow Matching
Kernel Fusion
Few-Step Distillation