Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On

📅 2026-07-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of existing virtual try-on methods, which are often confined to single garment categories or fixed scenarios and struggle to support flexible combinations across multiple clothing types, reference images, and real-world settings. We propose the first fashion-native unified foundation model that reframes virtual try-on as a multi-reference, understanding-driven generative task, departing from conventional mask-based inpainting paradigms. Trained through a three-stage pipeline—continuous pretraining, supervised fine-tuning, and reinforcement learning—the model leverages a custom try-on reward model and a high-quality data engine to generate photorealistic try-on results from a single person image and an arbitrary number of garment references (product or in-the-wild photos). It supports full- or upper-body rendering, cross-category multi-item composition, and pose editing. On public benchmarks and our newly introduced Oxygen-TryOn Bench, it achieves state-of-the-art performance for single-item try-on and significantly outperforms both open- and closed-source systems in multi-item scenarios, excelling in identity preservation and texture fidelity.
📝 Abstract
We present Oxygen-TryOn, a unified foundation model for any-item virtual try-on. Rather than repurposing a general-purpose image editor, Oxygen-TryOn is fashion-native, built for try-on through a dedicated data engine and try-on-specific training. Given one or more reference items (clean product shots or in-the-wild worn-on photos) and a single target subject image, it synthesizes a photorealistic image of the subject wearing the items across virtually any fashion category. Prior systems handle a single garment category in a studio setting, and recent multi-reference methods remain garment-centric; in contrast, Oxygen-TryOn supports diverse items and scenarios, including full- and half-body views, a variable number of references, and free multi-item composition, while faithfully preserving both subject identity and item appearance. Instead of mask-based inpainting, we reformulate try-on as a multi-reference, understanding-driven generation task. We build a data engine that collects, manufactures, annotates, and filters high-quality try-on data at scale, and design a three-stage recipe of continued pre-training (CPT), supervised fine-tuning (SFT), and reinforcement learning (RL). The RL stage uses a hybrid reward combining an in-house try-on reward model with a proprietary, rubric-guided general-purpose model, jointly supervising fine-grained consistency and instruction-level quality. It also follows general editing instructions (e.g., pose changes) in the same pass. Across public benchmarks and our in-house Oxygen-TryOn Bench, it achieves state-of-the-art consistency and realism on single-item try-on and leads on multi-item try-on, matching or surpassing both leading proprietary systems (Nano Banana Pro, GPT-Image-2, Seedream5 Lite) and open-source models (FLUX.2).
Problem

Research questions and friction points this paper is trying to address.

virtual try-on
fashion-native
multi-item composition
photorealistic synthesis
any-item try-on
Innovation

Methods, ideas, or system contributions that make the work stand out.

virtual try-on
foundation model
multi-reference generation
fashion-native
reinforcement learning