OmniTaskonomy: When Does Visual Generation Improve Visual Understanding?

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the conditions and mechanisms under which visual generative supervision enhances visual understanding capabilities. Methodologically, it establishes a unified task taxonomy and integrates multimodal joint training, gradient analysis, and large-scale task mapping to systematically evaluate the transfer effects of image generation on downstream understanding tasks, revealing non-intuitive beneficial inter-task connections. Building upon these findings, a gradient alignment-based training strategy is proposed to optimize cross-task knowledge transfer. The results demonstrate that, under appropriate configurations, generative data can significantly enhance visual understanding performance. Overall, this work provides a systematic roadmap for the synergistic optimization of visual generation and understanding.
📝 Abstract
Training a model to generate visual content can encourage it to learn rich perceptual capabilities related to geometry, spatial relationships, and objectness; yet, its benefits for visual understanding remain unclear. We ask: when and how does visual generation supervision improve visual understanding? We study controlled pairs of image-to-image (I2I) generation and image-to-text (I2T) understanding tasks that express the same underlying problem in different output modalities. We find that under the correct recipe, I2I training improves downstream I2T performance, with larger gains as the amount of I2I training data increases. We next ask which generation tasks benefit which understanding capabilities. To study transfer beyond paired tasks, we introduce OmniTaskonomy, a unified taxonomy spanning 19 I2I generation tasks and 25 I2T understanding capabilities. The resulting transfer map reveals selective, task-dependent benefits. Some follow intuitive correspondences, e.g., depth prediction improving metric 3D reasoning, object pointing improving counting, and jigsaw reconstruction improving 2D ordering. Interestingly, we also uncover surprising connections: 2.5D segmentation improving category recognition and Z-depth prediction improving localization. To probe these patterns, we analyze gradient alignment between generation and understanding tasks and find that stronger alignment is associated with larger downstream transfer gains. Together, our results highlight visual generation as a rich source of supervision for visual understanding and provide a roadmap for unlocking its benefits through the right training curriculum and task selection. Project page: https://omni-taskonomy.github.io/.
Problem

Research questions and friction points this paper is trying to address.

visual generation
visual understanding
task transfer
OmniTaskonomy
image-to-image
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Generation
Visual Understanding
OmniTaskonomy
Transfer Map
Gradient Alignment
🔎 Similar Papers
No similar papers found.