Score
Designs and implements systems that model and synthesize images using differentiable image-formation and rendering components, building neural rendering pipelines that combine differentiable reprojection, warping, transform implementations, and joint estimation of camera and scene parameters with generative model design, training, sampling, and evaluation. Analyzes image generation and image-to-image translation quality with appropriate image formation models and evaluation metrics, and integrates visible‑domain constraints into end-to-end synthesis and inference workflows.
3D scene generation suffers from limited diversity, low visual fidelity, and poor view consistency—hindering its deployment in immersive media, robotics, autonomous driving, and embodied AI. This paper presents a systematic survey of four dominant paradigms: procedural, neural 3D, image-driven, and video-driven generation. We introduce the first unified taxonomy to clarify technical evolution across these approaches. Three emerging frontiers are identified: physics-aware modeling, interactive generation, and perception-generation co-design. Leveraging NeRF, 3D Gaussian Splatting, diffusion models, GANs, and multimodal representations, we conduct a rigorous cross-paradigm evaluation on standard benchmarks, quantifying trade-offs among fidelity, diversity, and view consistency. To foster reproducibility and community advancement, we publicly release an open-source tracking platform that continuously monitors state-of-the-art progress.
This work presents the first comprehensive survey of mainstream image generation techniques developed over the past decade, including variational autoencoders (VAEs), generative adversarial networks (GANs), normalizing flows, autoregressive models, Transformers, and diffusion models. It systematically traces their evolution in terms of objective functions, architectural designs, and training algorithms, while also extending the discussion to applications such as video generation, deepfake detection, and watermarking. By constructing a clear technological roadmap, the paper synthesizes optimization strategies, failure modes, and inherent limitations across model families, highlighting the progression from static image synthesis to high-quality video generation. Furthermore, it underscores critical ethical and safety considerations regarding model robustness and responsible deployment, offering a systematic reference for future research and practical applications in generative modeling.
Image generation models frequently suffer from “hallucinations”—semantic distortions or structurally implausible outputs—arising from path deviations in flow matching (FM). To address this, we propose an iterative path correction and progressive refinement framework, the first to systematically integrate iterative optimization into the FM paradigm. Our method employs three core components: reweighted path optimization, gradient-guided dynamic correction, and multi-stage distribution alignment, enabling continuous trajectory refinement throughout generation. The framework is plug-and-play, fully compatible with existing FM models without architectural modification. Extensive experiments demonstrate significant hallucination suppression across multiple benchmarks, yielding 15–22% FID improvement over strong baselines. Crucially, our approach preserves high sampling efficiency and training stability. By explicitly modeling and correcting trajectory deviations, it establishes a new paradigm for enhancing generation fidelity and robustness in flow-based diffusion models.
Manual 3D modeling remains labor-intensive and time-consuming, failing to meet the rapidly growing demands of XR and metaverse applications. Method: This paper presents a systematic survey of state-of-the-art methods for static 3D object and scene generation, introducing— for the first time—a multidimensional classification and cross-comparative framework that jointly considers representation evolution (e.g., point clouds, meshes, NeRFs) and generative paradigms (e.g., supervised learning, diffusion models, 2D foundation model priors, procedural modeling). Contribution/Results: We identify key challenges—including geometric-semantic consistency and scalability—and establish a reusable, multi-axis evaluation framework. Our analysis clarifies technological trajectories and performance boundaries, providing both theoretical foundations and practical guidelines for industrial-grade 3D content generation, thereby advancing the paradigm shift in 3D content creation.
Traditional autoregressive models for image generation suffer from difficulties in modeling spatial dependencies, compromising semantic interpretability, resolution scalability, and generation controllability. To address this, we propose the Compositional Autoregressive Transformer (CAR-Transformer), which decomposes an image into a base map and multi-level detail factors. Generation proceeds hierarchically and iteratively via fine-grained incremental prediction, departing from conventional token-wise or scale-wise modeling paradigms. CAR-Transformer introduces the novel “compositional autoregression” paradigm, enabling zero-shot resolution scaling—arbitrary high-resolution images can be synthesized without retraining. While maintaining computational efficiency, the model achieves state-of-the-art performance on high-fidelity image synthesis benchmarks, with significant improvements in generation controllability and cross-resolution generalization.
This work addresses two key challenges in neural rendering: reliance on explicit 3D representations and scarcity of camera-annotated 3D data. To this end, we propose Kaleido—the first pure-decoder Transformer model that formulates 3D view synthesis as a sequence-to-sequence image generation task. Our core innovations are threefold: (1) treating 3D as a special case of video to unify object- and scene-level view synthesis; (2) eliminating explicit geometric or radiance field representations, instead directly modeling 6-DoF viewpoint transformations in pixel-sequence space via masked autoregressive modeling and refinement-flow Transformers; and (3) leveraging large-scale unlabeled video pretraining to drastically reduce dependence on 3D supervision. Experiments demonstrate state-of-the-art performance across multiple benchmarks: Kaleido achieves superior zero-shot generalization over existing generative methods in few-shot settings, and—uniquely among feedforward models—matches the quality of per-scene optimization approaches under many-view conditions.
Current single-image 3D generation methods suffer from insufficient geometric detail, over-smoothed surfaces, and structural discontinuities—particularly in thin-shell geometries—rendering them inadequate for industrial-grade applications. To address these limitations, we propose a multi-dimensional collaborative optimization framework: (1) a geometry-aware implicit 3D representation tailored for high-fidelity surface modeling; (2) a linear Transformer architecture to enhance long-range geometric coherence; and (3) a progressive super-resolution strategy integrated with strengthened multi-view 3D supervision. Our approach significantly improves geometric fidelity and structural integrity, achieving state-of-the-art performance on ShapeNet and Objaverse benchmarks. The generated models exhibit high-precision geometry, topologically consistent thin-wall structures, and immediate usability—enabling seamless integration into professional 3D production pipelines.
Existing vision models typically rely on disjoint modules for image understanding and generation, hindering coherent reasoning and efficient learning within a unified architecture. This work proposes CyCLeGen, a unified vision-language foundation model that jointly models comprehension and generation capabilities through a cyclic image↔layout generation mechanism within a single autoregressive framework. By integrating cycle-consistency learning with reinforcement learning–driven synthetic supervision, the model acquires introspective abilities and achieves data-efficient self-improvement. Experiments demonstrate that CyCLeGen significantly outperforms current methods across multiple benchmarks for both image understanding and generation, thereby validating the effectiveness and potential of a unified architectural approach.
Existing single-image 3D scene reconstruction methods struggle to produce editable, physically consistent textured meshes—suffering from erroneous object decomposition, inaccurate spatial relationships, and missing backgrounds—thus failing industrial requirements in film and game production. This paper introduces the first end-to-end framework for editable 3D asset generation: it enforces geometric plausibility via a novel 4-DoF differentiable ground-plane constraint; models occlusion recovery as a generative image editing task; and achieves, for the first time, background-driven spatially consistent reconstruction, yielding illumination-coherent, simulation-ready, fully textured meshes. The method synergistically integrates state-of-the-art modules—including object detection, monocular depth estimation, NeRF/3D Gaussian Splatting reconstruction, diffusion-based generation, and differentiable geometric optimization—across complementary domains. It establishes new state-of-the-art performance on single-image 3D scene reconstruction, producing structurally sound, texture-accurate, physically plausible, and post-editable 3D scenes compatible with standard pipelines.