Score
Designs and implements conditional image synthesis systems that translate structural inputs (edges, layouts, segmentation maps or skeletons) into photorealistic or high-resolution images. This work builds and analyses pix2pix-style generator/discriminator architectures and training pipelines that preserve scene layout and edges and produce coherent high-resolution textures using adversarial, perceptual, and reconstruction losses.
This work systematically evaluates the applicability of generative AI to scientific image understanding, focusing on text-to-image and image-to-image generation tasks. We propose the first horizontal evaluation framework tailored to scientific imaging scenarios, benchmarking three dominant generative architectures—VAEs, GANs, and diffusion models—across six quantitative dimensions: fidelity, controllability, physical consistency, noise robustness, fine-grained detail accuracy, and domain adaptation efficiency. To address domain-specific requirements, we introduce novel evaluation metrics for generative quality in scientific imaging. Our analysis reveals fundamental trade-offs among key performance indicators across architectures. Furthermore, we identify concrete technical pathways toward enhancing model interpretability. Collectively, these findings provide both theoretical foundations and practical guidelines for the reliable deployment of generative AI in computational imaging, microscopy analysis, and other scientific domains.
This paper addresses key deployment bottlenecks hindering practical adoption of generative AI for image synthesis—namely, high computational overhead, data bias, and poor alignment with user intent. To tackle these challenges, we propose a structured, input-modality–centric taxonomy that unifies modeling across GANs, diffusion models, and conditional generation paradigms. We systematically categorize core tasks—including image-to-image translation, text-to-image generation, domain adaptation, and multimodal alignment—and conduct an in-depth analysis of architectural design principles and applicability boundaries of representative models such as DALL·E, ControlNet, and DeepSeek Janus-Pro. Furthermore, we establish an industrial-deployment–oriented evaluation framework, explicitly delineating optimization pathways for computational efficiency, bias mitigation, and intent alignment. The resulting methodology provides researchers and practitioners with a theoretically grounded yet practically actionable guide for developing and deploying robust, equitable, and controllable generative image systems.
Generating high-quality urban street layouts that jointly model natural (e.g., terrain, hydrology) and socioeconomic (e.g., population, POIs, road density) factors remains challenging. Method: This paper proposes a conditional adversarial learning framework that jointly encodes multi-source natural and socioeconomic features into a conditional GAN architecture. A lightweight graph extraction module enables end-to-end mapping from synthesized images to topologically consistent street graphs. The method integrates autoencoder-based feature fusion with image-to-graph post-processing. Contribution/Results: The generated layouts achieve strong fidelity in both visual appearance and graph-theoretic metrics—including connectivity, degree distribution, and betweenness—closely matching real-world street networks. Experiments demonstrate semantic controllability and high-fidelity virtual urban scene generation: FID improves by 23.6% and graph structural similarity increases by 19.4% across multiple city datasets.
Conditional image synthesis with diffusion models suffers from a lack of systematic understanding due to architectural complexity, task heterogeneity, and diverse conditional mechanisms. Method: This paper introduces the first unified taxonomy for conditional diffusion modeling, categorizing approaches by *where* conditioning is injected—either into the denoising network architecture or the sampling process—and formalizes three paradigmatic stages: training, reuse, and specialization. It further classifies six mainstream sampling-time conditioning strategies. Contributions: Based on a structured analysis of over 100 works, the paper establishes a comprehensive knowledge framework and open-sources an authoritative resource repository (GitHub Awesome-Conditional-Diffusion-Models). It identifies persistent bottlenecks—including limited generalization, inefficient inference, and coarse-grained control—and proposes principled directions toward scalable, modular, and fine-grained conditional modeling.
Existing 3D generative models often struggle to achieve pixel-level fidelity when synthesizing 3D assets from images due to ambiguities in 2D–3D correspondences. This work proposes a pixel-aligned 3D generation paradigm that explicitly lifts multi-scale 2D image features into a 3D feature volume consistent with the input viewpoint via a pixel reprojection mechanism, integrated within an end-to-end trainable, native 3D generative model. The approach achieves, for the first time, large-scale, natively 3D pixel-aligned synthesis, significantly improving geometric and appearance fidelity under both single-image and multi-view settings—approaching the quality of traditional reconstruction methods—and successfully extends to high-fidelity, object-disentangled scene-level synthesis tasks.
研究通过提出逻辑驱动框架ImgCoder和评估基准SciGenBench,解决科学图像合成中的视觉-逻辑不一致问题,提高下游推理能力。
This work addresses the challenge of generating thermal imagery from RGB inputs in aerial scenes, where paired RGB–thermal datasets are scarce. To this end, the authors propose a conditional U-Net architecture that incorporates weather condition metadata embedded into the bottleneck layer. The method further enhances image fidelity by integrating saturation and contrast adjustments as preprocessing steps and applying Gaussian blur as postprocessing within a Pix2Pix GAN framework. Systematic experiments demonstrate the critical role of auxiliary environmental information and tailored image processing in improving generation quality. Evaluated via five-fold cross-validation on a dataset of 612 image pairs, the proposed model significantly outperforms the ThermalGen baseline, achieving a PSNR of 14.55, an SSIM of 0.8095, and an LPIPS score as low as 0.1666.
Existing layered image synthesis methods face limitations in foreground-background separation, data availability, synthesis quality, and scene diversity. This work proposes the BFS framework, which, for the first time, transfers knowledge from non-layered image synthesis to layered generation. BFS employs a dual-branch diffusion model that jointly synthesizes a foreground layer—complete with visual effects such as shadows and reflections—and a composite image, ensuring photorealism and coherence. To address data scarcity, the method introduces a two-stage training strategy that requires only high-quality non-layered images. Experimental results and user studies demonstrate that BFS significantly outperforms current approaches in terms of synthesis quality, visual consistency, and scene diversity.
This work addresses the challenge of multi-granularity modeling in image generation. We propose the Next Visual Granularity (NVG) generation framework, which decomposes an image into a structured sequence of tokens at identical spatial resolution but varying visual granularities, enabling hierarchical modeling—from global layout to local details—via coarse-to-fine progressive generation. Our method employs sequence-based modeling and trains cascaded class-conditional NVG models, demonstrating scalability on ImageNet. The core innovation is a controllable granularity transition mechanism that explicitly captures inter-granularity dependencies. Experiments show substantial improvements over the VAR series: on ImageNet, our method achieves FID scores of 3.03, 2.44, and 2.06 at three progressively finer scales, respectively. It yields higher-fidelity generations and exhibits greater stability under resolution scaling.
This work addresses the end-to-end reconstruction of complex 3D assets from a single RGB image. We propose a neural procedural graph generation framework that represents 3D structure via differentiable procedural graphs, introduces an edge-based tokenization strategy, leverages Transformers to model structural sequence priors, and—crucially—incorporates Monte Carlo Tree Search (MCTS) for guided sampling, significantly improving image-to-3D alignment accuracy. Unlike prior approaches, our method requires neither category-specific priors nor multi-view supervision, enabling direct generation of decodable and editable 3D assets from monocular images. Evaluated on diverse complex objects—including cacti, trees, and bridges—our approach outperforms existing generative 3D methods and domain-specific modeling techniques in both fidelity and generalizability. Notably, it demonstrates strong generalization to real-world images while preserving fine-grained geometric and topological structure.
To address the scarcity of high-quality pixel-level annotations for segmenting elongated fibrous structures—such as microtubules and actin—in biological images, this paper proposes a fiber-aware conditional generative adversarial framework. Built upon the Pix2Pix architecture, it is the first to adapt this GAN paradigm for controllable synthesis of slender biological structures. We introduce a structural loss function that jointly enforces skeleton consistency and orientation sensitivity, and incorporate microscopy-specific image priors to enable end-to-end training. Extensive evaluation across multiple biological datasets demonstrates that segmentation models trained on our synthetic data achieve a 4.2% improvement in mDice over the unenhanced baseline. Moreover, synthesized fibers attain 92% morphological similarity to ground-truth annotations, as quantified by standard metrics. This work establishes a generalizable, low-annotation-dependency data augmentation paradigm for segmenting elongated biological structures.