Score
Designs, implements, and evaluates generative systems that produce images from natural-language descriptions, encompassing text-guided or text-conditioned architectures (e.g., diffusion models, conditional generators), conditioning modules, training pipelines, and augmentation workflows. Measures and analyzes the fidelity of images to the input text, visual realism and diversity, and the usefulness of synthesized images for downstream objectives such as data augmentation or class-balancing.
This work systematically evaluates the applicability of generative AI to scientific image understanding, focusing on text-to-image and image-to-image generation tasks. We propose the first horizontal evaluation framework tailored to scientific imaging scenarios, benchmarking three dominant generative architectures—VAEs, GANs, and diffusion models—across six quantitative dimensions: fidelity, controllability, physical consistency, noise robustness, fine-grained detail accuracy, and domain adaptation efficiency. To address domain-specific requirements, we introduce novel evaluation metrics for generative quality in scientific imaging. Our analysis reveals fundamental trade-offs among key performance indicators across architectures. Furthermore, we identify concrete technical pathways toward enhancing model interpretability. Collectively, these findings provide both theoretical foundations and practical guidelines for the reliable deployment of generative AI in computational imaging, microscopy analysis, and other scientific domains.
This study addresses the current lack of systematic integration of image-generating generative AI in modeling and simulation. It presents the first comprehensive exploration of text-to-image generation techniques within this domain, proposing tool-agnostic, transferable principles and establishing a localized, reproducible generation pipeline that combines prompt engineering with simulation output mapping. The proposed approach supports diverse applications—including conceptual model representation, visualization of simulation results, generation of instructional materials, and construction of multi-scale model interfaces—thereby offering practitioners a structured knowledge framework to evaluate and adapt this emerging technology. By doing so, it significantly enhances the visual expressiveness and interactive capabilities of simulation systems.
This work addresses the challenge of efficiently generating high-quality training data required for supervised learning in text-to-image generation models. We propose the Guided Adversarial Prompts (GAP) framework—a closed-loop data generation system integrating three core mechanisms: (1) adversarial prompt optimization guided by supervised model loss, (2) target distribution alignment via feature matching or discriminator-based guidance, and (3) online feedback adaptation. GAP is the first method to synergistically couple adversarial generation with explicit distributional constraints, shifting data synthesis from open-loop, static prompting to closed-loop, adaptive refinement. Empirical evaluation across diverse settings—including multi-task learning, heterogeneous model architectures, and distribution shifts (e.g., spurious correlations, unseen domains)—demonstrates substantial improvements in downstream model generalization. Data utilization efficiency increases by up to 3.2× compared to baseline approaches.
This work identifies a systematic degradation phenomenon—termed “nepotistic training”—occurring when generative AI models are fine-tuned using images synthesized by themselves. Leveraging diffusion-based text-to-image frameworks guided by CLIP, the study conducts controlled retraining experiments, multi-dimensional quality assessments (e.g., FID), and distributional shift analyses. It empirically demonstrates that the degradation is (i) contagious—injecting merely 0.5% AI-generated images degrades FID by over 200%; (ii) generalizable—distortions persist across unseen prompts; and (iii) irrecoverable—subsequent fine-tuning on clean real data fails to restore performance. The paper formally defines and validates these three core properties of this novel training failure mode. By establishing both theoretical insight and empirical evidence, the findings provide critical implications for sustainable generative model training and responsible content governance.
This work addresses the attribute-object-relation compositional misalignment problem in text-to-image generation, revealing structural deficiencies in the CLIP text encoder: its final-layer token embeddings are susceptible to interference from irrelevant words, and its attention mechanism struggles to capture fine-grained semantic relations. We propose a lightweight solution—fine-tuning only the CLIP linear projection head—without modifying the encoder backbone or the diffusion model architecture. On benchmarks such as Compositional COCO, our method improves compositional accuracy by 23.6% while preserving FID. Through attention reweighting and representational space analysis, we empirically demonstrate that suboptimal geometry in the CLIP embedding space is the primary cause of compositional failure. To our knowledge, this is the first systematic study establishing that projection-head fine-tuning alone suffices to substantially enhance compositional generalization—introducing a novel, efficient paradigm for improving multimodal alignment.
Text-to-image diffusion models achieve high visual fidelity but struggle to faithfully render fine-grained subject details—such as precise text spelling—due to the limited semantic expressivity of standard text tokenizers. To address this, we propose a reference-guided generation framework based on lightweight expert plug-ins: leveraging a reference image as visual conditioning, jointly with textual input, to steer the Stable Diffusion architecture and overcome the representational bottlenecks of language-only conditioning. Our method integrates multiple task-specific, compact expert plug-ins (each with only 28.55M parameters), an auxiliary network, and dedicated loss functions tailored for cross-lingual text rendering and domain-specific applications like logo generation. Extensive experiments demonstrate consistent and significant improvements over existing state-of-the-art methods across English and multilingual text-to-image synthesis, as well as logo generation benchmarks.
Current generative models exhibit insufficient OCR capabilities in text-to-image generation and editing, particularly in producing legible, accurate, and layout-preserving textual content. Method: We propose the first systematic evaluation paradigm for “OCR-aware generation,” comprising 33 tasks across five real-world domains—document, handwritten, scene, artistic, and complex-layout images—and advocate photorealistic text generation as a foundational capability for general-purpose multimodal models. We introduce a customized input-prompt co-design mechanism and a multidimensional evaluation protocol integrating OCR accuracy metrics (e.g., CER, WER) with visual fidelity metrics (e.g., CLIP-Score, FID). Contribution/Results: Extensive benchmarking across six leading open- and closed-source models reveals critical deficiencies in character-level accuracy and structural layout preservation. Our analysis identifies language-vision alignment as the primary bottleneck limiting OCR generation performance, establishing a reproducible benchmark and concrete optimization directions for next-generation multimodal foundation models.
Although text-to-image (T2I) diffusion models from 2022 to 2025 have made remarkable progress in generating visually realistic and prompt-aligned images, their synthetic data consistently underperforms when used to train image classifiers. This study systematically evaluates the efficacy of data generated by successive generations of state-of-the-art T2I models through large-scale synthesis, standard classifier training protocols, and cross-model comparative analysis. The findings reveal that the pursuit of aesthetic quality has come at the cost of reduced data diversity and label consistency, leading to a disconnection between “generative realism” and “data utility.” Experiments demonstrate that classifiers trained on synthetic data from the latest T2I models exhibit significantly degraded accuracy on real-world test sets, indicating that current T2I-generated data is unsuitable as a reliable source for training robust image classifiers.
This work proposes a unified framework for understanding and developing generative artificial intelligence models capable of producing multimodal content, including images, text, video, and molecular structures. Addressing the current fragmentation in generative modeling, the study integrates core methodologies—such as variational autoencoders, generative adversarial networks, diffusion models, and large language models—into a cohesive theoretical system grounded in mathematical principles, architectural design, and mechanisms for controllable generation. This framework not only advances a systematic understanding of multimodal generative processes but also provides robust theoretical foundations and practical pathways for generating high-quality, controllable digital content, with direct implications for applications in scientific discovery and beyond.
Automated design of text-to-image multi-component pipelines faces two major bottlenecks: prohibitively high computational cost and poor generalization across tasks. Method: We propose the first end-to-end reinforcement learning framework that eliminates reliance on costly image generation for evaluation, introducing a novel image-generation-free ensemble reward model. Our approach employs a two-stage optimization strategy—lexical pretraining followed by Generalized Reinforcement Policy Optimization (GRPO)—and incorporates Classifier-Free Guidance (CFG)-guided model interpolation to enhance structural diversity. Contribution/Results: Without rendering any images, our method enables efficient workflow sequence modeling and search. It achieves state-of-the-art performance in image fidelity, structural novelty, and cross-task generalization, significantly outperforming existing baselines while reducing training overhead substantially.
Existing automated evaluation methods for generative content lack a systematic, cross-modal framework. Method: This paper conducts a large-scale literature review and cross-modal comparative analysis to establish, for the first time, a unified evaluation taxonomy covering text, image, and speech modalities. It identifies five fundamental evaluation paradigms and empirically validates their consistent applicability across three representative generative tasks. Furthermore, it introduces a comparability analysis framework to construct a structured knowledge graph that clarifies capability boundaries and limitations of existing methods per modality. Contributions/Results: (1) The first cross-modal unified classification system for generative evaluation; (2) abstraction of generalizable, transferable evaluation paradigms; and (3) a theoretical foundation and practical methodology for cross-modal consistent evaluation and joint metric design. This work bridges critical gaps in evaluating multimodal generative models and enables principled, interoperable assessment across modalities.