Score
Designs, builds, and evaluates systems (models, pipelines, and interfaces) that automatically produce textual or multimodal outputs from prompts, templates, or structured inputs, including components for conditioning, decoding/sampling, post-processing, and content selection. Analyses focus on output quality, coherence, factuality, style, safety, and efficiency using automated metrics and human evaluation, and on methods for constraint-control, personalization, and hallucination mitigation.
This work addresses the lack of standardized documentation and evaluation methodologies in prompt engineering, which hinders the reproducibility and interpretability of complex prompts. To remedy this, the authors propose “Prompt Cards,” a novel framework that adapts the model card concept to prompt engineering by introducing a structured template to explicitly document a prompt’s design objectives, contextual strategies, evaluation protocols, and ethical considerations. Demonstrated through a “wordification” task, the approach integrates natural language generation with qualitative assessment to enable systematic recording and analysis of the entire prompting pipeline. Prompt Cards substantially enhance transparency, reproducibility, and methodological rigor, offering the research community a scalable standard for prompt documentation and a new paradigm for benchmarking beyond conventional metrics.
Existing evaluation methods for audio captioning struggle to accurately assess the fidelity of multimodal semantics and acoustic attributes in structured audio descriptions. This work proposes the first multi-axis evaluation framework tailored for structured audio captioning, integrating large language model (LLM)-based semantic judgments with deterministic acoustic metrics across five orthogonal dimensions: label sets, descriptive content, logical reasoning, numerical measurements, and spectral contours. The framework incorporates a controlled perturbation protocol to validate its ability to distinguish between semantic preservation and acoustic distortion. Experiments on the AudioCards dataset demonstrate that the proposed approach effectively differentiates semantically consistent paraphrases from genuine errors, significantly outperforming existing methods in both reliability and sensitivity.
This study investigates how prompt structure and clarity affect large language model (LLM) output quality and user productivity. Drawing on empirical data from 243 users across education, professional, and creative domains, the research integrates questionnaire surveys, behavioral log analysis, and satisfaction assessments. It provides the first systematic validation that structured, context-sensitive, and semantically explicit prompts significantly improve task completion efficiency (+37.2%) and output adherence rate (+41.5%). Methodologically, the study introduces the first multi-context, user-driven evaluation framework for prompt engineering efficacy. Its key contributions include three generalizable, transferable principles for high-efficiency prompt design. Results demonstrate that prompt engineering transcends mere technical fine-tuning—it functions as a critical leverage point for unlocking LLMs’ practical productivity in real-world applications.
This study investigates how large language model (LLM) hallucination and cognitive forcing jointly influence user dependency behavior and data quality during human-LLM collaborative generation of customer service dialogues. Through a user behavior experiment (N=11, 88 tasks), LLM dialogue generation, qualitative coding, and quantitative analysis, we empirically demonstrate— for the first time—a significant interaction effect: hallucination substantially degrades data quality; while cognitive forcing does not universally mitigate hallucination’s adverse effects, it systematically reshapes user adoption strategies, giving rise to three distinct dependency patterns. Our core contributions are threefold: (1) uncovering the synergistic mechanism between hallucination and cognitive forcing; (2) proposing a novel data quality assessment paradigm anchored in observable user behavior; and (3) providing empirical foundations for designing robust, trustworthy human-AI collaborative data production pipelines resilient to hallucination-induced interference.
This paper addresses the challenge of operationalizing generative AI within collaborative software engineering teams. Drawing on a design study with 39 industry experts—including field observations, semi-structured interviews, and multi-role workshops—we systematically investigate how prompt engineering supports cross-functional AI prototyping and iterative co-design. Our study is the first to characterize three core phenomena in collaborative prompt prototyping: (1) the emergent construction of shared coordination norms, (2) dynamic role evolution across developers, domain experts, and AI specialists, and (3) context-sensitive evaluation mechanisms for prompt efficacy. We propose a generative-content-feature-driven rapid iteration paradigm and distill a reusable prompt prototyping strategy framework. Key technical challenges—including model opacity and example overfitting—are empirically identified. The findings provide both methodological grounding and actionable practice guidelines for industrial software teams, advancing the shift from generative AI as a technical capability to a collaborative design enabler.
This work addresses the need for automated generation of compliant, high-fidelity, and creatively diverse images in brand marketing contexts by proposing the first fully automatic text-to-image generation pipeline that jointly optimizes brand safety, image quality, and human preference. The system integrates state-of-the-art text-to-image models with a DINOv2-based image quality assessment module and incorporates a lightweight human feedback mechanism to dynamically refine outputs. Experimental results demonstrate that, compared to baseline methods, the proposed approach improves image fidelity by 30.77% (as measured by DINOv2) and increases human preference by 52.00%, effectively balancing automation efficiency with brand compliance in large-scale production environments.
This study addresses the frequent inefficiencies in human-AI collaboration caused by incomplete contextual information, which often leads to excessive iteration and suboptimal output quality. To mitigate this, the authors propose a structured context construction framework that integrates a five-role context package—comprising authority, exemplars, constraints, evaluation criteria, and metadata—within a four-stage workflow encompassing review, design, construction, and audit. Notably, this work pioneers the incorporation of information theory and reliability engineering principles into context quality assessment, yielding a reusable and auditable collaboration framework. Empirical results from 200 interaction trials demonstrate that the approach reduces the average number of iterations from 3.8 to 2.0, increases first-pass success rates from 32% to 55%, and achieves a final task success rate of 91.5%.
This work addresses the limitations of current automatic text simplification approaches, which overly rely on automated metrics that fail to capture users’ actual comprehension abilities and normative standards, thereby offering inadequate support for cognitive accessibility. To overcome this, the authors propose a human-in-the-loop hybrid framework that integrates real-time human guidance during large language model generation alongside post-hoc human oversight. For the first time, this framework systematically embeds human roles into both generation and evaluation phases, leveraging a standards-aligned checklist, an Event-Condition-Action (ECA) rule engine, and accessibility-oriented key performance indicators (KPIs) to enable a traceable, reproducible, and auditable text generation pipeline. Empirical results demonstrate that the approach effectively encodes human feedback, enhances model adaptability, and provides a structured, transparent, and inclusive pathway for evaluating and optimizing accessible texts.
Existing reference-based evaluation metrics struggle to capture the subjective and creative aspects of AI-generated stories. To address this limitation, this work proposes the first hierarchical, multidimensional framework for assessing creativity, encompassing four dimensions: novelty, value, adherence, and resonance. By employing Spike Prompting to control generated content and conducting a crowdsourced study with 115 participants, the research investigates how human judgments of creativity evolve across immediate and reflective evaluation phases. The findings reveal that reflective assessment significantly alters scoring outcomes and enhances inter-rater agreement. This framework effectively uncovers creativity dimensions overlooked by conventional metrics, substantially improving the granularity and reliability of creativity evaluation in AI-generated narratives.