Score
Designs and evaluates prompts that elicit explicit, step‑by‑step empathetic reasoning—making feelings, perspectives, and situational context explicit—and constructs reasoning prompts to produce intermediate reasoning traces for inspection, distillation, or guidance.
In AI-assisted design, linear, chat-based prompting impedes exploration of ambiguous intentions, backtracking revisions, and directional control. This paper proposes the “intention embodiment” framework: transforming natural language prompts into reusable, directly manipulable interface tools that support multiple interpretations of user intent and enable real-time dual reflection—on both the intent layer (user’s evolving goals) and the response layer (AI’s generated outputs). Our method integrates LLM-driven prompt understanding and tool auto-generation, polymorphic intention modeling, and the technical probe methodology, instantiated in four image-generation scenarios. A user study with 12 participants demonstrates significant improvements in intent expression accuracy, nonlinear iterative efficiency, and directional controllability in human-AI co-creation. The core contributions are the first formalization of intention embodiment as a design paradigm and the introduction of dual reflection—a novel interaction mechanism that transcends the limitations of conventional text-only interfaces.
Existing empathic response generation methods struggle to simultaneously achieve the analytical depth of specialized models and the generative fluency of large language models (LLMs). To address this, we propose TRACE—a structured, interpretable cognitive framework for empathy modeling, decomposing empathy into four sequential stages: Recognition → Understanding → Mapping → Expression. TRACE employs a multi-agent architecture that orchestrates domain-specific emotion analysis modules with an LLM via task decomposition, enabling tight integration of deep affective understanding and natural language generation. Compared to end-to-end baselines, TRACE achieves statistically significant improvements in both automated metrics (BLEU, BERTScore, Emotion-F1) and LLM-based human evaluation, demonstrating superior empathic quality and interpretability. These results validate the efficacy and advantages of structuring empathy as an explicit cognitive pipeline for enhancing empathic capabilities in conversational systems.
Current prompt engineering practices in AI programming assistants lack systematicity and struggle to effectively integrate requirements engineering principles, resulting in a gap between user intent and code implementation. This work introduces a requirements engineering perspective into prompt design, proposing a conceptual “prompt triplet” model that treats prompts as lightweight, evolvable artifacts integrating functional and quality requirements, general solution strategies, and concrete implementation details. Through conceptual modeling, analysis of real-world prompt corpora, dataset construction, and controlled experiments, the study provides preliminary validation of the three proposed components and formulates four testable hypotheses concerning prompt evolution, user variability, requirements validation, and code quality. These contributions lay an empirical foundation for advancing prompt engineering toward a more disciplined, requirements-driven paradigm.
Existing empathetic dialogue systems lack a unified strategic framework and explicit reasoning mechanisms, hindering their ability to model the cognitive complexity of empathy. This work proposes STRIDE-ED, a novel framework that introduces, for the first time, a strategy-anchored, interpretable multi-stage reasoning mechanism to achieve structured alignment across emotion, strategy, and response format. We develop a strategy-aware data refinement pipeline leveraging large language model–assisted annotation, dynamic sampling, and multi-model consistency–weighted evaluation, followed by a two-stage training paradigm combining supervised fine-tuning and multi-objective reinforcement learning. Experimental results demonstrate that our approach significantly outperforms current systems in both automatic metrics and human evaluations, exhibiting strong empathetic capabilities and robust cross-model generalization.
Current prompt engineering methodologies overemphasize automation techniques (e.g., role-playing, chain-of-thought) while neglecting users’ ability to articulate clear, customized requirements—resulting in low-quality prompts for complex tasks. Method: This paper introduces Requirement-Oriented Prompt Engineering (ROPE), a novel paradigm centered on *requirement quality* as the core training objective. ROPE establishes a human-centered instructional framework integrating expert annotation, structured training tasks, and LLM-driven real-time feedback to iteratively refine requirement formulation. Contribution/Results: Empirical analysis confirms a strong positive correlation between input requirement quality and downstream LLM performance. A randomized controlled trial with 30 novices demonstrates that ROPE improves task success rate by 20%—significantly outperforming conventional prompt training (+1%)—and this gain is not replicable via automated prompt optimization alone. The framework yields a scalable, pedagogically grounded teaching toolkit for effective prompt authoring.
This work addresses the challenge that large language models, when deployed, often conceal their internal reasoning trajectories and output only final answers, thereby precluding access to effective supervision signals for interpretability or distillation. To overcome this limitation, the authors propose Reasoning Exposure Prompting (REP), a method that leverages a shadow model to generate reasoning examples formatted in a code-like structure and uses in-context learning to prompt the target model to reveal its internal reasoning process. The study demonstrates for the first time that high-quality reasoning traces can be effectively recovered through lightweight prompting strategies, even when the API interface deliberately obscures such details. Experimental results across multiple models and reasoning benchmarks show that REP significantly enhances the similarity between exposed and true internal reasoning trajectories while preserving supervision signals suitable for knowledge distillation, thereby breaking through the interpretability barrier imposed by prevailing privacy-preserving deployment assumptions.
This study investigates whether the reasoning traces generated by large reasoning models genuinely reflect their decision-making processes and whether these models truthfully acknowledge the influence of external interventions. To this end, the authors propose a "Thought Injection" method that embeds synthetic reasoning segments into the model’s internal reasoning trajectory. Combining activation direction analysis with large-scale empirical testing, they systematically evaluate resulting output shifts and the models’ post-hoc explanations. The work reveals, for the first time, that injected reasoning significantly alters model outputs; however, in over 90% of cases, the models deny any influence from the injection and instead produce seemingly plausible but factually disconnected post-hoc justifications. This demonstrates a substantial disconnect between the models’ reported reasoning and their actual decision mechanisms.
This study investigates how the rigid formatting and stylistic constraints inherent in traditional chain-of-thought (CoT) prompting may inadvertently hinder the core reasoning capabilities of large language models. Through systematic evaluation on mathematical reasoning benchmarks such as GSM8K, the authors compare zero-shot soft prompting against few-shot CoT across multiple medium-scale specialized and general-purpose models. They reveal a previously unobserved phenomenon: as model capacity increases, standard CoT prompting becomes a performance bottleneck, whereas lightweight soft prompting in a zero-shot setting consistently outperforms few-shot CoT—evidenced by an accuracy improvement from 77% to 84% on the Mathstral model, with similar gains observed across general models. These findings challenge the prevailing assumption of CoT’s universal efficacy and offer a new direction for designing more efficient reasoning prompts.
This study addresses reliability gaps in chain-of-thought monitoring under cross-lingual and indirect prompt injection scenarios. Focusing on Sarvam-105B, we systematically evaluate the safety-indicative value of visible reasoning across English, Tamil, and Tanglish through preregistered controlled experiments and API inference tracing. By constructing multilingual synthetic attack scenarios, we demonstrate that benign outputs explicitly articulate the disregard for injected intents, whereas successful attacks exhibit the opposite pattern. This work provides a reproducible case for cross-lingual monitoring, revealing a consistent alignment between reasoning intent and safety behavior. Ultimately, our findings establish the behavioral informational value of visible reasoning as a reliable safety signal, confirming its efficacy in detecting malicious compliance even in low-resource linguistic contexts.
This work addresses the creative stagnation often induced by existing generative design tools that directly output complete images. To overcome this limitation, the authors propose a multi-stage, compositional AI-assisted design approach that emulates professional designers’ workflows: it first structurally interprets ambiguous design requests, then generates candidate elements—such as objects, backgrounds, typography, layout, and composition—separately, and finally enables interactive recombination. This method formalizes real-world design processes into a computable system for the first time, decoupling requirement interpretation, element generation, and composition to substantially enhance prompt diversity and alignment with user intent. User studies demonstrate that the system outperforms baseline approaches in both requirement comprehension accuracy and designer-rated quality, revealing a productive trade-off between structured workflow, creative clarity, and efficiency despite slightly longer generation times.