llm-based explanation generation

Designs, builds, and evaluates systems and pipelines that use large language models to generate human-readable explanations or justifications for model outputs or decisions, including prompt templates, few/zero-shot prompting strategies, and mechanisms for personalizing or aligning explanations with user- or item-level features and predicted scores. Analyzes explanation quality, transparency, faithfulness, and reasoning behavior, and implements methods to improve coherence, factual grounding, and interpretability of LLM-produced explanations.

llm-basedexplanationgeneration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.48
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

Manual annotation of textual explanations for interpretable NLP is costly and inherently unscalable. Method: We propose a multi-LLM collaborative framework for automated explanation generation to enhance natural language inference (NLI) classifiers. It integrates outputs from multiple state-of-the-art large language models to produce high-quality, faithful reasoning rationales; employs NLG evaluation metrics to assess explanation quality; and fine-tunes downstream NLI classifiers—specifically on the SNLI and MNLI benchmarks—using these generated explanations as auxiliary supervision. Contribution/Results: Explanations automatically generated by LLMs significantly improve pre-trained NLI model performance, matching the efficacy of human-annotated explanations. This work provides the first empirical validation of the effectiveness and scalability of *automatically generated* explanations for model enhancement. By eliminating reliance on manual annotation, it establishes a novel, scalable paradigm for interpretable NLP.

Assessing impact of automated explanations on model task performanceAutomating textual explanations to replace costly human annotationsEvaluating if LLM-generated explanations improve classification performance

Investigating Co-Constructive Behavior of Large Language Models in Explanation Dialogues

Apr 25, 2025
LF
L. Fichtel
🏛️ Leibniz University Hannover | LMU Munich | Paderborn University | Bielefeld University

This study investigates, for the first time, the capacity of large language models (LLMs) to function as “co-constructive explainers” in explanatory dialogues—specifically, their ability to dynamically adapt explanations to users’ background knowledge and cognitive needs. Method: A user study was conducted, integrating prompt-engineered dialogue interventions, pre-/post-test comprehension assessments, multidimensional user perception questionnaires, and behavioral coding analysis. Results: LLMs spontaneously generate verification questions, enhancing user engagement and comprehension; however, they exhibit significant limitations in modeling explanation pacing, real-time cognitive load, and knowledge gaps—lacking robust metacognitive monitoring and scaffolding capabilities. Contribution: The work establishes a novel evaluation framework for co-constructive explanation, empirically delineates LLMs’ emergent yet bounded capacities in interactive explanation guidance, and provides foundational evidence and design implications for the conversational evolution of explainable AI.

Assessing LLMs' effectiveness in monitoring explainee understandingEvaluating LLMs' ability to co-construct explanations dynamicallyMeasuring impact of LLMs' co-constructive behaviors on comprehension

This work addresses the limited accessibility of existing interpretability tools for language models, which are often too complex for non-expert users to effectively utilize. To bridge this gap, the authors introduce ELIA, an interactive web application that uniquely integrates multiple mechanistic analysis techniques—including Attribution Analysis, Function Vector Analysis, and Circuit Tracing—and innovatively incorporates a vision–language model to automatically generate natural language explanations for complex visualizations. User studies demonstrate that ELIA significantly lowers the barrier to understanding model behavior, with AI-generated explanations effectively mitigating knowledge disparities. Notably, users’ comprehension outcomes show no significant correlation with their prior experience using large language models, underscoring the system’s broad usability and general effectiveness across diverse user backgrounds.

accessibility gapLarge Language Modelsmechanistic interpretability

This work addresses the limitations of existing software documentation, which often suffers from poor consistency, weak relevance, and unclear expression, while purely automated generation methods struggle with insufficient reliability and controllability. To overcome these challenges, the authors propose a human-AI collaborative approach that integrates empirically derived quality guidelines with large language model (LLM) assistance through a generate–evaluate–iterate workflow. This method preserves developers’ domain control while enhancing explanation quality by embedding experience-driven quality criteria directly into the LLM-assisted writing process, enabling controllable, efficient, and high-quality output. Preliminary experiments demonstrate a 24.4% average improvement in authoring efficiency with tool support, and user studies reveal significantly higher satisfaction with the generated explanations compared to purely manual writing (p = 0.003, effect size = 0.86).

explanation qualityLLM-generated textnatural-language explanations

This study investigates whether self-explanations generated by large language models (LLMs) can improve the accuracy of human and LLM predictions regarding the models’ behavior in counterfactual scenarios. To this end, we introduce a novel integration of pragmatic perturbations with a counterfactual simulatability framework to construct test cases, and conduct a systematic evaluation using chain-of-thought and post-hoc explanation generation, joint human–LLM assessments, and qualitative analysis of free-text responses. Our findings demonstrate that self-explanations significantly enhance prediction accuracy, though this effect is moderated by the choice of perturbation strategy and the evaluators’ reasoning capabilities. Further analysis of user-generated rationales corroborates the constructive role of explanations in shaping human judgment.

counterfactual simulatabilitylarge language modelsmodel behavior prediction

Latest Papers

What's happening recently
View more

This work addresses the challenge that large language models often struggle to simultaneously satisfy content relevance and formal constraints, leading to procedural errors. To overcome this, the authors propose a multi-agent workflow that, for the first time, decouples the primary task description from fine-grained constraints and iteratively refines prompts through an evaluation-driven collaborative mechanism. By integrating automated scoring feedback, prompt rewriting, and multi-agent coordination, the approach significantly enhances adherence to formal constraints in model outputs. Experiments on Llama 3.1 8B and Mixtral-8x 7B demonstrate substantial improvements, validating the effectiveness of constraint decoupling and evaluation-guided refinement in boosting instruction-following performance.

complianceformal constraintsinstruction following

This work addresses the lack of standardized documentation and evaluation methodologies in prompt engineering, which hinders the reproducibility and interpretability of complex prompts. To remedy this, the authors propose “Prompt Cards,” a novel framework that adapts the model card concept to prompt engineering by introducing a structured template to explicitly document a prompt’s design objectives, contextual strategies, evaluation protocols, and ethical considerations. Demonstrated through a “wordification” task, the approach integrates natural language generation with qualitative assessment to enable systematic recording and analysis of the entire prompting pipeline. Prompt Cards substantially enhance transparency, reproducibility, and methodological rigor, offering the research community a scalable standard for prompt documentation and a new paradigm for benchmarking beyond conventional metrics.

documentationevaluationprompt engineering

Increasing AI Explainability by LLM Driven Standard Processes

Nov 10, 2025
MJ
Marc Jansen
🏛️ University of Applied Sciences Ruhr West

Large language models (LLMs) suffer from opaque and non-auditable decision-making processes, hindering trust and regulatory compliance. Method: This paper proposes a hierarchical, explainable AI architecture that integrates LLMs with structured decision frameworks—including QOC (Question-Options-Criteria), sensitivity analysis, game-theoretic modeling, and risk management—thereby decoupling reasoning and explanation spaces. Unlike post-hoc interpretability methods, it enables prospective modeling of inference paths through standardized analytical workflows. Contribution/Results: It is the first work to systematically co-model classical decision science paradigms with LLMs, supporting end-to-end traceability and formal verification of decision logic. Experiments demonstrate that the system replicates expert-level reasoning in complex domains—including decentralized governance, systems analysis, and strategic planning—while significantly enhancing transparency, auditability, and trustworthiness of AI-driven decisions.

Enhancing AI explainability through standardized LLM integrationEstablishing verifiable AI decision-making via structured reasoning frameworksTransforming opaque AI inference into transparent decision traces

This study addresses the challenge of supporting teachers in diagnosing student problem behaviors through trustworthy large language model (LLM)-based dialogue systems. While effective intervention requires integrating multidimensional information and formulating evidence-based strategies, existing LLM systems often lack explainability, undermining user trust and transparency. To bridge this gap, this work proposes an explainable dialogue system grounded in hierarchical attribution, which fine-tunes an LLM to support multi-turn diagnostic conversations and automatically generates natural language explanations that highlight the key dialogue evidence underlying each intervention recommendation. Technical evaluations demonstrate the method’s superiority over baseline approaches in identifying supportive evidence, and a user study with 22 pre-service teachers reveals that providing such explanations significantly enhances perceived system credibility. By integrating explainable AI, hierarchical attribution, and natural language generation, this research advances the development of trustworthy LLM applications in educational contexts.

dialogue systemexplainable AIlarge language models

This study investigates the discrepancies between large language models (LLMs) and human cognition in high-level chart understanding, with a focus on interpreting designer intent and extracting complex data patterns. Through qualitative user studies, it systematically compares the higher-order interpretation strategies employed by humans and LLMs on line charts, bar charts, and scatter plots, while analyzing LLM outputs and reasoning pathways under three distinct prompting conditions. The work reveals, for the first time, that LLMs consistently adopt a structured enumeration strategy rather than constructing coherent trend-based narratives, and their explanatory patterns remain remarkably stable across different prompts. In contrast, humans demonstrate a superior ability to synthesize holistic, narrative-driven interpretations. These findings highlight fundamental mechanistic limitations of LLMs in visual reasoning and offer critical insights for future model design.

communicative goalshigh-level patternshuman interpretation

Hot Scholars

WW

Weiping Wang

School of Information Science and Engineering, Central South University
Computer NetworkNetwork Security
AV

Andrea Visentin

Associate Professor, School of Computer Science & IT, University College Cork
NR

Nicola Rossberg

PhD Candidate, University College Cork
Artificial Intelligence
WM

Wentao Ma

Alibaba DAMO Academy
Dialog SystemLarge Language ModelQuestion Answering
SS

Siddharth Suresh

Graduate Student, University of Wisconsin, Madison
Cognitive ScienceHuman-AI alignmentLarge Language ModelsRepresentation Learning