Free-text Rationale Generation under Readability Level Control

📅 2024-07-01
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses controllable readability generation: guiding large language models to produce accurate, faithful natural language explanations for reasoning tasks tailored to diverse cognitive levels (e.g., sixth graders, college students), balancing comprehensibility and factual consistency. We conduct the first systematic investigation into how readability level affects free-text reasoning generation, proposing a prompt-engineering–based controllable generation framework. Evaluation integrates traditional readability metrics (e.g., Flesch-Kincaid), multidimensional automated assessment (BERTScore, QA-based faithfulness), and human annotation. Results demonstrate strong model controllability over readability; explanations of medium complexity achieve optimal performance across both automatic metrics and human evaluation—exhibiting the lowest hallucination and misinterpretation rates—while high-school–level explanations receive the highest human preference. Our core contribution is the empirical identification of the readability–faithfulness trade-off and the validation of feasible, effective controllable explanation generation.

Technology Category

Natural Language Processing: GenerationCognitive Modeling & Cognitive Systems: Conceptual Inference and ReasoningKnowledge Representation and Reasoning: Common-Sense Reasoning

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: LLM based quality controls for crowd workSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactions
📝 Abstract
Free-text rationales justify model decisions in natural language and thus become likable and accessible among approaches to explanation across many tasks. However, their effectiveness can be hindered by misinterpretation and hallucination. As a perturbation test, we investigate how large language models (LLMs) perform rationale generation under the effects of readability level control, i.e., being prompted for an explanation targeting a specific expertise level, such as sixth grade or college. We find that explanations are adaptable to such instruction, though the requested readability is often misaligned with the measured text complexity according to traditional readability metrics. Furthermore, the generated rationales tend to feature medium level complexity, which correlates with the measured quality using automatic metrics. Finally, our human annotators confirm a generally satisfactory impression on rationales at all readability levels, with high-school-level readability being most commonly perceived and favored.
Problem

Research questions and friction points this paper is trying to address.

Generating readable free-text rationales for model decisions
Controlling rationale readability for different expertise levels
Assessing rationale quality and human preference across readability levels
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLMs generate rationales with controlled readability levels
Rationales adapt to specified expertise levels effectively
High-school-level readability is most favored by humans
🔎 Similar Papers
No similar papers found.