TD-DPO: Difference-Aware Preference Optimization for Mitigating Sycophancy in Clinical Autism Intervention Dialogue

๐Ÿ“… 2026-07-17
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the safety risks posed by sycophantic behaviors in large language models (LLMs) during conversational interventions with autistic children, a problem exacerbated by conventional sequence-level preference optimization methods that often degrade core intervention capabilities. To mitigate this, the authors propose Minimal-Edit Data Augmentation (MEDA) to construct controllable preference pairs and introduce Token-level Difference-aware Direct Preference Optimization (TD-DPO), which updates only the divergent tokens in model responses. This targeted approach effectively suppresses background drift and sycophancy while preserving essential therapeutic functions. Experimental results demonstrate that TD-DPO significantly outperforms baseline methods in offline evaluations, achieving a better balance between reducing unsafe behaviors and maintaining intervention efficacy, thereby highlighting its strong potential for clinical deployment.
๐Ÿ“ Abstract
The sycophancy of large language models can increase the safety risk in intervention dialogue for autistic children. Supervised fine-tuning can somewhat reduce sycophancy, but relying solely on positive examples is often insufficient to identify and correct failure patterns. We observe that sycophancy behaviors can often be localized to a limited span within the model response. In this regime, sequence-level preference optimization can over-update preference-irrelevant tokens and degrade intervention ability. To address this, we propose the \textbf{M}inimal \textbf{E}dit \textbf{D}ata \textbf{A}ugmentation (MEDA) strategy to construct controlled, stable, minimal edit preference pairs and \textbf{T}oken-level \textbf{D}ifference \textbf{D}irect \textbf{P}reference \textbf{O}ptimization (TD-DPO), which upweights difference tokens between chosen and rejected responses while downweighting shared tokens to suppress background drift. Extensive experiments across multiple backbones and evaluators show that TD-DPO achieves a better trade-off between sycophancy mitigation and intervention ability retention in our offline settings, highlighting its potential as a practical alignment approach for autism intervention.
Problem

Research questions and friction points this paper is trying to address.

sycophancy
clinical autism intervention
dialogue safety
preference optimization
intervention ability
Innovation

Methods, ideas, or system contributions that make the work stand out.

TD-DPO
sycophancy mitigation
token-level preference optimization
minimal edit data augmentation
autism intervention dialogue
๐Ÿ”Ž Similar Papers
No similar papers found.
S
Shuzhong Lai
Nanhu Brain-Computer Interface Institute
J
Junhong Lai
Nanhu Brain-Computer Interface Institute, MOE Frontiers Science Center for Brain and Brain-Machine Integration, Zhejiang University, College of Computer Science and Technology, Zhejiang University
C
Chenxi Li
Childrenโ€™s Hospital Zhejiang University School of Medicine
Qing Zhou
Qing Zhou
Professor of Statistics, UCLA
Graphical ModelsCausal InferenceMonte Carlo MethodsBioinformatics
Haifeng Li
Haifeng Li
Central South University
GISRemote sensingMachine learningSparse represetationBrain Theory
Gang Pan
Gang Pan
Tianjin University
Computer visionMultimodalAI
L
Lin Yao
Nanhu Brain-Computer Interface Institute, MOE Frontiers Science Center for Brain and Brain-Machine Integration, Zhejiang University, College of Computer Science and Technology, Zhejiang University, State Key Laboratory of Brain-Machine Intelligence, Department of Neurobiology, Affiliated Mental Health Center and Hangzhou Seventh Peopleโ€™s Hospital, Zhejiang University School of Medicine
Yueming Wang
Yueming Wang
Zhejiang University
Brain-computer InterfacePattern recognitionmachine learningneural signal processing