🤖 AI Summary
This study addresses the misalignment in empathy-enhanced reinforcement learning, where support priorities evolve dynamically during conversations while existing reward specifications remain static. To overcome this limitation, we propose CARE, a framework that introduces a pioneering context-adaptive scoring rule evolution mechanism. This mechanism dynamically adjusts the weights and criteria across cognitive, affective, and proactive empathy dimensions to construct an adaptive reward interface. Furthermore, the rule generator is trained via supervised fine-tuning combined with preference-based reinforcement learning, while the RLVER and MICA algorithms are integrated to optimize multi-turn empathetic strategies. Experimental results demonstrate that CARE achieves state-of-the-art performance on benchmarks including SentientBench. Notably, it improves the EMPA metric by at least 13 points over the strongest baseline, elevating the maximum score from 28.11 to 83.54.
📝 Abstract
We identify a fundamental mismatch in empathetic reinforcement learning: support priorities evolve with the dialogue state, yet existing methods typically optimize predefined reward specifications that remain fixed across turns. To model these evolving support priorities, we organize empathetic support along cognitive, affective, and proactive empathy, and propose Context-Adaptive Rubric Evolution (CARE). At each turn, CARE generates a context-adaptive rubric by adjusting both the weights of these three empathy dimensions and their fine-grained evaluation criteria. The rubric generator is trained with turn-level rubric supervision and human preference data through supervised fine-tuning followed by preference-based reinforcement learning, and then serves as an adaptive reward interface for online empathetic RL. Integrated with both RLVER and MICA, CARE achieves state-of-the-art performance across SentientBench, EQBench3, and EMPA under three independent LLM judges. Notably, on EMPA, CARE improves EPM-Idx over the strongest baseline by at least 13 points under all three judges, including an increase from 28.11 to 83.54 under Gemini-2.5-Pro. Further analyses show that learned rubric priorities systematically vary across dialogue stages and user emotions, demonstrating that CARE adapts what is rewarded as support needs evolve.