Reinforcement Learning with Comparative Evidence for Social Intelligence

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance of social intelligence models on manual annotation and the absence of verification mechanisms in social prediction by proposing the RLCE method. RLCE introduces comparative evidence testing and a strategy co-evolution mechanism, constructing evidence tests from unlabeled data to identify optimal explanations. This enables reinforcement learning without ground-truth rewards, effectively overcoming social semantic ambiguity. By integrating unsupervised training, pairwise evidence aggregation, and dynamic test generation, the proposed approach achieves state-of-the-art performance across four benchmarks, surpassing the strongest baseline by up to 18.93 points.
📝 Abstract
Developing socially intelligent AI remains heavily dependent on human-annotated data, limiting the scale and breadth of social understanding models can acquire. Methods that derive training signals from unlabeled data offer a path beyond this dependence, but social predictions lack the verification oracles available in mathematics and coding. Moreover, core social targets such as affect, intent, preference, and pragmatic meaning are often ambiguous. The same behavior can support multiple plausible interpretations, making it difficult to verify which is best supported. To address this challenge, we introduce Reinforcement Learning with Comparative Evidence (RLCE), a reinforcement learning method that learns social understanding from unlabeled training data without constructing rewards from ground-truth annotations. Given distinct answers in a rollout group, RLCE constructs evidence tests that identify observable evidence favoring an answer over another, validates these tests against the input sample, and aggregates test outcomes to determine the best-supported interpretation. Tests are regenerated as the policy produces new answers, enabling them to evolve with the policy. Across four benchmarks spanning affect, pragmatics, communicative intent, and preference, RLCE attains the strongest performance among seven methods that use no ground-truth training labels for rewards, including consensus, policy LLM-judge verification, multimodal co-evolution, and rubric-based rewards. Gains over the strongest baseline reach up to +18.93 points. Analyses further show that RLCE exhibits a larger share of reward variation between correct and incorrect predictions than compared rubric methods, can overturn erroneous policy-derived preferences, and benefits from pairwise test construction, compositional test aggregation, and on-policy test evolution.
Problem

Research questions and friction points this paper is trying to address.

social intelligence
unlabeled data
ambiguity
reinforcement learning
verification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning with Comparative Evidence
Social Intelligence
Unlabeled Data
Evidence Tests
On-policy Evolution