Chinese Competitive Debating Dataset and Benchmark

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决辩论评估中缺乏精细记录与专业评判结合的问题,本文通过组织182场中文辩论赛并收集专业裁判评分,构建了包含多层级辩论数据的评测基准。
📝 Abstract
Debate adjudication requires tracking how arguments develop through interaction, yet existing datasets rarely combine fine-grained debate transcripts with professional judgments collected during real competitions under a shared rubric. We introduce a dataset and benchmark for evaluating large language models' understanding of competitive Chinese-language debate at the match, stage, and speaker levels. We organized 182 matches and recruited 120 professional judges, with each match independently adjudicated by three judges using a predefined rubric. After excluding matches with incomplete records, the dataset contains 148 matches, 2,698 stages, and 20,542 exchange units, with manually verified transcripts and segmentation. It preserves original stage scores, match votes, best-debater ballots, and adjudication rationales. We define three tasks: winner-tendency prediction, stage-score prediction, and best-debater prediction. Zero-shot evaluation of multiple large language models yields a highest winner-prediction accuracy of 66.2%, a highest Pearson correlation of 0.250 between model stage scores and mean human ratings, and a highest best-debater prediction accuracy of 56.8%. The dataset and benchmark provide a testbed for studying large language models' understanding of interactive argumentation and their agreement with professional judges.
Problem

Research questions and friction points this paper is trying to address.

Chinese Competitive Debating
Dataset
Benchmark
Professional Judgments
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Chinese Competitive Debating
Large Language Models
Benchmarking
Argumentation Understanding
Professional Judgments
🔎 Similar Papers
💼 Related Jobs
No related jobs found.