DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models

📅 2025-02-10
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitations of large language models (LLMs) in long-context reasoning—specifically, their weak capabilities in argument comprehension, multi-turn trade-off analysis, and alignment with human expert judgment. To this end, we introduce DebateBench, a novel benchmark grounded in 32 high signal-to-noise ratio British Parliamentary (BP) debate transcripts—each exceeding one hour in duration, averaging 32K input tokens, and comprising eight structured speeches. DebateBench is the first systematic effort to repurpose authentic competitive debate data as a long-context reasoning evaluation platform. It innovatively incorporates fine-grained speech-level scoring and team-ranking supervision signals, formalizing a new reasoning task requiring cross-paragraph dynamic stance tracking, implicit rule internalization, and multi-hop argument integration. Experiments reveal that state-of-the-art models—including GPT-4o, GPT-o1, and Claude Haiku—exhibit substantial performance gaps relative to human adjudicators at the 100K-token scale, exposing critical bottlenecks in long-horizon value alignment and adaptive reasoning.

Technology Category

Natural Language Processing: Discourse, Pragmatics & Argument MiningKnowledge Representation and Reasoning: ArgumentationMachine Learning: Large Multimodal Models (LMMs)

Application Category

Economics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingSemantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
We introduce DebateBench, a novel dataset consisting of an extensive collection of transcripts and metadata from some of the world's most prestigious competitive debates. The dataset consists of British Parliamentary debates from prestigious debating tournaments on diverse topics, annotated with detailed speech-level scores and house rankings sourced from official adjudication data. We curate 256 speeches across 32 debates with each debate being over 1 hour long with each input being an average of 32,000 tokens. Designed to capture long-context, large-scale reasoning tasks, DebateBench provides a benchmark for evaluating modern large language models (LLMs) on their ability to engage in argumentation, deliberation, and alignment with human experts. To do well on DebateBench, the LLMs must perform in-context learning to understand the rules and evaluation criteria of the debates, then analyze 8 seven minute long speeches and reason about the arguments presented by all speakers to give the final results. Our preliminary evaluation using GPT o1, GPT-4o, and Claude Haiku, shows that LLMs struggle to perform well on DebateBench, highlighting the need to develop more sophisticated techniques for improving their performance.
Problem

Research questions and friction points this paper is trying to address.

Evaluate LLMs on long-context reasoning tasks
Assess argumentation and deliberation skills in LLMs
Improve LLM alignment with human expert judgments
Innovation

Methods, ideas, or system contributions that make the work stand out.

Long-context reasoning benchmark
Speech-level scores annotation
In-context learning evaluation
💼 Related Jobs
No related jobs found.
Utkarsh Tiwari
Utkarsh Tiwari
Birla Institute of Technology and Science
computer visionmachine learning
Aryan Seth
Aryan Seth
Birla Institute of Technology and Science, Pilani
A
Adi Mukherjee
Birla Institute of Technology and Science, Pilani
K
Kaavya Mer
Birla Institute of Technology and Science, Pilani
K
Kavish
Birla Institute of Technology and Science, Pilani
D
Dhruv Kumar
Birla Institute of Technology and Science, Pilani