LongSocialBench: Do Long-Context LLMs Understand Online Discussion Threads?

πŸ“… 2026-10-02
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the difficulty long-context large language models face in comprehending reply structures and performing social reasoning within online discussions. To this end, we introduce the first long-text benchmark dedicated to structured social evidence retrieval. Specifically, we construct serialized conversation trees from real-world forums such as Hacker News and design a human-validated multiple-choice question framework with evidence span annotations. An experimental evaluation of eighteen models reveals that the best-performing system achieves only 63% accuracy, falling significantly short of the human baseline at 72.4%. These findings demonstrate that merely extending context windows is insufficient to overcome deficiencies in social reasoning. Ultimately, this work establishes a novel paradigm for evaluating and enhancing deep social understanding capabilities in large language models.
πŸ“ Abstract
Long-context LLMs can now ingest entire online discussion threads, but understanding their social discourse requires more than reading a long document: models must track parent-reply relations, turning points, scoped subtrees, cross-branch contrasts, and participant trajectories. To test this structure-aware social reasoning, we introduce LongSocialBench, a benchmark of 1,462 verified human-authored multiple-choice items drawn from 94 complete Hacker News, Stack Exchange, and Reddit r/ChangeMyView episodes, with a median length of approximately 73K tokens. Each item pairs a complete serialized discussion and reply structure with a four-option question, requiring models to recover structured social evidence. Released items are verified for answerability, option uniqueness, and evidence grounding. Across 18 models and 29 evaluation settings, current long-context workflows remain far below human performance. The best individual result comes from Claude-Opus-4.7, which reaches 63.0% when prompted to eliminate incorrect options before answering, compared with 72.4% for independent human readers. Averaged across all 18 models, the full-context Baseline scores 43.9%. Supplying the gold evidence scope raises this to 55.0%, showing that substantial errors remain even after the relevant thread region is identified. LongSocialBench shows that the missing capability is not context access or prompting alone, but social understanding over structured reply trees.
Problem

Research questions and friction points this paper is trying to address.

long-context LLMs
social reasoning
online discussion threads
structured reply trees
benchmark evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

LongSocialBench
long-context LLMs
structured social reasoning
online discussion threads
reply tree understanding