LLM Agents at the Roundtable: A Multi-Perspective and Dialectical Reasoning Framework for Essay Scoring

📅 2025-09-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address insufficient multi-perspective understanding and misalignment with human scoring in zero-shot automated essay scoring (AES), this paper proposes Roundtable Essay Scoring (RES), a multi-agent dialectical negotiation framework. RES employs specialized large language model (LLM) reviewers, each dedicated to a distinct scoring dimension, and orchestrates dialogue-based reasoning inspired by human roundtable discussions to enable multi-perspective evaluation, conflict identification, and consensus aggregation. Its key innovations include trait-specific rubric prompting, role-anchored agent design, and a dialogue-driven score integration mechanism. Evaluated on the ASAP dataset using ChatGPT and Claude as base models, RES achieves up to a 34.86% improvement in Quadratic Weighted Kappa (QWK) over baseline prompt-based methods—demonstrating substantially enhanced scoring accuracy and human alignment in zero-shot settings.

Technology Category

Multiagent Systems: Agreement, Argumentation & NegotiationMachine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Language Grounding & Multi-modal NLP

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsResponsible Web: Machine-in-the-loop, human agency and autonomyEconomics, Online Markets and Human Computation: LLM based quality controls for crowd work
📝 Abstract
The emergence of large language models (LLMs) has brought a new paradigm to automated essay scoring (AES), a long-standing and practical application of natural language processing in education. However, achieving human-level multi-perspective understanding and judgment remains a challenge. In this work, we propose Roundtable Essay Scoring (RES), a multi-agent evaluation framework designed to perform precise and human-aligned scoring under a zero-shot setting. RES constructs evaluator agents based on LLMs, each tailored to a specific prompt and topic context. Each agent independently generates a trait-based rubric and conducts a multi-perspective evaluation. Then, by simulating a roundtable-style discussion, RES consolidates individual evaluations through a dialectical reasoning process to produce a final holistic score that more closely aligns with human evaluation. By enabling collaboration and consensus among agents with diverse evaluation perspectives, RES outperforms prior zero-shot AES approaches. Experiments on the ASAP dataset using ChatGPT and Claude show that RES achieves up to a 34.86% improvement in average QWK over straightforward prompting (Vanilla) methods.
Problem

Research questions and friction points this paper is trying to address.

Achieving human-level multi-perspective understanding in essay scoring
Improving zero-shot automated essay scoring alignment with human evaluation
Enhancing consensus among diverse evaluation perspectives through dialectical reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-agent framework for essay scoring
Dialectical reasoning through roundtable discussion
Zero-shot trait-based rubric generation
💼 Related Jobs
No related jobs found.