To Consolidate or not to Consolidate? Evaluating the Impact of Consolidation in Multi-Reference Training using Peer Reviews

๐Ÿ“… 2026-09-19
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
็ ”็ฉถ้€š่ฟ‡ๅผ•ๅ…ฅMERC-36Kๆ•ฐๆฎ้›†๏ผŒ่ฏ„ไผฐไบ†ๅœจๅคšๅ‚่€ƒ่ฎญ็ปƒไธญๆ•ดๅˆๅ‚่€ƒ็š„้‡่ฆๆ€ง๏ผŒๅนถ่ฏๆ˜Žไบ†ไฝฟ็”จๆ•ดๅˆๅŽ็š„ๅ‚่€ƒ่ต„ๆ–™่ฎญ็ปƒๆจกๅž‹่ƒฝๆ˜พ่‘—ๆ้ซ˜ๆ€ง่ƒฝใ€‚
๐Ÿ“ Abstract
Natural language generation (NLG) tasks span the spectrum of conditional entropy, ranging from highly constrained machine translation to open-ended dialogue generation. Structured tasks like automated peer-review generation occupy the intermediate region, where a single input admits multiple valid, overlapping outputs. In this work, we demonstrate that traditional single- and multi-reference training paradigms are suboptimal for these intermediary tasks. We provide empirical evidence that consolidating diverse references into a unified training signal is crucial for developing effective systems. To facilitate this, we introduce MERC-36K, a large-scale corpus of over 36,000 papers paired with original and consolidated peer reviews. Using this dataset, we train specific architectures to isolate the impact of different reference paradigms and benchmark against existing state-of-the-art systems. Through extensive automatic and human evaluation, we demonstrate that models trained on consolidated references significantly outperform those trained on unconsolidated references. Dataset and code will be released upon acceptance.
Problem

Research questions and friction points this paper is trying to address.

Consolidation
Multi-Reference Training
Peer Reviews
Natural Language Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Consolidated References
Natural Language Generation
Peer Review Generation
MERC-36K
๐Ÿ”Ž Similar Papers
No similar papers found.