DREAM: Dynamic Resolution Assignment For Multimodal Multi-agent Debate

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of fixed visual resolution and erroneous consensus induced by groupthink in multimodal debate, proposing the DREAM framework. Methodologically, it introduces a pioneering zero-shot probing mechanism to dynamically allocate optimal visual resolutions and designs an uncertainty-guided rollback aggregation strategy to mitigate groupthink effects. Experimental evaluations demonstrate that this framework achieves accuracy improvements of 1.5%–3.2% across six benchmark datasets while effectively optimizing the trade-off between precision and token consumption, all without requiring task-specific fine-tuning.
📝 Abstract
Multi-agent debate (MAD) has emerged as an effective paradigm to improve the reasoning capabilities of large language models (LLMs) and is increasingly being extended to multimodal settings. However, existing multimodal MAD frameworks typically expose agents to the same fixed visual input, ignoring substantial variation in the visual scale needed across samples and agents. In addition, these frameworks frequently suffer from groupthink, a phenomenon where agents prematurely abandon correct deductions to conform with confident but hallucinated peer responses. To address these bottlenecks, we introduce DREAM (Dynamic Resolution Assignment For Multimodal Multi-Agent Debate), which operates via two core components: (1) Dynamic Resolution Assignment, a zero-shot probe round where agents test multiple resolutions, quantify uncertainty using Average Normalized Log-Likelihood (ANLL), and use an adaptive threshold to assign each agent to its empirically optimal resolution; (2) Uncertainty-Guided Rollback Aggregation counters groupthink by tracking each agent's uncertainty over rounds and restoring early low-uncertainty answers overridden by group pressure. On six multimodal datasets, DREAM improves the accuracy-token trade-off over multi-agent debate baselines by 1.5-3.2% accuracy without dataset-specific tuning.
Problem

Research questions and friction points this paper is trying to address.

multimodal multi-agent debate
dynamic visual resolution
groupthink
hallucination
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-agent Debate
Dynamic Resolution Assignment
Uncertainty-Guided Rollback
Groupthink Mitigation
Multimodal LLMs
🔎 Similar Papers
No similar papers found.