🤖 AI Summary
This study addresses the interpretability challenges in evaluating the effectiveness of self-critique and diversity within multi-agent hypothesis generation. It systematically investigates how critique rounds and equivalence rules influence mechanistic divergence, employing pairwise equivalence judgment as the core measurement criterion combined with LLM-as-a-Judge, TF-IDF, and dense embedding techniques for hypothesis pair scoring. Results demonstrate that a single round of critique induces 34.5% mechanistic divergence, while varying definitions of equivalence rules cause the identified discrepancy rate to fluctuate dramatically between 38% and 96%, revealing the decisive impact of evaluation criteria selection on outcomes. Overall, this work provides an interpretable analytical framework for understanding and quantifying hypothesis diversity in multi-agent systems.
📝 Abstract
Multi-agent systems built on large language models (LLMs) are increasingly applied to scientific discovery and hypothesis generation. Both the effect of refinement and the diversity of the delivered set are hard to interpret before experimental ground truth exists, and both are typically reported by deciding whether pairs of generated hypotheses describe the same underlying mechanism. We study two evaluation questions that rest on this pairwise equivalence judgment: (1) how much self-critique changes delivered hypotheses beyond run-to-run variability, and (2) how the equivalence rule used to group hypotheses affects measured diversity. Across four proprietary instances, we hold opening hypotheses fixed, rerun the downstream workflow with 0, 1, and 5 critique rounds, and score matched hypothesis pairs with an LLM-as-a-judge. Relative to matched same-depth reruns, moving from 0 to 1 round produces 34.5 percentage points (pp) of additional mechanism-level divergence, whereas 1 to 5 rounds adds 1.3 pp. We then compare three equivalence rules: term frequency--inverse document frequency (TF--IDF) similarity, dense embeddings, and the same LLM-as-a-judge. We construct controlled hypothesis pairs that either preserve the causal explanation through wording or biological-terminology changes, or replace one component of the causal chain while holding the rest fixed. All three rules are invariant to meaning-preserving edits, but when the initiating event is replaced, the LLM-as-a-judge identifies 83% of valid pairs as different mechanisms, versus 0% for TF--IDF and 8% for embeddings; varying only the rubric that defines same mechanism moves this figure from 38% to 96%. Together, these results show that pairwise equivalence judgments are a measurement choice: how mechanism equivalence is defined affects both the estimated effect of self-critique and the measured diversity of generated hypotheses.