FASE: Fast Adaptive Semantic Entropy for Code Quality

📅 2026-06-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of unreliable code generation in multi-agent systems, where hallucinations and error propagation in large language models undermine system robustness. Existing uncertainty estimation methods based on semantic entropy rely on costly equivalence judgments, limiting their practicality. To overcome this, the authors propose a novel proxy metric for functional correctness that requires neither ground-truth labels nor expensive model-based equivalence checks. Their approach constructs a discrepancy graph by fusing structural and semantic embeddings and efficiently quantifies uncertainty via minimum spanning trees. Evaluated on HumanEval and BigCodeBench, the method achieves an average 25% improvement in Spearman correlation and a 19% gain in ROCAUC over prior approaches, while reducing runtime to just 0.3% of conventional methods—demonstrating a substantial advance in both efficiency and assessment accuracy.
📝 Abstract
Multi-agent code generation offers a promising paradigm for autonomous software development by simulating the human software engineering lifecycle. However, system reliability remains hindered by LLM hallucinations and error propagation across interacting agents. While semantic entropy provides a principled way to quantify uncertainty without ground-truth answers, current methods often rely on costly LLM-driven equivalence checks. In this work, we introduce Fast Adaptive Semantic Entropy (FASE), a novel metric that approximates functional correctness based on the minimum spanning tree of structural and semantic dissimilarity graphs. Evaluations on HumanEval and BigCodeBench demonstrate that FASE outperforms state-of-the-art semantic entropy by LLM entailment, achieving a 25% average improvement in Spearman correlation and a 19% increase in ROCAUC score against Pass@1 from ground-truth test cases when using the Qwen3-Embedding-8B model. Furthermore, by eliminating costly LLM-driven equivalence evaluation, FASE incurs negligible computational overhead, requiring only approximately 0.3% of the runtime cost of traditional semantic entropy approaches. These results position FASE as a practical, cost-effective solution for optimizing uncertainty quantification in real-world multi-agent workflows.
Problem

Research questions and friction points this paper is trying to address.

multi-agent code generation
LLM hallucinations
semantic entropy
uncertainty quantification
code quality
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fast Adaptive Semantic Entropy
semantic entropy
multi-agent code generation
uncertainty quantification
minimum spanning tree