🤖 AI Summary
This study addresses the critical annotation errors in existing NL-to-SQL benchmarks that distort model evaluation. We propose an automated error correction framework based on multi-agent weak supervision, which reformulates benchmark auditing as an error detection task. By leveraging generative labeling models to extract high-confidence samples and training a decision plane that integrates reliability estimates, our approach achieves high-precision automatic correction without requiring human-annotated ground truth. Experiments demonstrate that the proposed method attains an F1 score of 0.9194 on BIRD-Clean-xs, significantly outperforming baselines. Furthermore, our analysis reveals annotation error rates as high as 37% and 27% in the development sets of BIRD and Spider, respectively, establishing a new paradigm for assessing benchmark reliability.
📝 Abstract
Natural language to SQL (NL-to-SQL) benchmarks are foundational to progress in data analysis research, yet recent work has shown that widely-used benchmarks contain significant annotation errors. These errors silently corrupt evaluation metrics, penalize correct model output, and distort the field's understanding of state-of-the-art performance. We present SALUS, a system that automatically detects annotation errors in NL-to-SQL benchmarks. SALUS frames benchmark auditing as a weakly supervised error detection: SQL generated by multiple LLM agents drive a suite of complementary weak-labeling functions. By passing this noisy vote matrix through a generative label model, we extract high-confidence training samples without requiring human ground truth. These samples train a decision plane that maps gold SQL query features to per-agent trustworthiness, allowing SALUS to intelligently fuse reliability estimates with raw verdicts for rigorous benchmark error detection. We evaluate on BIRD-Clean-xs, a benchmark of 298 BIRD development tasks with manually verified correctness labels. SALUS achieves F1 = 0.9194, significantly outperforming the state-of-the-art baselines. Applying SALUS to the full development sets, we estimate annotation error rates of approximately 37% on BIRD and 27% on Spider.