The Best of Both Worlds: Bridging Quality and Diversity in Data Selection with Bipartite Graph

๐Ÿ“… 2026-04-11
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
In supervised fine-tuning of large language models, data selection typically prioritizes either quality or diversity, yet simultaneously optimizing both remains challenging. Method: This paper proposes a unified optimization framework that models the dataset as a bipartite graph between sentences and n-grams, andโ€”noveltyโ€”the first formalization of data selection as a *coverage-constrained set cover problem*. It introduces a multiplicative priority function jointly encoding quality and diversity, integrated with a closed-loop greedy selection mechanism featuring dynamic graph updates and iterative re-scoring. Contribution/Results: Extensive experiments across three backbone models and six mainstream benchmarks demonstrate consistent superiority over nine baselines: the method improves downstream task performance while reducing computational cost. Ablation studies further validate the critical role of instruction diversity in enhancing model generalization.

Technology Category

Search and Optimization: Learning to SearchNatural Language Processing: Learning & Optimization for NLPMachine Learning: Optimization

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: Data quality aspects of human-annotated datasetsWeb Mining and Content Analysis: Large pretrained models with web data
๐Ÿ“ Abstract
The performance of large language models (LLMs) is strongly influenced by the quality and diversity of data used during supervised fine-tuning (SFT). However, current data selection methods often prioritize one aspect over the other, resulting in suboptimal training outcomes. To address this, we formulate data selection as a set cover problem and present GraphFilter, a novel approach that balances both quality and diversity in data selection. GraphFilter models the dataset as a bipartite graph connecting sentences to their constituent n-grams, then employs a priority function that combines quality and diversity metrics multiplicatively. GraphFilter iteratively selects sentences with the highest priority, removes covered n-grams from the bipartite graph, and recomputes priorities to reflect the changing data landscape. We validate GraphFilter using three model backbones across six widely-used benchmarks, demonstrating that it outperforms nine existing baselines in both model performance and computational efficiency. Further analysis shows that our design choices lead to more effective subset selection, underscores the value of instruction diversity, and provides insights into how quality and diversity interact with different subset sizes.
Problem

Research questions and friction points this paper is trying to address.

Balancing quality and diversity in data selection for LLMs
Formulating data selection as a set cover problem
Improving model performance and computational efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Models data as bipartite graph for selection
Multiplicatively combines quality and diversity metrics
Iteratively selects sentences and updates priorities
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
Monash University
M
Minghao Wu
Monash University
Thuy-Trang Vu
Thuy-Trang Vu
Monash University
Natural Language ProcessingMachine Learning
L
Lizhen Qu
Monash University
G
Gholamreza Haffari
Monash University