🤖 AI Summary
This study addresses the challenge of disentangling the specific contributions of code data, typically treated as a monolithic corpus, to different models and downstream tasks. To this end, this work proposes a decomposition perspective grounded in computational patterns, categorizing execution-verified code corpora accordingly. We construct a controlled comparative fine-tuning framework to systematically analyze the independent and combinatorial effects of each category on question answering, mathematical reasoning, and code generation tasks. Our findings reveal a "less is more" principle in data mixing: for specific model-task configurations, compact mixtures meticulously curated from merely 10%–15% of the full corpus consistently outperform both the best single-category baselines and training on the entire dataset. These results highlight the critical importance of strategic data composition over sheer volume in optimizing language model performance across diverse downstream applications.
📝 Abstract
Evaluating code data as a single corpus can obscure which types of code data benefit which models and downstream tasks. Effective data selection requires understanding both the benefits of individual categories and whether these benefits persist when categories are combined. We introduce a decompositional lens for studying these effects in LLM post-training. We first decompose an execution-verified code corpus into interpretable categories based on the computational patterns of its solutions. Through controlled fine-tuning experiments, we compare individual categories with a balanced mixture across instruction-tuned models on question answering, mathematics, and code generation. The resulting response maps reveal recurring gains in average question-answering performance, while the same category can improve one model or task and degrade another. The best-performing category also varies with the starting model and target task. We then compose compact mixtures guided by these results and examine whether benefits observed in individual categories persist under joint training. On selected model--task pairs, mixtures whose constituents each improve the target task outperform both their best constituent and full-corpus training while using roughly 10--15\% of the full corpus. These exploratory findings illustrate a \emph{less is more} pattern and highlight how the value of code data in post training depends on which categories are combined for which model and task.