CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data

📅 2026-07-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current evaluation methods for large language models (LLMs) primarily identify failing samples or categories but struggle to uncover underlying capability deficiencies, thereby limiting targeted model improvement. This work proposes CRAFT, a novel framework that diagnoses model weaknesses at the scoring-criterion level. CRAFT constructs a hierarchical capability tree by extracting capability descriptions and applying hierarchical clustering, then dynamically identifies low-performance nodes across multiple granularities to generate targeted fine-tuning data. Evaluated on financial and legal domains as well as 13 standard benchmarks, CRAFT significantly outperforms prompt-clustering and random data generation baselines. Fine-tuning four open-source LLMs with CRAFT-generated data consistently enhances their performance, demonstrating more precise localization of capability gaps and enabling efficient, targeted model refinement.
📝 Abstract
Evaluations should do more than measure a models current performance. They should tell us what to fix for the next model iteration and provide a way to generate targeted post training data. Most evaluation pipelines identify weak examples, topics, or categories, but they leave the underlying capability failure implicit: they say where a model fails, not why. We introduce CRAFT, a method that converts any rubric based evaluation dataset into a model specific diagnosis of weak capabilities. CRAFT treats each grading criterion as a capability probe: it extracts a capability description from every prompt rubric pair, clusters these descriptions into a hierarchical capability tree, scores the target model at every node, and selects low performing nodes dynamically across tree levels, at the granularity where each failure is clearest. The selected weak capabilities then direct the generation of targeted supervised finetuning data. Holding the data generation, finetuning, and evaluation setup fixed, we compare CRAFT against prompt level EvalTree clustering and untargeted random generation on four open source models, two professional domains (finance and legal), and 13 held out benchmarks disjoint from the diagnostic data. CRAFT achieves the strongest finance domain average for all four models under repeated temperature decoding; on legal domain, it is strongest for three of four models and remains within the decoding variance bands of the best baseline on the fourth. Diagnosing weaknesses at the level of rubric criteria, rather than prompts or categories, thus yields both a sharper picture of what a model cannot do and measurably better models after finetuning on that diagnosis.
Problem

Research questions and friction points this paper is trying to address.

evaluation
capability diagnosis
fine-tuning data
rubric-based assessment
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

capability diagnosis
rubric-based evaluation
hierarchical clustering
targeted fine-tuning
large language models
🔎 Similar Papers
No similar papers found.