Understanding Dataset Difficulty with V-Usable Information

📅 2021-10-16
🏛️ International Conference on Machine Learning
📈 Citations: 225
✨ Influential: 38
📄 PDF
🤖 AI Summary
This paper addresses the challenge of quantifying the intrinsic difficulty of a dataset for a given model. Method: We propose an information-theoretic framework—V-usable information—and introduce Pointwise V-Information (PVI), a fine-grained, model-aware metric that formalizes sample-level difficulty as the deficit of information usable by the model. PVI enables interpretable difficulty attribution across datasets, subpopulations, and input attributes, supports reverse evaluation (fixed model, multiple datasets), and facilitates diagnosis of annotation artifacts. Contribution/Results: Evaluated on NLP benchmarks, the framework successfully uncovers latent annotation biases. Experiments confirm its model-agnosticism, cross-task consistency, and strong interpretability. By grounding data difficulty in information usability, PVI provides a unified, computationally tractable theoretical tool for rigorous difficulty analysis.
📝 Abstract
Estimating the difficulty of a dataset typically involves comparing state-of-the-art models to humans; the bigger the performance gap, the harder the dataset is said to be. However, this comparison provides little understanding of how difficult each instance in a given distribution is, or what attributes make the dataset difficult for a given model. To address these questions, we frame dataset difficulty -- w.r.t. a model $mathcal{V}$ -- as the lack of $mathcal{V}$-$ extit{usable information}$ (Xu et al., 2019), where a lower value indicates a more difficult dataset for $mathcal{V}$. We further introduce $ extit{pointwise $mathcal{V}$-information}$ (PVI) for measuring the difficulty of individual instances w.r.t. a given distribution. While standard evaluation metrics typically only compare different models for the same dataset, $mathcal{V}$-$ extit{usable information}$ and PVI also permit the converse: for a given model $mathcal{V}$, we can compare different datasets, as well as different instances/slices of the same dataset. Furthermore, our framework allows for the interpretability of different input attributes via transformations of the input, which we use to discover annotation artefacts in widely-used NLP benchmarks.
Problem

Research questions and friction points this paper is trying to address.

Measures dataset difficulty using V-usable information
Introduces pointwise V-information for instance-level difficulty
Enables interpretability of input attributes via transformations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses $mathcal{V}$-usable information for dataset difficulty
Introduces pointwise $mathcal{V}$-information (PVI) for instances
Enables interpretability via input attribute transformations