Diversity-Based Active Learning: An Evaluation of Metric Spaces for Active Learning Selection

📅 2026-08-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文评估了基于多样性的主动学习方法,在不同度量空间下使用贪心K-中心算法选择最具代表性的样本,以减少标记数据需求。
📝 Abstract
With rapid advancement over the last few years, many different methods are now widely used for classification. However, training these models requires substantial labeled data. Active Learning is a potential solution to this problem. Pool-based active learning minimizes costs by querying only the most informative samples from an unlabeled dataset. Diversity-based approaches, on the other hand, attempt to select a representative subset of the data. There are many different objectives for determining the selection process, including exact K-center, exact K-median, and Greedy K-center. In this paper, we will focus on evaluating the performance of Greedy K-center across a variety of metric spaces: the raw feature space, a Linear Discriminant Analysis (LDA) space, and a model-derived probability space (with and without entropy-based weighting). Using Random Forest classifiers as a baseline evaluator, our empirical results on synthetic and real-world datasets demonstrate that mapping unlabeled instances into a predictive probability space and weighting the result by entropy often dominates the other options for active learning selection with Greedy K-center.
Problem

Research questions and friction points this paper is trying to address.

Active Learning
Diversity-based
Greedy K-center
Metric Spaces
Innovation

Methods, ideas, or system contributions that make the work stand out.

Greedy K-center
predictive probability space
entropy-based weighting
S
Siddharth Chilamkur
University of California, Berkeley
D
Dorit S. Hochbaum
University of California, Berkeley