Evaluation Sovereignty in Metadata-Driven Classification: A Multi-Track Framework for Weakly Supervised Information Systems

📅 2026-06-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
In weakly supervised learning, evaluation metrics are often confounded by the label generation mechanism, obscuring true model performance. This work introduces the concept of “evaluation sovereignty,” framing evaluation validity as a system-level property governed by label provenance, and proposes a multi-track evaluation framework that systematically varies the sources of training and evaluation labels—such as silver versus gold labels—to probe a model’s dependence on label authority. Through hierarchical multi-label classification experiments combining Micro-F1 and ranking-based metrics, the study reveals that models achieving strong performance under silver-label evaluation (Micro-F1 ≈ 0.54) collapse dramatically (to Micro-F1 ≈ 0.03) when assessed on independent gold labels, suggesting that current evaluation practices may primarily measure alignment with noisy supervision rather than genuine predictive capability.
📝 Abstract
Evaluation in machine learning is typically treated as a neutral measurement process. However, in operational information systems, evaluation outcomes are often conditioned by the processes used to generate labels. This paper does not seek to improve classification performance. Instead, it examines the validity of performance measurement under differing label-authority regimes. This issue is particularly relevant in large-scale metadata-driven systems, where labels are often incomplete, inconsistent, or weakly supervised. We introduce evaluation sovereignty, defined as the degree to which performance metrics are independent of label authority and supervision regime, and propose a multi-track evaluation framework that systematically varies training and evaluation label sources. Using hierarchical multi-label classification on large-scale scientific metadata, we demonstrate that models exhibiting strong performance under operational ("silver") evaluation degrade substantially under independent ("gold") evaluation, particularly for fine-grained classification. For example, Micro-F1 decreases from approximately 0.54 to 0.03. Notably, ranking-based metrics remain above baseline, revealing a divergence between latent model signal and classification validity. These findings suggest that commonly reported performance metrics may reflect alignment with labeling processes rather than true predictive capability. We therefore reconceptualize evaluation validity as a system-level property shaped by label governance and provide a practical methodology for auditing intelligent systems operating under weak supervision.
Problem

Research questions and friction points this paper is trying to address.

evaluation sovereignty
weak supervision
label authority
metadata-driven classification
performance validity
Innovation

Methods, ideas, or system contributions that make the work stand out.

evaluation sovereignty
weak supervision
multi-track evaluation
metadata-driven classification
label authority
🔎 Similar Papers
No similar papers found.