๐ค AI Summary
In multi-task learning, there is no standardized, test-set-free metric to uniformly assess the generalization capability of representations for downstream kernel ridge regression (KRR) tasks.
Method: We propose the Uniform Kernel Prober (UKP), the first task-agnostic, kernel-driven pseudo-metric for representation quality. UKP explicitly encodes desired invariances via a user-specified kernel and estimates representation quality solely from training data. It enjoys an $O(1/sqrt{n})$ convergence rate and admits efficient computation.
Contribution/Results: Unlike conventional metrics, UKP does not require task-specific labels or test sets, yielding comparable upper bounds on KRR generalization error across diverse features or representations. Extensive benchmark experiments demonstrate that UKP robustly discriminates representation quality and exhibits strong correlation with actual KRR generalization errorโvalidating its effectiveness, robustness, and plug-and-play utility.
๐ Abstract
The ability to identify useful features or representations of the input data based on training data that achieves low prediction error on test data across multiple prediction tasks is considered the key to multitask learning success. In practice, however, one faces the issue of the choice of prediction tasks and the availability of test data from the chosen tasks while comparing the relative performance of different features. In this work, we develop a class of pseudometrics called Uniform Kernel Prober (UKP) for comparing features or representations learned by different statistical models such as neural networks when the downstream prediction tasks involve kernel ridge regression. The proposed pseudometric, UKP, between any two representations, provides a uniform measure of prediction error on test data corresponding to a general class of kernel ridge regression tasks for a given choice of a kernel without access to test data. Additionally, desired invariances in representations can be successfully captured by UKP only through the choice of the kernel function and the pseudometric can be efficiently estimated from $n$ input data samples with $O(frac{1}{sqrt{n}})$ estimation error. We also experimentally demonstrate the ability of UKP to discriminate between different types of features or representations based on their generalization performance on downstream kernel ridge regression tasks.