🤖 AI Summary
Existing approaches to uncertainty quantification in machine learning inadequately identify and formalize the distinct sources of uncertainty, relying predominantly on the oversimplified aleatoric/epistemic dichotomy. Method: Departing from this binary paradigm, we ground our analysis in the fundamentals of statistical modeling, formally defining the conceptual distinctions and boundary conditions between uncertainty types, demonstrating their inherent non-decomposability and strong dependence on data characteristics. We establish a conceptual mapping framework linking statistical inference principles to ML uncertainty, integrating statistical theory, formal modeling, and conceptual analysis. Contribution/Results: This yields a traceable, multi-dimensional uncertainty taxonomy that resolves longstanding conceptual ambiguities and practical misapplications. The framework provides a rigorous statistical foundation for trustworthy AI and delivers an actionable methodology for principled uncertainty quantification.
📝 Abstract
Machine Learning and Deep Learning have achieved an impressive standard today, enabling us to answer questions that were inconceivable a few years ago. Besides these successes, it becomes clear, that beyond pure prediction, which is the primary strength of most supervised machine learning algorithms, the quantification of uncertainty is relevant and necessary as well. While first concepts and ideas in this direction have emerged in recent years, this paper adopts a conceptual perspective and examines possible sources of uncertainty. By adopting the viewpoint of a statistician, we discuss the concepts of aleatoric and epistemic uncertainty, which are more commonly associated with machine learning. The paper aims to formalize the two types of uncertainty and demonstrates that sources of uncertainty are miscellaneous and can not always be decomposed into aleatoric and epistemic. Drawing parallels between statistical concepts and uncertainty in machine learning, we also demonstrate the role of data and their influence on uncertainty.