RAISE: A Unified Framework for Responsible AI Scoring and Evaluation

📅 2025-10-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Current evaluations of high-risk AI systems lack unified, quantitative metrics for responsibility dimensions beyond predictive accuracy—namely, explainability, fairness, robustness, and sustainability. Method: This paper introduces RAISE, the first framework unifying these four responsibility dimensions into a computable, comparable, and aggregable scoring system. We conduct multidimensional empirical evaluation across financial, healthcare, and socioeconomic structured datasets, benchmarking models including MLPs, Tabular ResNets, and Feature Tokenizer Transformers. Results: We identify significant responsibility trade-offs across models—for instance, Transformers exhibit superior fairness but higher energy consumption, whereas MLPs demonstrate strong robustness yet limited explainability; no single model dominates all dimensions. RAISE enables cross-model responsibility profiling and ranking, advancing responsible AI from qualitative principles toward systematic, standardized, and quantitatively grounded assessment.

Technology Category

Philosophy and Ethics of AI: Accountability, Interpretability & ExplainabilityHumans and AI: Explainable AI (XAI) for Human UnderstandingMachine Learning: Ethics, Bias, and Fairness

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsResponsible Web: Algorithmic accountability and transparency on the webUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and ranking
📝 Abstract
As AI systems enter high-stakes domains, evaluation must extend beyond predictive accuracy to include explainability, fairness, robustness, and sustainability. We introduce RAISE (Responsible AI Scoring and Evaluation), a unified framework that quantifies model performance across these four dimensions and aggregates them into a single, holistic Responsibility Score. We evaluated three deep learning models: a Multilayer Perceptron (MLP), a Tabular ResNet, and a Feature Tokenizer Transformer, on structured datasets from finance, healthcare, and socioeconomics. Our findings reveal critical trade-offs: the MLP demonstrated strong sustainability and robustness, the Transformer excelled in explainability and fairness at a very high environmental cost, and the Tabular ResNet offered a balanced profile. These results underscore that no single model dominates across all responsibility criteria, highlighting the necessity of multi-dimensional evaluation for responsible model selection. Our implementation is available at: https://github.com/raise-framework/raise.
Problem

Research questions and friction points this paper is trying to address.

Extends AI evaluation beyond accuracy to include explainability, fairness, robustness, and sustainability
Quantifies model performance across multiple responsibility dimensions into unified scores
Reveals critical trade-offs between different responsibility criteria in model selection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified framework quantifies model performance across four dimensions
Aggregates explainability, fairness, robustness, sustainability into single score
Evaluated three deep learning models on structured datasets
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Loc Phuc Truong Nguyen
Friedrich-Alexander-Universität Erlangen-Nürnberg, 91054 Erlangen, Germany
H
Hung Thanh Do
Friedrich-Alexander-Universität Erlangen-Nürnberg, 91054 Erlangen, Germany