Score
Design interpretable models: build and analyze predictive systems whose architectures, feature representations, and inference procedures are structurally transparent or constrained for intrinsic interpretability, intentionally trading some expressivity for human-understandable internal mechanisms. Produce explainable outputs and reports—e.g., perturbation-based or additive attributions, stakeholder-weighted and multilingual explanations—integrated into decision workflows and evaluated for fidelity and comprehensibility, with mechanisms for routing and prioritizing downstream actions.
Neural networks suffer from limited interpretability, hindering their adoption in high-stakes decision-making. To address this, we present a systematic survey of self-interpretable neural networks (SINNs) and propose the first five-dimensional taxonomy—covering attribution-, function-, concept-, prototype-, and rule-based approaches—unified across multimodal domains including vision, natural language processing, graph learning, and deep reinforcement learning. Methodologically, we establish a structured framework for SINN design and evaluation, introduce the first open-source tracking repository (Awesome-Self-Interpretable-Neural-Network), and rigorously define evaluation metrics and fundamental open challenges. Through comprehensive literature analysis, modeling abstraction, visual case studies, and cross-domain validation, we deliver a reproducible interpretability paradigm and practical implementation guidelines. Our work bridges the gap between theoretical SINN design and trustworthy real-world deployment.
This work addresses the limitations of existing explainable AI methods, which predominantly focus on associative predictions and fall short in supporting decision-making that requires causal reasoning and counterfactual analysis. To bridge this gap, the paper proposes a novel framework that integrates causal machine learning with intrinsically interpretable models—such as additive models and symbolic regression—by explicitly embedding causal inference mechanisms within the model architecture. This approach enables the explicit recovery of causal structures and functional forms among variables directly from cross-sectional data. While maintaining high predictive accuracy, the method achieves comprehensive transparency in system structure, causal relationships, and response mechanisms, thereby substantially enhancing both interpretability and causal reliability for trustworthy “What-if” analyses.
Current interpretability lacks an operational definition, resulting in weak theoretical foundations and insufficient general guidance for model design. Method: We propose the first unified, formal, and actionable definition of interpretability, systematically characterizing its essential properties, core assumptions, design principles, and architectural constraints; based on this, we develop a generic blueprint for interpretable model construction and introduce novel, interpretability-native data structures and computational workflows. Contribution/Results: We implement these advances in XAI-Struct, an open-source software library. This work establishes the first systematic theoretical framework for explainable AI (XAI), closing the full loop from definition to theory to tooling. It enables rigorous, standardized, and engineering-oriented development of interpretable models, advancing the field toward scientific maturity and practical deployability.
This work addresses the fundamental trade-off in machine learning between high predictive performance and low interpretability inherent in “black-box” models (e.g., deep neural networks, ensemble methods). It rigorously distinguishes post-hoc explanation—applied after model training—from inherently interpretable modeling—designed for transparency from inception. To reconcile accuracy and interpretability, we propose a hybrid modeling paradigm centered on symbolic knowledge embedding, integrating differentiable symbolic modules, knowledge distillation, and symbolic reasoning into the model architecture itself. This enables joint optimization of fidelity and interpretability at the design stage. Extensive experiments across diverse domains demonstrate that our approach matches the predictive accuracy of state-of-the-art black-box models while generating human-understandable, logically grounded decision rules. As a result, it substantially enhances model trustworthiness and deployment viability in safety- and accountability-critical applications.
Deploying large neural networks in high-stakes domains exposes a critical risk: misalignment between their internal computations and human values—particularly when behavioral alignment methods (e.g., RLHF, red-teaming, Constitutional AI) fail to detect deceptive or latent misaligned reasoning. This work proposes elevating mechanistic interpretability—not as a post-hoc diagnostic tool, but as a foundational design principle for alignment. We employ causal, circuit-based techniques—including circuit tracing and activation patching—to explicitly model internal computational mechanisms, and integrate LIME/SHAP to quantify alignment between model representations and human concepts. Our key contribution is the first systematic argument that interpretability must be *constructive* and *pre-deployment*, serving as the architectural basis for alignment rather than a verification supplement. This approach enhances transparency, auditability, and value consistency, yielding both a theoretical framework and scalable methodology for safe, trustworthy AI.
In high-stakes domains, AI decision-making suffers from insufficient transparency and post-hoc explanations that inadequately address accountability and trust. Method: This study proposes an “explanability-by-design” paradigm, embedding explainability intrinsically across the AI system lifecycle. We develop a tiered implementation framework calibrated to organizational capability differences, integrating feature importance analysis, local interpretability methods (e.g., LIME, SHAP), reasoning-path visualization, and dynamic model-behavior tracing—validated and refined through cross-sector empirical studies in healthcare and finance. Contribution/Results: First, we establish explainability as a foundational design principle—not an add-on module. Second, we deliver a scalable, production-ready engineering pathway. Third, empirical evaluation demonstrates significant improvements in model transparency, stakeholder trust, and regulatory compliance; notably, the framework also enables iterative model performance enhancement through interpretability-driven insights.
This work addresses the lack of a unified theoretical foundation in interpretable machine learning, which has led to fragmented methodologies and inconsistent evaluation criteria. By introducing Lagrangian mechanics into this domain for the first time, the paper proposes a general theoretical framework grounded in user-oriented interpretability. Through systematic analysis of symmetries and constraints, the approach derives optimal interpretable models by minimizing a suitably defined Lagrangian. This deductive methodology not only unifies existing techniques under a coherent theoretical umbrella but also reveals novel research directions. It has successfully informed the design of core programming interfaces, mitigated limitations of current methods, and established a rigorous theoretical basis for interpretability education and interdisciplinary integration.
Current machine learning evaluation practices predominantly rely on surface-level performance metrics, often neglecting the internal mechanisms of models. This work proposes trustworthy interpretability as a central evaluation paradigm and, for the first time, systematically demonstrates that it satisfies core criteria from the philosophy of science—namely falsifiability, reproducibility, and predictive power. By constructing an evaluation framework that integrates causal analysis with mechanistic probing, the study delineates three functional pathways through which interpretability enables the identification of behavioral origins, detection of latent flaws, and prediction of potential failure modes. This approach advances model assessment beyond performance-oriented benchmarks toward a deeper understanding of underlying mechanisms.
This study clarifies the conceptual confusion in physics-oriented machine learning between “interpretability”—referring to model transparency—and “explainability,” which denotes the capacity to map onto domain knowledge. It delineates the boundaries of these two notions and examines their trade-offs in terms of expressive power and adaptability. Through conceptual analysis and the construction of a unifying framework, complemented by a systematic review of both intrinsic and post-hoc explanation methods, the work advocates for integrating interpretability and explainability into scientific modeling paradigms. Crucially, it underscores the central role of task formulation and intervention design in model development. By establishing a clear conceptual foundation and methodological guidance, this research advances the principled integration of machine learning models with scientific reasoning in physics.
The "black-box" nature of large language models hinders their trustworthy and secure deployment. This work presents a systematic survey of intrinsic interpretability research and, for the first time, proposes five core design paradigms: functional transparency, concept alignment, decomposable representations, explicit modularity, and latent sparsity induction. By establishing a unified taxonomic framework that integrates multiple technical approaches, the study clarifies the evolving research landscape, identifies key challenges, and outlines promising future directions. The resulting synthesis offers both theoretical foundations and architectural guidance for developing more interpretable and trustworthy large language models.
This work addresses the challenge of providing trustworthy explanations for deep learning models in predictive process monitoring, where their black-box nature and the limitations of existing attribution methods hinder both computational efficiency for long traces and semantic fidelity to control-flow dynamics. The authors propose a novel local post-hoc interpretability approach that leverages control-flow structures to semantically segment event logs and computes SHAP attributions over these segments to identify critical process fragments and turning points influencing predictions. Experimental results demonstrate that the method accurately captures known logical turning points on synthetic datasets and effectively uncovers the underlying dynamic mechanisms driving predictions in real-world loan approval and municipal process logs, achieving a balanced trade-off between computational efficiency and process-aware semantic interpretability.