feature importance analysis

Develops and applies quantitative methods to compute, score, and rank the importance or relevance of inputs, features, tokens, and model components (e.g., layers, parameters, experts) using layer-wise, parameter-wise, and relevance-scoring estimators. Uses those importance estimates to interpret contributions to model outputs and predictive accuracy, detect dominant drivers or removable/redundant components, and guide pruning, experiment reduction, or control decisions.

featureimportanceanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.1
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$188K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Variable Selection Using Relative Importance Rankings

Sep 13, 2025
TC
Tien-En Chang
🏛️ National Taiwan University

Conventional filter-based variable selection relies solely on marginal correlations, neglecting dependencies among predictors and thus failing to identify synergistic or context-dependent effects. Method: This paper pioneers the integration of relative importance (RI) analysis as a preprocessing step for variable ranking and screening. We propose CRI.Z—a computationally efficient method that combines generalized dominance (GD) analysis with composite relative importance (CRI) to decompose and quantify both direct and joint contributions of predictors within a multiple regression framework. Contribution/Results: CRI.Z effectively identifies key variables within highly correlated clusters and detects weak-margin but high-synergy predictors. Empirical evaluation demonstrates that RI-based filtering substantially outperforms Lasso and Relaxed Lasso in high-dimensional, multicollinear settings—yielding improved prediction accuracy and model stability. The approach establishes a novel, interpretable paradigm for variable selection grounded in effect decomposition rather than sparsity alone.

Addressing limitations of marginal correlation with dependency-aware measuresEvaluating performance against lasso methods in correlated predictor scenariosUsing relative importance for variable selection before modeling

Effective Data Pruning through Score Extrapolation

Jun 10, 2025
SS
Sebastian Schmidt
🏛️ Technical University of Munich | BMW Group | Pontificia Universidad Católica de Valparaíso | Friedrich-Alexander Universität Erlangen-Nürnberg | Pruna AI

Existing data pruning methods require full initial training to evaluate sample importance, undermining the efficiency benefits of single-stage training. Method: We propose the Score Extrapolation Framework (SEF), which accurately predicts global sample importance from only a small subset training run—enabling, for the first time, importance estimation without full training. SEF jointly leverages k-nearest-neighbor similarity modeling and graph neural network propagation, and is compatible with mainstream pruning strategies (e.g., Dynamic Uncertainty, TDDS) across supervised, unsupervised, and adversarial training paradigms. Results: Evaluated on CIFAR-10/100, Places-365, and ImageNet, SEF reduces pre-pruning computational overhead by up to 87% while preserving model accuracy with negligible degradation (<0.3%). This breaks the long-standing dependency of data pruning on full-model training, establishing a new efficiency frontier for scalable dataset pruning.

Predicting data importance without full trainingReducing computational costs in machine learning trainingScaling pruning methods across diverse datasets

This work addresses the challenge of deploying multi-component neural network controllers, which are often hindered by high computational complexity, and the inadequacy of conventional norm-based pruning methods in accurately capturing the functional importance of individual components. To this end, the paper introduces a component-aware structured pruning framework that, for the first time, integrates three gradient-driven importance metrics—gradient accumulation, Fisher information, and Bayesian uncertainty—into the pruning of multi-component controllers. These metrics enable dynamic assessment of component importance during training, uncovering structural dependencies and temporal variations overlooked by static heuristic approaches. Experiments on autoencoders and TD-MPC reinforcement learning agents demonstrate that the proposed method more accurately identifies critical components, achieving substantial model compression while effectively preserving performance.

computational complexitymodel compressionneural network controllers

Automatic Input Feature Relevance via Spectral Neural Networks

Jun 03, 2024
LC
Lorenzo Chicchi
🏛️ University of Florence | INFN

This work addresses the challenge of quantifying input feature importance in deep neural networks. We propose a training-embedded spectral reparameterization method that directly employs the eigenvalues associated with input nodes as robust proxies for feature relevance, enabling simultaneous feature importance estimation and model training—without post-hoc analysis or auxiliary supervision. Our key contribution is the first use of input-node eigenvalue sensitivity in spectral neural networks to characterize relative feature importance, coupled with spectral reparameterization during optimization to ensure numerical stability. Experiments on both synthetic and real-world datasets demonstrate that the method significantly improves feature selection efficiency and model interpretability while strictly preserving predictive accuracy—achieving zero performance degradation.

Estimate input importance via spectral neural networksIdentify relevant input features for compact datasetsRank input elements by relevance for decision making

Latest Papers

What's happening recently
View more

Conventional magnitude-based criteria in structured pruning often fail to identify redundant filters, leading to suboptimal compression. To address this, we propose IPPRO—a projection-space gradient dynamics analysis method for filter importance estimation. IPPRO abandons reliance on weight magnitudes and instead models the gradient descent trajectory in a linearly projected low-dimensional space, where it quantifies each filter’s directional contribution to optimization. Crucially, it introduces the PROscore, an amplitude-agnostic metric that enables fair and discriminative redundancy assessment. Extensive experiments across multiple CNN architectures (e.g., ResNet, VGG) and datasets (e.g., ImageNet, CIFAR-10/100) demonstrate that IPPRO achieves near-lossless compression at comparable sparsity levels—significantly reducing accuracy degradation over baselines. After fine-tuning, pruned models consistently outperform state-of-the-art structured pruning methods, validating both the effectiveness and generalizability of IPPRO’s importance evaluation mechanism.

Achieves near-lossless pruning with reduced performance dropChallenges magnitude dominance in filter pruning decisionsIntroduces projective space for fair filter pruning evaluation

This work addresses the lack of a unified R framework for computing conditional feature importance and conducting associated statistical inference, which hinders reliable interpretation of machine learning models. To bridge this gap, the authors introduce xplainfi, an R package built on the mlr3 ecosystem that features a novel modular conditional sampling architecture. It integrates diverse samplers—including Gaussian approximations, adversarial random forests, conditional inference trees, and knockoffs—making it suitable for both continuous and mixed-type data. The package supports multiple global importance measures such as permutation importance, conditional and marginal Shapley values, and leave-one-covariate-out methods. Rigorous statistical inference is enabled through variance-corrected confidence intervals and a conditional predictive impact framework. Empirical evaluations demonstrate that xplainfi yields importance scores consistent with existing approaches while maintaining competitive computational efficiency. The package is publicly available on CRAN.

conditional importancefeature importancemachine learning

Transformer inference suffers from low efficiency, and existing gradient-based Head Importance Score (HIS) pruning methods neglect attention pattern diversity, leading to unstable pruning. To address this, we propose a unified pruning criterion that jointly incorporates attention entropy and HIS: entropy quantifies the diversity of attention distributions across heads, while HIS captures task-specific, gradient-driven contribution. By integrating these complementary signals, our method enables a more comprehensive and robust head importance assessment. This work is the first to introduce an information-theoretic perspective—specifically, attention entropy—into attention head pruning. Extensive experiments on multiple NLP benchmarks demonstrate that our approach achieves up to a 15.2% improvement in post-pruning model quality while enhancing pruning stability by 2.04×, all without sacrificing accuracy. The proposed framework establishes a new paradigm for efficient and reliable Transformer compression.

Addresses limitations of gradient-only head importance evaluation methodsEnhances model compression while maintaining accuracy and stabilityImproves transformer pruning by combining importance and entropy scores

Existing attention visualization methods often rely on specific model architectures and incur high computational costs, lacking lightweight and general-purpose tools for token importance analysis. This work proposes a model-agnostic attribution method that incurs no additional overhead by perturbing inputs and introducing a three-matrix analytical framework: the Angular Deviation Matrix, Magnitude Deviation Matrix, and Dimensional Importance Matrix. These matrices respectively capture semantic directional shifts, magnitude changes, and dimensional contributions, enabling fine-grained and mathematically rigorous assessment of token importance. The approach demonstrates strong efficiency and interpretability across multiple large language models, and the authors release their code to support reproducible research.

attention visualizationinterpretabilitylarge language models

RAISE: A Unified Framework for Responsible AI Scoring and Evaluation

Oct 21, 2025
LP
Loc Phuc Truong Nguyen
🏛️ Friedrich-Alexander-Universität Erlangen-Nürnberg

Current evaluations of high-risk AI systems lack unified, quantitative metrics for responsibility dimensions beyond predictive accuracy—namely, explainability, fairness, robustness, and sustainability. Method: This paper introduces RAISE, the first framework unifying these four responsibility dimensions into a computable, comparable, and aggregable scoring system. We conduct multidimensional empirical evaluation across financial, healthcare, and socioeconomic structured datasets, benchmarking models including MLPs, Tabular ResNets, and Feature Tokenizer Transformers. Results: We identify significant responsibility trade-offs across models—for instance, Transformers exhibit superior fairness but higher energy consumption, whereas MLPs demonstrate strong robustness yet limited explainability; no single model dominates all dimensions. RAISE enables cross-model responsibility profiling and ranking, advancing responsible AI from qualitative principles toward systematic, standardized, and quantitatively grounded assessment.

Extends AI evaluation beyond accuracy to include explainability, fairness, robustness, and sustainabilityQuantifies model performance across multiple responsibility dimensions into unified scoresReveals critical trade-offs between different responsibility criteria in model selection

Hot Scholars

LZ

Linfeng Zhang

DP Technology; AI for Science Institute
AI for Sciencemulti-scale modelingmolecular simulationdrug/materials design
XL

Xuyang Liu

Sichuan University
Vision-language ModelsModel CompressionToken CompressionTransfer Learning
ET

Enzo Tartaglione

Associate Professor, Télécom Paris, Institut Polytechnique de Paris
deep learningcompressionpruningdebiasing
YM

Yuki Mitsufuji

Distinguished Engineer, Sony
Machine LearningAudioSource SeparationMusic Technology
JF

Junyi Fan

University of Southern California
machine learning