regression analysis

Applying and interpreting linear, nonlinear, panel, quantile and nonparametric regression methods (including surrogate-model adaptations like polynomial-chaos expansions) to quantify predictive relationships, explained variance and redundancy across multivariate outcomes and modalities.

regressionanalysis

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

On function-on-function linear quantile regression

Oct 12, 2025
MM
Muge Mutis
🏛️ Yildiz Technical University | Marmara University | Macquarie University

This paper addresses the challenge of quantile regression modeling between functional responses and functional predictors. We propose two novel functional partial quantile regression (FPQR) algorithms. Methodologically, we pioneer the integration of partial quantile regression into the function-on-function regression framework, combining functional principal component analysis (FPCA) for dimension reduction with basis function expansions to transform the infinite-dimensional quantile coefficient function estimation into a finite-dimensional multivariate quantile regression problem. Theoretical analysis and Monte Carlo simulations demonstrate that the proposed methods substantially improve estimation accuracy and robustness, particularly in small-sample settings. Empirical studies further confirm their practical effectiveness. The algorithms are implemented in the R package *ffpqr*, ensuring reproducibility and facilitating broad applicability in functional data analysis.

Approximating function-on-function quantile regression using basis expansion methodsDeveloping functional partial quantile regression algorithms for accurate estimationProjecting infinite-dimensional variables onto finite-dimensional space efficiently

This study addresses the dynamic modeling of conditional means, volatilities, and conditional geometric quantiles in multivariate financial time series by proposing a nonparametric regression framework that accommodates multidimensional responses and covariates. The approach captures temporal dynamics through functional dependence structures and integrates nonparametric kernel estimation with multivariate conditional geometric quantile techniques. Under temporal dependence, the paper establishes, for the first time, the consistency of the conditional geometric quantile estimator and demonstrates that the framework subsumes several classical parametric models as special cases. Theoretical analysis provides both strong and weak convergence results for the proposed estimators. Extensive simulations and empirical applications to stock returns of Maersk and Lockheed Martin confirm the method’s effectiveness and practical utility.

conditional quantilesfinancial econometricsmultivariate time series

Scenario Analysis with Multivariate Bayesian Machine Learning Models

Feb 12, 2025
MP
Michael Pfarrhofer
🏛️ WU Vienna University of Economics and Business | Oesterreichische Nationalbank

This paper addresses the challenge of modeling nonlinear and asymmetric dynamic relationships among macroeconomic and financial variables. We propose the first scenario-analysis-oriented, dynamic nonparametric multivariate Bayesian machine learning framework. Methodologically, we adapt classical econometric tools—including conditional forecasting and generalized impulse response analysis—to high-dimensional Bayesian nonparametric models, integrating dynamic factor extensions and Monte Carlo simulation to enable asymmetric shock response estimation and conditional scenario inference. Our key contribution is the first systematic integration of traditional scenario-analysis tools with nonlinear Bayesian machine learning, explicitly capturing structural asymmetry. The framework is validated across three empirical domains: financial stress testing, macroeconomic risk assessment, and cross-border spillover analysis. Results demonstrate substantial improvements in risk measurement accuracy and cross-jurisdictional early-warning capability, offering a novel paradigm for prudential regulation and policy evaluation.

Adapting scenario analysis tools for nonparametric econometric modelsDeveloping algorithms using predictive simulation and Monte Carlo methodsMeasuring nonlinear macroeconomic risks and financial shock spillovers

Linear Regression in a Nonlinear World

Dec 15, 2025
NK
Nadav Kunievsky
🏛️ University of Chicago

Standard interpretations of OLS coefficients in multiple linear regression presume linearity of the conditional expectation function (CEF), yet real-world data-generating processes are often nonlinear. Method: We show that when the CEF is nonlinear, OLS coefficients represent weighted averages of its partial derivatives, with bias arising systematically from the nonlinear structure of covariates. Leveraging Taylor expansion of the CEF, weighted average theory, and linear projection, we derive a closed-form expression for this bias—termed the “weighted derivative bias”—and prove its decomposition into functions of covariate nonlinearity. Contribution/Results: We unify this bias within an interpretable framework analogous to classical measurement error and omitted-variable bias, establishing a new theoretical foundation for coefficient interpretation under nonlinearity. Simulation and empirical analyses consistently validate the predicted direction and magnitude of the bias.

Assessing bias in coefficients when relationships are nonlinearComparing nonlinear bias to measurement error and omitted variable biasInterpreting regression coefficients under nonlinear data processes

Symbolic Quantile Regression for the Interpretable Prediction of Conditional Quantiles

Aug 11, 2025
CO
Cas Oude Hoekstra
🏛️ Reward Value | Vrije Universiteit Amsterdam

In high-stakes applications requiring simultaneous modeling of multiple conditional quantiles while preserving interpretability, this paper proposes Symbolic Quantile Regression (SQR)—the first framework extending symbolic regression to conditional quantile modeling. SQR employs a tree-based evolutionary algorithm to directly synthesize closed-form quantile functions and optimizes them via quantile loss, enabling distinct characterization of covariate effects on both distributional centers and tails. Its core innovation lies in the principled integration of symbolic regression and quantile regression, jointly ensuring global interpretability, sensitivity to extreme values, and cross-quantile comparability of effect estimates. Experiments demonstrate that SQR outperforms conventional interpretable models across multiple benchmarks, matches the predictive accuracy of state-of-the-art black-box methods, and successfully identifies heterogeneous influential factors at different quantiles in an aviation fuel demand forecasting case study.

Enhancing interpretability in high-stakes quantile estimationPredicting conditional quantiles with symbolic regressionUnderstanding feature influences across different distribution quantiles

Latest Papers

What's happening recently
View more

This study addresses the ill-posedness of integral equations arising from the curse of dimensionality in high-dimensional nonparametric instrumental variable (NPIV) estimation. For the first time, it introduces physical perturbation theory to this domain and proposes a kernel ridge regression–based higher-order perturbation correction method. By leveraging spectral analysis of integral operators and hybridization of eigenmodes, the approach systematically constructs a perturbation expansion that effectively mitigates high-dimensional ill-posedness while preserving computational tractability. Theoretical analysis and empirical experiments demonstrate that, with first-order correction, the method reduces prediction error by up to 99% under severe ill-posedness (β > 0.7), with performance gains intensifying as dimensionality increases—substantially outperforming existing approaches.

curse of dimensionalityhigh-dimensional estimationill-posed inverse problem

This study addresses the complex nonlinear influence of high-dimensional exogenous variables on the conditional variance of financial volatility by proposing the first nonparametric significance test and consistent variable selection method tailored for ARCH-X type models. The approach models covariate effects as unknown nonparametric functions, constructs a test statistic via an artificial one-way ANOVA framework, and controls the false discovery rate using the Benjamini–Yekutieli procedure. Theoretically, the test statistic is shown to be asymptotically standard normal, and the selection procedure consistently includes all truly relevant variables with probability approaching one. Extensive simulations demonstrate superior performance over existing methods, and an empirical application to S&P 500 volatility confirms its practical utility, offering both theoretical rigor and real-world applicability.

ARCH-m(X) modelconditional variancefalse discovery rate

This study addresses the lack of a unified formulation for scalar, multivariate, and functional regression models, which obscures their intrinsic connections. By leveraging an integral operator defined with respect to general measures, the authors propose a unified framework that subsumes all three regression types as special cases of the same operator under different input and output measures. This framework reveals classical regression forms as measure-dependent manifestations of a single operator, clarifies discretized modeling as operator estimation under specific measures, and explains the efficacy of vectorized multivariate regression in linear settings. Theoretically, the authors prove that discrete representations correspond exactly to operator evaluations under discrete measures and converge to the continuous case as the discretization grid refines; moreover, this estimator is equivalent to standard multivariate regression and inherits its classical statistical properties.

functional regressionintegral operatorsmeasure theory

This study addresses the limitations of traditional multivariate response regression methods, which often neglect inter-variable dependencies and struggle to simultaneously achieve smooth fitting and dimensionality reduction. The authors propose a novel unified framework that, for the first time, integrates P-spline smoothing into reduced-rank regression by employing B-spline basis expansions with penalized coefficients. This approach jointly models the correlations among multiple responses and their nonlinear trends. An efficient block-relaxation algorithm is developed for parameter estimation, while biplots and partial dependence plots are incorporated to enhance model interpretability. Extensive simulations and analyses of three real-world datasets demonstrate that the proposed method substantially improves both fitting smoothness and the ability to elucidate the underlying multivariate response structure.

Multiple OutcomesMultivariate RegressionP-splines

This study addresses the challenges of insufficient surrogate accuracy and low sampling efficiency in high-fidelity engineering models with multi-output responses, where conventional single-output modeling neglects inter-output correlations and incurs high computational costs. To overcome these limitations, this work proposes a multivariate active learning strategy based on polynomial chaos expansion for vector-valued outputs. The approach introduces a sequential sampling criterion that jointly exploits statistical correlations among outputs and balances exploration of the input space with exploitation of joint output variance information. By preserving data consistency across outputs, the method significantly enhances both sampling efficiency and global model accuracy. Numerical experiments on multiple engineering benchmarks demonstrate its superior performance in surrogate fidelity, numerical stability, and uncertainty quantification of second-order statistics.

Experimental DesignPolynomial Chaos ExpansionSurrogate Modeling

Hot Scholars

HL

Han Lin Shang

Department of Actuarial Studies and Business Analytics, Macquarie University
Functional data analysisnonparametric smoothingnonparametric statisticsmachine learning
SF

Stefan Feuerriegel

Professor, LMU Munich
AI in ManagementBusiness AnalyticsComputational Social ScienceAI for Good
LF

Long Feng

Professor of Nankai University
High Dimensional DataHigh Frequency Data
NL

Nils Lid Hjort

Professor of Mathematical Statistics, University of Oslo
Theoretical and applied statistics and probability theory
MK

Masahiro Kato

Mizuho-DL Financial Technology Co., Ltd. / The University of Tokyo
Economics