Score
Designs and implements optimization objectives, training schedules, or sampling procedures that adaptively change the Conditional Value-at-Risk (CVaR, also called expected shortfall) quantile during learning to concentrate computation on tail losses; builds mechanisms to prioritize worst‑case tail samples and progressively tighten the sampled tail. Analyzes the resulting algorithms’ behavior, convergence, and performance tradeoffs when using an adaptive CVaR (expected shortfall) schedule.
This work addresses the challenge of unifying the modeling of stochastic objectives such as risk, bias, regret, and error by proposing an optimization framework grounded in a generalized “risk quadrangle.” By incorporating advanced risk measures like superquantiles and expectiles, and by developing a “sub-regularity” axiom system that relaxes conventional regularity assumptions, the approach overcomes limitations of classical theory and enhances model flexibility. Leveraging duality analysis, generalized stochastic divergences, and robust optimization techniques, the framework demonstrates superior performance in portfolio optimization, regression, and classification tasks. The study highlights the central role of duality in risk-sensitive decision-making and significantly broadens the applicability of risk modeling in machine learning, finance, and related domains.
This work addresses the low sample efficiency of Conditional Value-at-Risk (CVaR) policy gradient methods, which stems from their exclusive reliance on tail trajectories. To overcome this limitation, the authors propose an augmented optimization objective that incorporates an expected quantile term, enabling full utilization of all sampled trajectories through quantile dynamic programming while preserving the original CVaR optimization goal. This approach constitutes the first method to achieve efficient, full-sample CVaR policy optimization within the class of Markov policies. Empirical evaluations across multiple environments with verifiable risk-averse behaviors demonstrate that the proposed method significantly outperforms conventional CVaR-PG and other state-of-the-art risk-sensitive reinforcement learning algorithms in terms of both sample efficiency and performance.
This work addresses the instability in value estimation, high Bellman residuals, and low sample efficiency that plague CVaR risk-sensitive Q-learning under fixed hyperparameters. To overcome these limitations while preserving the original CVaR optimization objective, the authors propose an adaptive training controller that integrates six synergistic mechanisms: adaptive inner-loop step-size adjustment, synchronized decay of inner and outer loops, early correction of VaR variables, coverage-prioritized re-greedy sampling, progressive estimate aggregation, and data-driven scale calibration. Evaluated on a Bitcoin trading task, the proposed method reduces average Bellman residuals by 85%, achieves a policy Sharpe ratio of 0.9281, limits maximum drawdown to merely 6.46%, and demonstrates significantly lower volatility compared to the buy-and-hold strategy.
This work addresses the lack of theoretical guarantees for Conditional Value-at-Risk (CVaR) learning under heavy-tailed and contaminated data, where endogenous quantile estimation often leads to threshold sensitivity and unstable decisions. Through a learning-theoretic analysis of CVaR empirical risk minimization in such settings, this study uncovers—for the first time—the error structure driven by the CVaR threshold and proposes a truncated median-of-means CVaR estimator. By integrating Bahadur–Kiefer-type expansions, minimax optimal rate analysis, techniques for β-mixing sequences, and robust statistical methods, the paper establishes high-probability generalization and excess risk bounds under weak moment conditions. The proposed estimator achieves minimax optimal convergence rates under adversarial contamination and precisely characterizes the boundary conditions that determine whether CVaR learning is generalizable or inherently unstable.
This study addresses the challenge that existing methods struggle to accurately characterize the influence of covariates on the tail distribution of a response variable, and that superquantile regression and Expected Shortfall (ES) regression often yield inconsistent estimates. To resolve this, the authors propose an optimization-based linear ES regression approach that avoids imposing additional assumptions on conditional quantiles and instead employs an implicit loss function to precisely model tail risk. The key innovation lies in explicitly distinguishing between superquantile and ES regression for the first time and introducing heterogeneity-adaptive weights to enhance estimation efficiency. By combining binning-based initial values with a tailored optimization algorithm, the method ensures consistency and asymptotic normality, while simulation studies demonstrate its superior performance over existing approaches across various settings.
Existing Conditional Value-at-Risk policy gradient (CVaR-PG) methods achieve risk-sensitive optimization by discarding low-return trajectories, resulting in severe sample inefficiency. This work proposes **Return Capping**, a novel mechanism that imposes a hard upper bound on trajectory returns during training and incorporates a reweighting estimator—provably equivalent to the original CVaR objective without information loss. Unlike prior approaches, Return Capping shifts CVaR-PG’s sample efficiency paradigm from *discarding* to *reusing* trajectories, substantially improving data utilization. Empirically, the method achieves faster convergence and enhanced robustness across multiple benchmark environments. It improves sample efficiency by several-fold over state-of-the-art baselines while maintaining theoretical equivalence to the CVaR optimization goal. This work thus provides the first provably equivalent, high-efficiency policy gradient implementation for risk-sensitive reinforcement learning.
This work addresses the challenge that when the loss function depends on decision variables, the regularity properties—such as continuity and differentiability—of Value-at-Risk (VaR) and Conditional Value-at-Risk (CVaR) are generally not guaranteed, thereby hindering theoretical and algorithmic advances in risk-aware optimization. Focusing on such decision-dependent losses, the paper establishes simple yet rigorous sufficient conditions under which CVaR is proven for the first time to be continuously differentiable, and provides an explicit expression for its gradient. By integrating tools from perturbation analysis of probability measures, path-differentiability, and real analysis, the study develops a unified theoretical framework that ensures the continuity of VaR and the continuous differentiability of CVaR, thereby furnishing reliable gradient information and convergence guarantees for optimization problems involving tail risk.
This study addresses the optimal joint allocation of put options and trend-following strategies for tail risk management under multifaceted adverse scenarios—including market crashes, volatility repricing, and prolonged drawdowns. The authors develop a continuous-time Conditional Value-at-Risk (CVaR) framework that unifies both mechanisms within a single optimization objective. By modeling wealth, spot price, stochastic variance, and an exponentially weighted log-return signal as Markovian state variables, they derive the viscosity solution to the associated Hamilton–Jacobi–Bellman equation. A key innovation is the temporal decoupling of protective mechanisms: put options deliver immediate convexity-based protection, while trend following enhances resilience during sustained drawdowns. The framework further incorporates a four-dimensional diagnostic layer assessing conditional convexity, tail-event reliability, holding costs, and drawdown persistence. Monte Carlo simulations demonstrate that the hybrid strategy substantially reduces terminal CVaR, with optimal allocations highly sensitive to parameter calibration.
This study addresses the significant underestimation of conditional value-at-risk (CVaR) at high confidence levels by traditional operational risk models, which neglect dynamic tail dependence between loss frequency and severity as well as parameter uncertainty. To overcome this limitation, the authors propose a Bayesian extreme value theory framework that innovatively integrates a Hawkes process to capture self-exciting clustering in loss frequency, an autoregressive latent variable to model persistence in stress regimes, and a Gumbel upper-tail copula to represent asymmetric tail dependence. Full Bayesian inference is implemented via Hamiltonian Monte Carlo using PyMC, and CVaR is estimated through posterior predictive simulation. At the 99.995% confidence level, the proposed approach effectively mitigates the approximately 40% CVaR underestimation inherent in standard Loss Distribution Approach (LDA) models and corrects structural flaws in shared-factor models, accurately reproducing empirical tail dependence and substantially improving the precision of extreme risk measurement.