Risk-Averse Online POMDP Planning via CVaR of the Immediate Cost with Performance Guarantees

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the limitation of expected costs in online POMDP planning, which can obscure hazardous states, while existing risk-sensitive methods require bespoke algorithms. We propose applying Conditional Value-at-Risk (CVaR) to immediate costs, directly capturing uncertainty within the belief state. This formulation preserves the standard MDP structure, eliminating the need to reconstruct value functions or develop new algorithms. Consequently, any expectation-based planner achieves risk-sensitive planning by merely modifying its cost computation, seamlessly integrating with particle filtering techniques. The core contribution lies in establishing finite-time theoretical guarantees independent of risk levels. Specifically, we rigorously bound the approximation error between the particle surrogate and the original POMDP, thereby providing end-to-end performance assurances for the proposed framework.
📝 Abstract
Online POMDP planners optimize the expected cumulative cost, which can mask dangerous states when the belief places significant mass on high-cost states. Existing risk-averse methods apply static or dynamic Conditional Value at Risk (CVaR) to the value function, capturing trajectory-level risk, but share two gaps: (i) by retaining the immediate cost as an expectation of a state-dependent cost over the belief, the risk \emph{within} the belief is left unaddressed; and (ii) by modifying the value function, they require new tailored algorithms rather than reusing existing expectation-based planners. We instead apply CVaR to the immediate cost over the belief at each step, directly targeting per-step uncertainty about the current state. The standard expected cumulative return is retained as the objective, so the resulting problem has a standard MDP structure: any expectation-based POMDP planner can be made risk-sensitive by changing only the cost computation. We inherit finite-time guarantees for policy evaluation and sparse sampling---with estimation error independent of the risk level---and, as our central theoretical result, prove a finite-time bound on the gap between the particle belief MDP surrogate and the original POMDP, which together yield an end-to-end guarantee from the true POMDP value to the algorithmic estimate. In the risk-neutral limit, the formulation recovers standard expectation-based planning.
Problem

Research questions and friction points this paper is trying to address.

Online POMDP Planning
Risk-Averse
Conditional Value at Risk (CVaR)
Immediate Cost
Performance Guarantees
Innovation

Methods, ideas, or system contributions that make the work stand out.

Risk-Averse POMDP
Conditional Value at Risk (CVaR)
Immediate Cost
Performance Guarantees
Particle Belief MDP
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yaacov Pariente
Faculty of Mathematics, Technion – Israel Institute of Technology, Haifa, Israel
Vadim Indelman
Vadim Indelman
Associate Professor, Technion
RoboticsPerception/SLAMPOMDP/Belief space planningAIMulti-robot systems