RSD-Poker: Structure-Adaptive and Shift-Robust Risk-Utility Certification for Residual Policies in Imperfect-Information Games

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the heterogeneous degradation and vulnerability of residual policies under distribution shift caused by fixed groupings. We propose an adaptive structural certification framework that freezes the residual family prior to label injection, learns policy-visible partitions, and introduces a hybrid robustness selection mapping. By integrating weighted simultaneous certificates, robust mapping selection over pre-declared uncertainty sets, and deterministic auditing techniques, the framework jointly guarantees risk upper bounds and utility. Experimental results demonstrate that the proposed approach significantly improves weak-response utility, reduces the violation rate to 0.0059, and achieves high-probability satisfaction of risk budgets.
📝 Abstract
Residual policy adaptation provides a lightweight way to modify a strong reference policy, but a shared scale and a fixed subgroup partition can hide heterogeneous degradation and become fragile when the deployment mixture of information states changes. We introduce RSD-Poker, a structure-adaptive and shift-robust certification framework that freezes a bank of residual families and scales, learns a policy-visible partition on an independent structure split, and freezes that partition before calibration labels are joined. Each candidate-group pair receives a weighted simultaneous upper certificate for anchor-relative risk and a lower certificate for weak-response utility. A robust group-to-candidate map is then selected over a predeclared uncertainty set of deployment group proportions. Under independent calibration units drawn from each frozen group's law, a candidate bank and partition fixed before calibration, and invariant within-group conditionals, the selected map satisfies its declared mixture-robust risk budget and utility certificate with probability at least $1-\zeta_{risk}-\zeta_{util}$. The information contract supports both a teacher-backed transform and a teacher-free observation-only student. The retained deterministic 24-state audit remains an exact replay diagnostic: empirical-zero selects $\alpha=0.08$, raising the weak-response proxy from 4.2082 to 4.2889 with $0/12$ held-out threshold crossings. On stratified held-out states, the learned-partition dual selector raises weak utility from 4.4074 under global dual certification to 4.4936 and lowers held-out violation from 0.0215 to 0.0078; its mixture-robust variant reaches violation 0.0059. Across five observation-only checkpoints, risk-calibrated residuals attain weak utility $4.3659\pm0.0177$ and violation rate $0.0178\pm0.0057$.
Problem

Research questions and friction points this paper is trying to address.

residual policy adaptation
imperfect-information games
distribution shift
risk-utility certification
heterogeneous degradation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Residual Policy Adaptation
Risk-Utility Certification
Structure-Adaptive Partitioning
Distribution Shift Robustness
Imperfect-Information Games
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Miaobo Hu
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China
S
Shuhao Hu
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
Xiaobo Guo
Xiaobo Guo
Dartmouth College
machine learningdeep learningnatural language processingsocia mediapropagantion
Xin Wang
Xin Wang
Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
Biomedical Engineering
Bokun Wang
Bokun Wang
Texas A&M University
Machine LearningArtificial IntelligenceMultimodal Machine Learning
Peng Zhang
Peng Zhang
Associate Professor, Nuclear Engineering and Radiological Sciences, University of Michigan
plasma physicscharged particle beamssurfaces and interfaces
D
Daren Zha
Institute of Information Engineering, Chinese Academy of Sciences, Beijing, China
J
Jun Xiao
School of Artificial Intelligence, University of Chinese Academy of Sciences, Beijing, China