Correcting Within-Group Self-Selection Bias in Prioritized Replay

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究解决了优先经验回放中的组内自选择偏差问题,通过SAMPLE、AVG和MODEL三种方法调整重放缓冲区,以提高学习效率。
📝 Abstract
Prioritized experience replay (PER) improves sample efficiency by replaying high-priority transitions, usually according to absolute temporal-difference error. In stochastic environments, PER can distort the distribution of realized outcomes replayed from transitions with the same state-action pair. We call this within-group self-selection. We quantify the resulting changes in within-group outcome frequencies and mean Bellman targets. We decompose PER into between-group allocation and conditional sibling selection, and derive fixed-buffer corrections that preserve current group-level priority mass: SAMPLE selects a group through PER and trains on a uniformly sampled sibling; AVG averages sibling Bellman targets; and MODEL samples from an empirical full-outcome model. In exact state-action environments with rare high-magnitude outcomes, sibling-aware replay improves learning efficiency over PER, although matched parameter sweeps show that tuning can narrow some gaps. In MinAtar, approximate VQ-VAE groups with SAMPLE mitigate degradation under mean-preserving reward tails in four of five games. Sibling-aware replay thus retains the focus on high-priority state-action regions while recovering their empirical outcome frequencies.
Problem

Research questions and friction points this paper is trying to address.

Prioritized Experience Replay
Within-Group Self-Selection
Sample Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prioritized Experience Replay
Within-Group Self-Selection Bias
Fixed-Buffer Corrections
Sibling-Aware Replay
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.