🤖 AI Summary
In the Cross-Entropy Method (CEM), scoring errors of world models are coupled across action selection and candidate distribution updates, limiting search performance. This work proposes decoupling the model’s dual roles as “scorer” and “proposer,” analyzing their impact on the search process via a cross-rescoring mechanism. By integrating multiple predictive models for comparison and constructing environment-referenced elite sets, this study is the first to quantify error propagation across both stages. Our findings reveal that the proposal distribution contracts progressively over iterations, and intervening in the initial update step significantly reduces the final execution cost. These insights offer a novel perspective for enhancing the efficiency of model-based planning.
📝 Abstract
The cross-entropy method (CEM) uses world-model scores to select action sequences and fit the distribution sampled in its next iteration. A scoring error can therefore change both the present decision and the candidates considered later. We evaluate these two roles separately. Four types of predictive model generate CEM traces, and every model rescores every saved candidate pool. Executing the same candidates in the environment provides a reference elite set and proposal update.
Across twelve independently trained task-seed units on Walker and Cheetah, the pre-specified proposal distance falls from the first to the final CEM iteration in every unit. Proposal widths contract and fitted means separate relative to the remaining search width. Pairwise ranking agreement stays near chance on Walker and declines on Cheetah; elite-set agreement does not improve. This comparison shows greater variation between scorers than between pool sources on Cheetah; Walker has variation in both and in their pairings. We use the original six units to select Random nonlinear for a one-update intervention, without inspecting intervention outcomes. Replacing its first model-ranked update with an environment-ranked update lowers final realised selected-sequence cost in those six units and in six further units held out from the selection.