When Offline Selectors Cannot Beat the Best Single Model: A Diagnostic Study on edX Dropout Prediction

📅 2026-06-02
📈 Citations: 0
Influential: 0
📄 PDF

career value

173K/year
🤖 AI Summary
Offline model selectors often fail to surpass the strongest individual model, yet the underlying reasons remain unclear. This work proposes a three-stage diagnostic framework: first estimating a local performance upper bound via k-nearest-neighbor consistency, then evaluating whether behavior cloning and offline reinforcement learning methods (DQN, CQL) approach this bound, and finally assessing feature expressiveness through state ablation. The framework uniquely attributes selection failure to three diagnosable factors—learner mismatch, distributional shift, or representational ambiguity—and provides a unified experimental pipeline to pinpoint bottlenecks. In an edX dropout prediction task, although an ideal selector could improve accuracy by 9.7%, all offline methods underperformed; diagnostics revealed that ambiguous state representations—not algorithmic limitations or distributional issues—were the primary cause.
📝 Abstract
Different predictors often excel on different inputs, so picking the best one per instance promises higher accuracy than committing to a single model. In practice, selectors trained from logged data routinely fail to beat the strongest single predictor. Three causes typically go unseparated before more tuning is applied: a mismatched learner, a state that does not predict which model wins, or buffer-to-deployment label shift. A three-stage diagnostic rules them out on a shared buffer. Stage~1 estimates a local ceiling on oracle recovery from $k$-NN label consistency. Stage~2 asks whether paired BC and offline-RL learners (BC, DQN, and CQL across penalty weights) reach that ceiling. Stage~3 ablates the selector state to test whether richer features would raise it. The combined verdict points to the most promising next step: tuning the learner, redesigning the state, or collecting new data. We apply it to selecting among five dropout-prediction models on edX clickstream data. Across 16 windows, the oracle beats the strongest single base model by 9.7 accuracy points on average, yet BC, DQN, and CQL land in the same test-accuracy band below it (robust to a tenfold buffer sweep and $N{=}2{,}000$ held-out examples). The bottleneck is local representational ambiguity: CQL closes the imitation gap without a deployment gain (not conservatism), regret clusters tightly across learners (not tie-breaking), and the three learners converge on test accuracy (not shift). The next iteration should change the state or collect new data, not tune the offline learner further.
Problem

Research questions and friction points this paper is trying to address.

offline learning
model selection
dropout prediction
label shift
representational ambiguity
Innovation

Methods, ideas, or system contributions that make the work stand out.

offline model selection
diagnostic framework
label shift
representational ambiguity
conservative Q-learning
🔎 Similar Papers
No similar papers found.