From Valid to Useful: Post-Verification Acquisition for Recursive Self-Improving Recommendation

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the training bias arising from uneven source contributions of synthetic sequences in recursive self-improving recommendation (RSIR). We propose the DA-RSIR framework, which establishes "post-validation acquisition" as an independent control point to decouple validation from selection without requiring additional labels or teacher models. Furthermore, it introduces Bayesian Active Learning by Disagreement (BALD) scoring with Monte Carlo Dropout estimation to enable unsupervised sample importance evaluation based on predictive disagreement, thereby optimizing training data selection for subsequent iterations. Experiments across four datasets demonstrate that our method comprehensively outperforms the retain-all strategy, achieving single-round performance that surpasses five-round cumulative gains with high statistical significance.
📝 Abstract
Sequential recommenders can generate synthetic interaction sequences and retrain on the augmented corpus in a recursive self-improvement loop. To limit error accumulation, current methods verify each generated sequence remains predictive of the user's real interactions and discard those that drift away from it. Verification does not, however, determine which verified sequences should train the next model. With every verified sequence used for training, source sequences yielding more verified sequences or longer continuations have more influence, although neither quantity indicates how much those sequences will help the next model. We formulate the decision of which verified sequences are used to train the next model as \emph{post-verification acquisition} and introduce {\bf Disagreement-Aware Recursive Self-Improving Recommendation (DA-RSIR)}. DA-RSIR caps each source sequence's contribution and ranks its verified sequences by how much the model's predictions disagree over their augmented interactions. It uses a score derived from Bayesian Active Learning by Disagreement (BALD) and estimated with Monte Carlo (MC) dropout. DA-RSIR requires no extra labels, teacher model, or quality scorer. Across four datasets, three recommender models, and two metrics, it improves on the retain-all approach in all $24$ comparisons and attains the highest mean in $23$ of $24$ overall; the aggregate improvement is statistically significant on both metrics. A single DA-RSIR round exceeds the retain-all approach's best gain over five recursive rounds. These findings establish post-verification acquisition as a separate control point in recursive self-improvement, separating which sequences pass verification from which verified sequences are used to train the next model.
Problem

Research questions and friction points this paper is trying to address.

Sequential Recommendation
Recursive Self-Improvement
Post-Verification Acquisition
Synthetic Interaction Sequences
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recursive Self-Improving Recommendation
Post-Verification Acquisition
Bayesian Active Learning by Disagreement (BALD)
Monte Carlo Dropout
Sequential Recommendation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.