🤖 AI Summary
To address the degradation of model robustness post-deployment caused by hardware/software environment shifts, this paper introduces Prom, an open-source framework that pioneers dynamic misprediction detection and lightweight feedback-driven adaptive repair at deployment time. The method integrates statistical significance testing, uncertainty quantification, and confidence calibration—enabling accuracy recovery without full retraining. Instead, it leverages an online feedback loop to incrementally annotate and learn from ≤5% of samples. Evaluated across 13 models and five code analysis and optimization tasks, Prom achieves an average misprediction identification rate of 96% (up to 100%), significantly enhancing cross-platform generalization and robustness against diverse hardware configurations and code patterns.
📝 Abstract
Supervised machine learning techniques have shown promising results in code analysis and optimization problems. However, a learning-based solution can be brittle because minor changes in hardware or application workloads -- such as facing a new CPU architecture or code pattern -- may jeopardize decision accuracy, ultimately undermining model robustness. We introduce Prom, an open-source library to enhance the robustness and performance of predictive models against such changes during deployment. Prom achieves this by using statistical assessments to identify test samples prone to mispredictions and using feedback on these samples to improve a deployed model. We showcase Prom by applying it to 13 representative machine learning models across 5 code analysis and optimization tasks. Our extensive evaluation demonstrates that Prom can successfully identify an average of 96% (up to 100%) of mispredictions. By relabeling up to 5% of the Prom-identified samples through incremental learning, Prom can help a deployed model achieve a performance comparable to that attained during its model training phase.