FlexiFlow: Bandit-based Model Switching in ML Workflows

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge that single models often fail to maximize accuracy in machine learning workflows, while existing systems lack adaptive switching capabilities. To overcome these limitations, this work proposes a dynamic data stream system based on multi-armed bandits. The core innovation lies in a novel bandit strategy that integrates runtime overhead with assertion probabilities, effectively mitigating the shortcomings of standard Thompson sampling and enabling intelligent model switching at runtime. Experimental results demonstrate that the proposed system improves accuracy by 23% over baseline methods and achieves a 48% efficiency gain compared to sequential execution.
📝 Abstract
Model optimizations help improve inference performance and accuracy of ML workflows. However, relying on a single model to perform inference across all data batches often fails to maximize accuracy and thus overall performance. In many cases, alternate models could perform better on specific subsets of data where a primary model underperforms. Our experiments with real ML workflows indeed show that switching models improves workflow accuracy by up to 23%. Yet, current systems lack the ability to adaptively switch between models based on performance, forcing users to manually test models in sequence. We present FlexiFlow, a dataflow system that dynamically switches between alternate models when the current model exhibits low accuracy. FlexiFlow learns to rank models using a novel multi-armed bandit approach that accounts for model runtimes, probability of passing user-defined assertions, and the computational structure of the ML workflow. We show that the standard Thompson sampling approach is insufficient for switching models in ML workflows. In contrast, our proposed approaches are effective and scales to complex real-world ML workflows. Experiments show that switching models at runtime while reusing intermediate results provides higher accuracy, but also 48% efficiency gain compared to sequential workflow runs.
Problem

Research questions and friction points this paper is trying to address.

Model Switching
ML Workflows
Inference Accuracy
Adaptive Switching
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-armed Bandit
Model Switching
Dataflow System
Thompson Sampling
Intermediate Result Reuse
🔎 Similar Papers
No similar papers found.