🤖 AI Summary
This study addresses the lack of modular open-source tools and standardized benchmarks for systematically analyzing the bullwhip effect in multi-echelon supply chains under asymmetric costs and realistic demand patterns. To bridge this gap, we propose the first open-source Python platform integrating a scalable simulation engine with a registry-driven benchmarking framework. The platform supports plug-in modules for demand generation, ordering policies, and cost functions, while unifying six bullwhip metrics and incorporating real-world datasets (e.g., AR(1), WSTS). Leveraging abstract base classes and vectorized Monte Carlo simulation, it achieves high computational efficiency—executing 20.8 million simulations in just 7 seconds—and reveals a cumulative amplification factor of 427× in a four-tier semiconductor supply chain. Our analysis uncovers stochastic filtering upstream and demonstrates that no single metric suffices to evaluate policy performance comprehensively: the bullwhip intensity of the Order-Up-To policy varies by up to 155× between synthetic and real demand data.
📝 Abstract
The bullwhip effect remains operationally persistent despite decades of analytical research. Two computational deficiencies hinder progress: the absence of modular open-source simulation tools for multi-echelon inventory dynamics with asymmetric costs, and the lack of a standardized benchmarking protocol for comparing mitigation strategies across shared metrics and datasets. This paper introduces deepbullwhip, an open-source Python package that integrates a simulation engine for serial supply chains (with pluggable demand generators, ordering policies, and cost functions via abstract base classes, and a vectorized Monte Carlo engine achieving 50 to 90 times speedup) with a registry-based benchmarking framework shipping a curated catalog of ordering policies, forecasting methods, six bullwhip metrics, and demand datasets including WSTS semiconductor billings. Five sets of experiments on a four-echelon semiconductor chain demonstrate cumulative amplification of 427x (Monte Carlo mean across 1,000 paths), a stochastic filtering phenomenon at upstream tiers (CV = 0.01), super-exponential lead time sensitivity, and scalability to 20.8 million simulation cells in under 7 seconds. Benchmark experiments reveal a 155x disparity between synthetic AR(1) and real WSTS bullwhip severity under the Order-Up-To policy, and quantify the BWR-NSAmp tradeoff across ordering policies, demonstrating that no single metric captures policy quality.