🤖 AI Summary
This study addresses the challenge of detecting highly elusive race conditions induced by asynchronous message passing in commercial network operating systems (NOS), for which existing detection techniques incur prohibitive overhead. To this end, this work proposes a two-phase dynamic analysis framework. The first phase leverages task locality to focus on critical components, thereby narrowing the search space and reducing analysis costs. The second phase integrates parallel grouping with an adaptive reinforcement strategy to proactively perturb message interleaving orders, efficiently exposing latent concurrency defects. The proposed approach has been deployed within an industrial-grade NOS testing platform, incurring less than 10% runtime system overhead. It successfully identified 21 critical bugs across five core components with a precision of 66%, substantially improving both the detection efficiency and coverage of concurrency defects.
📝 Abstract
Commercial on-device network operating systems (NOSes) run complex control planes in production routers and switches, where configuration update tasks are executed by multiple loosely coupled components through asynchronous message passing. Such executions are prone to race conditions: the same ordered input commands may produce different outcomes when messages are delivered in different orders. The race conditions are difficult to expose because they often manifest only as subtle, delayed malfunctions. Existing static analysis techniques lack scalability and precision for large-scale industrial NOSes, while dynamic ones incur substantial system-execution cost when attempting to cover the space of asynchronous message interleavings.
To address the challenges above, we present NosRacer, a dynamic analysis framework for race condition detection in industrial-grade on-device NOSes. NosRacer uses a two-phase design. The concentration phase reduces analysis scope by leveraging the locality of configuration update tasks. Race condition detection is then limited to a small subset of involved components and crucial state variables, thereby reducing the detection cost. The perturbation phase proactively perturbs message asynchrony to increase the likelihood of executions with race-condition-exposing message interleavings. It keeps perturbation practical by grouping compatible perturbations for parallel execution and adaptively strengthening perturbations.
We implement NosRacer in a commercial, actively developed NOS. NosRacer is integrated into the existing testing factory and used to detect race conditions in 5 key control-plane components. NosRacer achieves 66% precision and detects 21 race conditions confirmed by developers as severe bugs, while keeping the detection overhead below 10%.