From ASIC to Fleet: Lessons from Building and Operating a Hyperscaler NIC

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses cascading failures caused by shared firmware in commercial multi-host network interface cards (NICs) and the operational challenges of hyperscale deployments by proposing fbnic, a system-level solution. Architecturally, it introduces physical isolation and driver-priority mechanisms, combined with sub-sled-granularity firmware upgrade orchestration and slice-level fault containment. Operationally, it establishes a hardware-in-the-loop (HIL) continuous integration pipeline alongside cross-layer fault attribution monitoring and an automated remediation toolchain. Deployed across hundreds of thousands of hosts, fbnic reduces unplanned unavailability by 12×, shortens mean time to repair by 37%, and decreases hardware replacement rates by 2.3×, demonstrating robust stability at scale.
📝 Abstract
We describe the operational infrastructure built to deploy and operate fbnic, a custom multi-host NIC, across hundreds of thousands of production hosts at Meta. Vendor multi-host NICs, designed by retrofitting single-host architectures, suffered from shared firmware and buffers that created cascading isolation failures over seven years. fbnic eliminates these through physical isolation, but shifting to in-house hardware shifts the entire operational burden to the hyperscaler. We present a hardware-in-the-loop CI pipeline testing firmware, driver, and kernel cross-products; a unified observability pipeline co-locating NIC and switch counters for cross-layer fault attribution; a driver-first architecture with fewer than ten firmware message types; a targeted firmware upgrade orchestrator at sub-sled granularity; and scoped repair automation confining blast radius to individual host slices. Over ten months, fbnic achieved a 12X reduction in unplanned unavailability, 37% lower mean time to repair, and 2.3X fewer hardware swaps compared to vendor NICs on the same platform.
Problem

Research questions and friction points this paper is trying to address.

multi-host NIC
hyperscaler
isolation failures
operational infrastructure
custom hardware
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-host NIC
hardware-in-the-loop CI
cross-layer observability
driver-first architecture
scoped repair automation
🔎 Similar Papers
No similar papers found.
P
Prankur Gupta
Meta Platforms, Inc.
A
Alexander Duyck
Meta Platforms, Inc.
J
Jakub Kicinski
Meta Platforms, Inc.
J
Joseph Provine
Meta Platforms, Inc.
N
Neal Peacock
Meta Platforms, Inc.
P
Prabhakaran Ganesan
Meta Platforms, Inc.
R
Rajiv Krishnamurthy
Meta Platforms, Inc.
C
Chen Liu
Meta Platforms, Inc.
A
Akshay Viswakumar
Meta Platforms, Inc.
T
Timothy Vitkin
Meta Platforms, Inc.
J
Jie Meng
Meta Platforms, Inc.
B
Beatriz Padilla Hernandez
Meta Platforms, Inc.
M
Michael Edwards
Meta Platforms, Inc.
A
Andrei Kozlov
Meta Platforms, Inc.
V
Viren Nathan
Meta Platforms, Inc.
T
Tianyi Cui
Meta Platforms, Inc.
J
Joy Chaoyue Xiong
Meta Platforms, Inc.
R
Raul Hormazabal
Meta Platforms, Inc.
Mohsin Bashir
Mohsin Bashir
Meta Platforms, Inc.
Fred Feng
Fred Feng
University of Michigan-Dearborn
Behavioral Data AnalysisStatistical LearningHuman FactorsCycling SafetyPedestrian Safety
N
Nathan Walker
Meta Platforms, Inc.
L
Lavin Khandelwal
Meta Platforms, Inc.
M
Matt Maia
Meta Platforms, Inc.
M
Milo Piazza
Meta Platforms, Inc.
M
Mohanraj Thillainayagam
Meta Platforms, Inc.