Ofan: Optimal Load Balancing for AI Training

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of logical hallucinations and high computational overhead in large language models during complex reasoning tasks. To this end, we propose a joint optimization framework based on dynamic sparse attention and chain-of-thought distillation. Specifically, the method employs an adaptive routing mechanism to identify critical reasoning paths and leverages multi-granularity knowledge distillation to transfer the reasoning capabilities of teacher models into lightweight architectures. Experimental results demonstrate that the proposed approach improves reasoning accuracy by 4.2% across mainstream benchmarks while reducing computational latency by 37%. This work provides a novel paradigm for constructing efficient and reliable reasoning systems.
📝 Abstract
The extreme collective completion time (CCT) demands of AI workloads challenge existing packet spraying algorithms, which can have trouble efficiently load-balancing workloads that are sent at full line rates. We trace this to a structural cause: on a fat tree, once a packet picks its upward path, the downward path to its destination is unique, so destination-oblivious schemes cannot undo the imbalance it creates. We prove that such schemes can suffer from $Θ(\sqrt{m})$ queueing for messages of size $m$, thus eventually triggering rate reductions by the congestion control. Instead, we suggest Ofan, a switch-based destination-aware LB scheme that can reach $O(1)$ queueing. We also present its pOfan variant that fits the pipe-based architecture of current switches. Our P4 implementation shows that it consumes modest resources. An end-to-end FSDP2 evaluation with Llama-3 405B-parameter models shows that Ofan cuts CCT inflation by $16$--$39\times$ when compared to existing algorithms.
Problem

Research questions and friction points this paper is trying to address.

load balancing
AI training
collective completion time
packet spraying
fat tree
Innovation

Methods, ideas, or system contributions that make the work stand out.

Load Balancing
Destination-aware
AI Training
Collective Completion Time
Fat-tree
🔎 Similar Papers
2024-06-07International Symposium on High-Performance Computer ArchitectureCitations: 5