One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining

2026-06-29 • Machine Learning

Machine Learning

AI summaryⓘ

The authors study a method called asynchronous pipeline parallelism used to train large language models faster by keeping GPUs busy without waiting. They focus on a specific asynchronous schedule, PipeDream-2BW, which has a consistent one-step delay in updating gradients but was thought to perform poorly due to instability. Their work shows that this poor performance depends on the choice of optimizer: while AdamW struggles with delay, newer optimizers like Muon handle it well. They also introduce a correction technique to reduce delay problems, backed by theory and tested on large models, making asynchronous training more practical.

Pipeline ParallelismAsynchronous TrainingGradient StalenessPipeDream-2BWOptimizerAdamWMuon OptimizerError FeedbackConvergence AnalysisLarge Language Models

Authors

Philip Zmushko, Egor Petrov, Nursultan Abdullaev, Mikhail Khrushchev, Samuel Horváth

Abstract

Modern large-scale LLM pretraining benefits from utilizing Pipeline Parallelism; however, synchronous implementations leave GPUs idle during pipeline bubbles, wasting computational resources. Asynchronous Pipeline Parallelism eliminates these bubbles, maximizing throughput at the cost of gradient staleness. Among asynchronous schedules, PipeDream-2BW is particularly appealing: unlike the original PipeDream schedule, it ensures a constant one-step gradient delay regardless of pipeline depth. However, its adoption remains limited due to the common belief that optimizing under staleness is fundamentally unstable. In this work, we challenge this assumption, demonstrating that degradation under one-step delay depends strongly on optimizer choice rather than being an intrinsic limitation. We provide the first comprehensive empirical analysis showing that while AdamW, the predominant optimizer at the time when PipeDream-2BW was introduced, indeed suffers from severe degradation, recent methods like Muon exhibit strong robustness under a one-step delay. We introduce an optimizer-agnostic Error Feedback-inspired correction to further mitigate delay effects. We provide supporting theoretical analysis demonstrating convergence for Muon with and without this correction. Extensive evaluation on models up to 10B parameters confirms that our strategies bridge the performance gap with synchronous training, highlighting the practical potential of asynchronous pipeline parallelism at scale.

View PDFOpen arXiv