asynchronous-pipeline

Tag

Cards List
#asynchronous-pipeline

One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining

Hugging Face Daily Papers · 2026-06-29 Cached

This paper challenges the assumption that one-step gradient delay in asynchronous pipeline parallelism is inherently unstable, showing that degradation depends on optimizer choice. It demonstrates that optimizers like Muon are robust to one-step delay and introduces an error-feedback correction to further mitigate staleness, achieving near-synchronous performance in LLM pretraining up to 10B parameters.

0 favorites 0 likes
#asynchronous-pipeline

Breaking the Bubble: Asynchronous Pipeline Parallel Training with Bounded Weight Inconsistency

Hugging Face Daily Papers · 2026-06-05 Cached

Introduces PACI, a bubble-free asynchronous pipeline parallel training method that bounds forward/backward weight inconsistency using local gradient accumulation, achieving higher throughput and faster time-to-accuracy without sacrificing stability or memory usage.

0 favorites 0 likes
← Back to home

Submit Feedback