Tag
This article explores fault tolerance in distributed pre-training by simulating stage skipping in pipeline-parallel training, demonstrating that healthy workers can continue training with minimal validation loss impact when failures occur.
This paper presents a system to accelerate speculative decoding in large-scale RL post-training through online draft co-training, using context-parallel attention extensions and cross-stage feature transport.