Tag
The author shares experiences and insights from nnScaler to large-scale distributed training systems, discussing correctness, flexibility, boundary expansion, and the challenges brought by post-training and reinforcement learning.
LLMSys-PaperList is a curated reading list on GitHub that organizes LLM systems research papers and resources into practical categories such as training systems, serving systems, and multi-modal coverage, helping AI/ML engineers and researchers stay updated.