Tag
Nereus is a cost-aware runtime that adapts reinforcement learning post-training jobs for large language models to efficient execution plans, reducing latency by 27.7% and improving throughput up to 7.27 times over existing tools.
DistribAI v2 is a platform for distributed training of AI models across multiple consumer devices, featuring improved hosting options and robust handling of edge cases.
The paper introduces a federated learning framework for Averaged n-Dependence Estimators (AnDE) classifiers, enhancing privacy by avoiding transmission of semantically meaningful parameters and demonstrating effectiveness on discrete datasets.
This article explores fault tolerance in distributed pre-training by simulating stage skipping in pipeline-parallel training, demonstrating that healthy workers can continue training with minimal validation loss impact when failures occur.
Michael Yu shares his detailed implementation and learnings from Stanford's CS336 course on building LLMs from scratch, covering tokenizers, transformers, training systems, and alignment techniques.
This paper introduces block parallelism and context-sharded block parallelism (CSBP) to efficiently train long-context diffusion language models, achieving significant throughput improvements and better performance on benchmarks like SWE-bench Verified.
The paper proposes WFM, a Wiki Foundation Model for scalable, agent-native knowledge representation and retrieval, demonstrating improved performance in complex agentic reasoning tasks with a 10.5x training acceleration.
The article promotes the Training Track at PyTorch Conference North America, featuring practical sessions on distributed training, optimization techniques, and foundation models to enhance AI system efficiency.
HyperParallel is a PyTorch-based distributed acceleration library optimized for Ascend SuperPoD, demonstrating higher training throughput than standard PyTorch FSDP2 with Muon while maintaining loss convergence, as showcased in a demo at #PyTorchCon China.
This paper proposes an approach to enhance distributed training over wide area networks by leveraging multicast technology and in-line FPGAs, along with an optimization framework for synchronization schedules, to maximize information exchange and reduce performance gaps.
vLLM introduces a native sharded weight transfer engine using Ray Direct Transport (RDT) for efficient weight syncing in large-scale online RL setups, enhancing performance for models like Kimi K2.
A YouTube creator returns after a hiatus to teach building a distributed training framework from first principles, focusing on advanced AI topics like DeepSeek, MoE, and MLA, with an emphasis on developing problem-solving skills and self-confidence.
This paper investigates federated learning for CNC tool wear prediction, showing that federated models achieve performance close to centralized learning and surpass local client baselines in distributed manufacturing environments.
Introduces Data-Centric Parallel (DCP), a method for training deep learning models on variable long sequences by dynamically adjusting runtime settings per batch, achieving up to 2.88x speedup on 32 H200 GPUs with only 10 lines of code integration.
Presents SNI-GNN, a SmartNIC-assisted full-graph GNN training system that reduces inter-node communication by predicting remote embeddings in-network, achieving 1.3–3.6x speedups with negligible accuracy loss.
Presents DG-FedReuse, a federated learning mechanism that reuses age-decayed cached client updates under a proxy-gradient threshold to reduce uplink communication, achieving significant modeled savings with minimal accuracy loss on image classification benchmarks.
This paper proposes FedSLM, a parameter-centric framework for federated fine-tuning of foundation models with heterogeneous compressed clients, using SVD-based decomposition and a weak-to-strong elicitation step to handle resource asymmetry. Experiments show it outperforms existing federated baselines while reducing client GPU memory by ~50%.
PyTorch 2.13 release brings FlexAttention to Apple Silicon with up to 12x speedups and 4x peak memory reduction for large-vocabulary models via fused LinearCrossEntropyLoss, alongside updates to distributed training, compilation, and on-device inference.
This paper presents SLAI T-Rex, a full-parameter post-training optimization framework for trillion-parameter MoE models on Ascend NPU SuperPOD, achieving 34.22% MFU and outperforming GPT-5.4-Mini on Operations Research tasks by 3.98 percentage points.
Proposes DiLoCo, a distributed optimization algorithm that enables training large language models across poorly connected devices with 500x less communication while matching fully synchronous performance.