distributed-training

Tag

Cards List
#distributed-training

Nereus: Adaptive Parallelism for LLM Post-Training

Hugging Face Daily Papers ↗ · 2d ago Cached

Nereus is a cost-aware runtime that adapts reinforcement learning post-training jobs for large language models to efficient execution plans, reducing latency by 27.7% and improving throughput up to 7.27 times over existing tools.

0 favorites 0 likes
#distributed-training

DistribAI v2

Reddit r/LocalLLaMA ↗ · 4d ago

DistribAI v2 is a platform for distributed training of AI models across multiple consumer devices, featuring improved hosting options and robust handling of edge cases.

0 favorites 0 likes
#distributed-training

Federated Learning of AnDE Classifiers

arXiv cs.LG ↗ · 5d ago Cached

The paper introduces a federated learning framework for Averaged n-Dependence Estimators (AnDE) classifiers, enhancing privacy by avoiding transmission of semantically meaningful parameters and demonstrating effectiveness on discrete datasets.

0 favorites 0 likes
#distributed-training

Simulating fault tolerance with stage skipping in pipeline-parallel training [R]

Reddit r/MachineLearning ↗ · 2026-09-22

This article explores fault tolerance in distributed pre-training by simulating stage skipping in pipeline-parallel training, demonstrating that healthy workers can continue training with minimal validation loss impact when failures occur.

0 favorites 0 likes
#distributed-training

@_michael_yu_: I think Stanford's CS336 is one of the best materials to learn LLMs from scratch. I just finished working through it (t…

X AI KOLs Timeline ↗ · 2026-09-22 Cached

Michael Yu shares his detailed implementation and learnings from Stanford's CS336 course on building LLMs from scratch, covering tokenizers, transformers, training systems, and alignment techniques.

0 favorites 0 likes
#distributed-training

Block Parallelism For Efficient Distributed Long-Context Diffusion Language Model Training

arXiv cs.LG ↗ · 2026-09-18 Cached

This paper introduces block parallelism and context-sharded block parallelism (CSBP) to efficiently train long-context diffusion language models, achieving significant throughput improvements and better performance on benchmarks like SWE-bench Verified.

0 favorites 0 likes
#distributed-training

WFM: Wiki Foundation Model for Complex Agentic Reasoning

arXiv cs.AI ↗ · 2026-09-17 Cached

The paper proposes WFM, a Wiki Foundation Model for scalable, agent-native knowledge representation and retrieval, demonstrating improved performance in complex agentic reasoning tasks with a 10.5x training acceleration.

0 favorites 0 likes
#distributed-training

@PyTorch: Train smarter. From distributed training to optimization techniques and foundation models, the Training Track at #PyTor…

X AI KOLs Following ↗ · 2026-09-11 Cached

The article promotes the Training Track at PyTorch Conference North America, featuring practical sessions on distributed training, optimization techniques, and foundation models to enhance AI system efficiency.

0 favorites 0 likes
#distributed-training

@PyTorch: Under identical configurations, HyperParallel FSDP2 + Muon delivers substantially higher training throughput than PyTor…

X AI KOLs Following ↗ · 2026-09-09

HyperParallel is a PyTorch-based distributed acceleration library optimized for Ascend SuperPoD, demonstrating higher training throughput than standard PyTorch FSDP2 with Muon while maintaining loss convergence, as showcased in a demo at #PyTorchCon China.

0 favorites 0 likes
#distributed-training

Distributed Training using an Intelligent Network

arXiv cs.LG ↗ · 2026-08-28 Cached

This paper proposes an approach to enhance distributed training over wide area networks by leveraging multicast technology and in-line FPGAs, along with an optimization framework for synchronization schedules, to maximize information exchange and reduce performance gaps.

0 favorites 0 likes
#distributed-training

@vllm_project: This weight transfer engine is native to vLLM. Any Ray-based trainer can adopt it with a single WeightSource iterator. …

X AI KOLs Following ↗ · 2026-08-26 Cached

vLLM introduces a native sharded weight transfer engine using Ray Direct Transport (RDT) for efficient weight syncing in large-scale online RL setups, enhancing performance for models like Kimi K2.

0 favorites 0 likes
#distributed-training

The GOAT of local LLM youtube is back

Reddit r/LocalLLaMA ↗ · 2026-08-18 Cached

A YouTube creator returns after a hiatus to teach building a distributed training framework from first principles, focusing on advanced AI topics like DeepSeek, MoE, and MLA, with an emphasis on developing problem-solving skills and self-confidence.

0 favorites 0 likes
#distributed-training

Federated Learning for Distributed CNC Tool Wear Prediction

arXiv cs.LG ↗ · 2026-08-13 Cached

This paper investigates federated learning for CNC tool wear prediction, showing that federated models achieve performance close to centralized learning and surpass local client baselines in distributed manufacturing environments.

0 favorites 0 likes
#distributed-training

Training Variable Long Sequences with Data-Centric Parallel

arXiv cs.AI ↗ · 2026-08-11 Cached

Introduces Data-Centric Parallel (DCP), a method for training deep learning models on variable long sequences by dynamically adjusting runtime settings per batch, achieving up to 2.88x speedup on 32 H200 GPUs with only 10 lines of code integration.

0 favorites 0 likes
#distributed-training

SNI-GNN: SmartNIC-Assisted Full-Graph GNN Training with In-Network Embedding Prediction

arXiv cs.LG ↗ · 2026-08-10 Cached

Presents SNI-GNN, a SmartNIC-assisted full-graph GNN training system that reduces inter-node communication by predicting remote embeddings in-network, achieving 1.3–3.6x speedups with negligible accuracy loss.

0 favorites 0 likes
#distributed-training

DG-FedReuse: Proxy-Gradient-Gated Cached-Update Reuse with Matched Sparse Uplink Accounting

arXiv cs.LG ↗ · 2026-08-07 Cached

Presents DG-FedReuse, a federated learning mechanism that reuses age-decayed cached client updates under a proxy-gradient threshold to reduce uplink communication, achieving significant modeled savings with minimal accuracy loss on image classification benchmarks.

0 favorites 0 likes
#distributed-training

Federated Foundation Models Fine-Tuning with Heterogeneous Compressed Clients

arXiv cs.LG ↗ · 2026-08-03 Cached

This paper proposes FedSLM, a parameter-centric framework for federated fine-tuning of foundation models with heterogeneous compressed clients, using SVD-based decomposition and a weak-to-strong elicitation step to handle resource asymmetry. Experiments show it outperforms existing federated baselines while reducing client GPU memory by ~50%.

0 favorites 0 likes
#distributed-training

@PyTorch: PyTorch 2.13 brings FlexAttention to Apple Silicon, cuts peak memory by up to 4× for large-vocabulary models with nn.Li…

X AI KOLs Following ↗ · 2026-07-31 Cached

PyTorch 2.13 release brings FlexAttention to Apple Silicon with up to 12x speedups and 4x peak memory reduction for large-vocabulary models via fused LinearCrossEntropyLoss, alongside updates to distributed training, compilation, and on-device inference.

0 favorites 0 likes
#distributed-training

SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD

Hugging Face Daily Papers ↗ · 2026-07-22 Cached

This paper presents SLAI T-Rex, a full-parameter post-training optimization framework for trillion-parameter MoE models on Ascend NPU SuperPOD, achieving 34.22% MFU and outperforming GPT-5.4-Mini on Operations Research tasks by 3.98 percentage points.

0 favorites 0 likes
#distributed-training

DiLoCo: Distributed Low-Communication Training of Language Models

Reddit r/singularity ↗ · 2026-07-21 Cached

Proposes DiLoCo, a distributed optimization algorithm that enables training large language models across poorly connected devices with 500x less communication while matching fully synchronous performance.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback