Tag
Hugging Face Diffusers 上张量并行(tensor parallel)加载迎来重大优化:在 Flux.2-Dev DiT、TP=4(A10G)配置下,加载时间从 30.4s 降至 12.5s(约 2.4 倍提速),每 rank 峰值 CPU 内存从 64.1 GB 降至 6.8 GB(减少约 89%)。相关分布式推理(Accelerate 与 PyTorch Distributed)用法已更新到官方文档。
A user shares a workaround enabling vLLM tensor parallelism with tp=6 on six GPUs by padding model architecture dimensions with zeros until divisible, achieving higher KV cache utilization for Qwen 27B inference on Radeon 7900 XTX GPUs.
Tests of an updated DeepSeek V4.1 Flash model recipe on 4 DGX Sparks show significant performance improvements, with boosts in cold prefill, code generation, and text generation speeds.
This article is a handbook for the NVIDIA DGX Spark, a device designed for local AI inference, detailing its specifications, how to link multiple units for enhanced performance, and its optimization for running mixture-of-experts models.
A user successfully runs the Qwen 3.8 27B AI model on a mixed setup of RTX 3060 and 5060 Ti GPUs using tensor parallelism with exllamav3, achieving around 50 tokens per second with MTP enabled.
A developer forked NInfer, a C++20/CUDA inference engine, to add tensor-parallelism and YaRN rope scaling, enabling Qwen3.8-27B to run with a 1M token context on dual 5090 GPUs and outperforming vLLM in specific decode scenarios.
This paper compares tensor parallelism and KV-cache compression techniques for memory-bound LLM serving, finding that compression is generally cheaper and discusses decision rules based on model size relative to device memory.
FleetSieve introduces a decision-critical profiling method for SLO-aware LLM fleet configuration that optimizes resource allocation by reducing unnecessary measurements, achieving efficiency gains over uniform profiling.
A YouTube creator returns after a hiatus to teach building a distributed training framework from first principles, focusing on advanced AI topics like DeepSeek, MoE, and MLA, with an emphasis on developing problem-solving skills and self-confidence.
Enabling PCI-E peer-to-peer (P2P) for consumer Nvidia GPUs with patched drivers and vLLM environment variables yields roughly 25% prefill throughput improvement for free, as demonstrated by benchmarks.
Tests NVLink on dual RTX 3090s for AI inference and training, finding significant speedups for tensor parallel prompt processing (30%) and FSDP training (3x), but minimal effect on token generation or layer split inference.
The author shares six months of measurements on a four-RTX 3090 local setup, revealing that data parallelism often outperforms tensor parallelism for models fitting on fewer cards, with up to 3.4x throughput difference.
An analysis of PCIe transfer performance when running llama.cpp with dual GPUs using pipeline and tensor parallelism.
A guide on using Kaggle's free dual Tesla T4 GPUs (32GB VRAM) to run large LLMs with massive context windows, covering multi-GPU parallelism strategies in llama.cpp.
Detailed findings on PCIe bifurcation and P2P performance issues with 4x GPU setups, including workarounds and alternatives for tensor and pipeline parallelism.
This pull request makes tensor parallelism (TP) viable in llama.cpp when using the Vulkan backend, enabling distributed inference across multiple GPUs.
Summary of Lecture 19 on efficient AI distributed training, covering data, pipeline, tensor, and sequence parallelism methods with notes on memory and communication bottlenecks.
A user reports near-linear performance scaling when adding a second RTX 3090 for inference with a Qwen model, achieving roughly 1.8x decode TPS improvement without NVLink.
A tweet explains the correct answer to an ML performance interview question at Anthropic about the latency tradeoffs of splitting tensor-parallel linear layers by columns vs. rows when serving a 70B transformer model on 8 GPUs, highlighting that performance is not similar despite equal per-GPU weights.
llama.cpp maintainers and NVIDIA engineers collaborated to significantly improve multi-GPU performance in ggml, enabling hardware-agnostic tensor parallelism and major performance gains on RTX systems.