tensor-parallelism

Tag

Cards List
#tensor-parallelism

tp=6 can work on vLLM, with padding

Reddit r/LocalLLaMA ↗ · yesterday

A user shares a workaround enabling vLLM tensor parallelism with tp=6 on six GPUs by padding model architecture dimensions with zeros until divisible, achieving higher KV cache utilization for Qwen 27B inference on Radeon 7900 XTX GPUs.

0 favorites 0 likes
#tensor-parallelism

@YRSM_Simon: What kind of black magic is this! Tested @MiaAI_lab's updated DeepSeek V4.1 Flash recipe, 4 DGX Sparks, TP4, the improv…

X AI KOLs Timeline ↗ · 6d ago Cached

Tests of an updated DeepSeek V4.1 Flash model recipe on 4 DGX Sparks show significant performance improvements, with boosts in cold prefill, code generation, and text generation speeds.

0 favorites 0 likes
#tensor-parallelism

@exolabs: https://x.com/exolabs/status/2103617535765573959

X AI KOLs Timeline ↗ · 6d ago Cached

This article is a handbook for the NVIDIA DGX Spark, a device designed for local AI inference, detailing its specifications, how to link multiple units for enhanced performance, and its optimization for running mixture-of-experts models.

0 favorites 0 likes
#tensor-parallelism

Success running Qwen 3.8 27B EXL3 on RTX 3060 + 5060 Ti

Reddit r/LocalLLaMA ↗ · 2026-09-20

A user successfully runs the Qwen 3.8 27B AI model on a mixed setup of RTX 3060 and 5060 Ti GPUs using tensor parallelism with exllamav3, achieving around 50 tokens per second with MTP enabled.

0 favorites 0 likes
#tensor-parallelism

(NInfer Fork) I wanted to have a 1M context Qwen-3.8 27B, tp2, dual 5090s

Reddit r/LocalLLaMA ↗ · 2026-08-29

A developer forked NInfer, a C++20/CUDA inference engine, to add tensor-parallelism and YaRN rope scaling, enabling Qwen3.8-27B to run with a 1M token context on dual 5090 GPUs and outperforming vLLM in specific decode scenarios.

0 favorites 0 likes
#tensor-parallelism

More GPUs or a Smaller Cache? Tensor Parallelism versus KV Compression for Memory-Bound LLM Serving

arXiv cs.AI ↗ · 2026-08-26 Cached

This paper compares tensor parallelism and KV-cache compression techniques for memory-bound LLM serving, finding that compression is generally cheaper and discusses decision rules based on model size relative to device memory.

0 favorites 0 likes
#tensor-parallelism

FleetSieve: Decision-Critical Profiling for SLO-Aware LLM Fleet Configuration

arXiv cs.LG ↗ · 2026-08-21 Cached

FleetSieve introduces a decision-critical profiling method for SLO-aware LLM fleet configuration that optimizes resource allocation by reducing unnecessary measurements, achieving efficiency gains over uniform profiling.

0 favorites 0 likes
#tensor-parallelism

The GOAT of local LLM youtube is back

Reddit r/LocalLLaMA ↗ · 2026-08-18 Cached

A YouTube creator returns after a hiatus to teach building a distributed training framework from first principles, focusing on advanced AI topics like DeepSeek, MoE, and MLA, with an emphasis on developing problem-solving skills and self-confidence.

0 favorites 0 likes
#tensor-parallelism

enabling PCI-E p2p for consumer Nvidia cards will yield you more than you think

Reddit r/LocalLLaMA ↗ · 2026-08-08

Enabling PCI-E peer-to-peer (P2P) for consumer Nvidia GPUs with patched drivers and vLLM environment variables yields roughly 25% prefill throughput improvement for free, as demonstrated by benchmarks.

0 favorites 0 likes
#tensor-parallelism

When Is NVLink Worth It?

Hacker News Top ↗ · 2026-07-22 Cached

Tests NVLink on dual RTX 3090s for AI inference and training, finding significant speedups for tensor parallel prompt processing (30%) and FSDP training (3x), but minimal effect on token generation or layer split inference.

0 favorites 0 likes
#tensor-parallelism

@superalesha: https://x.com/superalesha/status/2077437741915312221

X AI KOLs Timeline ↗ · 2026-07-15 Cached

The author shares six months of measurements on a four-RTX 3090 local setup, revealing that data parallelism often outperforms tensor parallelism for models fitting on fewer cards, with up to 3.4x throughput difference.

0 favorites 0 likes
#tensor-parallelism

Measuring PCIe transfer under dual GPU with pipeline & tensor llama.cpp

Reddit r/LocalLLaMA ↗ · 2026-07-11

An analysis of PCIe transfer performance when running llama.cpp with dual GPUs using pipeline and tensor parallelism.

0 favorites 0 likes
#tensor-parallelism

@analogalok: I can't afford a $2,000 GPU is officially a dead excuse. yesterday I showed you how to unlock an enterprise grade 16GB …

X AI KOLs Timeline ↗ · 2026-07-06 Cached

A guide on using Kaggle's free dual Tesla T4 GPUs (32GB VRAM) to run large LLMs with massive context windows, covering multi-GPU parallelism strategies in llama.cpp.

0 favorites 0 likes
#tensor-parallelism

Findings from troubleshooting p2p on 4x5060 ti bifurcation.

Reddit r/LocalLLaMA ↗ · 2026-06-27

Detailed findings on PCIe bifurcation and P2P performance issues with 4x GPU setups, including workarounds and alternatives for tensor and pipeline parallelism.

0 favorites 0 likes
#tensor-parallelism

vulkan: make TP viable by pwilkin · Pull Request #25051 · ggml-org/llama.cpp

Reddit r/LocalLLaMA ↗ · 2026-06-26 Cached

This pull request makes tensor parallelism (TP) viable in llama.cpp when using the Vulkan backend, enabling distributed inference across multiple GPUs.

0 favorites 0 likes
#tensor-parallelism

@ickma2311: Efficient AI Lecture 19: Distributed Training (Part 1) This lecture gave me a much clearer picture of how self-attentio…

X AI KOLs Timeline ↗ · 2026-06-10 Cached

Summary of Lecture 19 on efficient AI distributed training, covering data, pipeline, tensor, and sequence parallelism methods with notes on memory and communication bottlenecks.

0 favorites 0 likes
#tensor-parallelism

Weird to get near linear scaling by adding another GPU?

Reddit r/LocalLLaMA ↗ · 2026-06-08

A user reports near-linear performance scaling when adding a second RTX 3090 for inference with a Qwen model, achieving roughly 1.8x decode TPS improvement without NVLink.

0 favorites 0 likes
#tensor-parallelism

@gpusteve: you're interviewing for an ml performance role at anthropic and they ask: "you're serving a 70b transformer model on 8 …

X AI KOLs Timeline ↗ · 2026-06-08 Cached

A tweet explains the correct answer to an ML performance interview question at Anthropic about the latency tradeoffs of splitting tensor-parallel linear layers by columns vs. rows when serving a 70B transformer model on 8 GPUs, highlighting that performance is not similar despite equal per-GPU weights.

0 favorites 0 likes
#tensor-parallelism

@ggerganov: Highlighting recent advances in multi-GPU and tensor parallel support in llama.cpp Over the last few months llama.cpp m…

X AI KOLs Following ↗ · 2026-06-04

llama.cpp maintainers and NVIDIA engineers collaborated to significantly improve multi-GPU performance in ggml, enabling hardware-agnostic tensor parallelism and major performance gains on RTX systems.

0 favorites 0 likes
#tensor-parallelism

Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism

arXiv cs.AI ↗ · 2026-05-26 Cached

This paper proposes PAT, an adaptive tensor parallelism method that dynamically reconfigures TP during the generation stage of synchronous RLHF training to mitigate long-tail generation bottlenecks. Evaluations on LLaMA3.1-8B and Qwen3-14B show reductions in generation latency by up to 34.6% and end-to-end iteration latency by up to 27.2%.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback