tensor-parallelism

Tag

Cards List
#tensor-parallelism

Accelerating Long-Tail Generation in Synchronous RLHF Training via Adaptive Tensor Parallelism

arXiv cs.AI ↗ · 2026-05-26 Cached

This paper proposes PAT, an adaptive tensor parallelism method that dynamically reconfigures TP during the generation stage of synchronous RLHF training to mitigate long-tail generation bottlenecks. Evaluations on LLaMA3.1-8B and Qwen3-14B show reductions in generation latency by up to 34.6% and end-to-end iteration latency by up to 27.2%.

0 favorites 0 likes
#tensor-parallelism

@Hikari_07_jp: Local LLM is incredibly complex. Hardware selection, quantization, harnesses, engines, tensor parallelism, unmodified m…

X AI KOLs Timeline ↗ · 2026-05-22 Cached

A user reflects on the complexity and fascination of running local LLMs, touching on hardware selection, quantization, and tensor parallelism.

0 favorites 0 likes
#tensor-parallelism

@levidiamode: Day 138/365 of GPU Programming One of my favorite lectures I've watched this year is Stanford's CS336 lecture 7 on GPU …

X AI KOLs Timeline ↗ · 2026-05-21 Cached

A learner shares enthusiasm for Stanford CS336 lecture 7 on GPU parallelism, which covers fundamental operations and connects them to multi-GPU setups and parallelism techniques like tensor, data, and pipeline parallelism.

0 favorites 0 likes
#tensor-parallelism

Dual GPU llama.cpp speedup

Reddit r/LocalLLaMA ↗ · 2026-05-17

A fork of llama.cpp fixes the --split-mode tensor issue with quantized KV caches, achieving up to 40% speed improvement on dual GPU setups without quality loss.

0 favorites 0 likes
#tensor-parallelism

@PyTorch: At #PyTorchCon Europe 2026, @ezyang (@Meta) explains why many developers find tensor parallelism difficult to work with…

X AI KOLs Following ↗ · 2026-05-14 Cached

At PyTorchCon Europe 2026, Edward Yang explains PyTorch's new pre-compilation support for distributed training and SPMD type system to help developers write correct tensor parallelism code, addressing common pitfalls in gradient correctness.

0 favorites 0 likes
#tensor-parallelism

NCCL-Free Tensor Parallelism on Dual Blackwell PCIe llama.cpp b9095 released!

Reddit r/LocalLLaMA ↗ · 2026-05-10

llama.cpp build b9095 introduces NCCL-free tensor parallelism for dual Blackwell PCIe GPUs, enabling efficient multi-GPU inference without relying on NCCL.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback