v100

Tag

Cards List
#v100

NVFP4 on VOLTA! Despite being built for Blackwell, I made four 2017 V100s run Qwen 3.8 NVFP4 natively and match my $6000 RTX 5090.

Reddit r/LocalLLaMA · 2d ago

The author achieved performance parity between four 2017 Tesla V100 GPUs and a modern RTX 5090 when running the Qwen 3.8 model with NVFP4 precision, using custom software optimizations like the QPN kernel for efficient inference.

0 favorites 0 likes
#v100

366 t/s Qwen3.6 27B NVFP4 on v100s

Reddit r/LocalLLaMA · 2026-08-11

Introduces v100-skinny, a custom kernel library enabling fast NVFP4 inference on V100 (sm70) GPUs, achieving 366 t/s for Qwen3.6 27B in best-case extraction, with lower speeds for structured generation and code.

0 favorites 0 likes
#v100

For V100 Users: SGLang running Qwen+Dflash and Laguna

Reddit r/LocalLLaMA · 2026-07-27

Modified SGLang to support Qwen and Laguna models on V100 GPUs using custom FlashAttention and Marlin kernels, achieving decent throughput on 4xV100 hardware.

0 favorites 0 likes
#v100

Chinese Hackers Latest Masterpiece with NVIDIA

Reddit r/LocalLLaMA · 2026-06-22 Cached

A DIY enthusiast created a single-slot, low-profile NVIDIA V100 GPU, showcasing custom hardware modding in the PC building community.

0 favorites 0 likes
#v100

Cheapest hardware for Qwen 3.6: both 27B and 35B-A3B

Reddit r/LocalLLaMA · 2026-06-15

Discusses the cheapest hardware options for running Qwen 3.6 models, comparing RTX 3090 and Tesla V100 GPUs, and provides a detailed cost breakdown for a system at around $2000.

0 favorites 0 likes
#v100

Cheap V100 32gb

Reddit r/LocalLLaMA · 2026-06-01

A deal for a used V100 32GB GPU on Aliexpress at approximately $526, including coupon codes.

0 favorites 0 likes
#v100

I Put a Datacenter GPU in My Gaming PC for £200

Lobsters Hottest · 2026-05-31 Cached

A blogger describes how they acquired a Tesla V100 SXM2 datacenter GPU for £150 and used a custom adapter to install it in their gaming PC alongside an RTX 4080, achieving 32GB of total VRAM and enabling local inference of 27B parameter models at 32 tokens per second.

0 favorites 0 likes
#v100

Anyone using Flash Attention 2 (ai-bond) on their V100's? How is the performance?

Reddit r/LocalLLaMA · 2026-05-29

A user benchmarks a V100-compatible port of Flash Attention 2, reporting 3x-17x speedups and up to 94% memory reduction over default PyTorch attention.

0 favorites 0 likes
#v100

1000 tps generation on Qwen3.6 27B with V100s

Reddit r/LocalLLaMA · 2026-05-25

Achieved 1000 tokens per second generation on Qwen3.6 27B using V100 GPUs with 128 concurrent requests, and 80 t/s for single user.

0 favorites 0 likes
← Back to home

Submit Feedback