@TheAhmadOsman: Laguna S 2.1 118B-A8B on DGX Station using - NVFP4 - FP8 KV Cache Can run 10 parallel agents > with 256k context each >…
Summary
Laguna S 2.1 118B-A8B model runs on DGX Station with NVFP4 and FP8 KV Cache, achieving ~1k tokens/second for 10 parallel agents with 256k context each.
View Cached Full Text
Cached at: 07/25/26, 06:01 AM
Laguna S 2.1 118B-A8B on DGX Station using
- NVFP4
- FP8 KV Cache
Can run 10 parallel agents
> with 256k context each > at ~1k tokens/second
Some other numbers below https://t.co/Kb7RaTxBQz
Similar Articles
Laguna-S-2.1 runs on my 6 years old gaming PC!
Laguna-S-2.1, a quantized AI model, runs on a 6-year-old gaming PC with an RTX 3080, achieving 10 t/s decode and using 8.3 GB VRAM and 52.2 GB host RAM.
@0xSero: GLM-5.1-478B-NVFP4 Running on: - 4x RTX Pro 6000 - Sglang - 370,000 max tokens (1.75x full context) - p10 27.7 | p90 45…
A quantized 478B-parameter GLM-5.1 model runs on 4×RTX Pro 6000 GPUs via SGLang, delivering 370k-token context at up to 45 tok/s decode and 1340 tok/s prefill, and is demoed driving Figma.
@onusoz: 16x parallel Gemma-4-26B-A4B-NVFP4 runs 18 output tokens/s, aggregate 300 tok/s 1 DGX Spark with 128 GB unified memo…
@onusoz demonstrates running 16 parallel instances of NVIDIA's quantized Gemma-4-26B-A4B-NVFP4 model on a single DGX Spark with 128GB unified memory, achieving 300 tok/s aggregate, showcasing high concurrency without flashinfer.
If you use Open Code or other agenting programs you are leaving a lot of t/s if you don't actually use agents in parallel. Benchmark : RTX5090, Qwen3.6 35B loaded via LM studio with parallel tasks set to 8
Benchmark shows that running 4-5 parallel agents with LM Studio on RTX 5090 maximizes throughput, while more agents yield diminishing returns due to VRAM and compute splitting.
NVFP4 kv cache quantization on sm120 will make 32GB VRAM systems very capable
NVFP4 KV cache quantization on sm120 significantly improves memory efficiency for large language models, enabling 32GB VRAM systems to achieve ~60 tok/sec inference at 196k context size with Qwen3.6-27B.