@TheAhmadOsman: Laguna S 2.1 118B-A8B on DGX Station using - NVFP4 - FP8 KV Cache Can run 10 parallel agents > with 256k context each >…

X AI KOLs Timeline Models

Summary

Laguna S 2.1 118B-A8B model runs on DGX Station with NVFP4 and FP8 KV Cache, achieving ~1k tokens/second for 10 parallel agents with 256k context each.

Laguna S 2.1 118B-A8B on DGX Station using - NVFP4 - FP8 KV Cache Can run 10 parallel agents > with 256k context each > at ~1k tokens/second Some other numbers below https://t.co/Kb7RaTxBQz
Original Article
View Cached Full Text

Cached at: 07/25/26, 06:01 AM

Laguna S 2.1 118B-A8B on DGX Station using

  • NVFP4
  • FP8 KV Cache

Can run 10 parallel agents

> with 256k context each > at ~1k tokens/second

Some other numbers below https://t.co/Kb7RaTxBQz

Similar Articles