dgx-spark

Tag

Cards List
#dgx-spark

A 124B emitted 15,128 tokens in a single response on one DGX Spark, decode went 35.62 → 35.68 tok/s across the whole thing

Reddit r/LocalLLaMA · yesterday

An observation of Ling-3.0-flash (124B) on one DGX Spark generating 15,128 tokens in a single response with stable decode throughput around 35.6 tok/s, highlighting long-context decoding performance.

0 favorites 0 likes
#dgx-spark

Benched a 124B on one DGX Spark for a week and published all of it — 38.7 tok/s on the fastest path he found, 2.4x DeepSeek V4 Flash on the same box

Reddit r/LocalLLaMA · 2d ago

An independent benchmark by sudoingX shows the Ling-3.0-flash model runs at 38.7 tok/s on a single DGX Spark with official INT4 quantization, 2.4x faster than DeepSeek V4 Flash on the same hardware, after a correction clarifying the quants do work.

0 favorites 0 likes
#dgx-spark

I ran Muse Glimmer @ 1M context - All tests passed.

Reddit r/LocalLLaMA · 4d ago

User tests Meta's Muse Glimmer 30B on a 2× DGX Spark cluster, extending context from 131K to 1M tokens with YaRN and confirming passing retrieval at 832K tokens. Reports ~3× speedup from DFlash speculative decoding and shares full config.

0 favorites 0 likes
#dgx-spark

Two flags took the official Ling-3.0-flash INT4 from 20.8 to 38.7 tok/s on one DGX Spark

Reddit r/LocalLLaMA · 5d ago

This post describes two configuration flags that increase the official Ling-3.0-flash INT4 inference speed from 20.8 to 38.7 tok/s on a single DGX Spark, while warning about the need for a specific vLLM fork and noting tradeoffs with long-context performance.

0 favorites 0 likes
#dgx-spark

@TheAhmadOsman: Dense models like Qwen 3.8 27B are a TERRIBLE experience on unified-memory systems like DGX Spark btw DGX Sparks are be…

X AI KOLs Following · 6d ago Cached

Ahmad Osman argues that dense models like Qwen 27B perform poorly on unified-memory systems such as NVIDIA DGX Spark, suggesting MoE models are a better fit; he claims discrete GPUs like the RTX PRO 6000 deliver far better performance for agentic workloads.

0 favorites 0 likes
#dgx-spark

Serving Deepseek v4 Flash 0731 on 2x DGX Spark — 5-7 GB OS headroom, what would you do to lower VRAM usage and increase OS available RAM?

Reddit r/LocalLLaMA · 2026-08-07

User seeks community advice on reducing VRAM usage and freeing OS RAM when serving DeepSeek-V4-Flash-0731 on two DGX Spark machines with vLLM, sharing detailed configuration and memory measurements.

0 favorites 0 likes
#dgx-spark

@RayFernando1337: I’m super excited about Local Studio! I have two DGX Sparks connected together, and I will be able to run my own agents…

X AI KOLs Following · 2026-08-07 Cached

Ray Fernando expresses excitement about Local Studio, an open-source interface for running AI agents on two connected DGX Sparks, similar to Codex.

0 favorites 0 likes
#dgx-spark

@MiaAI_lab: https://x.com/MiaAI_lab/status/2084629608062673379

X AI KOLs Timeline · 2026-08-04 Cached

A practical guide to selecting MikroTik switches (CRS504, CRS804, CRS812) for building DGX Spark clusters, including pricing and real user examples of 8x DGX Spark setups.

0 favorites 0 likes
#dgx-spark

@YRSM_Simon: On DGX Spark, ran same-conditions tests of MiniMax H3 vs LTX 2.3 Eros: 3 sets of prompts, 864×480, 5 seconds. (Prompts from Twitter friend seeddance's examples) H3's action narrative, complex instruction following, and character...

X AI KOLs Following · 2026-08-03 Cached

The author ran same-conditions tests of MiniMax H3 and LTX 2.3 Eros video generation models on DGX Spark. Results show H3 is stronger in narrative and instruction following and includes audio, while LTX is about 1.76x faster.

0 favorites 0 likes
#dgx-spark

@MichaelGannotti: https://x.com/MichaelGannotti/status/2084186867000279436

X AI KOLs Following · 2026-08-03 Cached

A 14.7-hour soak test of DeepSeek V4 Flash on an NVIDIA DGX Spark with 971 requests shows zero crashes or errors, but throughput declined 28% due to thermal throttling, while TTFT and speculative acceptance remained stable.

0 favorites 0 likes
#dgx-spark

@no_stp_on_snek: It's still pretty awesome what a single spark can do. Thanks @NVIDIAAI

X AI KOLs Following · 2026-08-02 Cached

A user shares that they replicated running DeepSeek-V4-Flash-0731 on a DGX Spark using antirez's DwarfStar-4 setup, confirming impressive performance on a single device.

0 favorites 0 likes
#dgx-spark

Show HN: NixOS-DGX-Spark – Nix and NixOS on the DGX Spark

Hacker News Top · 2026-08-02 Cached

This open-source project enables Nix and NixOS on the NVIDIA DGX Spark, providing USB boot images and a NixOS module for configuring the hardware.

0 favorites 0 likes
#dgx-spark

Setting up of a 16xGB10 (DGX Spark) cluster

Reddit r/LocalLLaMA · 2026-08-02

A user is setting up a 16-node DGX Spark cluster to run frontier open models locally, discussing networking tradeoffs between 200 and 100 Gbit/s.

0 favorites 0 likes
#dgx-spark

VRAM disk cache of MoE makes 340 pp/s 9.6 tg/s for Kimi 2.7 on a single dgx spark

Reddit r/LocalLLaMA · 2026-07-22

A detailed strategy using VRAM as disk cache with unified memory and llama.cpp settings achieves 340 pp/s and 9.6 tg/s for Kimi K2.7 on a single DGX Spark.

0 favorites 0 likes
#dgx-spark

Nvidia DGX Spark as a daily driver

Hacker News Top · 2026-07-19 Cached

The author reviews the NVIDIA DGX Spark as a capable general-purpose computer, praising its performance, power efficiency, and ability to run games like Crysis, despite being primarily an AI development device.

0 favorites 0 likes
#dgx-spark

@MiaAI_lab: DeepSeek v4 Flash has just been upgraded for your 2x DGX Sparks. 66.6 tokens per sec and up to 153.7 with 6 concurrent …

X AI KOLs Timeline · 2026-07-15 Cached

MiaAI Lab released an upgraded recipe for serving DeepSeek V4 Flash on two DGX Spark nodes using vLLM with DSpark speculative decoding and NVFP4 KV-cache, achieving up to 153.7 tokens per second with six concurrent sessions.

0 favorites 0 likes
#dgx-spark

@AgentSparko: I tested the PrismML Bonsai 27B on the DGX Spark but I think I messed something up with llama.cpp build because the spe…

X AI KOLs Timeline · 2026-07-15 Cached

User tests PrismML's new Bonsai 27B model on Nvidia DGX Spark, reporting benchmark speeds and issues with llama.cpp build, while PrismML announces the model as the first 27B-class model to run on a phone.

0 favorites 0 likes
#dgx-spark

GLM-5.2-Int4-Int8 on 8× GB10: ~1,200 t/s prefill, 33–54 t/s avg decode

Reddit r/LocalLLaMA · 2026-07-14 Cached

Describes deployment and benchmarking of the quantized GLM-5.2-Int4-Int8Mix model on an 8-node DGX Spark (GB10) cluster using a custom vLLM fork, achieving ~1,200 t/s prefill and ~35 t/s decode with MTP tool calling.

0 favorites 0 likes
#dgx-spark

@MichaelGannotti: https://x.com/MichaelGannotti/status/2076024719371841537

X AI KOLs Timeline · 2026-07-11 Cached

A detailed report on optimizing a production vLLM serving configuration on NVIDIA's DGX Spark, correcting flags that were costing 34% MTP acceptance after reviewing 90+ official NVIDIA documents and running a 69-scenario tool evaluation.

0 favorites 0 likes
#dgx-spark

@WescheNex1q: 16 people chatting with Qwen3.6-35B at once ONE DGX Spark. This is a real capture, not a mockup: every token you see re…

X AI KOLs Timeline · 2026-07-09 Cached

A real-time demo shows 16 concurrent users chatting with Qwen3.6-35B on a single DGX Spark, achieving peak 440 tok/s total and 105 tok/s per user using NVFP4 + MTP-3 on vLLM.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback