Tag
A detailed report on optimizing a production vLLM serving configuration on NVIDIA's DGX Spark, correcting flags that were costing 34% MTP acceptance after reviewing 90+ official NVIDIA documents and running a 69-scenario tool evaluation.
A real-time demo shows 16 concurrent users chatting with Qwen3.6-35B on a single DGX Spark, achieving peak 440 tok/s total and 105 tok/s per user using NVFP4 + MTP-3 on vLLM.
A full-stack evaluation of NVIDIA's Nemotron-3 Mamba-Transformer hybrid models on DGX Spark hardware, including architecture analysis, quantization (NVFP4), benchmark results, and introduction of the open-source smf-bench testing suite.
Benchmark results for Qwen3.6-35B-A3B-UD-Q8_K_XL on DGX Spark using llama.cpp script by Mia, showing fast token generation times across various context lengths.
Miro warns that most people will regret buying a Mac or DGX Spark for local LLMs, and recommends the RTX Pro 6000 for serious use.
A bug in vLLM's speculative decoding configuration for GLM-5.2 NVFP4 on four DGX Sparks was fixed, resolving a performance tradeoff and achieving ~24 tok/s at 128K context with MTP4.
A user describes their fully local AI stack using multiple hardware devices running Chinese models like GLM, Qwen, and Kimi, claiming 87% cost savings compared to frontier models like GPT-5.5 and Opus 4.8, while noting plans to self-host video generation.
Announcing Orinth 1.0 AEON ULTIMATE UNCENSORED, a model with BF16 and NVFP4 quantization for DGX Spark/Blackwell architecture, claiming 200-300% performance improvement with working DFlash.
A tweet detailing the best AI models to run on Nvidia's DGX Spark, including Qwen 3.6 and DeepSeek v4 Flash variants, with token speeds and context lengths for single and multi-unit setups.
A user thanks followers and promotes GitHub repos with SM121 optimized containers for running local LLMs on DGX Spark (GB10) systems.
A detailed analysis on whether to run AI models locally or via API, covering hardware options like RTX 5090, RTX PRO 6000, and DGX Spark, with emphasis on memory vs bandwidth trade-offs, cost considerations, and privacy needs.
A comparison between a single RTX Pro 6000 GPU and two DGX Spark systems for AI compute tasks.
The author successfully ran GLM-5.2 with MTP speculative decoding on a 4× DGX Spark (GB10) setup, revealing a missing component in the public build recipe.
Update on running a non-quantized DeepSeek-v4-Flash model at 11 tok/s on a single DGX Spark using sglang inference and a custom mega-kernel, progressing towards GLM-5.2.
Spark Doctor is an open-source diagnostic CLI for NVIDIA DGX Spark that collects system, GPU, memory, Docker, and recipe data, applies specific rules, and outputs the likely cause and next steps for common issues.
A user asks about the feasibility of running GLM-5.2 at 4-bit quantization on four Ascend GX10s or DGX Sparks, wondering about speed and memory for 100k context.
NVIDIA announces the third generation of its DGX Spark AI supercomputer.
AMD launches the Ryzen AI Halo Developer Platform, a $3,999 mini PC with 128GB unified memory and Windows 11 support, competing with Nvidia's DGX Spark for local AI workloads.
A Reddit user shares their experience running DeepSeek V4 Flash on a dual-ASUS GX10 DGX Spark setup, detailing performance metrics, configuration, and power consumption, with throughput benchmarks across various context lengths.
Dell confirms a new XPS laptop featuring the NVIDIA N1X chip, essentially a consumer version of the DGX Spark GB10, to be announced at Computex.