Tag
A buyer reports that NVIDIA DGX Spark units are hard to find and that prices appear to have jumped roughly $2,000 in a week, with local Micro Center stock vanishing overnight, raising questions about supply allocation.
A head-to-head benchmark of two community inference stacks running the 320B GLM-5.3-Flash on a pair of NVIDIA DGX Sparks, comparing Entrpi's EXL3 setup against TensorFold with DFlash2 drafters across prose, code, and prefill workloads.
This repository helps DGX Spark users utilize a spare GPU to free up memory for better context or quantization quality by offloading the draft model via remote inference with vllm modifications.
A tweet sharing a personal handbook guide for new users of NVIDIA's DGX Spark AI computing system.
The article compares the inference performance of Apple's Mac Studio M5 Ultra with two NVIDIA DGX Spark units when running the DeepSeek V4 Flash model, showing DGX Sparks are faster in prefill while Mac Studio is slightly faster in generation.
Nvidia's DGX Spark is out of stock for the first time on the marketplace, with prices rising at retailers, indicating strong AI hardware demand and potential economic implications.
A developer built a project in 8 hours using the Qwen3.8-Flash-Next model on a single NVIDIA DGX Spark, generating around 10k lines of code and consuming 800k tokens.
The author shares how an AI agent named Hermes, running on DGX Spark, automates business tasks like content management and client inquiries, saving 30-35 hours weekly, and argues that the agent market is at a pivotal moment akin to early mobile internet.
A user asks on a forum if a 2x DGX Spark cluster can run the DeepSeek 4.1 Flash model, or if more Sparks are required, seeking community insights.
DeepSeek-V4.1-Flash has been quantized to 4.75 bpw EXL3 for deployment on 4× DGX Spark, optimizing memory usage and enabling efficient local inference with plans for validation and further optimization.
NVIDIA's open kernel driver for DGX Spark received two fixes that return freed GPU memory to the OS when a process exits and enable huge pages for GPU page faults on system memory, boosting first-touch bandwidth from 0.4 to 19.6 GiB/s.
SGLang has added deployment recipes for DeepSeek-V4-Flash-Vision and DeepSeek-V4-Flash-0731 models on 2x DGX Spark hardware, with support for various configurations and optimizations.
User @ViC305 successfully runs DeepSeek-V4-Flash-Vision with EXL3 MixedK and DSpark speculative decoding on a single DGX Spark, achieving improved performance and fixing technical issues for multimodal AI deployment.
A bug has been identified in DGX Spark, with the author urging NVIDIA to ensure it receives the latest CUDA updates without delay.
User @QuixiAI reports difficulty installing the latest CUDA toolkit on their NVIDIA DGX Spark, suggesting that installing vanilla Ubuntu may be necessary to resolve the issue.
The Asus Ascent GX10 desktop AI supercomputer has experienced a substantial price increase, leading to speculation that the DGX Spark system might also see a price jump soon.
A user shares their optimized configuration for deploying the Qwen3.8-Flash-Next model on dual DGX Spark hardware, achieving up to 50t/s decode and 2,900t/s prefill speeds with technical patches and setup details.
Shantanu Goel shared a configuration recipe for optimizing Qwen 3.8 Flash Next on a single DGX Spark, tested for practical tasks and plans to benchmark it further.
Announcement that the repository for GLM 5.3 FLASH model is live, with initial benchmarks showing 59 tokens per second on DGX SPARK hardware and promises of more updates.
NVIDIA and Perplexity AI have partnered to bring AI processing to local desktops via the DGX Spark device, focusing on cost reduction and data privacy, with cloud offloading for advanced reasoning tasks.