Do you need some extra memory on your DGX Spark?
Summary
This repository helps DGX Spark users utilize a spare GPU to free up memory for better context or quantization quality by offloading the draft model via remote inference with vllm modifications.
Similar Articles
Serving Deepseek v4 Flash 0731 on 2x DGX Spark — 5-7 GB OS headroom, what would you do to lower VRAM usage and increase OS available RAM?
User seeks community advice on reducing VRAM usage and freeing OS RAM when serving DeepSeek-V4-Flash-0731 on two DGX Spark machines with vLLM, sharing detailed configuration and memory measurements.
DGX Spark agentic usage numbers
A user shares benchmark results and configuration for running Qwen3.6 models on NVIDIA DGX Spark using vLLM, focusing on agentic workloads with concurrent requests and tool calling.
@exolabs: https://x.com/exolabs/status/2103617535765573959
This article is a handbook for the NVIDIA DGX Spark, a device designed for local AI inference, detailing its specifications, how to link multiple units for enhanced performance, and its optimization for running mixture-of-experts models.
@QuixiAI: Two fixes for the NVIDIA open kernel driver on DGX Spark - Freed GPU memory now returns to the OS when a process exits …
NVIDIA's open kernel driver for DGX Spark received two fixes that return freed GPU memory to the OS when a process exits and enable huge pages for GPU page faults on system memory, boosting first-touch bandwidth from 0.4 to 19.6 GiB/s.
@antirez: DS4 running on DGX Spark (GB10 / CUDA), private branch for now. 12 tokens/sec, the memory bandwidth is limited in this …
Antirez reports benchmarking DS4 inference on the DGX Spark (GB10), noting 12 tokens/sec generation speed and high prefill performance, with plans to merge the codebase once mature.