New wave of miniboss models you can run on dual DGX Spark

Reddit r/LocalLLaMA News

Summary

A new wave of large language models including GLM 4.5, Qwen 3.5, MiniMax M2.7, Deepseek V4 Flash, Xiaomi MiMo 2.5, StepFun 3.7 Flash, and Tencent Hy3 can now be run locally on a dual DGX Spark setup with 250GB usable memory at 4-bit quantization, costing approximately $7,000–$8,000.

Two DGX Spark and a Connect-X7 cable give you about 250GB of usable memory for $7000 8000 USD. This allows using some interesting models at 4-bit. For what seemed like an eternity, the only serious models in that size were GLM 4.5/4.6/4.7 (194GB), Qwen 3.5 397B (210GB), and older MiniMax. Then 3 months ago we got MiniMax M2.7 (131GB), Deepseek V4 Flash (160GB) and Xiaomi MiMo 2.5 (181GB). Shortly later we got StepFun 3.7 Flash (129GB), and now we just got Tencent's Hy3 (182GB). I know this size is inaccessible to a lot of people on here, but 7k 8k seems reasonable acceptable to me to be able to run models of this power level. Have you been using them? Which is your favorite? [P.S. please don't read this post and run off to spend 7k without first spending $7 trying these models on OpenRouter first.]
Original Article

Similar Articles

dgx sparks and new models my tests and results

Reddit r/LocalLLaMA

This article presents test results for AI models like DeepSeek V4 Flash and Qwen3.8 on NVIDIA DGX Sparks hardware, detailing performance metrics, context lengths, and benchmark scores with operational insights.

Ling-3.0-flash MXFP4 released and running locally on one DGX Spark.

Reddit r/LocalLLaMA

Ling-3.0-flash MXFP4, a quantized model, has been released and runs locally on a single DGX Spark, achieving ~80 tok/s decoding and 2,500-3,500 tok/s long-input prefilling, enabling private on-device inference for coding, agents, and offline batch jobs.