@TeksEdge: Wow! New open source Computer Use model shows strong local performance on LLM Leaderboard using a single DGX Spark! Thi…
Summary
H Company released Holo-3.1-35B-A3B-NVFP4, an open-source computer-use model that achieves up to 195 tokens per second on a single DGX Spark node, outperforming larger models like Qwen3.5-397B and Kimi-K2.5.
View Cached Full Text
Cached at: 06/03/26, 05:53 PM
Wow! New open source Computer Use model shows strong local performance on LLM Leaderboard using a single DGX Spark!
This model can handle your agents Computer Use workload fast!
Holo-3.1-35B-A3B-NVFP4 (vLLM + NVFP4)
» 101 tokens/sec at low concurrency » Scales up to ~195 tokens/sec at higher concurrency (c10)
A 35B-class model delivering nearly 200 tokens per second on a single DGX Spark node for Computer Use is impressive. Great efficiency for a local agentic model.
how Mac/StrixHalo/RTX Card would do?
David Hendrickson (@TeksEdge): 🌞This is big Local AI news! A new open-source Computer-Use LLM has just launched.
Holo 3.1 is H Company’s (🇫🇷) new local computer-use agent model that beats Qwen3.5-397B, Kimi-K2.5, and Sonnet 4.6!
Since it is built for local deployment → ⬩ Runs fully on your machine
Similar Articles
@TeksEdge: This is big Local AI news! A new open-source Computer-Use LLM has just launched. Holo 3.1 is H Company’s () new local c…
H Company released Holo 3.1, an open-source computer-use LLM specialized for local deployment, achieving 79.3% on AndroidWorld benchmark, beating larger models like Qwen3.5-397B and Kimi-K2.5.
dgx sparks and new models my tests and results
This article presents test results for AI models like DeepSeek V4 Flash and Qwen3.8 on NVIDIA DGX Sparks hardware, detailing performance metrics, context lengths, and benchmark scores with operational insights.
DGX Spark agentic usage numbers
A user shares benchmark results and configuration for running Qwen3.6 models on NVIDIA DGX Spark using vLLM, focusing on agentic workloads with concurrent requests and tool calling.
@MiaAI_lab: What’s the best model you can run on your @NVIDIAAI DGX Spark? 1× DGX Spark * Qwen 3.6 35b NVFP4 - 256k ctx, 110 tok/s…
A tweet detailing the best AI models to run on Nvidia's DGX Spark, including Qwen 3.6 and DeepSeek v4 Flash variants, with token speeds and context lengths for single and multi-unit setups.
@DeRonin_: My current local AI setup: - 2x DGX Spark linked (256gb) > GLM 5.2 @ 2bit, reasoning + agent loops - Mac Studio M3 Ultr…
A user describes their fully local AI stack using multiple hardware devices running Chinese models like GLM, Qwen, and Kimi, claiming 87% cost savings compared to frontier models like GPT-5.5 and Opus 4.8, while noting plans to self-host video generation.