@TeksEdge: Wow! New open source Computer Use model shows strong local performance on LLM Leaderboard using a single DGX Spark! Thi…

X AI KOLs Timeline Models

Summary

H Company released Holo-3.1-35B-A3B-NVFP4, an open-source computer-use model that achieves up to 195 tokens per second on a single DGX Spark node, outperforming larger models like Qwen3.5-397B and Kimi-K2.5.

Wow! New open source Computer Use model shows strong local performance on LLM Leaderboard using a single DGX Spark! This model can handle your agents Computer Use workload fast! Holo-3.1-35B-A3B-NVFP4 (vLLM + NVFP4) » 101 tokens/sec at low concurrency » Scales up to ~195 tokens/sec at higher concurrency (c10) A 35B-class model delivering nearly 200 tokens per second on a single DGX Spark node for Computer Use is impressive. Great efficiency for a local agentic model. how Mac/StrixHalo/RTX Card would do?
Original Article
View Cached Full Text

Cached at: 06/03/26, 05:53 PM

Wow! New open source Computer Use model shows strong local performance on LLM Leaderboard using a single DGX Spark!

This model can handle your agents Computer Use workload fast!

Holo-3.1-35B-A3B-NVFP4 (vLLM + NVFP4)

» 101 tokens/sec at low concurrency » Scales up to ~195 tokens/sec at higher concurrency (c10)

A 35B-class model delivering nearly 200 tokens per second on a single DGX Spark node for Computer Use is impressive. Great efficiency for a local agentic model.

how Mac/StrixHalo/RTX Card would do?

David Hendrickson (@TeksEdge): 🌞This is big Local AI news! A new open-source Computer-Use LLM has just launched.

Holo 3.1 is H Company’s (🇫🇷) new local computer-use agent model that beats Qwen3.5-397B, Kimi-K2.5, and Sonnet 4.6!

Since it is built for local deployment → ⬩ Runs fully on your machine

Similar Articles

dgx sparks and new models my tests and results

Reddit r/LocalLLaMA

This article presents test results for AI models like DeepSeek V4 Flash and Qwen3.8 on NVIDIA DGX Sparks hardware, detailing performance metrics, context lengths, and benchmark scores with operational insights.

DGX Spark agentic usage numbers

Reddit r/LocalLLaMA

A user shares benchmark results and configuration for running Qwen3.6 models on NVIDIA DGX Spark using vLLM, focusing on agentic workloads with concurrent requests and tool calling.