@ciruai: Testing DeepSeek v4 Flash on the AMD Ryzen AI Max+ 395 Strix Halo with 128GB RAM. Getting ~15 TPS over a decently long …

X AI KOLs Timeline News

Summary

Testing DeepSeek v4 Flash on the AMD Ryzen AI Max+ 395 with 128GB RAM achieves ~15 TPS for a 284B MoE model (13B active) locally, costing $3,000 versus $25,000+ for a datacenter setup, highlighting the feasibility of running large models on consumer hardware.

Testing DeepSeek v4 Flash on the AMD Ryzen AI Max+ 395 Strix Halo with 128GB RAM. Getting ~15 TPS over a decently long context, which is honestly very usable for a model this smart. 284B parameter MoE, A13B active. Before anyone says “that’s slow,” remember: this is running on a $3,000 machine. Getting this kind of model to run fast normally means spending well over $25,000 (if you build it yourself). The accomplishment isn’t beating a datacenter GPU. The accomplishment is running it locally at all.
Original Article
View Cached Full Text

Cached at: 06/18/26, 04:19 PM

Testing DeepSeek v4 Flash on the AMD Ryzen AI Max+ 395 Strix Halo with 128GB RAM.

Getting ~15 TPS over a decently long context, which is honestly very usable for a model this smart.

284B parameter MoE, A13B active.

Before anyone says “that’s slow,” remember: this is running on a $3,000 machine. Getting this kind of model to run fast normally means spending well over $25,000 (if you build it yourself).

The accomplishment isn’t beating a datacenter GPU.

The accomplishment is running it locally at all.

Similar Articles

Deepseek V4 Flash running on RTX 5090 MoE

Reddit r/LocalLLaMA

User shares optimization benchmarks for DeepSeek-V4-Flash (Q2_K) running on an RTX 5090 using a fork of llama.cpp, achieving 21.3 tokens/s generation and 1 million context size.

DeepSeek V4 Flash on a Single AMD MI300X

Hacker News Top

This repository provides configuration, patches, and tuning to run the DeepSeek V4 Flash 304B checkpoint on a single AMD MI300X in production, achieving 168 tok/s decode without quantization. It includes correctness overlays for vLLM ROCm, AITER tuning tables, and a hybrid KV cache strategy.