@ciruai: Testing DeepSeek v4 Flash on the AMD Ryzen AI Max+ 395 Strix Halo with 128GB RAM. Getting ~15 TPS over a decently long …
Summary
Testing DeepSeek v4 Flash on the AMD Ryzen AI Max+ 395 with 128GB RAM achieves ~15 TPS for a 284B MoE model (13B active) locally, costing $3,000 versus $25,000+ for a datacenter setup, highlighting the feasibility of running large models on consumer hardware.
View Cached Full Text
Cached at: 06/18/26, 04:19 PM
Testing DeepSeek v4 Flash on the AMD Ryzen AI Max+ 395 Strix Halo with 128GB RAM.
Getting ~15 TPS over a decently long context, which is honestly very usable for a model this smart.
284B parameter MoE, A13B active.
Before anyone says “that’s slow,” remember: this is running on a $3,000 machine. Getting this kind of model to run fast normally means spending well over $25,000 (if you build it yourself).
The accomplishment isn’t beating a datacenter GPU.
The accomplishment is running it locally at all.
Similar Articles
@0x0SojalSec: You don’t need 256GB+ RAM to run a DeepSeek-V4.1-Flash 763B model locally, - Only need 64GB system RAM. - offloaded a 2…
The article explains how to run the DeepSeek-V4.1-Flash 763B model locally with only 64GB RAM by offloading a 200GB Engram to NVMe and using DSpark, achieving 200 TPS on 4 Max-Q cards.
DeepSeek v4 Flash on 4090 + DDR5, my experience
A user shares their experience running the DeepSeek v4 Flash model with a 24GB GPU and DDR5 RAM, including performance numbers and tips for optimization.
Deepseek V4 Flash running on RTX 5090 MoE
User shares optimization benchmarks for DeepSeek-V4-Flash (Q2_K) running on an RTX 5090 using a fork of llama.cpp, achieving 21.3 tokens/s generation and 1 million context size.
DeepSeek-V4-Flash-0731 at ~40–50 tok/s on 2× Radeon AI PRO R9700 with the affinity engine (prebuilt quant + fixes)
The article reports on running DeepSeek-V4-Flash-0731 on two Radeon AI PRO R9700 GPUs with the affinity inference engine, achieving 40-50 tok/s decode speed via a prebuilt quantization and stability fixes.
DeepSeek V4 Flash on a Single AMD MI300X
This repository provides configuration, patches, and tuning to run the DeepSeek V4 Flash 304B checkpoint on a single AMD MI300X in production, achieving 168 tok/s decode without quantization. It includes correctness overlays for vLLM ROCm, AITER tuning tables, and a hybrid KV cache strategy.