What's the best model you have running on strix halo 128GB?
Summary
The author shares their experience running local AI models on a Framework desktop with Strix Halo and 128GB unified memory, preferring Qwen models for coding, and asks for recommendations on better hardware utilization.
Similar Articles
Is it silly to get a 64GB Strix Halo (Framework Desktop) ~$2000?
A user queries whether a 64GB RAM Framework Desktop is sufficient for local AI tasks such as video generation with Minimax-H3 and Qwen 3.8 27B models, and asks for alternatives.
2x Strix Halo speed-up with an R9700
A user shares how they achieved a 2x performance boost in AI inference by splitting a large MoE model between a Strix Halo APU and an R9700 GPU, detailing configurations and code modifications.
Devs - you have 64gb of VRAM - which model do you use for coding?
A developer with 64GB VRAM shares their preference for an unsloth version of Qwen 3.5 122b-a10b for coding and asks the community for their recommendations.
High VRAM local coding model — still Qwen 3.6 27B?
The user discusses their experience with Qwen 3.6 27B for local coding tasks and asks for recommendations for larger models (100B+) suitable for systems with 224GB of VRAM.
Ran the same models across Strix Halo, RTX 3090, and RTX 5070 because I wanted my own numbers
The author ran 55 inference benchmark runs across Strix Halo, RTX 3090, and RTX 5070 with multiple backends, revealing that memory bandwidth dominates decode speed, the RTX 5070 beats the 3090 on small models, and reasoning models appear ~5x slower due to hidden reasoning content.