What's the best model you have running on strix halo 128GB?

Reddit r/LocalLLaMA News

Summary

The author shares their experience running local AI models on a Framework desktop with Strix Halo and 128GB unified memory, preferring Qwen models for coding, and asks for recommendations on better hardware utilization.

I have a framework desktop with 128gb of unified memory and have been having a blast learning local models with it. Ive noticed something, at least in my cases I always end up using smaller or moderately sized models like qwen3.6 35b or qwen 3.8 27b. Usually these models give me the best results for coding and im not complaining. But I feel like that extra 90GB is just sitting around being wasted, and the CPU isn't beefy enough to really run multiple agents at a decent speed. What have you all been running on your strix halo machines, and am I missing out on something?
Original Article

Similar Articles

2x Strix Halo speed-up with an R9700

Reddit r/LocalLLaMA

A user shares how they achieved a 2x performance boost in AI inference by splitting a large MoE model between a Strix Halo APU and an R9700 GPU, detailing configurations and code modifications.