Qwen3.5 122B is the best?

Reddit r/LocalLLaMA News

Summary

A user shares their experience comparing several large language models (Qwen, Gemma) on complex tool-calling tasks, finding Qwen3.5 122B the most reliable, while criticizing smaller MoE models for instability.

I’m using Opencode and a computer with 128gb. So maybe the results would be different on system. I’ve exhaustingly tried Qwen3.6 27B and Qwen3.6 33B. I have no idea why but they just fall apart when doing more complex tasks with many tool calls. They’re pretty aggressive, doing slightly more than asked, and end up digging themselves into problems. Gemma4 31B and the 26B are literally the opposite. They can’t simply get things done. I have to sit there babysitting them just saying ok, ok, ok. Tool calling on bot the Qwen and Gemma MoE models feel buggy. Consistently just getting blank responses. The one model that I just keep coming back to is Qwen3.5 122B. It seemingly just gets the job done. I spent all day trying to just extract a few specific data fields from about 160 PowerPoints using these models and just ran into issue after issue. I just gave Qwen3.5 122B a goal of what I wanted and it did it in about 2 hours. I feel like the fully dense ~30B dense models out there are alright but just aren’t worth how slow they are. The MoE models around this size are just trash. You’re just better off on a system with 16gb and using models by API. The 120B size MoE models really hit such a sweet of capability and speed. I really hope to see more at this size. Yeah not everyone has the ram for this but I really feel like I’m just wasting time and effort using anything g smaller. Anyone else feel the same?
Original Article

Similar Articles

Qwen 35b a3b surprises me

Reddit r/LocalLLaMA

User reports positive experience with Qwen 35b a3b for agentic coding tasks, noting it outperforms Gemma4 26b in their use case and works well for demo/data analytics, especially in agentic mode versus chat.