Ran a classic(medival europe) fantasy RP/agentic benchmark across 8 local models Qwen3.6-27B held up better than its size suggests
Summary
A benchmark comparing 8 local models on a classic medieval European fantasy role-playing and agentic task found that Qwen3.6-27B performed better than its size would suggest.
Similar Articles
Return of the local King! Qwen3.8-27B performance chart
Qwen3.8-27B is a new locally-runnable model that reportedly outperforms previous local models and rivals Opus4.6, with benchmark charts shared and links to ModelScope and HuggingFace.
@ModelScope2022: Qwen-AgentWorld just dropped two releases on ModelScope! An open 35B total / 3B active MoE world model with 256K contex…
Qwen-AgentWorld releases an open 35B total / 3B active MoE world model with 256K context, along with a 7-domain benchmark, achieving state-of-the-art performance on AgentWorldBench.
@AdinaYakup: Qwen released WebWorld an open world model series for web agents 8B/14B/32B+Dataset Apache2.0 +9.9% MiniWob++, +10.9% W…
Qwen released WebWorld, an open-source model series for web agents (8B/14B/32B) under Apache 2.0, which improves performance on MiniWob++ and WebArena benchmarks.
Qwen/Qwen-AgentWorld-35B-A3B
Qwen releases Qwen-AgentWorld-35B-A3B, a native language world model that simulates agentic environments across seven domains via long chain-of-thought reasoning. The model is trained with a three-stage pipeline and supports MCP, Search, Terminal, SWE, Android, Web, and OS interactions.
Is Qwen3.6 current king for local agentic use?
A user reports that Qwen3.6 35B A3B outperforms other local models like Gemma4 and GLM 4.7 Flash REAP for agentic tasks, though occasional loops still occur.