Tag
A benchmark comparing 8 local models on a classic medieval European fantasy role-playing and agentic task found that Qwen3.6-27B performed better than its size would suggest.