@MiaAI_lab: If you mainly use local LLMs for Hermes-style agentic loops, this might surprise you: Qwen 3.6 35B actually *beats* Dee…
Summary
Qwen 3.6 35B outperforms DeepSeek v4 Flash on tool-heavy and coding-adjacent workflows, according to benchmarks from MiaAI Lab.
View Cached Full Text
Cached at: 06/30/26, 07:37 AM
If you mainly use local LLMs for Hermes-style agentic loops, this might surprise you:
Qwen 3.6 35B actually beats DeepSeek v4 Flash — especially on tool-heavy & coding-adjacent workflows.
You’re not missing out.
Full results
https://github.com/MiaAI-Lab/Qwen3.6-35b-vs-DSF4-tool-eval-bench…
MiaAI-Lab/Qwen3.6-35b-vs-DSF4-tool-eval-bench
Source: https://github.com/MiaAI-Lab/Qwen3.6-35b-vs-DSF4-tool-eval-bench
Qwen3.6-35B vs DSF4 Tool Eval Bench
Evaluation benchmarks comparing Qwen3.6-35B (specifically the Qwen3.6-35B-A3B-UD-Q8_K_XL GGUF quant) against DeepSeek-V4-Flash-DSpark on tool-calling and function-calling tasks.
Test Setup
All benchmarks were run with the following command with THINKING ON:
tool-eval-bench --seed 42 --trials 8 --hardmode --base-url http://localhost:8888
Contents
| File | Description |
|---|---|
Qwen3.6-35B-A3B-UD-Q8_K_XL.html | Tool evaluation results for the Qwen3.6-35B model |
DeepSeek-V4-Flash-DSpark.html | Tool evaluation results for the DeepSeek V4 Flash DSpark model |
DeepSeek-V4-vs-Qwen3.6-35B_comparison.html | Side-by-side comparison of both models across all benchmark dimensions |
How to View
Open any of the .html files in your browser to see the full benchmark results with interactive tables and visualizations.
License
MIT
Similar Articles
Qwen 3.8 27b vs Deepseek Flash
The post compares the open-source AI models Qwen 3.8 (27B) and Deepseek Flash, discussing benchmarks and seeking user experiences to evaluate their performance.
DeepSeek v4 Flash vs. Qwen3.6-27B, 3.5-122B, and Gemma 4 31B Benchmark
User benchmarks DeepSeek v4 Flash against Qwen3.6-27B, Qwen3.5-122B, and Gemma 4 31B on a local coding benchmark, finding Flash wins overall but Qwen 122B performs surprisingly well with better first-try success and lower token usage.
Is Qwen3.6 current king for local agentic use?
A user reports that Qwen3.6 35B A3B outperforms other local models like Gemma4 and GLM 4.7 Flash REAP for agentic tasks, though occasional loops still occur.
Qwen 3.6 27B on DeepSWE
Qwen 3.6 27B scored 2% on the DeepSWE benchmark, placing 18/20 above Haiku 4.5 and Minimax M2.7, highlighting the gap between local and leading-edge models.
The Qwen 3.6 35B A3B hype is real!!!
The author benchmarks small local LLMs, highlighting Qwen 3.6 35B A3B for its superior ability to map academic code to research papers compared to models like Gemma 4 and Nemotron 3 Nano.