@xdotli: my friend @xeophon thinks coding is solved here's validation that a 3b model is trained with focus on algo efficiency a…
Summary
Nanbeige 4.1, a 3B model, outperforms Qwen3-30b-A3b and Qwen 3.5 4b in coding tasks with focus on algorithmic efficiency, achieving long horizon tasks with 600+ tool calls.
View Cached Full Text
Cached at: 06/20/26, 06:21 PM
my friend @xeophon thinks coding is solved
here’s validation that a 3b model is trained with focus on algo efficiency and not just correctness https://t.co/l16pWFC3dZ
Xiangyi Li (@xdotli): ICYMI
Nanbeige 4.1, a 3b model released by Chinese Indeed, outperforms Qwen3-30b-A3b + Qwen 3.5 4b. It can finish long horizon tasks with 600+ tool calls
We are working on something similar. thinking about doing a paper reading session. dm/comment for interest.
Similar Articles
@xdotli: ICYMI Nanbeige 4.1, a 3b model released by Chinese Indeed, outperforms Qwen3-30b-A3b + Qwen 3.5 4b. It can finish long …
Nanbeige 4.1, a 3B model from Chinese Indeed, outperforms larger Qwen models on tasks requiring 600+ tool calls.
@rasbt: Crazy model! It actually uses the old Qwen2.5-Coder-3B stack and got really great performance with their post-training …
A 3B parameter model using the Qwen2.5-Coder-3B stack achieves coding benchmark scores comparable to Claude Opus 4.5, with detailed post-training techniques including synthetic data, filtering, two-stage SFT, and a novel RL method (MGPO).
@f14bertolotti: Stellar performance from a 3B model. These results were achieved primarily through post-training refinements on Qwen2.5…
This technical report introduces VibeThinker-3B, a 3B parameter model that achieves frontier-level verifiable reasoning performance through post-training refinements on Qwen2.5-Coder, including curriculum-based supervised fine-tuning, multi-domain reinforcement learning, and offline self-distillation, matching or exceeding much larger models like DeepSeek V3.2.
Tested in Coding: Q8_K_XL Qwen3.8 27B vs BF16 Qwen3.6 27B
A user's detailed comparison of Qwen3.8 and Qwen3.6 models in coding tasks, highlighting improvements in instruction following and tracing for Qwen3.8, but with inefficiencies in reasoning.
Qwen 3.8 27b is strong even at Q3_xxs
The user finds Qwen 3.8 27b in Q3 quantization highly effective for coding tasks with fast inference speeds, outperforming previous models, despite minor issues in general conversations.