@rasbt: Crazy model! It actually uses the old Qwen2.5-Coder-3B stack and got really great performance with their post-training …
Summary
A 3B parameter model using the Qwen2.5-Coder-3B stack achieves coding benchmark scores comparable to Claude Opus 4.5, with detailed post-training techniques including synthetic data, filtering, two-stage SFT, and a novel RL method (MGPO).
View Cached Full Text
Cached at: 06/17/26, 01:42 AM
Crazy model! It actually uses the old Qwen2.5-Coder-3B stack and got really great performance with their post-training stack. Need to use it in the next days to see if vibes of VibeCoder actually check out in practice. But impressive first impression!
Based on the tech report, some of the important pieces of their post-training stack:
-
High-signal synthetic data (math problems with credible solutions, code with tests)
-
Multiple reasoning paths for each answer
-
Filtering, filtering, filtering
-
2-stage SFT (start with broad training, then train on hard long-reasoning samples)
-
Use target (pass@k) accuracy over validation loss for checkpoint selection
-
MGPO (MaxEnt-Guided Policy Optimization) for RLVR: basically a GRPO-style RL method with an extra weighting that favors examples that are neither too easy nor too hard for the current policy
-
Single 64k long-context RL (they found that the usual progressive context expansion hurt this model because early truncation damaged long-thinking behavior)
-
Training data order: they do Math RL, then Code RL, then STEM RL in this particular oder which they found helped overall
-
After optimizing for accuracy, they add a stage that rewards shorter correct trajectories; basically making the model more efficient without accuracy degradation
orcus108 (@orcus108): WHAT THE HELL is happening in AI?
A 3B parameter model just put up coding benchmark scores in the same league as Claude Opus 4.5.
3 BILLION.
The weights are on Hugging Face, anyone can test it.
I genuinely don’t know if this is a breakthrough or if the benchmarks are broken.
Similar Articles
@xdotli: my friend @xeophon thinks coding is solved here's validation that a 3b model is trained with focus on algo efficiency a…
Nanbeige 4.1, a 3B model, outperforms Qwen3-30b-A3b and Qwen 3.5 4b in coding tasks with focus on algorithmic efficiency, achieving long horizon tasks with 600+ tool calls.
@f14bertolotti: Stellar performance from a 3B model. These results were achieved primarily through post-training refinements on Qwen2.5…
This technical report introduces VibeThinker-3B, a 3B parameter model that achieves frontier-level verifiable reasoning performance through post-training refinements on Qwen2.5-Coder, including curriculum-based supervised fine-tuning, multi-domain reinforcement learning, and offline self-distillation, matching or exceeding much larger models like DeepSeek V3.2.
Qwen3.6-27B: Flagship-Level Coding in a 27B Dense Model
Qwen releases Qwen3.6-27B, a 27B dense model claiming flagship-level coding performance surpassing the larger Qwen3.5-397B-A17B MoE, with impressive SVG generation demos.
@KyleHessling1: Hello again, everyone! We've got another really fun 9b, this one specifically trained for tool calling and agentic codi…
A new 9B fine-tuned model called Qwopus3.5-9B-Coder is released, optimized for tool calling and agentic coding workflows, achieving strong SWE-bench and HermesAgent-20 scores while running on affordable hardware.
Wow! Qwen 3.6:35b-a3b on a 3090... pretty amazing.
A user shares impressive results running a quantized Qwen 3.6:35b-a3b model on a used RTX 3090, achieving 160 tokens per second output after fitting the model into VRAM, and demonstrates vision capabilities with a 75-second video processing time.