I RL-trained Qwen3.6-35B-A3B to RL-train small task-specific Qwen models. Fully open source! 🤓
Summary
The author trained a Qwen3.6-35B-A3B model using reinforcement learning to then RL-train small task-specific Qwen models, and has released everything fully open source.
Similar Articles
Qwen 3.7 Max
Qwen 3.7 is an impressive new AI model from Chinese labs, with discussion on whether weights will be available for download.
Qwen is never going to open source Qwen 3.7, aren't they?
After firing Junyang Lin, Qwen has locked down its large models and is no longer releasing open source models, while other Chinese AI labs continue to open source their latest models. Rumors suggest the small model team is gone and Qwen 3.6/3.7 may be the last open source models.
I trained Qwen3.5 to jailbreak itself with RL, then used the failures to improve its defenses
The author trained Qwen3.5 to jailbreak itself with reinforcement learning, using diversity rewards to surface multiple attack strategies, then improved the defender's robustness from 64% to 92% defense rate with a slight drop in benign accuracy.
Qwen-Image-2.0-RL Technical Report
This technical report presents Qwen-Image-2.0-RL, a post-training pipeline using reinforcement learning from human feedback and on-policy distillation to enhance visual quality and instruction-following in image generation and editing tasks.
Qwen/Qwen3.6-35B-A3B
Qwen releases Qwen3.6-35B-A3B, an open-weight Mixture-of-Experts model with 35B total parameters and 3B active parameters, featuring significant improvements in agentic coding and reasoning preservation.