Tag
Mercor details the reinforcement learning post-training of Qwen3.5-397B-A17B and Qwen3.6-35B-A3B for knowledge work agents, achieving a 70% relative improvement on the APEX-Agents benchmark, and open-sources the full training recipe.