Tag
Model routing is a hot trend to reduce inference costs, but the best routing is deeply task-specific. Teams like Harvey and Factory achieve significant cost savings by focusing on single workflows rather than generic routers.
The author trained a Qwen3.6-35B-A3B model using reinforcement learning to then RL-train small task-specific Qwen models, and has released everything fully open source.
A 6-person team built task-specific AI models that are 4-8x faster than OpenAI or Anthropic models, with 500K downloads on HuggingFace.