Do you think a few Qwen3.8-27B models working together could score as well as Fable-5 on LiveCodeBench Hard?

Reddit r/LocalLLaMA Papers

Summary

A new paper claims that an ensemble of Qwen3.8-27B models achieves coding performance comparable to Fable-5 on LiveCodeBench, potentially at a significantly lower cost.

Has anyone tested this? Ensemble of small Qwen models claiming Fable 5-level coding performance. A new paper claims that running several Qwen3.8-27B models together matches Fable 5’s accuracy on LiveCodeBench. The authors also say their setup paired with GPT Terra reaches Fable 5-level coding accuracy on LiveCodeBench at roughly a fifth of the cost. Curious what people here think; is this worth actually trying out? https://github.com/slee-persis/GVS5H https://arxiv.org/abs/2608.26480
Original Article

Similar Articles