Been running Qwen3.6-27B through a 3-critic harness. The harness matters more than I thought

Reddit r/LocalLLaMA Models

Summary

Reports on running Qwen3.6-27B (8-bit) through a 3-critic coding harness, finding the harness effectively catches errors and makes final output quality comparable to frontier models, with a proposed workflow of frontier for planning and Qwen for execution.

Been running Qwen3.6-27B (8-bit) through my coding harness for a few days, alongside GLM5.2. The harness uses 3 critics — code review, test review, Playwright e2e — each with fresh context before accepting output. Qwen3.6 is legit for a 27B dense model. Benchmarks weren't lying. It handles repo-level reasoning, produces decent code. But yeah it makes more mistakes than frontier models. Expected. What I didn't expect was that the 3-critic pipeline I built for frontier models turns out to be a great fit here. Critics catch the extra mistakes. Harness handles the retry overhead without breaking flow. The output after critics have done their work is good enough that I can't really tell the difference from a frontier run in terms of final quality. The path is just noisier. One thing though, the plan for this run is executing was written by GLM5.2, not Qwen3.6. My guess is the optimal split is frontier for planning + Qwen3.6 for execution. Strong model where reasoning matters most, cheap model for high-volume implementation where the harness catches errors.
Original Article

Similar Articles

Most powerful harness for Qwen 3.8?

Reddit r/LocalLLaMA

The post discusses which harness is most powerful for the Qwen 3.8 model, comparing Qwen code and open code in terms of features and usability.

Are you running Qwen 3.8 27b or Qwen Flash Next?

Reddit r/LocalLLaMA

The user discusses preferences between Qwen 3.8 27b and Qwen Flash Next models on Apple hardware, comparing speeds, and inquires about improving performance with MLX and harnesses without reasoning.