Tag
The author details benching a quad 5060Ti setup for running Qwen3.6-27B at Q8 with FP16 KV and MTP for code generation, concluding it offers good performance per dollar compared to alternatives like dual 3090s or modded 3080s.
A user reports that the Qwen3.6-27B model performs better and more reliably with llama.cpp than with vLLM, citing tool call errors and 'lobotomized' behavior in vLLM despite extensive configuration.
A developer reports that using Qwen3.6-27B for vibe coding introduced subtle bugs in their codebase and asks the community for best practices to reduce errors when working with the model.
This post presents Abliterlitics, an open-source toolkit for analyzing abliteration techniques, and compares five abliteration variants of Qwen3.6-27B using 85 GPU-hours of benchmarks, safety evaluations, and weight forensics. Heretic and Huihui show best capability preservation while all achieve near-complete safety removal.