Tag
This article presents a comprehensive benchmark of 8 abliterated variants of the Qwen 3.8 27B model against the base model, using weight analysis, KL divergence, 13 benchmarks, and HarmBench refusal tests over 167 GPU hours. The analysis reveals that surgical edits significantly outperform heavy modifications, with aggressive abliteration causing thinking loops in up to 45% of adversarial responses and chat template manipulation detected in some variants.
CLEAR introduces a continuous latent adapter routing framework for LLM safety alignment, using a hidden-state gate to modulate safety adapters and improve robustness on HarmBench while preserving utility on benign inputs.