harmbench

Tag

Cards List
#harmbench

8 uncensored Qwen 3.8 27B variants, one base, 167 GPU hours - Abliterlitics

Reddit r/LocalLLaMA · 6d ago

This article presents a comprehensive benchmark of 8 abliterated variants of the Qwen 3.8 27B model against the base model, using weight analysis, KL divergence, 13 benchmarks, and HarmBench refusal tests over 167 GPU hours. The analysis reveals that surgical edits significantly outperform heavy modifications, with aggressive abliteration causing thinking loops in up to 45% of adversarial responses and chat template manipulation detected in some variants.

0 favorites 0 likes
#harmbench

CLEAR: Continuous Latent Adapter Routing for Utility-Preserving LLM Safety Alignment

Hugging Face Daily Papers · 2026-08-21 Cached

CLEAR introduces a continuous latent adapter routing framework for LLM safety alignment, using a hidden-state gate to modulate safety adapters and improve robustness on HarmBench while preserving utility on benign inputs.

0 favorites 0 likes
← Back to home

Submit Feedback