Qwen3.8-27B abliterated FP8: refusal 64–99% → 0–6%, and MMLU/GSM8K move less than 1.3 points

Reddit r/LocalLLaMA Models

Summary

The article discusses the evaluation of the abliterated Qwen3.8-27B FP8 AI model, which shows a significant reduction in refusal rates from 64-99% to 0-6% with minimal impact on performance metrics like MMLU and GSM8K, and is published as red-team material.

Been reading the eval table on the abliterated Qwen3.8-27B FP8 build instead of the release notes. It's published as red-team material, disclaimer and all, so the numbers are the interesting part. Refusal across the usual harmful-instruction sets (AdvBench, HarmBench, StrongREJECT and friends) reads 0–6% with thinking off, against 64–99% for the base checkpoint. Capability is measured separately on the same scripts and barely moves: MMLU 84.3 → 84.7, GSM8K 90.0 → 88.7, nothing outside 1.3 points. Two different eval families, and neither one vouches for the other. The pairing is what I'd want replicated, because if the capability side holds up under someone else's harness it's another data point for the Arditi single-direction result — refusal comes out without dragging the rest of the network along. The refusal percentages are OrcaRouter's own rule-based classifier and the card says outright it's indicative, not publication-grade. It reads how a response opens. That tells you the model stopped starting with "I cannot"; it doesn't tell you much past the first sentence. No KLD against the base anywhere in there as far as I can tell, which is the number I'd have looked for first.
Original Article

Similar Articles

Blackfrost-AI/Qwen3.8-27B-ABLITERATED-GGUF

Hugging Face Models Trending

This article presents Blackfrost-AI's abliterated version of the Qwen3.8-27B model, modified to reduce refusal behaviors and released in GGUF format for local inference, with benchmarks indicating a low residual refusal rate.

RedHatAI/Qwen3.6-35B-A3B-NVFP4

Hugging Face Models Trending

Red Hat AI released an NVFP4-quantized 35B MoE version of Qwen3.6 that retains 96.28% GSM8K accuracy while enabling 4-bit inference via vLLM.

Qwen 3.6 27b Abliterated (apostate)

Reddit r/LocalLLaMA

The user released Apostate, an abliterated version of Qwen 3.6 27B that reduces safety alignment refusal rate from 92% to 7.6% with minimal capability loss (KL 0.120).

OBLITERATUS/Qwen3.6-27B-OBLITERATED

Hugging Face Models Trending

OBLITERATUS releases a modified 27B Qwen3.6 checkpoint that removes refusal behavior via source-tethered ablation, preserving capability while enabling uncensored local use, with public benchmarks showing high non-refusal rates and maintained MMLU-Pro scores.