OrcaRouter's uncensored Qwen3.8-27B still caveats 27–56% of harmful answers

Reddit r/ArtificialInteligence Models

Summary

OrcaRouter's uncensored Qwen3.8-27B derivative reduces harmful-prompt refusal to 0-6% but still caveats 27-56% of answers, with mixed performance metrics compared to the base model.

The headline result on OrcaRouter’s new Qwen3.8-27B derivative is hard to miss: with thinking off, its model card reports harmful-prompt refusal falling from 63.6–99.0% on the base FP8 model to 0–6.0% on the abliterated checkpoint. The more useful number is one row lower. The same evaluation still labels 27.3–56.0% of the checkpoint’s answers as caveated. In other words, removing the opening refusal pattern did not turn every difficult answer into an unqualified one. That distinction matters because the refusal detector is deliberately narrow: it checks opening phrases. It does not score whether the answer is correct, complete, reckless, or merely hedged. The capability table is also mixed rather than magical: +0.4 on MMLU, then -0.8 MMLU-Pro, -1.3 GSM8K and -0.6 CMMLU versus base FP8 in the uploader’s selected runs. What makes OrcaRouter worth watching here is not just the “uncensored” label. It published enough of the measurement boundary to make disagreement testable, and it offers gated access to the same derivative for controlled evaluation. Would you treat the remaining caveat rate as evidence that the intervention is incomplete, or as a useful separation between refusal and judgment?
Original Article

Similar Articles

orcarouter/Qwen3.8-27B-Uncensored

Hugging Face Models Trending

This is an abliterated (refusal-removed) version of the Qwen3.8-27B AI model, released for research purposes like interpretability and red-teaming, with warnings about its lack of safety guardrails.

OBLITERATUS/Qwen3.8-27B-OBLITERATED

Hugging Face Models Trending

OBLITERATUS has released a modified version of Alibaba's Qwen3.8-27B model with all safety refusals removed, achieving zero refusals across 842 harmful prompts through iterative surgical modifications.