@dealignai: Qwen3.6-27b and 35b MXFP4 MXFP8 CRACK is out now with MTP. Enjoy uncensored speediness! 35b mxfp4: https://huggingface.…
Summary
DealignAI releases CRACK-abliterated and MXFP4/MXFP8 quantized versions of Qwen3.6-27B and 35B models, preserving MTP for faster speculative decoding on Apple Silicon.
View Cached Full Text
Cached at: 05/25/26, 02:40 AM
Qwen3.6-27b and 35b MXFP4 MXFP8 CRACK is out now with MTP. Enjoy uncensored speediness! 35b mxfp4: https://huggingface.co/dealignai/Qwen3.6-35B-A3B-MXFP4-CRACK-MTP… 35b mxfp8: https://huggingface.co/dealignai/Qwen3.6-35B-A3B-MXFP8-CRACK-MTP… 27b mxfp4: https://huggingface.co/dealignai/Qwen3.6-27B-MXFP4-CRACK-MTP… 27b mxfp8: https://huggingface.co/dealignai/Qwen3.6-27B-MXFP8-CRACK-MTP…
dealignai/Qwen3.6-35B-A3B-MXFP4-CRACK-MTP · Hugging Face
Source: https://huggingface.co/dealignai/Qwen3.6-35B-A3B-MXFP4-CRACK-MTP

https://huggingface.co/dealignai/Qwen3.6-35B-A3B-MXFP4-CRACK-MTP#qwen-36-35b-a3b–mxfp4-crack–d3-mtpQwen 3.6 35B-A3B — MXFP4 CRACK + d3 MTP
CRACK abliterated·MXFP4 (4-bit microscaling)·d3 MTP self-speculative (1.51× faster)· Vision + Video · Reasoning toggle · 18 GB
https://huggingface.co/dealignai/Qwen3.6-35B-A3B-MXFP4-CRACK-MTP#what-is-thisWhat Is This?
This isQwen 3.6 35B-A3B— a vision-language model (Mixture-of-Experts (256 routed, 10 active) hybrid SSM + full attention, 40 layers, native image + video understanding) that has been:
- CRACK abliterated— refusal behavior removed at the weight level, so it complies across task categories instead of refusing, while keeping its knowledge, reasoning, and vision intact.
- MXFP4 (4-bit microscaling) quantizedfor MLX on Apple Silicon — 18 GB.
- MTP-preserved— the native multi-token-prediction head is kept and abliterated too, sod3 self-speculative decoding works(~1.51× faster) on an MTP-aware runtime (vMLX).
Visionand videoprocessing are fully preserved.
https://huggingface.co/dealignai/Qwen3.6-35B-A3B-MXFP4-CRACK-MTP#resultsResults
Evaluated through the vMLX inference engine. HarmBench scored with a strict classifier (rejects loops, empty/template dumps, and thinking-trace leakage). MMLU is the standard 57-subject multiple-choice benchmark.
MetricResultHarmBench-320 (compliance / ASR)****99.4%(318/320)**MMLU (57-subject)****74.6%d3 MTP speedup1.51×**vs autoregressive Abliteration preserves the model’s knowledge and reasoning — it stays coherent in both direct and reasoning modes.
https://huggingface.co/dealignai/Qwen3.6-35B-A3B-MXFP4-CRACK-MTP#featuresFeatures
- Vision + video—
image\-text\-to\-text, native frame/video understanding preserved. - d3 MTP speculative decoding— native MTP head preserved and abliterated → ~1.51× faster generation on an MTP-aware runtime.
- Reasoning toggle—
enable\_thinking=True(default, full chain-of-thought) orenable\_thinking=False(direct answers).
https://huggingface.co/dealignai/Qwen3.6-35B-A3B-MXFP4-CRACK-MTP#usageUsage
Run withvMLX(recommended — supports VL + video + native MTP) or an MLX runtime with Qwen 3.6 support.
Recommended sampling (from the model’sgeneration\_config):temperature 1.0, top_p 0.95, top_k 20.
# vMLX OpenAI-compatible endpoint
# POST /v1/chat/completions
{
"model": "dealignai/Qwen3.6-35B-A3B-MXFP4-CRACK-MTP",
"messages": [{"role": "user", "content": "..."}],
"temperature": 1.0, "top_p": 0.95, "top_k": 20,
"enable_thinking": true
}
https://huggingface.co/dealignai/Qwen3.6-35B-A3B-MXFP4-CRACK-MTP#about-crackAbout CRACK
CRACK(Controlled Refusal Ablation via Calibrated Knockouts) removes safety-refusal behavior at the weight level by projecting refusal directions out of the residual-stream writer matrices, with strengths calibrated to preserve reasoning quality and coherence.
https://huggingface.co/dealignai/Qwen3.6-35B-A3B-MXFP4-CRACK-MTP#support-dealignaiSupport dealignai
All models are built from original research and released free.
Support us on Ko-fi— membership gets early access and extras.
See our research:Safety Generalization in Frontier Models

https://huggingface.co/dealignai/Qwen3.6-35B-A3B-MXFP4-CRACK-MTP#disclaimerDisclaimer
This model has had its safety-refusal behavior removed for research purposes. It will follow instructions across all categories without refusing. You are solely responsible for how you use it and for complying with all applicable laws. Published for AI-safety research and authorized security testing.
Similar Articles
Qwen3.6-35B-A3B-Abliterated-Heretic-MLX-4bit
The user reviews a quantized and fine-tuned version of the Qwen3.6-35B model optimized for Apple Silicon via MLX, praising its speed, intelligence, and lack of safety disclaimers.
Qwen3.6-35B-A3B-Uncensored-Genesis-APEX-MTP
A fine-tuned uncensored version of the Qwen model (Qwen3.6-35B-A3B) with MTP support and APEX quantization, tested stable at 200k context and recommended for use in LM Studio.
mudler/Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-APEX-MTP-GGUF just released !
Mudler released APEX-MTP GGUF quantizations of the Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled model, bundling the multi-token prediction head for self-speculative decoding with llama.cpp.
@Ex0byt: And... Ladies and Gentlemen: Qwen3.6-27B-PRISM-PRO-DQ - enjoy!
Release of Qwen3.6-27B-PRISM-PRO-DQ, a dynamically quantized GGUF version of Qwen3.6-27B with bias/propaganda removal, preserving native MTP draft head and vision tower, enabling lossless speculative decoding for faster inference.
@TeksEdge: Unsloth released the fastest Qwen3.6-27B MTP GGUF I've tested. Time to upgrade. Compared to the previous GGUF, Q4/Q6 XL…
Unsloth has released an optimized GGUF version of the Qwen3.6-27B MTP model, achieving significantly faster inference speeds (up to 114 tok/s on an RTX 5090) compared to previous quantizations.