@Italianclownz: Converted Qwen 3.6 35b a3b to ROCmfp4 and this is flying. Used the mtp version bc this ROCmfp4 can also incorporate the…

X AI KOLs Timeline Tools

Summary

Converted the Qwen 3.6 35b a3b model to ROCmfp4 format, leveraging MTP benefits for improved performance on AMD hardware.

Converted Qwen 3.6 35b a3b to ROCmfp4 and this is flying. Used the mtp version bc this ROCmfp4 can also incorporate the merged benefits of MTP. 262k. Reasoning On. @FrameworkPuter AMD 💪 🔥 https://t.co/K8j1KIZ8Ee
Original Article
View Cached Full Text

Cached at: 05/25/26, 08:48 AM

Converted Qwen 3.6 35b a3b to ROCmfp4 and this is flying. Used the mtp version bc this ROCmfp4 can also incorporate the merged benefits of MTP. 262k. Reasoning On. @FrameworkPuter AMD 💪 🔥 https://t.co/K8j1KIZ8Ee

Similar Articles

Qwen 3.5 122B Heretic ROCmFP4 iMatrix

Reddit r/LocalLLaMA

A compact, importance-calibrated ROCmFP4 quantization of Qwen 3.5 122B model for high-memory AMD systems, achieving improved quality (14% lower KLD) and performance (28.45 tok/s). Requires ROCmFPX runtime; not compatible with stock llama.cpp.

@Snixtp: https://x.com/Snixtp/status/2055734339346768225

X AI KOLs Timeline

A user benchmarks the MTP variant of Qwen3.6 27B against the normal version on a single RTX 3090 using llama.cpp, finding MTP offers up to 2.37x faster generation at long contexts (32k-64k) but with slower prefill and no concurrency support yet.