@TeksEdge: Solved! Qwen3.6-27B-FP8 is now running on Intel Arc Pro B70! LocalMaxxing shows a working 4× Arc Pro B70 32GB run at ~5…
Summary
Qwen3.6-27B-FP8 model is now running on Intel Arc Pro B70 GPUs at ~50 tok/s with a vLLM bug fix, marking a significant milestone for Intel GPU local AI inference.
View Cached Full Text
Cached at: 05/16/26, 07:15 AM
Solved! Qwen3.6-27B-FP8 is now running on Intel Arc Pro B70!
LocalMaxxing shows a working 4× Arc Pro B70 32GB run at ~50 tok/s solved by @xyster
That’s a BIG deal for Intel GPU local AI
27B-class FP8 inference is no longer just “maybe with the right patch in the future“
It’s now benchmarked running real on Intel Arc Pro B70
The software stack is still early , but this is exactly the kind of proof point Arc Pro B70 needed
David Hendrickson (@TeksEdge): 🪲 New vLLM bug fix for Qwen3.5-27B is available for Intel Pro ARC B70. Anyone tested it yet?
✅ New bug fix landed 2 days ago (commit b3169b8): → Fixed
gdn_conv_fused_seqcrash for Qwen3.5-27B at TP=1 (H=16, HV=48) ⚠️ Still no major fix confirmed for higher TP (like TP=4)
Similar Articles
@Tono_Ken3: On GPU, driving Qwen3.6-27b (100 TPS) and on CPU, Hy3-299B (25 TPS) simultaneously The dawn of a new era in local LLM i…
Demonstration of running Qwen3.6-27b on GPU at 100 TPS and Hy3-299B on CPU at 25 TPS simultaneously, marking a milestone in local LLM inference.
Qwen 3.6-35B-A3B with 977 tk/s prompt processing and 262k context window on Intel Arc B70 Pro
This article describes how to use the SYCL backend with llama.cpp to achieve over 60 tokens per second on the Qwen 3.6-35B-A3B model using an Intel Arc Pro B70 GPU, with the entire model and KV cache in VRAM.
Running Qwen3.6 35b a3b on 8gb vram and 32gb ram ~190k context
The author shares a high-performance local inference configuration for running Qwen3.6 35B A3B on limited hardware (8GB VRAM, 32GB RAM) using a modified llama.cpp with TurboQuant support, achieving ~37-51 tok/sec with ~190k context.
Tried Qwen3.6-27B-UD-Q6_K_XL.gguf with CloudeCode, well I can't believe but it is usable
User reports surprisingly usable coding performance from Qwen3-27B-UD-Q6_K_XL.gguf running locally on RTX 5090 at ~50 tok/s with 200K context, marking a significant leap in local model quality.
@DeepTechTR: Qwen 3.6 27B is incredibly fast with 16 GB VRAM! The impact of Pure Quant The era of the 27B model that runs seamlessly…
Qwen 3.6 27B runs fast on 16 GB VRAM thanks to 'Pure Quant' technology, achieving 40 tokens/s with MTP and supporting 64k contexts, enabling local AI on consumer GPUs like RTX 4060 Ti.