@TeksEdge: Solved! Qwen3.6-27B-FP8 is now running on Intel Arc Pro B70! LocalMaxxing shows a working 4× Arc Pro B70 32GB run at ~5…

X AI KOLs Following News

Summary

Qwen3.6-27B-FP8 model is now running on Intel Arc Pro B70 GPUs at ~50 tok/s with a vLLM bug fix, marking a significant milestone for Intel GPU local AI inference.

Solved! Qwen3.6-27B-FP8 is now running on Intel Arc Pro B70! LocalMaxxing shows a working 4× Arc Pro B70 32GB run at ~50 tok/s solved by @xyster That’s a BIG deal for Intel GPU local AI 27B-class FP8 inference is no longer just “maybe with the right patch in the future" It’s now benchmarked running real on Intel Arc Pro B70 The software stack is still early , but this is exactly the kind of proof point Arc Pro B70 needed
Original Article
View Cached Full Text

Cached at: 05/16/26, 07:15 AM

Solved! Qwen3.6-27B-FP8 is now running on Intel Arc Pro B70!

LocalMaxxing shows a working 4× Arc Pro B70 32GB run at ~50 tok/s solved by @xyster

That’s a BIG deal for Intel GPU local AI

27B-class FP8 inference is no longer just “maybe with the right patch in the future“

It’s now benchmarked running real on Intel Arc Pro B70

The software stack is still early , but this is exactly the kind of proof point Arc Pro B70 needed

David Hendrickson (@TeksEdge): 🪲 New vLLM bug fix for Qwen3.5-27B is available for Intel Pro ARC B70. Anyone tested it yet?

✅ New bug fix landed 2 days ago (commit b3169b8): → Fixed gdn_conv_fused_seq crash for Qwen3.5-27B at TP=1 (H=16, HV=48) ⚠️ Still no major fix confirmed for higher TP (like TP=4)

Similar Articles

Running Qwen3.6 35b a3b on 8gb vram and 32gb ram ~190k context

Reddit r/LocalLLaMA

The author shares a high-performance local inference configuration for running Qwen3.6 35B A3B on limited hardware (8GB VRAM, 32GB RAM) using a modified llama.cpp with TurboQuant support, achieving ~37-51 tok/sec with ~190k context.