Call for compute - help optimize Qwen inference speed on local hardware

Reddit r/LocalLLaMA Tools

Summary

The user has renamed a GitHub repository to HyperQwen to focus on optimizing Qwen model inference speeds on local hardware and is seeking testers with 4090s and 5090s GPUs for both Windows and Linux.

Yoyo I renamed https://github.com/syv-ai/qwen38-27b-rtx3090 to https://github.com/syv-ai/HyperQwen/ to focus more on Qwen models and "local" hardware in general, than just a single card and a single model. What I am looking for now, is people with 4090s, 5090s (1 or more is fine), for both Windows and Linux. Qwen4 is launching soon, and HyperQwen will most likely be the spot to get the maximum decode and prefill speeds. So if you have available hardware to help test (I only have a 3090), then please hit me up.
Original Article

Similar Articles

Running Qwen3.6 35b a3b on 8gb vram and 32gb ram ~190k context

Reddit r/LocalLLaMA

The author shares a high-performance local inference configuration for running Qwen3.6 35B A3B on limited hardware (8GB VRAM, 32GB RAM) using a modified llama.cpp with TurboQuant support, achieving ~37-51 tok/sec with ~190k context.

Testing Qwen 3.8 27B running locally on a single 5090

Reddit r/ArtificialInteligence

The article demonstrates the capabilities of running the Qwen 3.8 27B AI model locally on a single 5090 GPU, using Row-Bot to generate a rich animation showcasing tasks from language synthesis to physics simulation.

5090: Windows or Linux for Qwen3.8.27b

Reddit r/LocalLLaMA

User seeks advice on the best operating system (Windows or Linux) and inference server to run the Qwen3.8.27b model on a dedicated AI rig with RTX 5090 and 96GB RAM for optimal performance.