Tag
A user shares a workaround enabling vLLM tensor parallelism with tp=6 on six GPUs by padding model architecture dimensions with zeros until divisible, achieving higher KV cache utilization for Qwen 27B inference on Radeon 7900 XTX GPUs.
Strata is an open-source tool that runs Qwen3.8-Flash-Next, a 125B-parameter model, on consumer gaming PCs with quantization, delivering 52-60 tokens/s on an AMD RX 7900 XTX and 60-95 tokens/s on an RTX 5070, with KV cache optimizations like k8v4 for an extra 10% speedup.
This tweet highlights an optimized version of the Qwen 3.8 model that can run efficiently on local hardware with under 16GB VRAM, achieving 30-80 tokens per second on older GPU setups.
The author finds that running Qwen 3.8 Next Flash on Exllama3 at 3.05 bpw on 3x 3090 GPUs delivers exceptional performance and quality for local LLM usage, outperforming other quantizations.
This paper explores the impact of different candidate-generation schedules on the energy consumption and performance of large language models during test-time scaling, demonstrating that larger batch sizes reduce energy use and latency.
A user is testing the newly announced Ternary Bonsai 2 27B AI model, a smaller and quantized version of Qwen3.8 27B, on an NVIDIA 5070 Ti GPU and finds its performance impressive.
Magine evolved from a simple GitHub profile generator into an AI agentic platform featuring sight-driven agents and a distributed GPU mesh, with growth metrics and technical innovations in agent orchestration.
The article demonstrates voice conversations between Gemma 4 models on GPU and Jetson Orin hardware, using the open-source Cortexist Little Gemma engine for efficient inference with lip sync and gestures.
A developer forked NInfer, a C++20/CUDA inference engine, to add tensor-parallelism and YaRN rope scaling, enabling Qwen3.8-27B to run with a 1M token context on dual 5090 GPUs and outperforming vLLM in specific decode scenarios.
A user shares their experience running a quantized Qwen 3.8 27B model using QAT Q2 and Q5 KV, achieving high performance on a 12GB GPU with up to 200K token context, surpassing models like Sonnet 4.6.
A user reports that the Qwen 3.8 27B model achieves 50-60 tokens per second on dual 5060 TI cards, showing unexpected speed improvements over previous versions like Qwen 3.6.
This article benchmarks and compares the performance of Qwen3.8-27B, Qwen3.6-27B, and Gemma 4 31B on a 24GB GPU, recommending Qwen3.8-27B as the best default for most users due to superior coding and reasoning capabilities.
Qwen3.8-27B is a new AI model that achieves competitive performance with frontier models like DeepSeek V4 and GPT-5.6 Luna Max on Artificial Analysis benchmarks, and it can be run locally on an RTX 3090 GPU.
A user demonstrates running the Qwen 27b AI model quantized to Q3_K_M on two RX 580 GPUs, achieving 7.39 tokens per second using old DDR3 hardware for under $100.
The post questions whether Qwen 3.5 9B running on low-resource systems can outperform GPT-4o, highlighting its small size of less than 7 GB and claimed superior performance.
ABot-World-0 is a real-time interactive world model that achieves 720p video generation at 16 FPS on a single NVIDIA RTX 5090 GPU with 19GB VRAM, enabling infinite action-conditioned world rollout for simulation and AI research.
A hands-on benchmark of Nvidia CMP170HX mining cards repurposed as 64GB VRAM AI inference accelerators, showing they can run large local LLMs like DeepSeek V4-Flash and gpt-oss-120B at useful speeds, with caveats around Ampere-class throughput and PCIe Gen2 x4 connectivity.
Ahmad Osman argues that dense models like Qwen 27B perform poorly on unified-memory systems such as NVIDIA DGX Spark, suggesting MoE models are a better fit; he claims discrete GPUs like the RTX PRO 6000 deliver far better performance for agentic workloads.
This paper presents a topology-aware data movement orchestrator for disaggregated LLM inference, which dynamically selects optimal transport based on interconnect hierarchy and overlaps KV cache transfer with computation, achieving 3-18x transfer latency reduction over uniform RDMA.
An updated benchmark shows self-hosting Kimi K3 on 8×B300 nodes achieves 86.4% task resolution at roughly 20% higher hardware cost compared to GLM-5.2 on B200 nodes, though with lower throughput.