Tag
Achieved 85.6 tokens per second using the Qwen3.8:27b model on a single RTX 5090 GPU.
Tests of Qwen-Image-2.1 on an NVIDIA 5070 Ti GPU demonstrate it fits within 16GB VRAM, with generation times around 25s for 1024² images, and highlight excellent prompt adherence and typography capabilities.
A draft model for Qwen 3.8 27B is highly effective on 16 GB GPUs, achieving around 60 tokens per second on an RX 9070 XT and offering better VRAM efficiency than built-in MTP.
Showcases the performance of the Qwen3.8-27B AI model running on a 12GB RTX 3080 Ti GPU, achieving 47 tokens per second at 128K context using llama.cpp with specific quantization settings.
A user reports achieving high token throughput with the Ninfer tool on an NVIDIA RTX 5090 GPU using a Qwen 3.8B model, significantly outperforming llama.cpp.
A benchmark comparison of community quantized Qwen 3.8 27B models on RTX 6000 GPUs versus Claude Opus 4.6, evaluating token generation speed and efficiency for creating HTML games.
LLM4LLM introduces a deployment-aware closed-loop optimization framework to bridge kernel benchmarks and real LLM inference, achieving up to 6.98x speedups on H100 GPUs.
The article describes hosting the Kimi K3 AI model with 2.8 trillion parameters using 8 B300 GPUs, achieving 92 tokens per second and costing $190 per million tokens, while comparing it with Unsloth's dynamic GGUF quantization method.
UC Berkeley has open-sourced FreeToken, a tool that significantly speeds up local AI inference on consumer GPUs, achieving up to 39.3 tok/s on an 8GB RTX 4060 laptop.
A user tested the Qwen3.8 27B model on SVG generation with a complex prompt and found the results exceeded expectations, highlighting the model's capability in creative tasks.
The author optimized the Qwen3.8-27B model inference on an RTX 3090 GPU, achieving up to 99 tokens per second for single requests and 1150 tps with batch processing through various quantization and optimization techniques, and released the updated code on GitHub.
The article explains how to determine if a GPU workload is compute-bound or memory-bound by analyzing operations per byte fetched from HBM, using NVIDIA's H100 as an example, and discusses how batching and prompt length affect performance.
A developer describes letting their AI agent autonomously test its own model upgrade by running controlled probes and measuring performance, revealing issues that throttled itself.
NVIDIA introduces a series on AI Model Co-Design, explaining how model dimensions affect GPU performance and the trade-offs between throughput and interactivity for LLM deployment. The first post provides a practical primer on designing hardware-friendly LLMs to improve system throughput and user responsiveness.
The user reports that the Qwen3.6 27B NVFP4 quantization is unreliable for coding, with inconsistent quality despite high throughput, and suggests that Q4_K_M may be more consistent.
Unsloth has released an optimized GGUF version of the Qwen3.6-27B MTP model, achieving significantly faster inference speeds (up to 114 tok/s on an RTX 5090) compared to previous quantizations.