Tag
A user shares and asks for feedback on running Qwen3.8-Flash-Next in EXL3 5.05bpw on a single Radeon RX 9700 (gfx12) with mixed GPU/CPU expert offloading, reporting ~863 t/s prefill and ~35 t/s decode at 230k context.
Appreciation post for tcclaviger adding expert RAM offloading support to vLLM, enabling running large AI models like DeepSeek-V4-Flash-Vision-Exp on local setups with multiple GPUs.
BeeLlama v0.4.7 is released, offering gigabytes of VRAM savings through independent drafter ubatch sizing for MTP and DFlash, along with major ROCm and CUDA performance enhancements for AI inference.
This article demonstrates a setup where an AMD Radeon AI Pro R9700 GPU is used on a MacBook Pro via Thunderbolt 5 with a custom driver and LemonSeed-Engine for running LLMs like Qwen3.6-3.8.
Optimized performance for running the Qwen3.8 27b NVFP4 model on a single AMD Radeon R9700 GPU, achieving up to 153 tokens per second in decode and 3,619 tokens per second in prefill with improved concurrency.
The article introduces AMDKernelVault, an open dataset and training framework for AMD GPU kernel optimization, featuring large-scale HIP and Triton kernels and agent-driven pipelines for generating and validating kernels.
This tool provides a reproducible Windows setup to run CUDA-targeted applications on AMD GPUs using ZLUDA and AMD HIP/ROCm, validated with RX 9060 XT and LibTorch workloads.
This pull request adds missing AMD GCN MMQ configuration to ggml-cuda for HIP, enhancing prefill performance for RDNA2 GPUs such as MI50 and MI60 in the llama.cpp inference library.
VoxGen is a new lightweight native inference engine for VoxCPM2 models, optimized for AMD cards using Rust and Vulkan compute to enhance performance and remove Python/PyTorch dependencies.
The article describes the optimization of AMD MI350X GPUs for running the Qwen3.6-35B-A3B LLM, achieving high output token throughput and open-sourcing the kernel to improve performance.
The user fixed crashes when running large AI models on an AMD GPU in headless Linux by disabling runtime power management, preventing model weights from being dumped into limited RAM and causing OOM errors.
AsmEvo is an agentic assembly-level optimizer for AMD GPU kernels that improves performance by proposing low-level edits and verifying functional equivalence against original binaries, achieving speedups up to 3.88x on MI308X GPUs.
A fork of llama.cpp has been created and optimized for AMD GFX906 GPUs, improving performance on Mi50, Mi60, Radeon VII, and GCN HIP for machine learning and LLM applications.
A user demonstrates running the Qwen 27b AI model quantized to Q3_K_M on two RX 580 GPUs, achieving 7.39 tokens per second using old DDR3 hardware for under $100.
Linux kernel patches introduce KNOD, a mechanism for in-kernel network packet offloading directly to AMD GPUs, enabling accelerated packet processing without user-space dependencies like ROCm. The code manages GPU queues, JIT compiles per-packet programs, and dispatches work entirely from the kernel.
USAF (Ultra Sparse Adaptive Fine-Tuning) is a new method that allows fine-tuning MoE models on consumer GPUs with as little as 12GB VRAM, including on AMD hardware, by training only the most important sparse weights and the router, unlike LoRA/QLoRA which cannot.
Discusses running a Q6 quantized version of the Gemma 4 31B model on a dual 9060 XT GPU configuration, likely for local inference.
A guide on optimizing VRAM usage on an AMD 7900XTX to run a 27B Qwen model with Q6K quantization and 131k context by compiling llama.cpp with OpenBLAS and CUDA_FA_ALL_QUANTS, and using kvcache quantization at q5_0/q4_0.
llama.cpp b9158 has been released with a fix for Flash Attention on RDNA3 GPUs, improving performance for AMD users.
A user shares a fix for performance bottlenecks when running AI models on AMD GPUs in Windows 11 by disabling memory compression via the command 'Disable-mmagent -mc'.