amd-gpu

Tag

Cards List
#amd-gpu

Qwen3.8-Flash-Next (5.05bpw + ngram at bf16) exl3 on one r9700: 863 t/s prefill and 35 t/s decode at 230k context (256k max), is that ok or am i missing something?

Reddit r/LocalLLaMA ↗ · 12h ago

A user shares and asks for feedback on running Qwen3.8-Flash-Next in EXL3 5.05bpw on a single Radeon RX 9700 (gfx12) with mixed GPU/CPU expert offloading, reporting ~863 t/s prefill and ~35 t/s decode at 230k context.

0 favorites 0 likes
#amd-gpu

RAM Offloading with vLLM - tcclaviger appreciation post

Reddit r/LocalLLaMA ↗ · 3d ago

Appreciation post for tcclaviger adding expert RAM offloading support to vLLM, enabling running large AI models like DeepSeek-V4-Flash-Vision-Exp on local setups with multiple GPUs.

0 favorites 0 likes
#amd-gpu

@Anbeeld: BeeLlama v0.4.7 is out! You can now save gigabytes of VRAM when using MTP and DFlash! In the example shown in the image…

X AI KOLs Timeline ↗ · 2026-09-25

BeeLlama v0.4.7 is released, offering gigabytes of VRAM savings through independent drafter ubatch sizing for MTP and DFlash, along with major ROCm and CUDA performance enhancements for AI inference.

0 favorites 0 likes
#amd-gpu

MacBook Pro M5 Max LSE LLM running an AMD Radeon AI PRO R9700 over Thunderbolt 5 in a Razer enclosure

Reddit r/LocalLLaMA ↗ · 2026-09-24

This article demonstrates a setup where an AMD Radeon AI Pro R9700 GPU is used on a MacBook Pro via Thunderbolt 5 with a custom driver and LemonSeed-Engine for running LLMs like Qwen3.6-3.8.

0 favorites 0 likes
#amd-gpu

153 tok/s on 1x AMD Radeon R9700 running Qwen3.8 27b NVFP4, 470 tok/s @ 8 conc requests, Prefill @ 3,619 tok/s

Reddit r/LocalLLaMA ↗ · 2026-09-17

Optimized performance for running the Qwen3.8 27b NVFP4 model on a single AMD Radeon R9700 GPU, achieving up to 153 tokens per second in decode and 3,619 tokens per second in prefill with improved concurrency.

0 favorites 0 likes
#amd-gpu

AMDKernelVault: Large-Scale Datasets and Agentic Training for AMD GPU Kernel Optimization

arXiv cs.CL ↗ · 2026-09-14 Cached

The article introduces AMDKernelVault, an open dataset and training framework for AMD GPU kernel optimization, featuring large-scale HIP and Triton kernels and agent-driven pipelines for generating and validating kernels.

0 favorites 0 likes
#amd-gpu

CUDA for AMD on Windows

Hacker News Top ↗ · 2026-09-13 Cached

This tool provides a reproducible Windows setup to run CUDA-targeted applications on AMD GPUs using ZLUDA and AMD HIP/ROCm, validated with RX 9060 XT and LibTorch workloads.

0 favorites 0 likes
#amd-gpu

ggml-cuda: hip: add missing AMD GCN MMQ config by thelittlefireman · Pull Request #27841 · ggml-org/llama.cpp - PP improvements for RDNA2(MI50, MI60)

Reddit r/LocalLLaMA ↗ · 2026-09-12 Cached

This pull request adds missing AMD GCN MMQ configuration to ggml-cuda for HIP, enhancing prefill performance for RDNA2 GPUs such as MI50 and MI60 in the llama.cpp inference library.

0 favorites 0 likes
#amd-gpu

VoxGen, an AMD-optimized TTS inference engine for VoxCPM 2 models

Reddit r/LocalLLaMA ↗ · 2026-09-02

VoxGen is a new lightweight native inference engine for VoxCPM2 models, optimized for AMD cards using Rust and Vulkan compute to enhance performance and remove Python/PyTorch dependencies.

0 favorites 0 likes
#amd-gpu

Open Source Kernel in Qwen3.6-35B-A3B for AMD MI350X: 78,498 output tok/s on 8 GPUs

Reddit r/LocalLLaMA ↗ · 2026-08-26

The article describes the optimization of AMD MI350X GPUs for running the Qwen3.6-35B-A3B LLM, achieving high output token throughput and open-sourcing the kernel to improve performance.

0 favorites 0 likes
#amd-gpu

FIXED 7900 xtx + headless Linux crashes (Low RAM OOM) amdgpu.runpm=0

Reddit r/LocalLLaMA ↗ · 2026-08-25

The user fixed crashes when running large AI models on an AMD GPU in headless Linux by disabling runtime power management, preventing model weights from being dumped into limited RAM and causing OOM errors.

0 favorites 0 likes
#amd-gpu

AsmEvo: Agentic Assembly-Level Optimization of AMD GPU Kernels with Functional Equivalence Verification

arXiv cs.CL ↗ · 2026-08-24 Cached

AsmEvo is an agentic assembly-level optimizer for AMD GPU kernels that improves performance by proposing low-level edits and verifying functional equivalence against original binaries, achieving speedups up to 3.88x on MI308X GPUs.

0 favorites 0 likes
#amd-gpu

GLM and I created a llama.cpp fork optimized for AMD GFX906 (Mi50, Mi60, Radeon VII, GCN HIP) - Machine Learning, LLMs, & AI

Reddit r/LocalLLaMA ↗ · 2026-08-22

A fork of llama.cpp has been created and optimized for AMD GFX906 GPUs, improving performance on Mi50, Mi60, Radeon VII, and GCN HIP for machine learning and LLM applications.

0 favorites 0 likes
#amd-gpu

100$ worth of gpu runs qwen 3.8 27b at 7.39 t/s

Reddit r/LocalLLaMA ↗ · 2026-08-17

A user demonstrates running the Qwen 27b AI model quantized to Q3_K_M on two RX 580 GPUs, achieving 7.39 tokens per second using old DDR3 hardware for under $100.

0 favorites 0 likes
#amd-gpu

Linux patches introduce "KNOD" for in-kernel network offloading directly to AMD GPUs

Reddit r/LocalLLaMA ↗ · 2026-07-20 Cached

Linux kernel patches introduce KNOD, a mechanism for in-kernel network packet offloading directly to AMD GPUs, enabling accelerated packet processing without user-space dependencies like ROCm. The code manages GPU queues, JIT compiles per-packet programs, and dispatches work entirely from the kernel.

0 favorites 0 likes
#amd-gpu

If your GPU can run inference, it should be able to fine-tune too. [P]

Reddit r/MachineLearning ↗ · 2026-07-04 Cached

USAF (Ultra Sparse Adaptive Fine-Tuning) is a new method that allows fine-tuning MoE models on consumer GPUs with as little as 12GB VRAM, including on AMD hardware, by training only the most important sparse weights and the router, unlike LoRA/QLoRA which cannot.

0 favorites 0 likes
#amd-gpu

Gemma 4 31B Q6 on Dual 9060 XT

Reddit r/LocalLLaMA ↗ · 2026-06-22

Discusses running a Q6 quantized version of the Gemma 4 31B model on a dual 9060 XT GPU configuration, likely for local inference.

0 favorites 0 likes
#amd-gpu

7900XTX 24GB vram, can finally fit Q6K+MTP with Qwen 3.6 27B at 131k context

Reddit r/LocalLLaMA ↗ · 2026-06-20

A guide on optimizing VRAM usage on an AMD 7900XTX to run a 27B Qwen model with Q6K quantization and 131k context by compiling llama.cpp with OpenBLAS and CUDA_FA_ALL_QUANTS, and using kvcache quantization at q5_0/q4_0.

0 favorites 0 likes
#amd-gpu

RDNA3 Flash Attention fix just dropped by llama.cpp b9158

Reddit r/LocalLLaMA ↗ · 2026-05-15

llama.cpp b9158 has been released with a fix for Flash Attention on RDNA3 GPUs, improving performance for AMD users.

0 favorites 0 likes
#amd-gpu

If you're using Windows, disable memory compression to stop bottlenecks!

Reddit r/LocalLLaMA ↗ · 2026-05-14

A user shares a fix for performance bottlenecks when running AI models on AMD GPUs in Windows 11 by disabling memory compression via the command 'Disable-mmagent -mc'.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback