rocm

Tag

Cards List
#rocm

Two ~300B MoE models, each on ONE 128 GB mini PC (AMD Strix Halo): GLM-5.3-Flash at ~580 tok/s prefill, MiMo-V2.6-Flash up to 44 tok/s decode. EXL3 weights + open ROCm engine

Reddit r/LocalLLaMA ↗ · yesterday

Yamz Labs released Kyojin, a ROCm inference engine built on ExLlamaV3 for AMD Strix Halo, enabling two ~300B-class MoE models (GLM-5.3-Flash and MiMo-V2.6-Flash) to each fit and run on a single 128 GB mini PC, with EXL3 quantized weights hitting up to 580 tok/s prefill and 44 tok/s speculative decode.

0 favorites 0 likes
#rocm

AMD Ryzen AI Developer Platform OS updated with ROCm 10.0, Linux 7.2

Reddit r/artificial ↗ · 2d ago

AMD's Ryzen AI Developer Platform OS has been updated to ROCm 10.0 with Linux 7.2, advancing its open software stack for AI development on Ryzen AI hardware.

0 favorites 0 likes
#rocm

@PyTorch: RL post-training has become a critical stage in modern LLM development, but deploying an end-to-end pipeline requires m…

X AI KOLs Timeline ↗ · 2026-09-21 Cached

The article promotes a poster presentation at PyTorch Conference North America on enabling the open-source vime RL post-training framework on AMD Instinct GPUs using ROCm, and provides registration details for the conference.

0 favorites 0 likes
#rocm

Getting slower speeds WITH MTP on Gemma 4 12B QAT than without...

Reddit r/LocalLLaMA ↗ · 2026-09-01

A user reported slower inference speeds with MTP on the Gemma 4 12B QAT model when using the Vulkan backend, but discovered that switching to ROCm significantly improved performance, achieving 60t/s with MTP.

0 favorites 0 likes
#rocm

A very confusing report from Puget Systems

Reddit r/LocalLLaMA ↗ · 2026-09-01 Cached

This article evaluates the AI inference performance of dual AMD Radeon AI PRO R9700 GPUs, comparing them to Intel Arc Pro B70 and NVIDIA RTX 5090, highlighting cost-effectiveness and software challenges.

0 favorites 0 likes
#rocm

ROCm 10.0: A Decade of Open Compute, Built for the Age of Agentic AI

Reddit r/LocalLLaMA ↗ · 2026-08-28 Cached

AMD releases ROCm 10.0, a major update to its open-source GPU compute platform, marking a decade and introducing native agentic AI developer experience with ROCm.AI.

0 favorites 0 likes
#rocm

4xR9700, 2xMi210 or 4x4080S 32G

Reddit r/LocalLLaMA ↗ · 2026-08-24

The user is comparing GPU options like R9700, Mi210, and 4080S to achieve 128GB VRAM for running multiple AI models in parallel, considering factors like cost, performance, and compatibility.

0 favorites 0 likes
#rocm

AMD Users: Have you tried the llamma.cpp AMD-Ecosystem branch? Up to 2x PP Speed

Reddit r/LocalLLaMA ↗ · 2026-08-23

The article discusses an AMD-specific branch of llama.cpp that significantly boosts prompt processing speed for AMD users, with up to 2x faster performance on dense models using ROCm/Hip, though with some trade-offs in other metrics.

0 favorites 0 likes
#rocm

Qwen 3.8 27B KV f16 vs q8_0 are not equivalents

Reddit r/LocalLLaMA ↗ · 2026-08-20

A user shares test results indicating that KV cache types f16 and q8_0 are not equivalent for the Qwen 3.8 27B model, with f16 showing better detail and consistency, and provides configuration details for AMD ROCm hardware.

0 favorites 0 likes
#rocm

Running Qwen 3.6 35B A3B-Q8_0 gguf on a cheap radeon 7600 at 18 token/s * update increased to 21 t/s

Reddit r/LocalLLaMA ↗ · 2026-08-12

User reports running Qwen 3.6 35B A3B-Q8_0 gguf on a Radeon 7600 with llama.cpp and ROCm, achieving 21 tokens per second after VRAM overclocking, with a note about a display-related performance bug.

0 favorites 0 likes
#rocm

Ling 3.0 Flash on Strix Halo

Reddit r/LocalLLaMA ↗ · 2026-08-10

Tweet reports that Ling 3.0 Flash on AMD Strix Halo is significantly faster than Qwen-122b using ROCm-optimized formats, but notes tool calls are broken in certain harnesses.

0 favorites 0 likes
#rocm

DeepSeek V4 Flash on a Single AMD MI300X

Hacker News Top ↗ · 2026-08-04 Cached

This repository provides configuration, patches, and tuning to run the DeepSeek V4 Flash 304B checkpoint on a single AMD MI300X in production, achieving 168 tok/s decode without quantization. It includes correctness overlays for vLLM ROCm, AITER tuning tables, and a hybrid KV cache strategy.

0 favorites 0 likes
#rocm

MSLK kernel reference (Website)

TLDR AI ↗ · 2026-08-03 Cached

Documentation reference for MSLK 1.3.0, a library of fused GPU kernels for transformer workloads including attention, quantization, GEMM, and MoE routing, supporting CUDA and ROCm with PyTorch integration.

0 favorites 0 likes
#rocm

@PyTorch: As the officially recommended framework for the @AMD AI DevMaster Hackathon, PyTorch enables developers to build AI app…

X AI KOLs Timeline ↗ · 2026-07-22 Cached

The AMD AI DevMaster Hackathon, running July 10–August 6, 2026, features three tracks (multimodal AI, agentic AI, physical AI) and a $30,000 prize pool, with PyTorch as the official framework on ROCm-enabled AMD GPUs.

0 favorites 0 likes
#rocm

There's a new PR for llamacpp claiming to boost prompt processing with rocm by around 15%, also fixes a bug which makes Q2_K 28x faster

Reddit r/LocalLLaMA ↗ · 2026-07-21 Cached

A new PR for llama.cpp boosts prompt processing on ROCm by ~15% and fixes a bug making Q2_K quantization 28x faster.

0 favorites 0 likes
#rocm

AMD ROCm 7.14 "TheRock" tech preview tagged for latest AMD GPU compute stack

Reddit r/LocalLLaMA ↗ · 2026-07-16 Cached

AMD has tagged the ROCm 7.14 'TheRock' tech preview, bringing AI training enhancements, performance improvements up to 16% for select AI workloads like Comfy UI, and ongoing Windows support for the open-source GPU compute stack.

0 favorites 0 likes
#rocm

Qwen 3.5 122B Heretic ROCmFP4 iMatrix

Reddit r/LocalLLaMA ↗ · 2026-07-15 Cached

A compact, importance-calibrated ROCmFP4 quantization of Qwen 3.5 122B model for high-memory AMD systems, achieving improved quality (14% lower KLD) and performance (28.45 tok/s). Requires ROCmFPX runtime; not compatible with stock llama.cpp.

0 favorites 0 likes
#rocm

Modded RTX 4090 48GB vs Radeon AI Pro R9700 vs Arc Pro B70 for local coding LLMs?

Reddit r/LocalLLaMA ↗ · 2026-07-09

A user seeks advice on choosing between a modded RTX 4090 48GB, dual AMD Radeon AI Pro R9700, or dual Intel Arc Pro B70 for running local coding LLMs, highlighting trade-offs in price, VRAM, software ecosystem, and inference speed.

0 favorites 0 likes
#rocm

@PyTorch: New on the PyTorch Foundation blog: @AMD and @Meta contributors share how PyTorch Monarch was brought to AMD Instinct G…

X AI KOLs Following ↗ · 2026-07-07 Cached

AMD and Meta contributors ported PyTorch Monarch to AMD Instinct GPUs with ROCm, enabling fault-tolerant distributed training at scale. The blog details the engineering work and validation on large clusters.

0 favorites 0 likes
#rocm

Bringing PyTorch Monarch to AMD GPUs: Single-Controller Distributed Training on ROCm (13 minute read)

TLDR AI ↗ · 2026-07-07 Cached

The article describes the porting of PyTorch Monarch, a distributed training runtime, to AMD GPUs with ROCm, enabling single-controller fault-tolerant training at scale and addressing reliability challenges in large-scale LLM training.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback