rocm

Tag

Cards List
#rocm

Qwen 3.8 27B KV f16 vs q8_0 are not equivalents

Reddit r/LocalLLaMA · 4h ago

A user shares test results indicating that KV cache types f16 and q8_0 are not equivalent for the Qwen 3.8 27B model, with f16 showing better detail and consistency, and provides configuration details for AMD ROCm hardware.

0 favorites 0 likes
#rocm

Running Qwen 3.6 35B A3B-Q8_0 gguf on a cheap radeon 7600 at 18 token/s * update increased to 21 t/s

Reddit r/LocalLLaMA · 2026-08-12

User reports running Qwen 3.6 35B A3B-Q8_0 gguf on a Radeon 7600 with llama.cpp and ROCm, achieving 21 tokens per second after VRAM overclocking, with a note about a display-related performance bug.

0 favorites 0 likes
#rocm

Ling 3.0 Flash on Strix Halo

Reddit r/LocalLLaMA · 2026-08-10

Tweet reports that Ling 3.0 Flash on AMD Strix Halo is significantly faster than Qwen-122b using ROCm-optimized formats, but notes tool calls are broken in certain harnesses.

0 favorites 0 likes
#rocm

DeepSeek V4 Flash on a Single AMD MI300X

Hacker News Top · 2026-08-04 Cached

This repository provides configuration, patches, and tuning to run the DeepSeek V4 Flash 304B checkpoint on a single AMD MI300X in production, achieving 168 tok/s decode without quantization. It includes correctness overlays for vLLM ROCm, AITER tuning tables, and a hybrid KV cache strategy.

0 favorites 0 likes
#rocm

MSLK kernel reference (Website)

TLDR AI · 2026-08-03 Cached

Documentation reference for MSLK 1.3.0, a library of fused GPU kernels for transformer workloads including attention, quantization, GEMM, and MoE routing, supporting CUDA and ROCm with PyTorch integration.

0 favorites 0 likes
#rocm

@PyTorch: As the officially recommended framework for the @AMD AI DevMaster Hackathon, PyTorch enables developers to build AI app…

X AI KOLs Timeline · 2026-07-22 Cached

The AMD AI DevMaster Hackathon, running July 10–August 6, 2026, features three tracks (multimodal AI, agentic AI, physical AI) and a $30,000 prize pool, with PyTorch as the official framework on ROCm-enabled AMD GPUs.

0 favorites 0 likes
#rocm

There's a new PR for llamacpp claiming to boost prompt processing with rocm by around 15%, also fixes a bug which makes Q2_K 28x faster

Reddit r/LocalLLaMA · 2026-07-21 Cached

A new PR for llama.cpp boosts prompt processing on ROCm by ~15% and fixes a bug making Q2_K quantization 28x faster.

0 favorites 0 likes
#rocm

AMD ROCm 7.14 "TheRock" tech preview tagged for latest AMD GPU compute stack

Reddit r/LocalLLaMA · 2026-07-16 Cached

AMD has tagged the ROCm 7.14 'TheRock' tech preview, bringing AI training enhancements, performance improvements up to 16% for select AI workloads like Comfy UI, and ongoing Windows support for the open-source GPU compute stack.

0 favorites 0 likes
#rocm

Qwen 3.5 122B Heretic ROCmFP4 iMatrix

Reddit r/LocalLLaMA · 2026-07-15 Cached

A compact, importance-calibrated ROCmFP4 quantization of Qwen 3.5 122B model for high-memory AMD systems, achieving improved quality (14% lower KLD) and performance (28.45 tok/s). Requires ROCmFPX runtime; not compatible with stock llama.cpp.

0 favorites 0 likes
#rocm

Modded RTX 4090 48GB vs Radeon AI Pro R9700 vs Arc Pro B70 for local coding LLMs?

Reddit r/LocalLLaMA · 2026-07-09

A user seeks advice on choosing between a modded RTX 4090 48GB, dual AMD Radeon AI Pro R9700, or dual Intel Arc Pro B70 for running local coding LLMs, highlighting trade-offs in price, VRAM, software ecosystem, and inference speed.

0 favorites 0 likes
#rocm

@PyTorch: New on the PyTorch Foundation blog: @AMD and @Meta contributors share how PyTorch Monarch was brought to AMD Instinct G…

X AI KOLs Following · 2026-07-07 Cached

AMD and Meta contributors ported PyTorch Monarch to AMD Instinct GPUs with ROCm, enabling fault-tolerant distributed training at scale. The blog details the engineering work and validation on large clusters.

0 favorites 0 likes
#rocm

Bringing PyTorch Monarch to AMD GPUs: Single-Controller Distributed Training on ROCm (13 minute read)

TLDR AI · 2026-07-07 Cached

The article describes the porting of PyTorch Monarch, a distributed training runtime, to AMD GPUs with ROCm, enabling single-controller fault-tolerant training at scale and addressing reliability challenges in large-scale LLM training.

0 favorites 0 likes
#rocm

@QuixiAI: https://x.com/QuixiAI/status/2073936537213915611

X AI KOLs Following · 2026-07-06 Cached

QuixiAI released QuixiCore, a family of native high-performance AI kernel libraries for modern accelerators, with standalone implementations for CUDA, Metal, ROCm, XPU, and Gaudi backends, all sharing a common contract but no shared code.

0 favorites 0 likes
#rocm

@0x0SojalSec: Fuck your paid courses, Master GPU engineering for AI systems. From foundational books and CUDA/ROCm programming to low…

X AI KOLs Timeline · 2026-07-02 Cached

A curated list of resources for mastering GPU engineering for AI systems, covering CUDA, ROCm, optimization tools, multi-GPU orchestration, and distributed training.

0 favorites 0 likes
#rocm

@DanKornas: GPU engineering is too broad to learn from random tabs. Awesome GPU Engineering is a curated GitHub list of resources f…

X AI KOLs Timeline · 2026-06-28 Cached

A curated GitHub list of resources for learning GPU engineering, covering architecture, kernel programming, optimization, distributed systems, and AI acceleration with books, frameworks, profilers, and interview prep.

0 favorites 0 likes
#rocm

AMD Strix Halo RDMA Cluster Setup Guide

Hacker News Top · 2026-06-28 Cached

A setup guide for using a custom Docker/Podman toolbox with ROCm/RCCL RDMA support to cluster two AMD Strix Halo nodes, enabling vLLM with tensor parallelism across 256GB unified memory.

0 favorites 0 likes
#rocm

If LLMs are so good at coding…

Reddit r/LocalLLaMA · 2026-06-25

A discussion questioning why LLMs haven't helped ROCm and Intel's software ecosystems catch up to CUDA, highlighting NVIDIA's premium pricing and the need for genuine market competition.

0 favorites 0 likes
#rocm

Big News for AMD / Strix Halo+ Owners

Reddit r/LocalLLaMA · 2026-06-24

The NPU on AMD Strix Halo devices is now usable for AI inference, enabling hybrid mode that combines NPU and iGPU for faster prompt processing. Tools like Lemonade and AMD's ROCm software make this possible.

0 favorites 0 likes
#rocm

ROCm vs Vulkan vs vLLM on Dual R9700's

Reddit r/LocalLLaMA · 2026-06-21

A comparison of AI inference frameworks ROCm, Vulkan, and vLLM running on dual AMD Radeon 9700 GPUs, likely benchmarking performance for large language models.

0 favorites 0 likes
#rocm

Benchmarks from the latest eBay special: W6800 (modded V620)

Reddit r/LocalLLaMA · 2026-06-17

A user benchmarks a modded AMD V620 GPU flashed with W6800 firmware and a custom blower fan for running LLMs via Vulkan and ROCm backends, comparing performance on Qwen2.5-27B at various quantization levels.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback