Tag
Yamz Labs released Kyojin, a ROCm inference engine built on ExLlamaV3 for AMD Strix Halo, enabling two ~300B-class MoE models (GLM-5.3-Flash and MiMo-V2.6-Flash) to each fit and run on a single 128 GB mini PC, with EXL3 quantized weights hitting up to 580 tok/s prefill and 44 tok/s speculative decode.
AMD's Ryzen AI Developer Platform OS has been updated to ROCm 10.0 with Linux 7.2, advancing its open software stack for AI development on Ryzen AI hardware.
The article promotes a poster presentation at PyTorch Conference North America on enabling the open-source vime RL post-training framework on AMD Instinct GPUs using ROCm, and provides registration details for the conference.
A user reported slower inference speeds with MTP on the Gemma 4 12B QAT model when using the Vulkan backend, but discovered that switching to ROCm significantly improved performance, achieving 60t/s with MTP.
This article evaluates the AI inference performance of dual AMD Radeon AI PRO R9700 GPUs, comparing them to Intel Arc Pro B70 and NVIDIA RTX 5090, highlighting cost-effectiveness and software challenges.
AMD releases ROCm 10.0, a major update to its open-source GPU compute platform, marking a decade and introducing native agentic AI developer experience with ROCm.AI.
The user is comparing GPU options like R9700, Mi210, and 4080S to achieve 128GB VRAM for running multiple AI models in parallel, considering factors like cost, performance, and compatibility.
The article discusses an AMD-specific branch of llama.cpp that significantly boosts prompt processing speed for AMD users, with up to 2x faster performance on dense models using ROCm/Hip, though with some trade-offs in other metrics.
A user shares test results indicating that KV cache types f16 and q8_0 are not equivalent for the Qwen 3.8 27B model, with f16 showing better detail and consistency, and provides configuration details for AMD ROCm hardware.
User reports running Qwen 3.6 35B A3B-Q8_0 gguf on a Radeon 7600 with llama.cpp and ROCm, achieving 21 tokens per second after VRAM overclocking, with a note about a display-related performance bug.
Tweet reports that Ling 3.0 Flash on AMD Strix Halo is significantly faster than Qwen-122b using ROCm-optimized formats, but notes tool calls are broken in certain harnesses.
This repository provides configuration, patches, and tuning to run the DeepSeek V4 Flash 304B checkpoint on a single AMD MI300X in production, achieving 168 tok/s decode without quantization. It includes correctness overlays for vLLM ROCm, AITER tuning tables, and a hybrid KV cache strategy.
Documentation reference for MSLK 1.3.0, a library of fused GPU kernels for transformer workloads including attention, quantization, GEMM, and MoE routing, supporting CUDA and ROCm with PyTorch integration.
The AMD AI DevMaster Hackathon, running July 10–August 6, 2026, features three tracks (multimodal AI, agentic AI, physical AI) and a $30,000 prize pool, with PyTorch as the official framework on ROCm-enabled AMD GPUs.
A new PR for llama.cpp boosts prompt processing on ROCm by ~15% and fixes a bug making Q2_K quantization 28x faster.
AMD has tagged the ROCm 7.14 'TheRock' tech preview, bringing AI training enhancements, performance improvements up to 16% for select AI workloads like Comfy UI, and ongoing Windows support for the open-source GPU compute stack.
A compact, importance-calibrated ROCmFP4 quantization of Qwen 3.5 122B model for high-memory AMD systems, achieving improved quality (14% lower KLD) and performance (28.45 tok/s). Requires ROCmFPX runtime; not compatible with stock llama.cpp.
A user seeks advice on choosing between a modded RTX 4090 48GB, dual AMD Radeon AI Pro R9700, or dual Intel Arc Pro B70 for running local coding LLMs, highlighting trade-offs in price, VRAM, software ecosystem, and inference speed.
AMD and Meta contributors ported PyTorch Monarch to AMD Instinct GPUs with ROCm, enabling fault-tolerant distributed training at scale. The blog details the engineering work and validation on large clusters.
The article describes the porting of PyTorch Monarch, a distributed training runtime, to AMD GPUs with ROCm, enabling single-controller fault-tolerant training at scale and addressing reliability challenges in large-scale LLM training.