ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++
Summary
ExLlamav3 has released major updates including CPU offload for MoE experts, support for new AI models like GLM-5.3-Flash and Qwen-3.8-Flash, and various performance optimizations.
Similar Articles
ExLlamaV3 Major Updates!
ExLlamaV3 has released a series of major updates including Gemma 4 support, improved caching efficiency, and the new DFlash technology for significantly faster inference speeds across various model categories.
ExLlamaV3 v1.0.0 - Major Performance Upgrades
ExLlamaV3 v1.0.0 brings major performance upgrades for running large language models locally.
BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)
BeeLlama.cpp is a performance-focused fork of llama.cpp that introduces DFlash speculative decoding and TurboQuant KV-cache compression, enabling high-speed local inference of large models like Qwen 3.6 27B on consumer hardware.
GLM-5.3-Flash
Release of GLM-5.3-Flash, an AI language model optimized for fast inference and performance updates.
llama.cpp support for Qwen3.8-Flash-Next has been merged
Support for the Qwen3.8-Flash-Next model has been merged into llama.cpp, enhancing its capabilities for local LLM inference in C/C++.