ExLlamav3 Recent Updates : CPU offload, GLM-5.3-FLASH, Qwen3.8-Flash, SC Quants ++

Reddit r/LocalLLaMA Tools

Summary

ExLlamav3 has released major updates including CPU offload for MoE experts, support for new AI models like GLM-5.3-Flash and Qwen-3.8-Flash, and various performance optimizations.

More new massive updates from turboderp: - CPU offload of MoE experts - Qwen-3.8-Flash-Next ngram disk offload - GLM-5.3-Flash - New self-calibrated optimization technique - Countless other optimizations and improvements If you have an NVIDIA card and haven't tried it lately, you might be missing out. The attached cat image was made with Qwen-3.8-Flash-Next-3.05bpw-exl3 and this prompt: Create a detailed SVG image of a cute kitten riding a magic turtle into space. Come join the crew at the exllama discord More frequent news on the exllama sub
Original Article

Similar Articles

ExLlamaV3 Major Updates!

Reddit r/LocalLLaMA

ExLlamaV3 has released a series of major updates including Gemma 4 support, improved caching efficiency, and the new DFlash technology for significantly faster inference speeds across various model categories.

GLM-5.3-Flash

Hacker News Top

Release of GLM-5.3-Flash, an AI language model optimized for fast inference and performance updates.