MiniMax-M3-EAGLE3-GGUF - Llama.cpp compatible MiniMax M3 EAGLE draft model!
Summary
A GGUF conversion of MiniMax M3's EAGLE draft model for llama.cpp is now available, enabling speculative decoding speedups on compatible hardware.
Similar Articles
EAGLE3 has landed in llama.cpp
EAGLE3, a speculative decoding method, has been integrated into llama.cpp, enabling faster inference.
GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-Thinking-GGUF
A GGUF quantized version of MiniCPM5-1B-Claude-Opus-Fable5-Thinking model is released on Hugging Face, with usage instructions for llama.cpp, vLLM, and Ollama.
unsloth/MiniMax-M3-GGUF
Unsloth releases a GGUF quantized version of the MiniMax-M3 multimodal model, enabling image-text-to-text tasks with support for Transformers, llama.cpp, vLLM, and other inference engines.
whats happening on llama.cpp
A significant update to llama.cpp requires all previously generated GGUF files to be regenerated, indicating a major breaking change to the model format.
unsloth/North-Mini-Code-1.0-GGUF · Hugging Face
This page hosts GGUF quantized versions of Cohere's North-Mini-Code-1.0 model, a 30B-A3B MoE model optimized for code generation and agentic tasks. Instructions are provided for building llama.cpp from a specific PR to support the cohere2moe architecture.