Tag
Vision support for the Minimax-M3 model has been merged into the llama.cpp project, enabling multimodal inference for this model locally.
Minimax M3 support with MSA has been merged into llama.cpp, enabling inference for the Minimax M3 model using the MSA architecture.
Reports a peak throughput of 19 tokens per second for the Minimax M3 model running on 8-16 MI50 GPUs.
Open-weights models have caught up with proprietary ones, with GLM 5.2 achieving near Opus-level scores in browser agent tasks at low cost. Other models like Minimax M3 and Kimi k2.7 also show notable improvements.
A pruned and quantized version of MiniMax-M3 (MiniMax-M3-Medium-JANG_2L) optimized to run on 128GB Macs using vMLX, featuring 32% expert pruning and JANG_2L mixed-precision quantization to fit within ~105 GB.
A product manager shares hands-on testing of Minimax M3's 1M context window on a real Q3 strategic brief, noting strong source attribution up to ~200K tokens but synthesis degradation beyond that.
A test of the open-weight MiniMax M3 model using MLX-VLM on a Mac Studio shows it can autonomously fill out a US customs form from a driver's license photo and a scanned document, using tool calls for fields, checkboxes, and signature.
Minimax-M3 is demonstrated running on 4x RTX Pro 6000 GPUs with 800k context, achieving 70-120 tok/s inference and 2000 tok/s prefill at 4x concurrency using 376GB VRAM in mxfp4 format.
A hands-on comparison of Kimi K2.6 and Minimax M3 in real agent workflows shows M3 costs roughly 5x less while delivering nearly identical quality, making it more cost-effective for production systems.
Unsloth is uploading a GGUF quantized version of the MiniMax M3 model to Hugging Face.
Unsloth releases a GGUF quantized version of the MiniMax-M3 multimodal model, enabling image-text-to-text tasks with support for Transformers, llama.cpp, vLLM, and other inference engines.
A user inquires about the upcoming open-source Minimax M3 model's performance in agentic tasks and coding, asking how it compares to older GPT models like GPT 5.2.
MiniMax released M3, a model with a 1M-token context window and native multimodal input, via API. The company promises open-weight release and a technical report within 10 days.
MiniMax announced MiniMax-M3, an open-weights model combining frontier coding and agentic capabilities with sparse attention scaling to 1M context, set to arrive on HuggingFace next week.