Tag
Announces the release of a Config-I quantization of MiniMax-M3 on MLX, using 2-bit experts and 4-bit attention to reduce the 427B MoE model from 869GB to ~167GB, though the quant is untested and requires a patch for mlx_lm.
Modular's kernel team is optimizing serving for MiniMax M3's 1M-token context and native multimodality, with open weights dropping soon for immediate deployment on Modular.
MiniMax unveils MiniMax M3, the first open-weights AI model combining frontier capabilities in coding and agentic tasks, achieving strong benchmark scores with sparse attention scaling to 1M context.