MLX 16/8/4/2-bit quants of nvidia/llama-embed-nemotron-8b
Summary
The user converted Nvidia's Llama-Embed-Nemotron-8B model to MLX format with fp16, 8-bit, 4-bit, and 2-bit quantizations, enabling in-process embedding loading on Apple Silicon via mlx-embeddings.
Similar Articles
SOTA ImageGen Locally NVIDIA Cosmos3(64B) INT4 quants CUDA/MLX
NVIDIA Cosmos3, a 64B parameter image generation model, is released with INT4 quantization for local deployment on CUDA and MLX, with code and weights available and performance demonstrated on Apple Silicon.
Qwen3.8-27B Hybrid IQ4_XS quantization for 16GB gang
This is a quantized version of the Qwen3.8-27B AI model using IQ4_XS quantization, optimized for 16GB RAM systems, with instructions for local deployment using various tools like llama.cpp and Ollama.
Qwen3.8-27B-Uncensored-MLX (4 minute read)
An uncensored MLX build of Qwen's Qwen3.8-27B model, quantized for Apple Silicon, with safety alignment removed for research purposes.
Qwen3.6-35B-A3B-Abliterated-Heretic-MLX-4bit
The user reviews a quantized and fine-tuned version of the Qwen3.6-35B model optimized for Apple Silicon via MLX, praising its speed, intelligence, and lack of safety disclaimers.
@no_stp_on_snek: Config-I quant of MiniMax-M3 is up on MLX. 2-bit experts, 4-bit attention, 8-bit boundaries + embeddings, f16 router. ~…
Announces the release of a Config-I quantization of MiniMax-M3 on MLX, using 2-bit experts and 4-bit attention to reduce the 427B MoE model from 869GB to ~167GB, though the quant is untested and requires a patch for mlx_lm.