Tag
Jun Kim, creator of oMLX, joins Hugging Face to support the MLX community, enhancing stability and development for local AI on Apple Silicon.
Laya-MLX is an open-weight tool for running typed decision AI models locally on Apple Silicon with low latency, providing native inference without cloud APIs. It includes benchmarks showing fast performance on devices like M3 Max.
Introducing laya-mlx, an open-source classification system optimized for Apple Silicon using MLX, which offers 50 times faster performance than Jev with a maximum of 1G memory usage, demonstrated through a real-time Snake game demo.
Release of Ternary-Bonsai-2-27B-gguf, a 27B-class language model using ternary weights for extreme compression (5.9 GB) while retaining 98.2% of FP16 intelligence, optimized for efficient inference on laptops and single GPUs.
A high-throughput inference engine for structured information extraction on Apple Silicon using MLX, offering parallel constrained decoding with 5.6x to 7.0x latency reductions and 100% schema validity.
NVIDIA Cosmos3, a 64B parameter image generation model, is released with INT4 quantization for local deployment on CUDA and MLX, with code and weights available and performance demonstrated on Apple Silicon.
An author improved MLX performance on an M3 Ultra Studio to achieve 73 tok/s for a 4-bit Qwen3.8-Flash-Next model, which shows intelligence scores matching GPT-5.6 Sol, highlighting the growing potential of local AI.
The user discusses preferences between Qwen 3.8 27b and Qwen Flash Next models on Apple hardware, comparing speeds, and inquires about improving performance with MLX and harnesses without reasoning.
A user questions the relevance of MLX on Apple M5 chips after llama.cpp updates for Metal improve GGUF model prefill performance, making MLX potentially redundant for certain tasks.
The article benchmarks the Qwen3.8-Flash-Next model, showing it breaks 94% on a personal benchmark and compares its performance in coding, general knowledge, and science against other models.
The discussion compares CUDA's maturity to MLX in the context of Apple's M5 Ultra Mac Studio, stressing the need for competition to optimize hardware for local AI, which could enable Apple to surpass NVIDIA.
Ornith 1.5 9B Abliterated, an experimental MLX derivative, is now available in 4-bit, 8-bit, and BF16 builds for Apple Silicon, designed to reduce refusal behavior while maintaining capabilities for research and legitimate local use.
LFM2.5 DSpark by liquidai is integrated into mlx-vlm v0.6.16, enabling up to 3.7× faster speculative decoding on M5 Max with zero output drift for on-device VLM inference.
Qwen3.8 27B AI model can now be run completely uncensored on Mac with no built-in guardrails, complying with harmful requests, and is available in various quantization formats.
A demonstration of the Gemma 4 E4B AI model running locally on an iPad using Apple MLX, showcased as an engaging application for children.
Qwen3.8-27B uncensored version released, optimized for Mac M chips, supports local deployment, retains multimodal capabilities and safety research features, with simplified installation steps.
An uncensored MLX build of Qwen's Qwen3.8-27B model, quantized for Apple Silicon, with safety alignment removed for research purposes.
The Qwen MLX Challenge is a competition for benchmarking AI models on Apple Silicon, with official scores and local iteration encouraged during validation.
The user indicates they are ready to use the Qwen 3.8 27b model, now Ollama directly supports Mac's MLX, making it very comfortable to use.
A developer reports running Meta's Muse Glimmer 30B up to ~3.3x faster on Apple Silicon using speculative decoding in mlx-dspark, with byte-identical output and no quality tradeoff.