mlx

Tag

Cards List
#mlx

Jun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX community

Hugging Face Blog · yesterday Cached

Jun Kim, creator of oMLX, joins Hugging Face to support the MLX community, enhancing stability and development for local AI on Apple Silicon.

0 favorites 0 likes
#mlx

@LinQ444: jev 和laya的对比 https://github.com/mizorewww/laya-mlx…

X AI KOLs Timeline · 2d ago Cached

Laya-MLX is an open-weight tool for running typed decision AI models locally on Apple Silicon with low latency, providing native inference without cloud APIs. It includes benchmarks showing fast performance on devices like M3 Max.

0 favorites 0 likes
#mlx

@mizorewww: Introducing a version 50 times faster than Jev, running on your device: laya-mlx! Occupies a maximum of 1G memory only …

X AI KOLs Timeline · 3d ago Cached

Introducing laya-mlx, an open-source classification system optimized for Apple Silicon using MLX, which offers 50 times faster performance than Jev with a maximum of 1G memory usage, demonstrated through a real-time Snake game demo.

0 favorites 0 likes
#mlx

prism-ml/Ternary-Bonsai-2-27B-gguf

Hugging Face Models Trending · 6d ago Cached

Release of Ternary-Bonsai-2-27B-gguf, a 27B-class language model using ternary weights for extreme compression (5.9 GB) while retaining 98.2% of FP16 intelligence, optimized for efficient inference on laptops and single GPUs.

0 favorites 0 likes
#mlx

harshatheg/Qwen-2.5-1B-RLCD

Hugging Face Models Trending · 2026-09-16 Cached

A high-throughput inference engine for structured information extraction on Apple Silicon using MLX, offering parallel constrained decoding with 5.6x to 7.0x latency reductions and 100% schema validity.

0 favorites 0 likes
#mlx

SOTA ImageGen Locally NVIDIA Cosmos3(64B) INT4 quants CUDA/MLX

Reddit r/LocalLLaMA · 2026-09-09

NVIDIA Cosmos3, a 64B parameter image generation model, is released with INT4 quantization for local deployment on CUDA and MLX, with code and weights available and performance demonstrated on Apple Silicon.

0 favorites 0 likes
#mlx

@ashxhart: Been giving my M3 Ultra Studio some love and tinkering with MLX. Started the afternoon at 42 tok/s. Now: 73 tok/s, and …

X AI KOLs Timeline · 2026-09-08 Cached

An author improved MLX performance on an M3 Ultra Studio to achieve 73 tok/s for a 4-bit Qwen3.8-Flash-Next model, which shows intelligence scores matching GPT-5.6 Sol, highlighting the growing potential of local AI.

0 favorites 0 likes
#mlx

Are you running Qwen 3.8 27b or Qwen Flash Next?

Reddit r/LocalLLaMA · 2026-09-07

The user discusses preferences between Qwen 3.8 27b and Qwen Flash Next models on Apple hardware, comparing speeds, and inquires about improving performance with MLX and harnesses without reasoning.

0 favorites 0 likes
#mlx

Mac Heads: Is there any point to MLX in September 2026?

Reddit r/LocalLLaMA · 2026-09-02

A user questions the relevance of MLX on Apple M5 chips after llama.cpp updates for Metal improve GGUF model prefill performance, making MLX potentially redundant for certain tasks.

0 favorites 0 likes
#mlx

Qwen3.8-Flash-Next: Time to Update Those Benchmarks

Reddit r/LocalLLaMA · 2026-08-27

The article benchmarks the Qwen3.8-Flash-Next model, showing it breaks 94% on a personal benchmark and compares its performance in coding, general knowledge, and science against other models.

0 favorites 0 likes
#mlx

@TheAhmadOsman: To be clear, CUDA is still more mature than MLX if you’re comparing GPUs to the new M5 Ultra Mac Studio However, the ec…

X AI KOLs Timeline · 2026-08-26 Cached

The discussion compares CUDA's maturity to MLX in the context of Apple's M5 Ultra Mac Studio, stressing the need for competition to optimize hardware for local AI, which could enable Apple to surpass NVIDIA.

0 favorites 0 likes
#mlx

@trevorwood222: Ornith 1.5 9B Abliterated is now available in MLX for Apple Silicon 4-bit, 8-bit and BF16 builds are live. Tuned 4-bit …

X AI KOLs Timeline · 2026-08-22 Cached

Ornith 1.5 9B Abliterated, an experimental MLX derivative, is now available in 4-bit, 8-bit, and BF16 builds for Apple Silicon, designed to reduce refusal behavior while maintaining capabilities for research and legitimate local use.

0 favorites 0 likes
#mlx

@Prince_Canuma: LFM2.5 DSpark by @liquidai is coming to mlx-vlm in v0.6.16 Exact speculative decoding on M5 Max, delivering up to 3.7× …

X AI KOLs Following · 2026-08-21 Cached

LFM2.5 DSpark by liquidai is integrated into mlx-vlm v0.6.16, enabling up to 3.7× faster speculative decoding on M5 Max with zero output drift for on-device VLM inference.

0 favorites 0 likes
#mlx

@LuminaBench: Qwen3.8 27B can now be run completely uncensored It will apparently comply with harmful, unethical, offensive or even i…

X AI KOLs Timeline · 2026-08-18 Cached

Qwen3.8 27B AI model can now be run completely uncensored on Mac with no built-in guardrails, complying with harmful requests, and is available in various quantization formats.

0 favorites 0 likes
#mlx

@googlegemma: Check out original post by @ivanfioravanti here:

X AI KOLs Following · 2026-08-18 Cached

A demonstration of the Gemma 4 E4B AI model running locally on an iPad using Apple MLX, showcased as an engaging application for children.

0 favorites 0 likes
#mlx

@Lonely__MH: Unleashed! The uncensored version of Qwen3.8-27B with safety restrictions removed is here! Kudos to the community for the speed! Deeply optimized for Mac M chips! I see everyone discussing the DGX Spark deployment for ling-3.0-flash, and many people's first reaction is that the compute power is too expensive to buy. Since cloud costs are high...

X AI KOLs Timeline · 2026-08-18 Cached

Qwen3.8-27B uncensored version released, optimized for Mac M chips, supports local deployment, retains multimodal capabilities and safety research features, with simplified installation steps.

0 favorites 0 likes
#mlx

Qwen3.8-27B-Uncensored-MLX (4 minute read)

TLDR AI · 2026-08-18 Cached

An uncensored MLX build of Qwen's Qwen3.8-27B model, quantized for Apple Silicon, with safety alignment removed for research purposes.

0 favorites 0 likes
#mlx

@no_stp_on_snek: wow look at the gains on metal!!!

X AI KOLs Following · 2026-08-17 Cached

The Qwen MLX Challenge is a competition for benchmarking AI models on Apple Silicon, with official scores and local iteration encouraged during validation.

0 favorites 0 likes
#mlx

@akazwz_: Ready to start using it, Qwen 3.8 27b, now Ollama directly supports Mac's MLX which is very nice.

X AI KOLs Following · 2026-08-15 Cached

The user indicates they are ready to use the Qwen 3.8 27b model, now Ollama directly supports Mac's MLX, making it very comfortable to use.

0 favorites 0 likes
#mlx

Meta's Muse Glimmer 30B now runs up to ~3.3x faster on Mac with mlx-dspark

Reddit r/LocalLLaMA · 2026-08-12

A developer reports running Meta's Muse Glimmer 30B up to ~3.3x faster on Apple Silicon using speculative decoding in mlx-dspark, with byte-identical output and no quality tradeoff.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback