Tag
The user requests serious benchmarks for the Mac M5 Ultra 256GB to evaluate its performance against Nvidia GPUs, criticizing the available influencer-driven content.
Jun Kim, creator of oMLX, joins Hugging Face to support the MLX community, enhancing stability and development for local AI on Apple Silicon.
The article reviews the Apple Mac Mini with the M6 chip, highlighting its performance enhancements and appeal for AI enthusiasts while discussing design and hardware details.
Extended Splash Engine to support native 8-bit Qwen3.8-27B on Apple Silicon, achieving 37-55 tok/s without quantization degradation and scaling up to 256k context.
Laya-MLX is an open-weight tool for running typed decision AI models locally on Apple Silicon with low latency, providing native inference without cloud APIs. It includes benchmarks showing fast performance on devices like M3 Max.
Introducing laya-mlx, an open-source classification system optimized for Apple Silicon using MLX, which offers 50 times faster performance than Jev with a maximum of 1G memory usage, demonstrated through a real-time Snake game demo.
Inco Splash is an open-source inference engine optimized for Apple silicon, offering significant speed improvements for running AI models like Qwen3.8-27B on M-series MacBooks.
Inco AI releases Splash, an open-source inference engine optimized for Apple silicon, claiming up to 3× faster decode speeds for local model serving, enabling agentic workloads on devices like M5 Max MacBook Pro.
Benchmark results for Apple's M5 Ultra and M6 chips show significant GPU performance gains, with the M5 Ultra achieving up to 59% higher Metal scores on Geekbench compared to the M3 Ultra, and the M6 chip scoring 34% higher than the M5.
Vinix is a modern, lightweight operating system written in V, featuring low RAM and disk usage, compatibility with Alpine Linux binaries, and optional app sandboxing, currently in alpha.
Prism ML released a ternary weight 27B-class AI model optimized for on-device use on Apple laptops, retaining 98.2% of full-precision intelligence with an 8.60 GB footprint and ~47 tok/s performance.
GameToMac leverages GPT-6 Astra AI to make Windows games playable on Apple Silicon Macs with high performance, supporting initial titles and promising broader compatibility soon.
The Ars Technica review of macOS 27 Golden Gate highlights its deep integration of Apple Intelligence, making AI unavoidable, while also delivering incremental improvements to the OS. It marks the end of Intel Mac support, requiring Apple Silicon with M1 or newer.
A high-throughput inference engine for structured information extraction on Apple Silicon using MLX, offering parallel constrained decoding with 5.6x to 7.0x latency reductions and 100% schema validity.
Cody Ho and Niklas built a fully OpenGL ES 3.0 compliant GPU driver for the M4 Mac Mini in one month through clean-room reverse engineering, enabling Linux to run graphics-intensive applications like Minecraft at high frame rates.
QuietHint is a meeting assistant for Mac that suggests real-time replies during calls, using on-device speech transcription and local processing to maintain privacy.
macOS 27 Golden Gate is now available for Apple Silicon Macs, introducing the Siri AI assistant, Liquid Glass UI improvements, and enhanced performance.
Kinesis is an open-source tool that allows users to control their Mac using gestures from Meta's Neural Band, including swiping between desktops and adjusting volume or brightness.
The article retrospectively reverse-engineers Apple's Neural Engine on the M1 chip, analyzing its architecture and discussing its evolution with the M5 chip's integration into GPU cores.
SiliconBench evaluates nine Apple Silicon LLM serving engines on speed, memory, and fidelity, finding that explicit memory budgets don't ensure headroom and only a few stacks meet all criteria for concurrency scaling and model coverage.