@bstnxbt: dflash-mlx v0.1.6 is out. Biggest agentic update so far: ► much more usable for real OpenCode / coding-agent sessions ►…
Summary
dflash-mlx v0.1.6 is released with major agentic improvements, including adaptive verification, custom kernels, prefix cache improvements, and broader compatibility with agentic coding tools like OpenCode, aider, and Continue.
Similar Articles
@Prince_Canuma: Today we're shipping our biggest MLX-VLM release yet: v0.6.0 ...and we are raising This one's about turning your Apple …
MLX-VLM v0.6.0 is released, adding speculative decoding, an agent-ready server compatible with Anthropic's API, new models (DeepSeek V4, ZAYA1-VL, etc.), image generation/editing, and audio input support, enabling local AI agents on Apple devices.
@Raullen: Rapid-MLX 0.11.0 is out! Making local models on Apple Silicon reliable enough to run your agent workflows, not just dem…
Rapid-MLX 0.11.0 brings major performance gains with prefix-cache and response caching, supports new model families including HY3 295B MoE and Qwen3-Coder-Next 80B, introduces structured output with guaranteed valid tool calls, and adds seamless integration with MCP servers for autonomous agent workflows.
@bstnxbt: DFlash v0.1.4 : custom Metal verify kernels for quantized Qwen3 hybrid models, plus significant peak memory reduction a…
DFlash v0.1.4 releases custom Metal verify kernels for quantized Qwen3 hybrid models with significant peak memory reduction and 2.2x throughput improvements at long context on M5 Max GPUs.
@jundotkim: oMLX 0.3.9rc1 released. Highlights: - Low-memory Macs stay stable instead of getting killed by the OS - DFlash bumped t…
oMLX 0.3.9rc1, an LLM inference server optimized for Apple Silicon Macs, adds low-memory stability, chunked prefill, multi-tasking admin chat, and more.
@jundotkim: oMLX 0.3.9.dev2 released. Highlights: - Gemma 4 MTP on the vision path (thanks to @Prince_Canuma's mlx-vlm). Image+text…
oMLX 0.3.9.dev2 is released with improved Gemma 4 support, DFlash engine integration, and ParoQuant capabilities for local LLM inference on Apple Silicon.