Tag
The article announces the release of Qwen3.8-Flash-Next on MLX-serve, supporting 1 million token context with efficient performance on M5 Max hardware using quantized weights.
A developer updates MLX-Serve, a fast local inference server for Apple Silicon, to support recent models like LiquidAI 2.6B, MiniMax H3 video generation, and DeepSeek V4 Flash, with AntLing 3.0-flash coming soon.