标签
一位作者改进了 M3 Ultra Studio 上的 MLX 性能,使 4 位量化的 Qwen3.8-Flash-Next 模型达到 73 tok/s,其智能评分与 GPT-5.6 Sol 相当,凸显了本地 AI 日益增长的潜力。