@jianchen1799: Local models can now handle agentic workloads. Local inference engines need to catch up. Today we’re releasing Splash: …

X AI KOLs Timeline Tools

Summary

Inco AI releases Splash, an open-source inference engine optimized for Apple silicon, claiming up to 3× faster decode speeds for local model serving, enabling agentic workloads on devices like M5 Max MacBook Pro.

Local models can now handle agentic workloads. Local inference engines need to catch up. Today we’re releasing Splash: model-optimized Apple silicon serving, powered by DFlash2. 2× faster than the next-fastest engine. Qwen3.8-27B: up to 140 tokens/s on M5 Max in real agentic serving.
Original Article
View Cached Full Text

Cached at: 09/19/26, 07:11 PM

Local models can now handle agentic workloads. Local inference engines need to catch up.

Today we’re releasing Splash: model-optimized Apple silicon serving, powered by DFlash2.

2× faster than the next-fastest engine. Qwen3.8-27B: up to 140 tokens/s on M5 Max in real agentic serving.

Inco AI (@inco_ai): Qwen3.8-27B at 144 tok/s on an M5 Max MacBook Pro ⚡

Meet Inco Splash: our open-source inference engine, built around the model and around Apple silicon.

Up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents.

Similar Articles

Localmaxxing (3 minute read)

TLDR AI

The article analyzes the viability of running AI inference locally on a MacBook Pro, comparing a local Qwen 35B model against the cloud-based Claude Opus 4.5. It concludes that local models are 2x faster for routine tasks, making them a practical choice for half of daily workloads despite a slight capability gap.