DeepSeek V4 @ IQ3XXS on M1 Ultra 128GB- 16 tok/s in LM Studio after patch

Reddit r/LocalLLaMA Tools

Summary

A GitHub patch allows running DeepSeek V4 Flash in LM Studio on 128GB Macs by sideloading antirez's llama.cpp fork, working around struct-layout drift, decoding splits, and code-signing issues.

No content available
Original Article
View Cached Full Text

Cached at: 08/03/26, 01:34 AM

noreff/lmstudio-dsv4-patch

Source: https://github.com/noreff/lmstudio-dsv4-patch

DeepSeek V4 Flash in LM Studio (Mac)

Sideload antirez’s deepseek_v4 llama.cpp fork into LM Studio so you can run DeepSeek V4 Flash — 256-expert MoE, ~80 GB resident, up to 1M token context — on a Mac with 128 GB unified memory.

This is a workaround. Upstream llama.cpp doesn’t yet support the deepseek_v4 architecture (discussion #22376). LM Studio’s MLX runtime rejects V4 GGUFs (bug #1872). Until those land, this is the only working path to drive V4-Flash from LM Studio’s UI / OpenAI-compatible server.

Tested on MacBook Pro M5 Max, 128 GB, macOS 26.4. Should work on M2/M3/M4 Max with 128 GB too.

What you get

  • DeepSeek V4 Flash loaded as a normal model in LM Studio
  • OpenAI-compatible API on http://localhost:1234
  • Full 1 048 576 token context
  • ~190 tok/s prompt processing, ~26 tok/s generation
  • Works with any OpenAI-API client (incl. pi, Open WebUI, your own scripts)

Why a patch is needed

LM Studio’s closed-source bridge (libllm_engine.dylib) was built against an older llama.h. When you swap in antirez’s newer libllama.dylib, the bridge crashes or produces garbage. The patch in this repo handles:

  1. llama_context_params struct-layout drift — bridge writes to old offsets, we read at new offsets, so most ctx fields arrive as garbage (the load-bearing one is yarn_attn_factor, which silently breaks RoPE on YARN-scaled models). We detect bridge calls by isnan(rope_freq_base) and override the struct.
  2. Bridge chunks prompts into multiple llama_decode() calls. V4’s compressed-attention indexer keeps state across calls and corrupts on a split, so we buffer chunks and re-submit them as one batch (with the right merge for image-embedding batches too, so vision models like Gemma 4 31B keep working).
  3. Probe decodes that the bridge issues (common_context_can_seq_rm) SIGBUS on V4’s hybrid memory — early-returned.
  4. Metal mpp::tensor_ops kernels don’t resolve inside the LM Studio process — hardcoded has_tensor=false so the simdgroup fallback compiles.
  5. PAC mismatch on cplan->abort_callback between Element Labs–signed arm64e bridge and ad-hoc-signed arm64 ggml-cpu — bypassed.
  6. Missing ABI symbols that the bridge expects (llama_memory_breakdown_print, llama_n_rs_seq, llama_get_embeddings_pre_norm*, mtmd_get_memory_usage) — added as no-op stubs.

Plus: code signing. LM Studio is hardened-runtime, Element Labs Developer ID, no disable-library-validation — refuses to load our dylibs out of the box. Re-sign LM Studio’s binaries with that entitlement added.

Recipe

You need: Xcode CLT, cmake, ~100 GB free disk, LM Studio installed.

# 1. Clone antirez's fork
git clone https://github.com/antirez/llama.cpp ~/llama-ds4
cd ~/llama-ds4

# 2. Apply this patch + drop in the two stub files
git apply /path/to/this/repo/patches/01-lms-bridge-shims.diff
cp /path/to/this/repo/patches/llama-lms-stubs.cpp     src/
cp /path/to/this/repo/patches/mtmd-lms-stubs.cpp      tools/mtmd/

# 3. Build shared libs with Metal
cmake -B build-shared \
  -DBUILD_SHARED_LIBS=ON -DGGML_METAL=ON -DGGML_METAL_EMBED_LIBRARY=ON \
  -DLLAMA_CURL=OFF
cmake --build build-shared --target llama llama-server -j

# 4. Clone the LM Studio runtime to a new sibling so original stays intact
RUNTIME_SRC=~/.lmstudio/extensions/backends/llama.cpp-mac-arm64-apple-metal-advsimd-2.17.0
RUNTIME=$RUNTIME_SRC   # use the same folder if you're fine pinning this version
# (or copy to <suffix>-antirez and bump version in backend-manifest.json — see notes)

# 5. Overlay our dylibs into the runtime
for f in libllama libggml-base libggml-cpu libggml-blas libggml-metal libmtmd libllama-common; do
  cp ~/llama-ds4/build-shared/bin/${f}.dylib "$RUNTIME/${f}.dylib"
  install_name_tool -id @rpath/${f}.dylib    "$RUNTIME/${f}.dylib"
done
# Rewrite version-suffixed deps (e.g. libfoo.0.dylib → libfoo.dylib)
for d in "$RUNTIME"/lib*.dylib; do
  otool -L "$d" | awk '/@rpath.*\.0\..*\.dylib/ {print $1}' | while read dep; do
    new=$(echo "$dep" | sed 's/\.0\.0\.1\.dylib/.dylib/;s/\.0\.dylib/.dylib/')
    install_name_tool -change "$dep" "$new" "$d" 2>/dev/null
  done
  codesign --force --sign - "$d"
done

# 6. Replace llama-server binary in the runtime (LM Studio 2.16+ uses subprocess mode)
cp ~/llama-ds4/build-shared/bin/llama-server "$RUNTIME/llama-server"
codesign --force --sign - "$RUNTIME/llama-server"

# 7. Re-sign LM Studio binaries with disable-library-validation
#    (one-time; survives across LM Studio app updates if you re-do after each update)
APP=/Applications/LM\ Studio.app
xattr -dr com.apple.quarantine "$APP"
# For each of: main "LM Studio", four Helpers, ~/.lmstudio/.internal/utils/{node,deno}
# and /Applications/LM Studio.app/Contents/Resources/app/.webpack/bin/{node,deno}:
#   - extract entitlements:  codesign -d --entitlements - --xml <bin> > ent.plist
#   - add the key:           PlistBuddy -c "Add :com.apple.security.cs.disable-library-validation bool true" ent.plist
#   - resign:                codesign --force --sign - --options runtime --entitlements ent.plist <bin>

# 8. Download the GGUF
huggingface-cli download antirez/deepseek-v4-gguf \
  DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2.gguf \
  --local-dir ~/.lmstudio/models/antirez/deepseek-v4-gguf

Then in LM Studio’s runtime picker (⌘⇧R) make sure the 2.17.0 runtime is selected, and:

~/.lmstudio/bin/lms load antirez/deepseek-v4-flash \
  --gpu max --context-length 1048576 --identifier dsv4 --exact

You’re done. The model answers via http://localhost:1234/v1/chat/completions like any other LM Studio model.

Tuning notes

The patch hardcodes these cparams (override via env vars LMS_UBATCH, LMS_THREADS_BATCH, LMS_KV_TYPE):

  • n_ubatch = 512 — sweet spot on M5 Max 128 GB. 1024+ crashes sched_reserve at 1M ctx (compute scratch).
  • n_batch = min(max(n_ctx, 4096), 16384) — must cap; output buffer for prompt-eval logits is n_batch × vocab × 4 B (would be ~480 GB uncapped).
  • n_threads_batch = 6 — Metal is GPU-bound; doesn’t matter.
  • type_k = type_v = F16 — F32 is ~8% slower on generation (memory-bandwidth bound).
  • kv_unified = true, swa_full = false, flash_attn = AUTO.

Caveats

  • LM Studio auto-updates its bundled runtime. When the runtime version moves past 2.17.0, repeat steps 5–6 against the new folder, or pin 2.17.0 in ⌘⇧R.
  • Long prompts (≫ 16 k tokens) hit the n_batch cap and get flushed in 16 k chunks. V4’s hybrid state degrades slightly; output is still coherent. Most chat use is well below that.
  • This is a hack, not an officially supported extension path. See bug-tracker #1843 for the request to expose a real backend SDK.

Credit

This only exists because @antirez wrote the deepseek_v4 arch implementation in his llama.cpp fork and shipped the IQ2XXS GGUF. The patches here are just glue between his fork and LM Studio’s closed bridge.

License

MIT.

Similar Articles

How I made DeepSeek V4 Flash 12x faster on an M3 Ultra

Reddit r/LocalLLaMA

The author achieved a 12x speedup for DeepSeek V4 Flash on a Mac Studio M3 Ultra by optimizing kernels and implementing effective caching strategies, reducing chat turn latency from 6-20 seconds to 1.6 seconds.