DeepSeek V4 @ IQ3XXS on M1 Ultra 128GB- 16 tok/s in LM Studio after patch
Summary
A GitHub patch allows running DeepSeek V4 Flash in LM Studio on 128GB Macs by sideloading antirez's llama.cpp fork, working around struct-layout drift, decoding splits, and code-signing issues.
View Cached Full Text
Cached at: 08/03/26, 01:34 AM
noreff/lmstudio-dsv4-patch
Source: https://github.com/noreff/lmstudio-dsv4-patch
DeepSeek V4 Flash in LM Studio (Mac)
Sideload antirez’s deepseek_v4 llama.cpp fork into LM Studio so you can run DeepSeek V4 Flash — 256-expert MoE, ~80 GB resident, up to 1M token context — on a Mac with 128 GB unified memory.
This is a workaround. Upstream llama.cpp doesn’t yet support the deepseek_v4 architecture (discussion #22376). LM Studio’s MLX runtime rejects V4 GGUFs (bug #1872). Until those land, this is the only working path to drive V4-Flash from LM Studio’s UI / OpenAI-compatible server.
Tested on MacBook Pro M5 Max, 128 GB, macOS 26.4. Should work on M2/M3/M4 Max with 128 GB too.
What you get
- DeepSeek V4 Flash loaded as a normal model in LM Studio
- OpenAI-compatible API on
http://localhost:1234 - Full 1 048 576 token context
- ~190 tok/s prompt processing, ~26 tok/s generation
- Works with any OpenAI-API client (incl. pi, Open WebUI, your own scripts)
Why a patch is needed
LM Studio’s closed-source bridge (libllm_engine.dylib) was built against an older llama.h. When you swap in antirez’s newer libllama.dylib, the bridge crashes or produces garbage. The patch in this repo handles:
llama_context_paramsstruct-layout drift — bridge writes to old offsets, we read at new offsets, so most ctx fields arrive as garbage (the load-bearing one isyarn_attn_factor, which silently breaks RoPE on YARN-scaled models). We detect bridge calls byisnan(rope_freq_base)and override the struct.- Bridge chunks prompts into multiple
llama_decode()calls. V4’s compressed-attention indexer keeps state across calls and corrupts on a split, so we buffer chunks and re-submit them as one batch (with the right merge for image-embedding batches too, so vision models like Gemma 4 31B keep working). - Probe decodes that the bridge issues (
common_context_can_seq_rm) SIGBUS on V4’s hybrid memory — early-returned. - Metal
mpp::tensor_opskernels don’t resolve inside the LM Studio process — hardcodedhas_tensor=falseso the simdgroup fallback compiles. - PAC mismatch on
cplan->abort_callbackbetween Element Labs–signed arm64e bridge and ad-hoc-signed arm64 ggml-cpu — bypassed. - Missing ABI symbols that the bridge expects (
llama_memory_breakdown_print,llama_n_rs_seq,llama_get_embeddings_pre_norm*,mtmd_get_memory_usage) — added as no-op stubs.
Plus: code signing. LM Studio is hardened-runtime, Element Labs Developer ID, no disable-library-validation — refuses to load our dylibs out of the box. Re-sign LM Studio’s binaries with that entitlement added.
Recipe
You need: Xcode CLT, cmake, ~100 GB free disk, LM Studio installed.
# 1. Clone antirez's fork
git clone https://github.com/antirez/llama.cpp ~/llama-ds4
cd ~/llama-ds4
# 2. Apply this patch + drop in the two stub files
git apply /path/to/this/repo/patches/01-lms-bridge-shims.diff
cp /path/to/this/repo/patches/llama-lms-stubs.cpp src/
cp /path/to/this/repo/patches/mtmd-lms-stubs.cpp tools/mtmd/
# 3. Build shared libs with Metal
cmake -B build-shared \
-DBUILD_SHARED_LIBS=ON -DGGML_METAL=ON -DGGML_METAL_EMBED_LIBRARY=ON \
-DLLAMA_CURL=OFF
cmake --build build-shared --target llama llama-server -j
# 4. Clone the LM Studio runtime to a new sibling so original stays intact
RUNTIME_SRC=~/.lmstudio/extensions/backends/llama.cpp-mac-arm64-apple-metal-advsimd-2.17.0
RUNTIME=$RUNTIME_SRC # use the same folder if you're fine pinning this version
# (or copy to <suffix>-antirez and bump version in backend-manifest.json — see notes)
# 5. Overlay our dylibs into the runtime
for f in libllama libggml-base libggml-cpu libggml-blas libggml-metal libmtmd libllama-common; do
cp ~/llama-ds4/build-shared/bin/${f}.dylib "$RUNTIME/${f}.dylib"
install_name_tool -id @rpath/${f}.dylib "$RUNTIME/${f}.dylib"
done
# Rewrite version-suffixed deps (e.g. libfoo.0.dylib → libfoo.dylib)
for d in "$RUNTIME"/lib*.dylib; do
otool -L "$d" | awk '/@rpath.*\.0\..*\.dylib/ {print $1}' | while read dep; do
new=$(echo "$dep" | sed 's/\.0\.0\.1\.dylib/.dylib/;s/\.0\.dylib/.dylib/')
install_name_tool -change "$dep" "$new" "$d" 2>/dev/null
done
codesign --force --sign - "$d"
done
# 6. Replace llama-server binary in the runtime (LM Studio 2.16+ uses subprocess mode)
cp ~/llama-ds4/build-shared/bin/llama-server "$RUNTIME/llama-server"
codesign --force --sign - "$RUNTIME/llama-server"
# 7. Re-sign LM Studio binaries with disable-library-validation
# (one-time; survives across LM Studio app updates if you re-do after each update)
APP=/Applications/LM\ Studio.app
xattr -dr com.apple.quarantine "$APP"
# For each of: main "LM Studio", four Helpers, ~/.lmstudio/.internal/utils/{node,deno}
# and /Applications/LM Studio.app/Contents/Resources/app/.webpack/bin/{node,deno}:
# - extract entitlements: codesign -d --entitlements - --xml <bin> > ent.plist
# - add the key: PlistBuddy -c "Add :com.apple.security.cs.disable-library-validation bool true" ent.plist
# - resign: codesign --force --sign - --options runtime --entitlements ent.plist <bin>
# 8. Download the GGUF
huggingface-cli download antirez/deepseek-v4-gguf \
DeepSeek-V4-Flash-IQ2XXS-w2Q2K-AProjQ8-SExpQ8-OutQ8-chat-v2.gguf \
--local-dir ~/.lmstudio/models/antirez/deepseek-v4-gguf
Then in LM Studio’s runtime picker (⌘⇧R) make sure the 2.17.0 runtime is selected, and:
~/.lmstudio/bin/lms load antirez/deepseek-v4-flash \
--gpu max --context-length 1048576 --identifier dsv4 --exact
You’re done. The model answers via http://localhost:1234/v1/chat/completions like any other LM Studio model.
Tuning notes
The patch hardcodes these cparams (override via env vars LMS_UBATCH, LMS_THREADS_BATCH, LMS_KV_TYPE):
n_ubatch = 512— sweet spot on M5 Max 128 GB.1024+crashessched_reserveat 1M ctx (compute scratch).n_batch = min(max(n_ctx, 4096), 16384)— must cap; output buffer for prompt-eval logits isn_batch × vocab × 4 B(would be ~480 GB uncapped).n_threads_batch = 6— Metal is GPU-bound; doesn’t matter.type_k = type_v = F16— F32 is ~8% slower on generation (memory-bandwidth bound).kv_unified = true,swa_full = false,flash_attn = AUTO.
Caveats
- LM Studio auto-updates its bundled runtime. When the runtime version moves past
2.17.0, repeat steps 5–6 against the new folder, or pin2.17.0in⌘⇧R. - Long prompts (≫ 16 k tokens) hit the
n_batchcap and get flushed in 16 k chunks. V4’s hybrid state degrades slightly; output is still coherent. Most chat use is well below that. - This is a hack, not an officially supported extension path. See bug-tracker #1843 for the request to expose a real backend SDK.
Credit
This only exists because @antirez wrote the deepseek_v4 arch implementation in his llama.cpp fork and shipped the IQ2XXS GGUF. The patches here are just glue between his fork and LM Studio’s closed bridge.
License
MIT.
Similar Articles
DeepSeek V4.1F Q4 on M3 Ultra with native DSpark MTP (40tps / 800tps)
The author optimized DeepSeek V4.1 Flash for Apple M3 Ultra, achieving up to 40 t/s decode speed with DSpark speculative decoding while maintaining byte-identical accuracy to the upstream model.
How I made DeepSeek V4 Flash 12x faster on an M3 Ultra
The author achieved a 12x speedup for DeepSeek V4 Flash on a Mac Studio M3 Ultra by optimizing kernels and implementing effective caching strategies, reducing chat turn latency from 6-20 seconds to 1.6 seconds.
@antirez: I didn't expect DeepSeek v4 PRO (not Flash) to run well on the Mac Studio M3 Ultra with 512GB of RAM. This is 2 bit qua…
Antirez reports that DeepSeek v4 PRO runs well on a Mac Studio M3 Ultra with 512GB RAM using 2-bit quantization, achieving 130 t/s prefill and 13 t/s generation.
You can run Deepseek 4 flash on mac (M3 Max, 96gb)
A guide on running DeepSeek 4 flash on a Mac M3 Max with 96GB RAM using Antirez's ds4 engine and SSD streaming, achieving ~12 tokens/second inference speed.
DeepSeek v4 Flash 0731 4bit ~50tps prefill, ~1tps decode on M5 Air 32gb
A user shares experiments running a 4-bit quantized DeepSeek v4 Flash on a 32GB M5 MacBook Air, achieving roughly 50 tokens/s prefill and 1 token/s decode using streamed experts and other tricks.