@BrianRoemmele: Another day and another full frontier model running on your computer. Been teething DeepSeek V4 Flash on over 60 employ…
摘要
Brian Roemmele reports that DeepSeek V4 Flash (304B, 1M context) now runs locally on Apple Silicon via the ds4 engine, sharing GGUF quantized builds with a fresh imatrix. The Hugging Face repo provides installation instructions and notes that these files are ds4-specific, not for llama.cpp.
查看缓存全文
缓存时间: 2026/08/03 07:39
Another day and another full frontier model running on your computer.
Been teething DeepSeek V4 Flash on over 60 employees at The Zero Human Company for a few hours and it is stunning. Between Kimi K3 and this nearly 80% of most folks needs is covered locally.
ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4 · Hugging Face
Source: https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#deepseek-v4-flash-0731–gguf-for-ds4-mixed-24-bitDeepSeek-V4-Flash-0731 — GGUF for ds4 (mixed 2+4 bit)
GGUF builds of**deepseek-ai/DeepSeek-V4-Flash-0731— the official release of DeepSeek V4 Flash (304B total parameters, hybrid CSA+HCA attention, 1M context) — for theds4 / DwarfStar**inference engine. Runs fully resident on a single 128 GB Apple Silicon machine.
Quantized in the asymmetric style ofantirez/deepseek-v4-gguf: crush the routed experts (they are almost all of the weights), keep the decision-making parts high precision.The filename is the spec.
📊The imatrix is published here too, not just the models it produced:
imatrix/DeepSeek\-V4\-Flash\-0731\-chat\-v2\-routed\-moe\-ds4\-1p5m\.dat— 430 MB, collectedon the0731weights themselves, 2729 prompts / 1.5M tokens / 387M routed-expert observations, 129 entries (43 layers × gate/up/down, full coverage, zero missing tensors). Reuse it withdeepseek4\-quantize \-\-imatrixto build your own mix, or to check mine.
**This is not my recipe.**The recipe, the quantizer (
gguf\-tools/deepseek4\-quantize), the imatrix pipeline and the engine are allantirezand the DwarfStar contributors. All I did was run that toolchain against the newer0731checkpoint and collect a fresh imatrix on those weights.
✅antirez has now shipped official
0731builds(2026-08-01), including\.\.\.Layers37\-42Q4KExperts\-\.\.\.\-imatrix\-fixed\-0731\.gguf— the same recipe as the file here.Get them fromantirez/deepseek-v4-gguf, that is the reference. What this repo still adds:the imatrix collected on the0731weights, published as a file. As of writing, hisimatrix/folder publishes only the preview-checkpoint\.dat(last touched 2026-05-12; thefixedin his filenames refers to a May “fixed routed-mid imatrix build”, not to anything0731-specific). What he actually built his0731files with is not something I can verify — he may well have collected one without publishing it. So: this is the only0731-collected imatrix I know to beavailable as a file, not a claim that his builds lack one. And readthe analysis belowbefore assuming it matters much: measured on general text, using it changes nothing significant, and there is a structural reason why.
⚠️**Needs ds4 / DwarfStar.**This is not a generic GGUF. It will not load in llama.cpp, Ollama or LM Studio — the tensor layout, quant mix and metadata are specific to the DS4 engine.
Not affiliated with DeepSeek or antirez. Weights © DeepSeek, released under MIT.
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#installationInstallation
git clone https://github.com/antirez/ds4
cd ds4
make # macOS Metal
Then download one of the files below intogguf/and point\-mat it.
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#filesFiles
FileSizeimatrixstatusDeepSeek\-V4\-Flash\-0731\-Layers37\-42Q4KExperts\-OtherExpertLayersIQ2XXSGateUp\-Q2KDown\-AProjQ8\-SExpQ8\-OutQ8\-chat\-v2\.gguf97.6 GB—✅ availableDeepSeek\-V4\-Flash\-0731\-Layers37\-42Q4KExperts\-OtherExpertLayersIQ2XXSGateUp\-Q2KDown\-AProjQ8\-SExpQ8\-OutQ8\-chat\-v2\-imatrix\.gguf97.6 GB✅ collected on0731✅ availableimatrix/DeepSeek\-V4\-Flash\-0731\-chat\-v2\-routed\-moe\-ds4\-1p5m\.dat430 MB—✅ available
**Either works.**On held-out wikitext the two are statistically indistinguishable (−1.36 %, p ≈ 0.098 —full numbers). The imatrix build is the one to prefera priori, since its statistics come from the0731weights, but I have no measurement proving it on general text. The plain build is the intermediate the imatrix was collected on, published so the comparison can be reproduced rather than taken on trust.
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#usageUsage
./ds4 -m gguf/DeepSeek-V4-Flash-0731-...-imatrix.gguf \
-p "Explain Redis streams in one paragraph."
./ds4-server --ctx 100000 --kv-disk-dir /tmp/ds4-kv --kv-disk-space-mb 8192
Sampling, per DeepSeek:temperature = 1\.0,top\_p = 0\.95for agentic scenarios,top\_p = 1\.0otherwise.
0731introduces areasoning\_effortparameterwith three levels —low,high,max— controlling how much the model deliberates before answering. The preview’s card documented no such parameter (it exposedthinking\_modeonly), so this is an addition.0731is also reported to spend substantially more tokens thinking than the preview; budget context and output limits accordingly, or use\-\-nothink.
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#quantization-recipeQuantization recipe
Unchanged from antirez’s mixed 2+4 bit build:
Tensor classQuantNotesblk\.37\.\.42\.ffn\_\{gate,up,down\}\_exps``Q4\_Klast 6 layers, most sensitiveblk\.\*\.ffn\_\{gate,up\}\_exps``IQ2\_XXSrouted expert gate/upblk\.\*\.ffn\_down\_exps``Q2\_KK-quant:downis far more sensitiveblk\.\*\.ffn\_\{gate,up,down\}\_shexp``Q8\_0shared expertsblk\.\*\.attn\_q\_a/q\_b/kv/output\_a/output\_b``Q8\_0MLA + low-rank outputoutput\.weight``Q8\_0``token\_embd\.weight``F16``blk\.\*\.ffn\_gate\_inp``F16router — never touch thisrouter bias,attn\_sinks, all\*\_norm``F32``blk\.\*\.ffn\_gate\_tid2eid``I32hash-routing tables, first 3 layersattn\_compressor\_\*,indexer\_\*,hc\_\*``F16/F32DSv4-specific blocks
Routed experts are the overwhelming majority of the parameters, but each individual expert only sees a fraction of the tokens — so aggressive quantization there costs little on average. Everything every token passes through stays high precision.
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#how-this-was-builtHow this was built
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#the-source-weights-are-already-4-bitThe source weights are already 4-bit
Worth stating plainly, because it is easy to assume otherwise:DeepSeek never published a BF16 version of this model.The routed experts ship at FP4 because that is how they weretrained— the paper (2606.19348, §5.2.1) appliesFP4 quantization-aware trainingto the MoE expert weights and the indexer QK path during post-training, and states that “the routed expert parameters utilize FP4 precision”.
Reading the safetensors headers across all 48 shards:
dtypetensorsbyteswhat it isI835 328148.18 GBFP4 weights,two per byteF8\_E8M035 7189.26 GBblock scales, one per 32 weightsF8\_E4M33906.30 GBBF164452.97 GBattention / shared / embeddingsF324330.15 GBnorms, biasestotal72 317166.88 GB
Declared shapes are thepackedshapes. Taken literally they sum to 165B parameters. Unpacking gives the real count — note these arebytes on the left, parameters on the right:
148.18 GB of I8 -> 148.18e9 bytes x 2 weights/byte = 296.36e9 FP4 weights
+ 1.48e9 BF16 + 6.30e9 F8_E4M3
= 304.14e9 parameters
which is the advertised count.
**The scale count confirms the unpacking independently.**IfI8really holds two FP4 weights per byte in blocks of 32, there must be296\.36e9 / 32 = 9\.26e9block scales — andF8\_E8M0is exactly 9.26 GB, i.e. 9.26e9 one-byte scales. The two numbers are derived from different fields of the header and agree, so this isn’t a guess about the layout.
Two different bit figures follow, and they should not be confused:
FP4 expert path4.25 bits/weight— 4 bits + one 8-bit scale per 32 weightswhole model on disk4.39 bits/param—166\.88 GB × 8 / 304\.1B, including the BF16/F32 tensors that were never quantized
So this is a4-bit → 2-bit requantization: one lossy step, not two. The tensors crushed toIQ2\_XXSare exactly the ones DeepSeek trained at FP4. The quantizer unpacks FP4/FP8 before re-encoding.
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#two-passes-because-the-imatrix-needs-a-running-modelTwo passes, because the imatrix needs a running model
Collecting an imatrix meansrunningthe model — the DwarfStar collector hooks the layer-major Metal prefill graph and accumulatessum\(x\[column\]^2\)per routed expert. That needs a loadable GGUF, which doesn’t exist yet before the first quantization:
- Quantize without imatrix→ the plain file above.
- Run itover the calibration corpus to collect the imatrix.
- Re-quantize with
\-\-imatrix→ the final file.
Budget disk for it: the 167 GB of source safetensors have to stay available for step 3, and each GGUF is 97.6 GB, so plan for ~360 GB free if you want to keep both builds around. Drop the intermediate after step 2 if you don’t.
antirez collects on the Q4 build for Flash, but Q4 for0731is ~165 GB and won’t stay resident on 128 GB. So this imatrix was collected on the2-bitbuild — which is exactly what the upstream README does for DeepSeek-V4-Pro, for the same reason.
That’s defensible because of the recipe itself: the router isF16, attention projections and shared experts areQ8\_0. Only the routed experts are degraded, so therouting decisionsobserved during collection are essentially the full-precision model’s.
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#why-a-fresh-imatrix-and-not-the-previewsWhy a fresh imatrix and not the preview’s
The collector records, fordowntensors, the routed SwiGLU rowafter route weighting— so the statistics depend directly on which experts fire.0731moves a lot versus the preview on exactly that axis (DeepSWE 7.3 → 54.4, Terminal-Bench 61.8 → 82.7), and the calibration corpus is heavily agent/code weighted (1106 of 4690 prompts areagent, 2074 aresource). A stale imatrix would be most wrong precisely onffn\_down\_exps— the tensor the recipe protects withQ2\_Kbecause it’s the most fragile. So it was recollected.
**That reasoning turned out to be only half right.**It holds forffn\_down\_exps, but the collector records something else entirely forgate/up— seewhy the imatrix buys so little, which measures the consequence and largely deflates this argument.
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#calibration-corpusCalibration corpus
Upstream’s, unmodified —gguf\-tools/imatrix/dataset/rendered\_prompts\.txtfromantirez/ds4. Not mine, and not regenerated:
ds4 commit54b36edlast commit touching the datasetb166a73``rendered\_prompts\.txt12 MB, sha2561159b0e7f1eff1c9…prompts4690tokens~2.92M (bytes/4 estimate, permanifest\.json)
Category mix, straight from upstream’smanifest\.json:source2074,agent1106,language1024,translation180,eval\_reasoning150,programming48,general40,long\_context36,algorithms32.
Use the tracked file, don’t regenerate it.
build\_ds4\_imatrix\_dataset\.pybuilds part of the corpusfrom the ds4 repository’s own C/Metal sources, so regenerating at a different commit yields a different corpus and a non-comparable imatrix. The tracked file was used verbatim, which is what makes this reproducible: check out54b36ed, verify the sha256, and you have the same input.
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#why-the-hub-sidebar-says-284bWhy the Hub sidebar says 284B
The model tree widget reads the GGUF metadata, which comes from the\-\-templatefile — a preview-checkpoint build. So it reports284B paramsand architecturedeepseek4, the preview’s numbers. The tensor data is0731(that is what\-\-compare\-tensorverifies), and the 304.1B figure derived from the safetensors headers above is the correct one for this checkpoint. Fixing the sidebar would mean rewriting the metadata block; it does not affect inference.
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#template-ggufTemplate GGUF
deepseek4\-quantizeregenerates tensor bytes from safetensors but takes metadata, tokenizer, tensor order and logical shapes from an existing DS4 GGUF passed as\-\-template. The preview-checkpoint GGUF was used. Safe here because the tokenizer isbyte-identicalacross the two releases —tokenizer\.json,tokenizer\_config\.jsonandgeneration\_config\.jsonshare the same blob hashes on the Hub. The onlyconfig\.jsondifferences are four added DSpark keys (dspark\_block\_size,dspark\_noise\_token\_id,dspark\_target\_layer\_ids,dspark\_markov\_rank).
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#pre-flight-checksPre-flight checks
Run before writing anything, as the upstream README requires:
--dspark-manifest dspark_stages=3 unknown_dspark_tensors=0
--dry-run n_tensors=1328 type_changes=0 97 591 747 168 bytes
--compare-tensor blk.0.attn_q_a.weight → bytes match, hashes differ
type\_changes=0proves the recipe was reproduced exactly from the template. The\-\-compare\-tensormismatch is theexpectedresult and the whole point of the check: identical byte counts prove the shape and target type are right, differing hashes prove the new0731weights are actually being read and not the template’s.
0731also carries4 705mtp\.\*tensorsfor the DSpark speculative decoding module that the preview template knows nothing about. The dry-run confirms they’re correctly excluded from the main model (1328 tensors out). Convert them separately with\-\-dspark\-supportif you want\-\-dspark.
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#does-the-imatrix-actually-helpDoes the imatrix actually help?
**Honest answer: not measurably, on general text.**Three runs, each widening the scored window on the same held-out corpus:
ctxtokens scoredppl plainppl imatrixdeltasignificance5124805.5804905.250970−5.90 %0.9 σ819281605.0202384.847609−3.44 %2.1 σ32768327364.8480564.782034−1.36 %**1.7 σ, p ≈ 0.098 The gapshrinks as the sample grows**— −5.90 → −3.44 → −1.36 %. That is the signature of an effect collapsing toward zero, not of a real one being measured more precisely. At 32k tokens it is not significant at the conventional threshold.
So the imatrix build isnot demonstrably better on Wikipedia prose. It is not worse either. If you were hoping for a number that justifies picking one file over the other on general text, this measurement does not provide it, and I am not going to dress it up.
Method, so it can be repeated:
./ds4 -m <build>.gguf --perplexity-file wikitext2_test_300kb.txt -c 32768
- Corpus: wikitext-2raw testsplit, first 300 KB. Deliberatelynot
rendered\_prompts\.txt— the imatrix was collected on that, measuring there would be circular. Verified: zero overlap with the calibration corpus. - Identical file, context length and tokenizer on both runs, 32736 tokens scored out of 68186 read. Teacher-forced NLL, no sampling, so no seed to fix.
- Significance uses a conservative per-token σ ≈ 1.5 and treats the two runs as independent. They are actuallypaired— same corpus, same tokens, same order — so a proper paired test on per-token log-probs would have more power.
ds4 \-\-perplexity\-fileonly returns the aggregate NLL, so that test could not be run here. The figures above are therefore a lower bound on significance, not an upper one.
⚠️Not comparable to published wikitext perplexities.ds4 \-\-perplexity\-filescores a single context window, not a sliding window over the whole test set likellama\-perplexity. These numbers are validagainst each otherand nothing else — and this GGUF cannot be loaded byllama\-perplexityat all.
**What this does not test.**The imatrix targets the routed-expert distribution seen in agent and code work, and it records the routed SwiGLU rowafter route weighting. General prose exercises a different mix of experts. A held-out code/agent slice would be the measurement that matters for the case this build is actually meant for — it has not been run yet. Until it is, treat the two builds as equivalent and pick either.
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#why-the-imatrix-buys-so-little–reading-the-dat-itselfWhy the imatrix buys so little — reading the .dat itself
The perplexity result above says the imatrix changes almost nothing on general text. Rather than leave that as a shrug, I opened the\.datand looked at what it actually contains. All of the following is reproducible from the published file with ~30 lines of numpy.
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#the-imatrix-does-not-discriminate-on-two-thirds-of-the-tensorsThe imatrix does not discriminate on two thirds of the tensors
For each expert, the imatrix holds one importance value per input column. If that vector is flat, the quantizer has no information to act on. Measuring its effective width (expof the Shannon entropy, so it is scale-free):
tensorcolumnsmedian effective widthreadingffn\_gate\_exps40963487 – 392285–96 % flat — near-uselessffn\_down\_exps2048370 – 129518–63 % — genuinely structured
This follows from how the collector is built. Upstream documents it plainly: for gate/up it records*“the squared FFN-normalized input activation”— the input to the FFN block, which isidentical for all 256 experts because it precedes routing— while fordownit records the routed SwiGLU rowafter route weighting*, i.e. what each expert actually saw.
The consequence is arithmetic.gateanduparetwo thirdsof the routed-expert weights and they are the ones crushed hardest, toIQ2\_XXS. That is exactly where the imatrix has nothing to say. The remaining third,down, is the only place it informs — and it is already the best-protected, atQ2\_K.
**So the imatrix carries usable signal on one third of the weights, and that third is the one quantized least aggressively.**A small measured gain is the expected outcome, not an anomaly. This is a property of the collector, not of the checkpoint or the recipe.
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#per-expert-importance-is-dominated-by-single-outliersPer-expert importance is dominated by single outliers
Summing importance per expert onffn\_down\_exps, the deep layers look extraordinarily concentrated — until you look at why:
layerdominant expertshare of importancemax / median30#17699.9 %3.0e529#12699.8 %2.9e536#10499.8 %2.0e534#3399.8 %1.7e5 One expert with 300 000× the median. Remove it and the distribution is unremarkable — effective width climbs back to 150–171 out of 255. This is themassive activationspattern, not specialization. It does not mislead the quantizer, which slices each expert’s own segment, but it does mean per-expert importance sums are not a usable basis for allocating bits.
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#there-is-no-expert-highwayThere is no expert “highway”
If some experts mattered everywhere, they could be protected globally. They do not. Overlap of the top-26 experts between layers runs 4–19 %, against10.2 % expected from chance. Experts never appearing in any top-26 across the 17 deep layers:43 of 256, versus 41.5 predicted by chance. Adjacent layers 29 and 30 share exactly one expert out of 26.
Specialization is strictly per-layer; the router redistributes completely at every level. Good news about the model — its 256 experts are genuinely used — and bad news for anyone hoping to find a global subset worth extra bits.
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#what-this-impliesWhat this implies
- **The
37\-42 → Q4\_Kboundary is a convention, not a measurement.**Nothing in the imatrix singles those six layers out. Note that cross-layer importance totals arenotcomparable (activation scale grows with depth), so this analysis cannot propose a better boundary either — it can only say the current one is not derived from data. - **Finer allocation is blocked by the format, not the tooling.**A GGUF tensor carries one type, written once; all 256 experts of a layer live in a single tensor. Per-expert types would require changing the formatandthe Metal dequantization kernels, not just the quantizer.
- The tractable improvement is upstream, in the collector: recording, for gate/up, the activation each expert actually receivesafterrouting, instead of the shared pre-routing input. That would give the quantizer real per-expert signal on two thirds of the weights, where today it has almost none.
Complementary to this,nazeshinjitemeasured weight drift between the preview and0731checkpoints and found correlation above 0.99 at every depth, with three quarters of 4-bit values bit-identical — a continued-training refresh rather than a retrain. Two independent routes to the same conclusion: there was little for a freshly collected imatrix to recover.
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#performancePerformance
Apple M3 Max, 128 GB. Single-run Metal CLI,\-\-ctx 32768 \-\-nothink \-\-temp 0 \-n 256, short prompt — the same conditions upstream uses for its own table:
imatrix buildprefill45.08 t/sgeneration27.01 t/s
For reference, upstream reports 58.52 / 26.68 t/s for theuniform q2preview build on the same machine class. Generation matches; prefill is lower here because this is the mixed 2+4 bit recipe — layers 37-42 stay atQ4\_K, so there are more bytes to move per token during prompt processing.
Memory, measured at load:
resident model90.88 GiBKV @ 32k ctx0.61 GiB (raw 0.36 + compressed 0.25)total planned91.74 GiBmodel residency~35-40 s from SSD Build cost on the same machine, for anyone reproducing it:
steptimequantization pass (\-\-threads 16, CPU-bound, GPU idle)~1 h 05imatrix collection, 1.5M tokens (Metal, GPU-bound)~4 hsecond quantization pass with\-\-imatrix~1 h 10
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#caveatsCaveats
- Community build, not endorsedby antirez or DeepSeek.
- Not scoredagainst the official DeepSeek continuation vectors that antirez uses as a release gate (
tests/test\-vectors). The perplexity delta above is a same-engine A/B between these two files, not an absolute quality claim against the FP4 original. - The imatrix was collected on the 2-bit build, not a Q4 one — see above for why that is a reasonable compromise, but itisa compromise.
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#creditsCredits
- DeepSeek— base model and weights (MIT)
- **antirez**and the DwarfStar contributors — recipe, quantizer, imatrix pipeline, inference engine. This repo is a straight application of their work to a newer checkpoint.
- llama.cpp / GGML— quant formats and the groundwork all of the above rests on
https://huggingface.co/ox-ox/DeepSeek-V4-Flash-0731-gguf-ds4#licenseLicense
MIT, following the base model’s release terms.
相似文章
我简直不敢相信,我居然在家用PC上运行了前沿模型DeepSeek-V4-Flash-0731。太疯狂了!
一位用户对在配备24GB显存的中端Windows PC上通过量化运行前沿模型DeepSeek-V4-Flash-0731表示震惊,凸显了本地AI的飞速进步。
@Snixtp: DeepSeek V4 Flash 能否在单张 RTX Pro 6000 上运行?
antirez 已发布 DeepSeek V4 Flash 的 GGUF 量化版本,使该模型能够在单张 GPU(如 RTX Pro 6000)以及 128GB 以上内存的 Mac 上运行。量化文件已上传至 Hugging Face,并附有 DS4 推理引擎的使用说明。
@danveloper: 简直不敢相信,我竟然在树莓派 5(8GB 版)上以超过1 tok/s的速度运行了 DeepSeek-V4-Flash(284B 参数)……
一位开发者经过大量实验,成功在树莓派 5 上以超过1 tok/s的速度运行了284B参数的DeepSeek-V4-Flash模型,使用的是来自 antirez 的未经修改的 GGUF 文件。
DeepSeek V4 @ IQ3XXS 在 M1 Ultra 128GB 上的 LM Studio 中经补丁后可达 16 tok/s
一个 GitHub 补丁通过侧载 antirez 的 llama.cpp 分支,解决了结构布局漂移、解码拆分和代码签名问题,从而允许在 128GB Mac 上的 LM Studio 中运行 DeepSeek V4 Flash。
@Saboo_Shubham_:开源 AI 势头强劲。DeepSeek v4 Flash 是一款准前沿模型,拥有高达 100 万的上下文窗口。它可本地…
文章重点介绍了 DeepSeek v4 Flash,这是一款拥有 100 万上下文窗口的准前沿开源模型,并指出其能够通过 2 比特量化在 128GB 内存的 Mac 上本地运行。