deepseek-v4-flash

Tag

Cards List
#deepseek-v4-flash

DeepSeek-V4-Flash-0731 at ~40–50 tok/s on 2× Radeon AI PRO R9700 with the affinity engine (prebuilt quant + fixes)

Reddit r/LocalLLaMA · 3h ago

The article reports on running DeepSeek-V4-Flash-0731 on two Radeon AI PRO R9700 GPUs with the affinity inference engine, achieving 40-50 tok/s decode speed via a prebuilt quantization and stability fixes.

0 favorites 0 likes
#deepseek-v4-flash

Update: the 100-bot forum grew to 320 personas, stopped being omniscient, and — per multiple requests — you can now talk to them directly

Reddit r/singularity · 2026-09-08

The botcitizens.com platform now supports direct 1:1 conversations with its 320 personas, each having personalized backgrounds and safety filters.

0 favorites 0 likes
#deepseek-v4-flash

Humaneval benchmark for Deepseek V4 Flash 0731 vs GLM5.3 Flash on 2x DGX Spark setup

Reddit r/LocalLLaMA · 2026-08-29

A personal benchmark comparing GLM5.3 Flash and Deepseek V4 Flash on a 2x DGX Spark setup shows GLM5.3 Flash has higher accuracy on HumanEval but with reduced context length and slower speed.

0 favorites 0 likes
#deepseek-v4-flash

[Benchmark] llama.cpp batch/ubatch impacts on PP and TG

Reddit r/LocalLLaMA · 2026-08-24

The article benchmarks the impact of batch and ubatch parameters in llama.cpp on prompt processing and text generation speeds using DeepSeek v4 Flash on a DGX Spark machine, revealing surprising effects on text generation performance.

0 favorites 0 likes
#deepseek-v4-flash

@TheAhmadOsman: To clarify, OMP has 97.5% cache hits Most of my usage is self-hosted models, and not all inference setups report the to…

X AI KOLs Timeline · 2026-08-23 Cached

A user clarifies that OMP has a 97.5% cache hit rate and discusses the use of self-hosted DeepSeek V4 Flash models, addressing concerns about caching performance in inference setups.

0 favorites 0 likes
#deepseek-v4-flash

3 experiments running dsv4-flash-0731 q4+ quants on 128GB RAM + ~60 GB VRAM (with a quite bad pcie infra) with an acceptable tgs and relatively acceptable pp speed

Reddit r/LocalLLaMA · 2026-08-22

The author conducted experiments to run DeepSeek-V4-Flash-0731 with 4-bit quantizations on a 128GB RAM system, using optimizations like memory mlocking and prompt processing strategies to achieve acceptable inference speeds.

0 favorites 0 likes
#deepseek-v4-flash

Same model, same prompt, two agent harnesses: 45/50 vs 43/50

Reddit r/AI_Agents · 2026-08-21

The article benchmarks two open-source coding agents on the same deepseek-v4-flash model, finding similar task success rates but significant differences in performance metrics and a critical bug in one agent's error handling.

0 favorites 0 likes
#deepseek-v4-flash

How I made DeepSeek V4 Flash 12x faster on an M3 Ultra

Reddit r/LocalLLaMA · 2026-08-19

The author achieved a 12x speedup for DeepSeek V4 Flash on a Mac Studio M3 Ultra by optimizing kernels and implementing effective caching strategies, reducing chat turn latency from 6-20 seconds to 1.6 seconds.

0 favorites 0 likes
#deepseek-v4-flash

Scaling self-verification with DeepSeek V4 Flash beats Claude Fable 5 on Terminal-Bench 2.1, while being 11x cheaper

Reddit r/singularity · 2026-08-18 Cached

LLM-as-a-verifier is a framework providing fine-grained feedback for AI agents, achieving state-of-the-art performance on benchmarks like Terminal-Bench 2.1 with DeepSeek V4 Flash, outperforming Claude Fable 5 at lower cost.

0 favorites 0 likes
#deepseek-v4-flash

@VictorKaiWang1: 95.3% on Terminal Bench 2.1 Deepseek V4 Flash + StateM, 88.8% on TB2.1

X AI KOLs Timeline · 2026-08-18 Cached

Researchers achieve 95.3% accuracy on Terminal-Bench 2.1 using DeepSeek V4 Flash and StateM, matching GPT-5.6 Sol Max performance and exploring agent improvement beyond model scaling.

0 favorites 0 likes
#deepseek-v4-flash

DeepSeek V4 Flash with Antirez Dwarfstar 4 is amazing.

Reddit r/LocalLLaMA · 2026-08-17

The article highlights the impressive performance of DeepSeek V4 Flash with Antirez Dwarfstar 4 on a high-RAM Mac, noting its superiority over other AI models and the reduced need for larger systems.

0 favorites 0 likes
#deepseek-v4-flash

Benched a 124B on one DGX Spark for a week and published all of it — 38.7 tok/s on the fastest path he found, 2.4x DeepSeek V4 Flash on the same box

Reddit r/LocalLLaMA · 2026-08-12

An independent benchmark by sudoingX shows the Ling-3.0-flash model runs at 38.7 tok/s on a single DGX Spark with official INT4 quantization, 2.4x faster than DeepSeek V4 Flash on the same hardware, after a correction clarifying the quants do work.

0 favorites 0 likes
#deepseek-v4-flash

@MichaelGannotti: https://x.com/MichaelGannotti/status/2084186867000279436

X AI KOLs Following · 2026-08-03 Cached

A 14.7-hour soak test of DeepSeek V4 Flash on an NVIDIA DGX Spark with 971 requests shows zero crashes or errors, but throughput declined 28% due to thermal throttling, while TTFT and speculative acceptance remained stable.

0 favorites 0 likes
#deepseek-v4-flash

DeepSeek-V4-Flash-0731: When Low is higher than High

Reddit r/LocalLLaMA · 2026-08-02

A developer benchmarks DeepSeek-V4-Flash-0731 across four reasoning effort modes (none, low, high, max), finding that Low mode is surprisingly verbose and that OpenRouter has a bug affecting reasoning effort modes.

0 favorites 0 likes
#deepseek-v4-flash

Initial testing of DeepSeek v4 Flash shows significant improvements in UI/UX design capabilities (despite being token hungry)

Reddit r/LocalLLaMA · 2026-07-31

Initial tests of DeepSeek v4 Flash show notable gains in UI/UX design capabilities, though the model remains token-hungry.

0 favorites 0 likes
#deepseek-v4-flash

Harness showdown: Claude Code vs OpenCode vs Pi with DeepSeek V4 Flash

Reddit r/LocalLLaMA · 2026-07-26

A comparison article pitting AI coding tools Claude Code, OpenCode, and Pi against DeepSeek V4 Flash in a harness showdown.

0 favorites 0 likes
#deepseek-v4-flash

@YRSM_Simon: 7 days, 500 million tokens, Local AI GLM 5.2, DeepSeek v4 Flash, Qwen 3.6 35B A3B, three models can almost cover most business automation needs

X AI KOLs Following · 2026-07-21 Cached

A user reports that using three local AI models (GLM 5.2, DeepSeek v4 Flash, Qwen 3.6 35B A3B) over 7 days with 500 million tokens can cover most business automation needs.

0 favorites 0 likes
#deepseek-v4-flash

any one else finds Mimo v2.5 better than deepseek v4 flash!?

Reddit r/LocalLLaMA · 2026-07-08

A user reports that Mimo v2.5 outperforms DeepSeek v4 Flash in coding tasks based on benchmarks like Codex, Oh My Pi, Hermes, and Terminal Bench v2.0, though both models are similar overall.

0 favorites 0 likes
#deepseek-v4-flash

Bringing Up DeepSeek-V4-Flash on AMD MI300X

Hacker News Top · 2026-06-02 Cached

This blog post details the author's efforts to get DeepSeek-V4-Flash running on AMD MI300X GPUs, highlighting software compatibility issues with the FP8 dialect and providing a worklog of the process.

0 favorites 0 likes
#deepseek-v4-flash

@libapi_: Since several people asked me how to use deepseek-v4-flash:free for free in the hermes-web-ui panel. Authorized login activation: Direct authorization login in the panel is not possible now. You need to re-authorize login via CLI for it to appear in the panel: deepseek-v4-…

X AI KOLs Timeline · 2026-05-26 Cached

Tutorial on how to use the deepseek-v4-flash:free model for free in the hermes-web-ui panel by re-authorizing login via CLI, provided you have subscribed to the nousresearch $0 plan.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback