deepseek-v4

Tag

Cards List
#deepseek-v4

MAI (Microsoft AI) is very far behind on coding

Reddit r/ArtificialInteligence · 2026-07-22

The article criticizes Microsoft AI's lackluster coding model performance compared to rivals like Kimi K3 and Deepseek V4, suggesting MAI is far behind despite vast resources.

0 favorites 0 likes
#deepseek-v4

DeepSeek V4 Flash (98GB) on 1x 4060ti + CPU got 300% faster this week [ 2->7t/s]

Reddit r/LocalLLaMA · 2026-07-16

DeepSeek V4 Flash (98GB) now runs up to 7 tokens per second on a single RTX 4060 Ti with CPU offloading, a 3x speed improvement over the previous week's 2 t/s.

0 favorites 0 likes
#deepseek-v4

DeepSeek v4 Flash on 4090 + DDR5, my experience

Reddit r/LocalLLaMA · 2026-07-10

A user shares their experience running the DeepSeek v4 Flash model with a 24GB GPU and DDR5 RAM, including performance numbers and tips for optimization.

0 favorites 0 likes
#deepseek-v4

I merged fixes for quantized KV cache into my DeepSeek V4 branch

Reddit r/LocalLLaMA · 2026-07-04

Merged fixes for quantized KV cache into the DeepSeek V4 branch of llama.cpp, with benchmark perplexity results for f16, q8_0, and q4_0 cache types.

0 favorites 0 likes
#deepseek-v4

@PyTorch: While SGLang provided Day-0 support for DeepSeek-V4, the collaboration between the @lmsysorg and @NVIDIAAI engineering …

X AI KOLs Following · 2026-06-23 Cached

SGLang provided Day-0 support for DeepSeek-V4, and collaboration between LMSys and NVIDIA engineering teams achieved up to 5x throughput increase in production, with improvements shown on the SemiAnalysis InferenceX dashboard.

0 favorites 0 likes
#deepseek-v4

@QuixiAI: https://x.com/QuixiAI/status/2068776183102067086

X AI KOLs Following · 2026-06-21 Cached

DwarfStar is a self-contained native inference engine optimized for DeepSeek V4 Flash and PRO models, supporting Metal, CUDA, and ROCm backends, with a focus on high-end personal machines and Mac Studios.

0 favorites 0 likes
#deepseek-v4

@karminski3: Magic! DeepSeekV4 context memory compressed to 1/10! Everyone knows DeepSeekV4 supports 1M context and is heavily optimized. To actually use 1M context, VRAM usage is only about 10GB (compared to DeepSeek-V3.2 which needs about…

X AI KOLs Following · 2026-06-12 Cached

FlashMemory-DeepSeek-V4 proposes a novel inference paradigm called Lookahead Sparse Attention (LSA), which uses a neural memory indexer to actively predict future context needs, compressing physical KV cache usage to 13.5% of full context baseline while improving average accuracy by 0.6%. This method adopts a decoupled training strategy that allows independent training of the indexer without loading the base model, significantly reducing training cost.

0 favorites 0 likes
#deepseek-v4

@danveloper: https://x.com/danveloper/status/2064387956387758206

X AI KOLs Timeline · 2026-06-09 Cached

A developer ran DeepSeek-V4-Flash on a Raspberry Pi 5 by streaming model weights from an NVMe SSD, achieving 1.3 tokens/second at 8 watts, demonstrating the feasibility of frontier-adjacent open-weight models on low-cost, offline hardware.

0 favorites 0 likes
#deepseek-v4

FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention

Hugging Face Daily Papers · 2026-06-08 Cached

Proposes Lookahead Sparse Attention with a Neural Memory Indexer on DeepSeek-V4, reducing GPU memory usage to ~13.5% of full-context baseline while maintaining or slightly improving accuracy.

0 favorites 0 likes
#deepseek-v4

@jakevin7: An interesting thing. The DeepSeek V4 technical report conducted a comprehensive evaluation of all major LLMs, concluding that Gemini 3.1 Pro has the strongest world knowledge among all models. Not GPT, not Claude, but Gemini. But when people use Gemini...

X AI KOLs Following · 2026-06-07 Cached

According to the DeepSeek V4 technical report's evaluation of mainstream LLMs, Gemini 3.1 Pro is considered to have the strongest world knowledge, but users generally find it hard to use because the model does not proactively use search tools.

0 favorites 0 likes
#deepseek-v4

@Lonely__MH: Announcing something: DeepSeek v4 is the undisputed god of Chinese writing! Personally tested Opus 4.8 and DeepSeek-v4-pro head-to-head. Result: Opus 4.8 got completely crushed. This proves that for Chinese writing, you've got to go with the homegrown champion! Paying respects to Master Liang

X AI KOLs Timeline · 2026-05-29 Cached

A personal test claims DeepSeek v4 surpasses Opus 4.8 in Chinese writing, asserting that DeepSeek v4 is currently the strongest model for Chinese writing.

0 favorites 0 likes
#deepseek-v4

@Teknium: DeepSeek V4 Flash IS BACK on Nous Portal for FREE for use in Hermes Agent! Check it out at

X AI KOLs Following · 2026-05-25 Cached

DeepSeek V4 Flash is now available again for free on Nous Portal for use in Hermes Agent.

0 favorites 0 likes
#deepseek-v4

Running DeepSeek-V4 locally with 4x legacy RTX 2080 Ti ($2k budget setup). Custom Turing kernels, W8A8 quantization, and 255 prefill tok/s!

Reddit r/LocalLLaMA · 2026-05-20

A developer successfully runs DeepSeek-V4-Flash (284B total, 13B active) locally on four RTX 2080 Ti GPUs with a $2,500 budget, achieving 255 prefill tokens/s using custom Turing CUDA kernels, W8A8 quantization, and heterogeneous inference. The implementation is open-sourced.

0 favorites 0 likes
#deepseek-v4

Deepseek V4's 1M context window: the breaking point

Reddit r/LocalLLaMA · 2026-05-17

A detailed evaluation of Deepseek V4's 1M token context window across production codebases reveals optimal performance at 150-250k tokens, with degradation past 300k and significant latency in reasoning mode. The model exhibits high hallucination rates on unknown tasks, requiring validation layers for production use.

0 favorites 0 likes
#deepseek-v4

@PrajwalTomar_: https://x.com/PrajwalTomar_/status/2055263873348124904

X AI KOLs Following · 2026-05-15 Cached

A guide on using DeepSeek V4 as a cheaper alternative to Claude Opus 4.7 for agentic coding in Claude Code, including setup steps and cost comparison.

0 favorites 0 likes
#deepseek-v4

@cjzafir: 359M Tokens burned in 72 hours. Cost: $78~ Results: New 240M fine-tuning dataset. Process: > Codex 5.5 as Orchestrator.…

X AI KOLs Timeline · 2026-05-14 Cached

A developer used Codex 5.5 as an orchestrator and Deepseek v4 pro as an executor to generate a 240M token fine-tuning dataset, burning 359M tokens at a cost of only $78.

0 favorites 0 likes
#deepseek-v4

We Tested DeepSeek V4 Pro and Flash Against Claude Opus 4.7 and Kimi K2.6 (11 minute read)

TLDR AI · 2026-05-14 Cached

DeepSeek released V4 Pro and V4 Flash under MIT license on April 24, 2026. In benchmarks against Claude Opus 4.7 and Kimi K2.6, V4 Pro scored 77/100 at $2.25, placing between Opus 4.7 (91) and Kimi K2.6 (68), while V4 Flash scored 60/100 at $0.02, the cheapest in the comparison, with a 75% discount on V4 Pro through May 31.

0 favorites 0 likes
#deepseek-v4

Claude Mythos, Deepseek v4, HappyHorse, Meta’s new AI, realtime video games: AI NEWS

YouTube AI Channels · 2026-04-21 Cached

Anthropic unveils a withheld Claude Mythos model that autonomously finds thousands of 0-days, ZAI open-sources the 1.5 TB GLM-5.1 that tops open-weight benchmarks, Alibaba’s unreleased HappyHorse video model hits #1 on public leaderboards, and Deepseek teases an “Expert Mode” v4 preview.

0 favorites 0 likes
← Back to home

Submit Feedback