local-llm

Tag

Cards List
#local-llm

@YRSM_Simon: Let Claude Fable 5 play chess against locally running DeepSeek V4 Flash (DGX Spark, thinking maxed out) — the year's most absurd comparison: Fable: ~2,600 tokens per move, 50 seconds to move; V4 Flas…

X AI KOLs Following · 6d ago Cached

The author had Claude Fable 5 play chess against a locally running DeepSeek V4 Flash. As a result, DeepSeek consumed tens of thousands of tokens per move and even timed out under extended thinking, yet it ended up with a better position. The game lasted 5 hours and was still unfinished.

0 favorites 0 likes
#local-llm

Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark

Reddit r/LocalLLaMA · 6d ago

A user reports that DeepSeek V4 Flash, running as a 2-bit quantized GGUF on dual RTX 3080s, is the first local model to score 100% on a real-world SQL benchmark, matching frontier models like Opus 4.7 and GPT-5.5.

0 favorites 0 likes
#local-llm

[Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]

Reddit r/LocalLLaMA · 6d ago

Technical write-up on running DeepSeek-V4-Flash-0731 with full 1M context on a single RTX 5090 + 256GB DDR5 desktop using vLLM with CPU/RAM offloading, achieving ~800 tps prefill and ~15 tps decode.

0 favorites 0 likes
#local-llm

Homebench – Benchmark local LLMs for speed, memory, and quality

Hacker News Top · 6d ago Cached

Homebench is a zero-config terminal tool that benchmarks locally-run LLMs for speed, memory, and quality, presenting a live leaderboard. It supports Ollama, LM Studio, llama.cpp, vLLM, and OpenAI-compatible servers.

0 favorites 0 likes
#local-llm

Is LM Studio abandoning their core product?

Reddit r/LocalLLaMA · 2026-08-04

LM Studio users are concerned that the company is de-emphasizing its original local LLM app by redirecting attention and downloads to its new Bionic agent, while the main app receives few updates and is harder to find on the website.

0 favorites 0 likes
#local-llm

@rohanpaul_ai: Qwen 3.8 Max just built better 3D physics scenes than Fable 5 while costing about 7x less to run. Test was done by @ato…

X AI KOLs Following · 2026-08-03 Cached

A test by atomic.chat shows Qwen 3.8 Max outperforming Fable 5 at generating self-contained 3D physics scenes while costing about 7x less per run.

0 favorites 0 likes
#local-llm

@no_stp_on_snek: Very nice. Huge for team 3090. And TurboQuant+ is already implemented in a bunch of inference engines.

X AI KOLs Following · 2026-08-03 Cached

A reply celebrates Unsloth AI's upcoming Qwen3.8-27B model, which will run on 17GB RAM/VRAM setups, and notes TurboQuant+ is already integrated into many inference engines — great news for RTX 3090 users.

0 favorites 0 likes
#local-llm

KAT Coder 2.5 dev: Do yourself a favor and try it!

Reddit r/LocalLLaMA · 2026-08-03

A developer enthusiastically recommends KAT Coder 2.5 dev, claiming it is faster, more accurate, and uses fewer tokens than Qwen 3.6 35b a3b, and outperforms Gemma 4 models on their setup, with a GitHub repo containing detailed benchmarks.

0 favorites 0 likes
#local-llm

@UnslothAI: Qwen3.8-27B is coming! Will run locally on 17GB RAM/VRAM setups.

X AI KOLs Timeline · 2026-08-03 Cached

Alibaba announces Qwen3.8-27B open-weights release, capable of running locally on 17GB RAM/VRAM, alongside the larger Qwen3.8-Max.

0 favorites 0 likes
#local-llm

@MikeBradleyAI: TLDR on @deepseek_ai 0731 V4 Flash. It is comfortably the current SOTA for 190GB VRAM or unified memory based systems. …

X AI KOLs Following · 2026-08-02 Cached

Mike Bradley shares benchmark results claiming DeepSeek V4 Flash 0731 is the current state-of-the-art for 190GB VRAM systems, matching or exceeding an Unsloth 3-bit Qwen3.5-397B in quality while running about 3x faster.

0 favorites 0 likes
#local-llm

Deepseek v4 flash - 100-150 faster t/s in prefill/pp.

Reddit r/LocalLLaMA · 2026-08-02

This post shares fixes to improve DeepSeek v4 Flash prefill/PP speed: downgrading CUDA from 13.3 to 13.1 or using a custom fork, achieving up to 1.3K prompt processing tokens/s.

0 favorites 0 likes
#local-llm

DeepSeek-V4-Flash 284B on 5.3GB of memory

Reddit r/LocalLLaMA · 2026-08-02

A developer showcases Mference, a new inference engine that runs MoE models like DeepSeek-V4-Flash on just ~5.3GB of memory by streaming experts from SSD, with a native Mac app and OpenAI-compatible server.

0 favorites 0 likes
#local-llm

@no_stp_on_snek: https://huggingface.co/thetom-ai/DeepSeek-V4-Flash-ConfigI-MLX… Fyi. Fits on 128GB of ram for metal. GGUF is coming jus…

X AI KOLs Following · 2026-08-01 Cached

TheTom releases an MLX quantized version of DeepSeek V4 Flash (284B MoE, 21B active) at 3.05 bpw, fitting in 101 GiB to run on 128GB Apple Silicon, with GGUF sibling also available.

0 favorites 0 likes
#local-llm

@no_stp_on_snek: Ooh nice. This one gonna be fun

X AI KOLs Timeline · 2026-07-31 Cached

A reply to Unsloth AI expressing excitement about the rumored DeepSeek-V4-Flash model and the possibility of running it locally via quantized versions.

0 favorites 0 likes
#local-llm

Is a second local LLM actually a security boundary, or just another probabilistic opinion?

Reddit r/LocalLLaMA · 2026-07-30

A critical analysis questioning whether a second local LLM as a guard creates a reliable security boundary for agentic systems, advocating for deterministic policy enforcement over probabilistic guardrails.

0 favorites 0 likes
#local-llm

4090 + 5060 Ti + 64GB RAM: 206 t/s on a 35B-A3B, and a 122B at 37 t/s

Reddit r/LocalLLaMA · 2026-07-30

A user shares benchmark results for running large language models (Qwen 27B-122B) on a dual-GPU setup with RTX 4090 and RTX 5060 Ti, achieving high token generation speeds (e.g., 206 t/s on 35B-A3B, 37-41 t/s on 122B). The post includes setup details and a link to a GitHub repo with scripts and raw data.

0 favorites 0 likes
#local-llm

Are you guys not scared of where we're heading? A year ago, GPT-5 was considered one of the best models in the world. Today, we have open-weight models like Qwen3.6-27B that are competitive enough to run locally on high-end consumer hardware. The pace of progress is absolutely brutal.

Reddit r/LocalLLaMA · 2026-07-29

Commentary on the rapid pace of AI progress, noting that open-weight models like Qwen3.6-27B are now competitive enough to run locally on consumer hardware, a year after GPT-5 was among the best.

0 favorites 0 likes
#local-llm

I tested proven orchestration techniques on small local models. 90% failed. The 10% that survived roughly doubled task completion.

Reddit r/LocalLLaMA · 2026-07-29

A Reddit user shares results from testing proven orchestration techniques on small local LLMs, finding that 90% failed but the surviving 10% roughly doubled task completion across models like LFM 1.2B and Gemma 4 26B-A4B.

0 favorites 0 likes
#local-llm

@Hacksterio: Run a local LLM on Raspberry Pi’s bare metal — Linux not necessary.

X AI KOLs Timeline · 2026-07-29 Cached

Hackster.io shares a method to run a local LLM on a Raspberry Pi without requiring Linux, enabling bare-metal execution.

0 favorites 0 likes
#local-llm

Mix local LLMs, Claude Code, Codex, Gemini and more in one SDLC pipeline (open source) [P]

Reddit r/MachineLearning · 2026-07-28

AutoDev Studio is an open-source tool that enables mixing different AI models across stages of the software development lifecycle, allowing one to use local LLMs for planning, Claude Code for implementation, and different models for review, all while driving tools headlessly.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback