local-llm

Tag

Cards List
#local-llm

Kimi K3 (Unsloth) IQ2-XXS from 711GB down to 478GB!!! Only Multi-language was removed to trim the size

Reddit r/LocalLLaMA · 19h ago

A trimmed English-only GGUF version of Kimi K3 (IQ2-XXS) reduces model size from 711GB to 478GB by removing multi-language components, with early tests suggesting it may match or outperform the standard 2-bit version on coding tasks.

0 favorites 0 likes
#local-llm

Anyone else amped up over Qwen 3.8?

Reddit r/LocalLLaMA · yesterday

The author shares excitement for the upcoming Qwen 3.8 model, highlighting their experience with Qwen 3.6 27B for local LLM use, and discusses the potential of self-hosted AI to replace subscription-based frontier models.

0 favorites 0 likes
#local-llm

@mohitwt_: Day 20/30 of Inference Engineering building a speculative decoding runtime that drafts multiple tokens ahead with a sma…

X AI KOLs Following · 2d ago Cached

A developer is building Octane, a speculative decoding runtime for local LLM inference on consumer hardware, aiming for 2-3x speedup with exact output quality. Currently in active development with paged KV cache, continuous batching, and batched attention implemented.

0 favorites 0 likes
#local-llm

LabyrinthBench: a local-focused, judge-free LLM benchmark that measures context recall under interference for multi-step agentic tasks.

Reddit r/LocalLLaMA · 2d ago

LabyrinthBench is a local, judge-free LLM benchmark that deterministically measures context recall under interference for multi-step agentic tasks. Initial experiments show that wiping context and re-injecting gate answers helped 7 of 9 models but backfired on 2, highlighting the complexity of context management strategies.

0 favorites 0 likes
#local-llm

@YRSM_Simon: 120 t/s ! Good job, @UnslothAI

X AI KOLs Following · 3d ago Cached

Unsloth AI announces DSpark, enabling DeepSeek-V4-Flash GGUF models to run ~1.4–2× faster locally, reaching 120 tokens/s with no accuracy change.

0 favorites 0 likes
#local-llm

@CycleDecoded: Olares (from the beclab team), an open-source personal cloud OS designed for "always-on 7x24 AI Agents". Built on Kubernetes, it turns your retired old computer, N1 box, or idle NAS computing power into your own "...

X AI KOLs Timeline · 3d ago Cached

Olares is an open-source personal cloud OS from the beclab team, built on Kubernetes, specifically designed for 7x24 always-on AI Agents. It can convert old computers or NAS into a self-hosted AI platform, supporting local data storage and natural language operation.

0 favorites 0 likes
#local-llm

Deepseek V4 Flash just hit Colibri, does anyone have numbers?

Reddit r/LocalLLaMA · 3d ago

User asks for performance numbers on Deepseek V4 Flash running via Colibri, focusing on high VRAM setups, long context prefill, and token generation speed for agentic workloads.

0 favorites 0 likes
#local-llm

Graph engineering ? Or we can say agents on steroids....

Reddit r/artificial · 3d ago

Introduces GraphARC, an MIT-licensed open-source tool that lets a model author agent graph topologies at runtime, with a deterministic admission gate for auditable execution, built on LangGraph and running locally via ollama or against cloud APIs.

0 favorites 0 likes
#local-llm

Building a Rust Inference Engine That Matches Llama.cpp

Hacker News Top · 4d ago Cached

Ferrox is a pure-Rust inference engine that loads GGUF models and runs local LLMs on CPU, Metal, or CUDA, with a CLI and an OpenAI-compatible server. It aims to match llama.cpp's performance while being written from scratch with no bindings.

0 favorites 0 likes
#local-llm

DeepSeek-v4-Flash-Mini 54GB GGUF running at ~20.5 t/s

Reddit r/LocalLLaMA · 4d ago

A community build crushes DeepSeek-V4-Flash down to a 54GB IQ2_XXS GGUF variant with aggressive 2-bit quantization, achieving ~20.5 tokens/s on local hardware while drastically reducing memory footprint.

0 favorites 0 likes
#local-llm

DeepSeek v4 Flash vs. Qwen3.6-27B, 3.5-122B, and Gemma 4 31B Benchmark

Reddit r/LocalLLaMA · 5d ago

User benchmarks DeepSeek v4 Flash against Qwen3.6-27B, Qwen3.5-122B, and Gemma 4 31B on a local coding benchmark, finding Flash wins overall but Qwen 122B performs surprisingly well with better first-try success and lower token usage.

0 favorites 0 likes
#local-llm

@YRSM_Simon: Let Claude Fable 5 play chess against locally running DeepSeek V4 Flash (DGX Spark, thinking maxed out) — the year's most absurd comparison: Fable: ~2,600 tokens per move, 50 seconds to move; V4 Flas…

X AI KOLs Following · 5d ago Cached

The author had Claude Fable 5 play chess against a locally running DeepSeek V4 Flash. As a result, DeepSeek consumed tens of thousands of tokens per move and even timed out under extended thinking, yet it ended up with a better position. The game lasted 5 hours and was still unfinished.

0 favorites 0 likes
#local-llm

Deepseek V4 Flash 2-bit quant is the first model I can run locally that achieves 100% in this SQL benchmark

Reddit r/LocalLLaMA · 5d ago

A user reports that DeepSeek V4 Flash, running as a 2-bit quantized GGUF on dual RTX 3080s, is the first local model to score 100% on a real-world SQL benchmark, matching frontier models like Opus 4.7 and GPT-5.5.

0 favorites 0 likes
#local-llm

[Deepseek-V4-Flash-0731] Full 1M context on a single RTX5090 + DDR5 Desktop Setup with VLLM CPU/Ram Offloading, ~800 tps pp & 15+ tps decode [Agentic Coding]

Reddit r/LocalLLaMA · 5d ago

Technical write-up on running DeepSeek-V4-Flash-0731 with full 1M context on a single RTX 5090 + 256GB DDR5 desktop using vLLM with CPU/RAM offloading, achieving ~800 tps prefill and ~15 tps decode.

0 favorites 0 likes
#local-llm

Homebench – Benchmark local LLMs for speed, memory, and quality

Hacker News Top · 5d ago Cached

Homebench is a zero-config terminal tool that benchmarks locally-run LLMs for speed, memory, and quality, presenting a live leaderboard. It supports Ollama, LM Studio, llama.cpp, vLLM, and OpenAI-compatible servers.

0 favorites 0 likes
#local-llm

Is LM Studio abandoning their core product?

Reddit r/LocalLLaMA · 5d ago

LM Studio users are concerned that the company is de-emphasizing its original local LLM app by redirecting attention and downloads to its new Bionic agent, while the main app receives few updates and is harder to find on the website.

0 favorites 0 likes
#local-llm

@rohanpaul_ai: Qwen 3.8 Max just built better 3D physics scenes than Fable 5 while costing about 7x less to run. Test was done by @ato…

X AI KOLs Following · 6d ago Cached

A test by atomic.chat shows Qwen 3.8 Max outperforming Fable 5 at generating self-contained 3D physics scenes while costing about 7x less per run.

0 favorites 0 likes
#local-llm

@no_stp_on_snek: Very nice. Huge for team 3090. And TurboQuant+ is already implemented in a bunch of inference engines.

X AI KOLs Following · 6d ago Cached

A reply celebrates Unsloth AI's upcoming Qwen3.8-27B model, which will run on 17GB RAM/VRAM setups, and notes TurboQuant+ is already integrated into many inference engines — great news for RTX 3090 users.

0 favorites 0 likes
#local-llm

KAT Coder 2.5 dev: Do yourself a favor and try it!

Reddit r/LocalLLaMA · 6d ago

A developer enthusiastically recommends KAT Coder 2.5 dev, claiming it is faster, more accurate, and uses fewer tokens than Qwen 3.6 35b a3b, and outperforms Gemma 4 models on their setup, with a GitHub repo containing detailed benchmarks.

0 favorites 0 likes
#local-llm

@UnslothAI: Qwen3.8-27B is coming! Will run locally on 17GB RAM/VRAM setups.

X AI KOLs Timeline · 6d ago Cached

Alibaba announces Qwen3.8-27B open-weights release, capable of running locally on 17GB RAM/VRAM, alongside the larger Qwen3.8-Max.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback