@ItsmeAjayKV: So last 7 day analysis of local LLM's i ran on my home-lab. MoE models now rule my 3090. Top used Model: @thinkymachine…
Summary
A user shares 7-day stats of local LLMs running on a home-lab, noting MoE models dominate their RTX 3090, with Inkling-small and Ling-3.0-flash performing well, and plans to open-source their dashboard.
View Cached Full Text
Cached at: 08/09/26, 03:19 PM
So last 7 day analysis of local LLM’s i ran on my home-lab.
MoE models now rule my 3090.
Top used Model: @thinkymachines Inkling-small-UD-Q4_K_M @AntLingAGI Ling-3.0-flash Q4 has now climbed to second position @deepseek_ai DSV4-flash-0731 is third
That being said, there were few sessions i forgot to record, but it will not really change result as those sessions were on Inkling-small.
From my testing Inkling-small is a really solid model, but Ling-3.0 is slowly becoming my fav, it’s a good balance between speed and quality.
DSV4-flash is good, but really slow to be useful, i hate waiting.
Man my dashboard looks awesome! Will be open-sourcing it soon
Similar Articles
@leopardracer: https://x.com/leopardracer/status/2055341758523883631
A user shares their experience setting up a dual-GPU local AI lab with RTX 4080 Super and 5060 Ti, running Qwen 3.6 models via llama.cpp and llama-swap to reduce API costs and enable unrestricted experimentation.
@TheAhmadOsman: You can run local models at home and use any agent harness like Codex or Claude Code with them
Ahmad built a simple tool that makes Claude Code work with any local LLM, demonstrated using vLLM serving GLM-4.5 Air on 4x RTX 3090s.
@Snixtp: More efficiency tests on a single 3090 TL;DR: - I tested 8 local LLMs on a single RTX 3090, power limit from 100W to 45…
The article presents benchmark results for 8 local LLMs on an RTX 3090, showing that power efficiency peaks around 225W, with diminishing returns at maximum power.
@ItsmeAjayKV: Update on 3090: Now with Qwen 3.6-35b-a3b moe (q6_k_xl). Crossed 90 t/s for the very first time, no MTP yet, prefill sp…
A user reports achieving over 90 tokens per second inference speed with Qwen 3.6-35b-a3b MoE model on an RTX 3090 using llama.cpp, with prefill speeds exceeding 1000 t/s, indicating practical local deployment of large language models on consumer hardware.
@TheAhmadOsman: Local AI is now good btw
Ahmad announces a Local AI Hardware Arena using ODS to benchmark LLMs on hardware like RTX PRO 6000, DGX Spark, Strix Halo, M5 MacBook Pro, and ChatGPT, inviting community input for future comparisons.