@theemozilla: We're working on making the local model experience better in Hermes, what are the best local models at each weight clas…
Summary
The user asks for recommendations on the best local AI models for different VRAM classes (8-16GB, 24-32GB, 128GB), mentioning Gemma4, Qwen, and DeepSeek variants, as they work on improving local model support in Hermes.
View Cached Full Text
Cached at: 07/10/26, 08:07 AM
We’re working on making the local model experience better in Hermes, what are the best local models at each weight class?
My blind guess, please correct:
8-16 GB VRAM Gemma4 12B
24-32 GB VRAM Qwen3.6 27B Qwen3.6 35B
128 GB VRAM (Spark, M3 Max) ??? Can you do DSv4-Flash?
Similar Articles
@svpino: Hermes with Gemma 4 or Qwen 3.5 is literally the best combo you can run locally on your computer. You've got to give th…
Developer claims Hermes fine-tunes of Gemma 4 and Qwen 3.5 deliver the best local LLM performance, suggesting they rival paid BigAI models.
@0xSero: Best models for your hardware - 4gb to 12gb vram - VibeThinker-3B - smokes everything remotely close to its weight clas…
This thread recommends AI models optimized for different VRAM levels, highlighting VibeThinker-3B for its strong reasoning performance at 3B parameters, along with other models for coding and general use.
@itsolelehmann: The best model setups to run on Hermes (by price tier): 1. If you have infinite budget: Go with GPT 5.5 or Claude Opus …
This post outlines budget-tiered AI model configurations for the Hermes application, recommending premium options like GPT 5.5 and Claude Opus 4.7 for unlimited budgets, cost-effective fallbacks like DeepSeek V4 Flash for tighter budgets, and local deployment via Qwen 3.6 for zero-cost inference.
@analogalok: I just got Gemma 4 26B A4B MoE model running fully locally with Hermes agent on an 8GB RTX 4060 and it's now backtestin…
A developer demonstrates running Gemma 4 26B MoE model locally on an 8GB RTX 4060 with Hermes agent to fully automate backtesting of trading strategies, highlighting the growing capability of local LLMs as autonomous agents.
@TraffAlex: AI MODELS FOR 32GB VRAM — TOP 17 CHEAT SHEET Hit the HuggingFace API, grabbed real .gguf Q4 sizes. Every link = direct …
A cheat sheet listing top AI models optimized for 32GB VRAM using GGUF Q4 quantization, with direct download links from HuggingFace. Includes models from Qwen, DeepSeek, Llama, and Mistral families, with tips on quantization and context settings.