My RULE of Thumb of choosing a models
Summary
The author shares personal experience showing how LLMs like Qwen 27B drastically reduce programming task time, offering rules of thumb for model selection.
Similar Articles
Medium sized MoE LLM models
A user asks for recommendations on medium-sized Mixture-of-Experts LLMs (up to 60B params in float8/110B in mxfp4), listing Qwen 3.5 35B, Gemma 4 A4B, and Nemotron 3 Nano, and inquiring about niche options beyond these.
Stop asking what model to run. There are literally only two.
A tech enthusiast argues that only two local AI models (Qwen 3.6 35b a3b and Qwen 3.6 27b) are worth running, dismissing smaller models and recommending heavy quantization of larger models.
I'm (mostly) picking models on speed now, not intelligence
The author argues that frontier LLMs have reached a 'good enough' intelligence threshold, so they now prioritize speed over raw intelligence when choosing models, citing fast open-weights models like GLM5.2 and DeepSeek V4 Flash as daily drivers.
LLM planner - pick a rig for your use-case/model/budget, or pick models for your rig. 60+ builds, 50+ models, 130+ cited t/s sources, 150+ reviewer YouTube videos, idle+active watts, multi-region prices, regular updates.
A comprehensive web tool and public dataset that helps users choose the right hardware for running LLMs, featuring 60+ builds, 50+ models, performance benchmarks, and reviewer videos, with two-way matching between models and hardware.
I tested 9 local models on the same flight sim prompt, all Q8, different Q providers, MLX
Benchmark of 9 quantized local LLMs running MLX on a flight-combat HTML prompt shows quant provider choice and model quirks matter more than parameter count or bit-width for usable code output.