A deep benchmark of 8 tiny LLMs (135M to 1B parameters) on a $250 Jetson Orin Nano Super across four power modes finds 25W to be Pareto-optimal, with SmolLM2-135M achieving 165.1 tok/s and best efficiency.
Just released a deep benchmark of 8 tiny LLMs (135M → \~1B) on a $250 Jetson Orin Nano Super 8GB using llama.cpp CUDA - across all 4 power modes: 7W, 15W, 25W, and MAXN Hardware: * NVIDIA Ampere GPU - 1024 CUDA cores, 32 Tensor cores * 6× Arm Cortex-A78AE CPU @ 1.728 GHz * 8 GB LPDDR5 @ 204.8 GB/s (unified CPU + GPU - no VRAM split) * Active fan cooling - peak junction temp stayed ≤ 73 °C across every run Stack: * JetPack R36.4.7 (Ubuntu 22.04), CUDA 12.6 * llama.cpp CUDA backend, all layers on GPU (-ngl 99) * Load: NVIDIA aiperf — 20 requests per combo, 12 prompt × gen combos per model * Power measured via tegrastats VDD\_CPU\_GPU\_CV rail at 500ms intervals Brief methodology: * Sweep: prompt ∈ {128, 512, 1024, 2048} tokens × gen ∈ {64, 128, 256} tokens × 4 power modes = 384 benchmark cells per model, 8 models. * Key metric: output tok/J = tokens generated per joule of compute energy Findings: - Key finding: 25W is the Pareto-optimal mode for every model we have tested. - 36–47% more tok/s than 15W - 3–26% better output tok/J than 15W - 8–35% better output tok/J than even MAXN (highest power mode) - More clocks ≠ more efficiency. MAXN costs ~17% more power for marginal throughput gains. Sub-1B standouts at 25W (ctx=2048, gen=256): - SmolLM2-135M - 165.1 tok/s, 22.6 output tok/J (best in suite), 101 MB, ~5.4W - LFM2.5-350M - 115.1 tok/s in 219 MB. Matches SmolLM2-360M (369 MB) at less than half the size ~1B class at 25W (ctx=2048, gen=256): - LFM2.5-1.2B: 54.1 tok/s, 5.26 output tok/J, 698 MB - fastest + best output tok/J in class - Gemma3-1B: edges ahead on total tok/J (118.5 vs LFM's 116.2) - lower power draw (6.87W vs 8.46W) compensates for slower decode - Llama3.2-1B: 47.0 tok/s, 4.67 output tok/J Full blog with all charts, heatmaps, latency tables, and raw HuggingFace datasets (384 cells × 4 modes) linked in the blog! Do check it out — and if you have a Jetson, what are you running on it? Would love to know! [Blog](https://www.smolhub.com/posts/jetson-nano-super-benchmark-non-reasoning/)
PrismML's Bonsai 27B model runs on the Jetson Orin Nano 8GB with 4.31 tokens/s and 27 t/s prompt processing, using 6.2GB RAM and about 25W power. It indicates surprisingly usable edge AI performance.
The article discusses the debate between small, highly capable local LLMs and large models optimized for efficiency, comparing performance on CPU using MiniCPM5 2B and MoE Qwen3.6 35B.
A developer successfully ran a 5.2 million parameter MoE LLM quantized to INT4 on an ESP32 Dev Kit V1 using only 81KB of SRAM by streaming experts from flash, achieving about 5 tokens per second.
Benchmark results for the Nex-N2.5-mini-MLX-4bit model on Apple M5 Max hardware, achieving 133.6 tokens per second generation speed and quality scores up to 85.80 in research tasks.
A developer benchmarked 21 local LLMs on MacBook Air M5 using HumanEval+ and found Qwen 3.6 35B-A3B (MoE) leads at 89.6% with 16.9 tok/s, while Qwen 2.5 Coder 7B offers the best RAM-to-performance ratio at 84.2% in 4.5 GB. Notably, Gemma 4 models significantly underperformed expectations (31.1% for 31B), possibly due to Q4_K_M quantization effects.