All currently popular local models in one table + Opus 4.8 results

Reddit r/LocalLLaMA News

Summary

This article provides a comparative table of popular local AI models across various benchmarks, including agentic, coding, general, and multimodal tasks, to help users choose models based on hardware specs and use cases.

If you are thinking what model will fit best your HW specs and tasks you are doing here is one table with all currently popular models that still can be considered as local. LLM Test Scores Feature DeepSeek-V4-Flash-Vision-Exp DeepSeek-V4-Flash-0731 Qwen3.8-Flash-Next GLM-5.3-Flash Qwen3.8-27B Opus-4.8 Total parameters ≈285B 284B 125B 320B 27B not published Active parameters 13B 13B 6B 18B 27B not published Agentic benchmarks Benchmark DeepSeek-V4-Flash-Vision-Exp DeepSeek-V4-Flash-0731 Qwen3.8-Flash-Next GLM-5.3-Flash Qwen3.8-27B Opus-4.8 Terminal Bench 2.1 83.9 82.7 – 82.6 73.0 85.0 NL2Repo 57.7 54.2 48.1 52.1 42.3 69.7 DeepSWE 59.3 54.4 58.7 61.1 42.2 58.0 Toolathlon-Verified 75.9 70.3 73.5 72.1 – 76.2 Agents' Last Exam 27.3 25.2⁷ 24.3 28.1 20.4 25.7 AutomationBench (Public) 25.7 25.1 – 25.3 – 27.2 GDPval-AA v2 – 68.1 – 72.3 – 75.1 Cybergym 75.3 76.7 – – – 78.3 DSBench-Hard 63.6 59.6 – – – 71.7 DSBench-FullStack – 68.7 – – – 71.6 ApexBench (Pass@1) 36.5 26.2⁷ – – – 39.4 HLE with tools (full set) – 16.8 – 22.9 – 25.4 Coding benchmarks Benchmark DeepSeek-V4-Flash-Vision-Exp DeepSeek-V4-Flash-0731 Qwen3.8-Flash-Next GLM-5.3-Flash Qwen3.8-27B Opus-4.8 SWE-bench Pro – 56.0 62.5 – 61.7 – SWE-bench Multilingual – – 81.0 – 73.8 – CoWorkBench – 45.1 73.9 – 70.7 – JobBench – 41.3 55.7 – 33.4 – General benchmarks Benchmark DeepSeek-V4-Flash-Vision-Exp DeepSeek-V4-Flash-0731 Qwen3.8-Flash-Next GLM-5.3-Flash Qwen3.8-27B Opus-4.8 GPQA Diamond – 90.8 91.7 – 89.2 – HLE (without tools) – 33.8 35.9 – 30.8 – LiveCodeBench v6 – 90.6 91.9 – 90.3 – IFBench – 79.2 81.3 – 79.5 – Multimodal benchmarks Benchmark DeepSeek-V4-Flash-Vision-Exp DeepSeek-V4-Flash-0731 Qwen3.8-Flash-Next GLM-5.3-Flash Qwen3.8-27B Opus-4.8 Chartography 64.3 – – – – 65.0 ZeroBench (Pass@5) 35.0 – – – – 34.0 BabyVision¹⁰ – – – 73.0 65.7 / 85.6 34.1 MathVision¹⁰ – – 90.6 / 95.7 – 90.0 / 94.6 – RealWorldQA – – 88.5 – 85.9 – AndroidWorld – – 84.5 – 81.9 – OSWorld 2.0 (partial credit) – – 52.3 – 48.0 – Vision2Web – – 64.0 – 62.9 – ClawEval-MM (Pass@3) – – 64.4 – 57.4 – RecreationBench – – 49.9 – 47.1 – ERQA – – 72.3 – 65.5 – Note: I used GLM-5.3 to compose the table from official HF pages of the models. Note2: Opus-4.8 results are presented only for illustration and are omitted from selecting the best model in a row.
Original Article

Similar Articles

Running local models is good now

Hacker News Top

The author reports that running local AI models has become surprisingly good, with recent releases like GPT-OSS and Gemma 4 enabling agentic coding locally at about 75% accuracy of frontier models, a significant improvement from just months ago.

Local text to image model comparaison: The ultimate test.

Reddit r/LocalLLaMA

User presents a comprehensive comparison of local text-to-image models using 192 prompts, evaluating capabilities like text rendering, faces, anatomy, and spatial composition, with results and prompts publicly available at imagebench.ai.