@YRSM_Simon: 7 days, 500 million tokens, Local AI GLM 5.2, DeepSeek v4 Flash, Qwen 3.6 35B A3B, three models can almost cover most business automation needs
Summary
A user reports that using three local AI models (GLM 5.2, DeepSeek v4 Flash, Qwen 3.6 35B A3B) over 7 days with 500 million tokens can cover most business automation needs.
View Cached Full Text
Cached at: 07/21/26, 02:45 PM
7 days, 500M tokens, Local AI
GLM 5.2, deepseek v4 flash, Qwen 3.6 35B A3B — these three models can meet almost all business automation needs. https://t.co/qULfBAOg13
Similar Articles
@DeRonin_: My current local AI setup: - 2x DGX Spark linked (256gb) > GLM 5.2 @ 2bit, reasoning + agent loops - Mac Studio M3 Ultr…
A user describes their fully local AI stack using multiple hardware devices running Chinese models like GLM, Qwen, and Kimi, claiming 87% cost savings compared to frontier models like GPT-5.5 and Opus 4.8, while noting plans to self-host video generation.
@cjzafir: Models that I'm using daily: > Codex 5.5 high (fast) > Deepseek v4 pro via API > Kimi 2.6 via API Models that I am fine…
User shares a personal list of AI models they use daily (Codex 5.5, Deepseek v4 pro, Kimi 2.6) and for fine-tuning (Qwen 3.5 variants, Gemma4 E4B, GPT-oss 20B), aiming to fine-tune Small Language Models into Expert Language Models.
@cryptoresetlife: 本地无审核版 GLM5.2 754B 参数模型 231GB 在我的MAC studio M3 ultra 512gb 上部署成功了 @support_huihui 948 tokens / 4分25秒 = 948 / 265 ≈ 3.6 …
An uncensored version of the GLM5.2 754B parameter model (231GB GGUF) was successfully deployed on a Mac Studio M3 Ultra with 512GB RAM, achieving approximately 3.6 tokens/s.
@jakevin7: DeepSeek cache hit rate 95%, feels great. Maka's performance under the latest round of long-context tasks with the Deepseek model is outstanding. Total runtime close to 18 hours, nearly 400 million tokens, cost 33 bucks. The Make builders are amazing…
DeepSeek cache hit up to 95%, Maka desktop AI workstation performs excellently in long-context tasks, supports multiple models and tools, open source and local-first.
@DeRonin_: i ran Fable 5 the whole day and still haven't touched my limits why? i stopped paying surgeon rates for small talk here…
A user shares a detailed workflow strategy for efficiently using multiple AI models (Fable, Opus, Codex, DeepSeek, GLM, Qwen, Kimi) by delegating tasks based on cost and capability, using a single CLAUDE.md routing table, and avoiding small talk to reduce token usage.