@0xSero: Opus at home, at 200+ tok/s I love ZAI
Summary
The user @0xSero shares running Anthropic's Opus model at home with over 200 tokens per second using ZAI, expressing enthusiasm for ZAI.
View Cached Full Text
Cached at: 08/28/26, 11:50 AM
Opus at home, at 200+ tok/s
I love ZAI https://t.co/vuQSDPRR0m
Similar Articles
@RoundtableSpace: Qwen 3.8 27B running locally on an RTX 5090 beats Opus 4.8 in personal benchmarks at up to 200 tokens per second with n…
Qwen 3.8 27B, a 27-billion parameter AI model, outperforms Opus 4.8 in personal benchmarks when running locally on an RTX 5090 GPU, achieving up to 200 tokens per second without internet or API access.
@analogalok: Operating System powered by Qwen 3.8 27B at 1950 tokens/sec! here is what 1,950 tokens/second @Alibaba_Qwen's 3.8 27b a…
A demonstration of an operating system powered by Qwen 3.8 27B running at 1950 tokens per second on Cerebras hardware, where the model weights act as the runtime for real-time software generation.
@rauchg: I animated the token spend race, from ~lifetime Vercel AI Gateway usage, which aggregates trillions of tokens from mill…
The tweet presents an animation of token spend from Vercel AI Gateway, illustrating shifts in usage among AI labs, Anthropic's dominance, and the growth of open weight AI.
Anthropic's Opus 5 is about token efficiency, not a capability leap
Anthropic released Opus 5, focusing on token efficiency and cost reduction rather than a major capability leap, offering performance close to Fable at half the cost.
Help - I'm spending $100/day on AI
A user shares their experience spending over $100/day on AI tokens via OpenRouter, primarily for coding tasks, and asks for cost reduction advice.