Tag
A user claims that OpenAI's GPT-6.1 Sol generates text about 2.5x slower than GPT-6 Sol and 2.3x slower than GPT-5.6 Sol, suggesting OpenAI is throttling compute for paying customers to reallocate resources to internal experiments.
Achieved 85.6 tokens per second using the Qwen3.8:27b model on a single RTX 5090 GPU.
The Mercury 2.5 LLM achieves a speed of 770 tokens per second, as evaluated by Artificial Analysis through various intelligence benchmarks and capability indexes.
Achieving 600 tokens per second on the Qwen3.6 model using Ninfer on an RTX Pro 6000, noted as useful for brute-force tasks despite not being the most advanced model.
The tweet highlights the performance of the GLM 5.3 Flash AI model, achieving over 40 tokens per second for prose generation.
Benchmark results for the Nex-N2.5-mini-MLX-4bit model on Apple M5 Max hardware, achieving 133.6 tokens per second generation speed and quality scores up to 85.80 in research tasks.
A user states that after experiencing high token throughput of thousands per second, they can no longer use cloud providers due to the significant performance difference.
The user @0xSero shares running Anthropic's Opus model at home with over 200 tokens per second using ZAI, expressing enthusiasm for ZAI.
Discusses the token speed requirements for running local large models and compares API output speeds of multiple top AI models.
The author argues that frontier LLMs have reached a 'good enough' intelligence threshold, so they now prioritize speed over raw intelligence when choosing models, citing fast open-weights models like GLM5.2 and DeepSeek V4 Flash as daily drivers.
A developer created an offline, single-file GPU build picker that estimates which local AI models a system can run and at what token generation speed.
A web tool that lets users visually experience different LLM token generation rates (e.g., 5–800 tok/s) across code, text, reasoning, and agent modes, helping internalize performance numbers from benchmarks.