token-speed

Tag

Cards List
#token-speed

GPT-6.1 Sol feels unlimited, because it runs at 20 tokens per second

Reddit r/singularity ↗ · 2d ago

A user claims that OpenAI's GPT-6.1 Sol generates text about 2.5x slower than GPT-6 Sol and 2.3x slower than GPT-5.6 Sol, suggesting OpenAI is throttling compute for paying customers to reallocate resources to internal experiments.

0 favorites 0 likes
#token-speed

Got 85.6tok/s MTP with Qwen3.8:27b. Single RTX 5090

Reddit r/LocalLLaMA ↗ · 6d ago

Achieved 85.6 tokens per second using the Qwen3.8:27b model on a single RTX 5090 GPU.

0 favorites 0 likes
#token-speed

Mercury 2.5 LLM hits 770 tokens per second

Hacker News Top ↗ · 2026-09-23 Cached

The Mercury 2.5 LLM achieves a speed of 770 tokens per second, as evaluated by Artificial Analysis through various intelligence benchmarks and capability indexes.

0 favorites 0 likes
#token-speed

600tok/s single request on qwen3.6 35ba3b with Ninfer on an RTX Pro 6000. Anybody remember that Comcast ad "stupid fast"?

Reddit r/LocalLLaMA ↗ · 2026-09-17

Achieving 600 tokens per second on the Qwen3.6 model using Ninfer on an RTX Pro 6000, noted as useful for brute-force tasks despite not being the most advanced model.

0 favorites 0 likes
#token-speed

@MiaAI_lab: Who wants 40+ tok/s for prose on GLM 5.3 Flash?

X AI KOLs Following ↗ · 2026-09-17 Cached

The tweet highlights the performance of the GLM 5.3 Flash AI model, achieving over 40 tokens per second for prose generation.

0 favorites 0 likes
#token-speed

Nex-N2.5-mini-MLX-4bit on Apple M5 Max — 133.6 tok/s — llm-bench.io

Reddit r/LocalLLaMA ↗ · 2026-09-12 Cached

Benchmark results for the Nex-N2.5-mini-MLX-4bit model on Apple M5 Max hardware, achieving 133.6 tokens per second generation speed and quality scores up to 85.80 in research tasks.

0 favorites 0 likes
#token-speed

@TheAhmadOsman: I genuinely cannot use cloud providers for tokens anymore My 1000s of tokens per second are a privilege and once you ge…

X AI KOLs Timeline ↗ · 2026-09-02 Cached

A user states that after experiencing high token throughput of thousands per second, they can no longer use cloud providers due to the significant performance difference.

0 favorites 0 likes
#token-speed

@0xSero: Opus at home, at 200+ tok/s I love ZAI

X AI KOLs Timeline ↗ · 2026-08-26 Cached

The user @0xSero shares running Anthropic's Opus model at home with over 200 tokens per second using ZAI, expressing enthusiasm for ZAI.

0 favorites 0 likes
#token-speed

@xueyu1125: Running local large models requires at least 50 tokens/s for usability. Here are API output speeds for top models (Gemini Flash 300+) DeepSeek V4 Flash: 99.5 tokens/s GPT-5.6 Sol: 68.1 to…

X AI KOLs Following ↗ · 2026-08-18 Cached

Discusses the token speed requirements for running local large models and compares API output speeds of multiple top AI models.

0 favorites 0 likes
#token-speed

I'm (mostly) picking models on speed now, not intelligence

Lobsters Hottest ↗ · 2026-08-02 Cached

The author argues that frontier LLMs have reached a 'good enough' intelligence threshold, so they now prioritize speed over raw intelligence when choosing models, citing fast open-weights models like GLM5.2 and DeepSeek V4 Flash as daily drivers.

0 favorites 0 likes
#token-speed

I made an offline, single-file GPU build picker that estimates what local models a rig will run — and at what tok/s

Reddit r/LocalLLaMA ↗ · 2026-06-27

A developer created an offline, single-file GPU build picker that estimates which local AI models a system can run and at what token generation speed.

0 favorites 0 likes
#token-speed

How fast is N tokens per second really?

Hacker News Top ↗ · 2026-05-18 Cached

A web tool that lets users visually experience different LLM token generation rates (e.g., 5–800 tok/s) across code, text, reasoning, and agent modes, helping internalize performance numbers from benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback