tokens-per-second

Tag

Cards List
#tokens-per-second

GPT-5.6 Sol can run now at an incredible rate of ~750 tokens per second

Reddit r/singularity · 5h ago

GPT-5.6 Sol now runs at an impressive inference speed of about 750 tokens per second.

0 favorites 0 likes
#tokens-per-second

Qwen3.6 27B on a 5090, 6.4k sample tok/s distribution after tuning MTP/cache settings

Reddit r/LocalLLaMA · 2026-07-04

Running Qwen3.6 27B on an RTX 5090, achieving 6.4k tokens per second after tuning MTP and cache settings, demonstrating optimization techniques for inference.

0 favorites 0 likes
#tokens-per-second

RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8

Hacker News Top · 2026-06-13

A setup using RTX 5080 and RTX 3090 GPUs achieves 80 tokens per second on the Qwen 3.6 27B Q8 model.

0 favorites 0 likes
#tokens-per-second

How fast is 10 tokens per second really?

Simon Willison's Blog · 2026-05-20 Cached

Simon Willison explores the practical meaning of 10 tokens per second speed for large language models, offering context on how fast that feels and its implications for usability.

0 favorites 0 likes
← Back to home

Submit Feedback