speed-benchmark

Tag

Cards List
#speed-benchmark

I pushed Qwen3.8-27B to 99 tps single request and 1150 tps with a batch request on a RTX 3090

Reddit r/LocalLLaMA · yesterday

The author optimized the Qwen3.8-27B model inference on an RTX 3090 GPU, achieving up to 99 tokens per second for single requests and 1150 tps with batch processing through various quantization and optimization techniques, and released the updated code on GitHub.

0 favorites 0 likes
← Back to home

Submit Feedback