gpu-performance

Tag

Cards List
#gpu-performance

Got 85.6tok/s MTP with Qwen3.8:27b. Single RTX 5090

Reddit r/LocalLLaMA ↗ · 2d ago

Achieved 85.6 tokens per second using the Qwen3.8:27b model on a single RTX 5090 GPU.

0 favorites 0 likes
#gpu-performance

@BenjaminDEKR: Here are my tests of Qwen-Image-2.1 on a NVIDIA 5070 Ti GPU (16gb vram) It fits, and it's fairly quick. ~25s for a 1024…

X AI KOLs Timeline ↗ · 2026-09-20 Cached

Tests of Qwen-Image-2.1 on an NVIDIA 5070 Ti GPU demonstrate it fits within 16GB VRAM, with generation times around 25s for 1024² images, and highlight excellent prompt adherence and typography capabilities.

0 favorites 0 likes
#gpu-performance

This draft model is OP on 16 GB cards for Qwen 3.8 27b

Reddit r/LocalLLaMA ↗ · 2026-09-12

A draft model for Qwen 3.8 27B is highly effective on 16 GB GPUs, achieving around 60 tokens per second on an RX 9070 XT and offering better VRAM efficiency than built-in MTP.

0 favorites 0 likes
#gpu-performance

@superalesha: Most of you dont know what an 8.4GB model can do on a 12GB RTX 3080 Ti. Qwen3.8-27B made this voxel Japanese pagoda gar…

X AI KOLs Timeline ↗ · 2026-09-02 Cached

Showcases the performance of the Qwen3.8-27B AI model running on a 12GB RTX 3080 Ti GPU, achieving 47 tokens per second at 128K context using llama.cpp with specific quantization settings.

0 favorites 0 likes
#gpu-performance

Ninfer and a 5090 with 3.8 27B is making me cry tears of joy it's so good.

Reddit r/LocalLLaMA ↗ · 2026-08-28

A user reports achieving high token throughput with the Ninfer tool on an NVIDIA RTX 5090 GPU using a Qwen 3.8B model, significantly outperforming llama.cpp.

0 favorites 0 likes
#gpu-performance

Compared Qwen 3.8 27B community quants on RTX 6000 vs Claude Opus 4.6

Reddit r/LocalLLaMA ↗ · 2026-08-26

A benchmark comparison of community quantized Qwen 3.8 27B models on RTX 6000 GPUs versus Claude Opus 4.6, evaluating token generation speed and efficiency for creating HTML games.

0 favorites 0 likes
#gpu-performance

LLM4LLM: Bridging Kernel Benchmarks and Real Deployment via Closed-Loop Agentic Optimization

arXiv cs.AI ↗ · 2026-08-25 Cached

LLM4LLM introduces a deployment-aware closed-loop optimization framework to bridge kernel benchmarks and real LLM inference, achieving up to 6.98x speedups on H100 GPUs.

0 favorites 0 likes
#gpu-performance

I hosted Kimi K3 (2.8T parameters) using 8 B300s. 92 tok/s, $190 per million tokens

Reddit r/LocalLLaMA ↗ · 2026-08-23

The article describes hosting the Kimi K3 AI model with 2.8 trillion parameters using 8 B300 GPUs, achieving 92 tokens per second and costing $190 per million tokens, while comparing it with Unsloth's dynamic GGUF quantization method.

0 favorites 0 likes
#gpu-performance

@Yuchenj_UW: UC Berkeley open-sourced FreeToken. Wild results: A single RTX PRO 6000 runs the 753B GLM-5.2 at 14.9 tok/s! An 8GB RTX…

X AI KOLs Following ↗ · 2026-08-21 Cached

UC Berkeley has open-sourced FreeToken, a tool that significantly speeds up local AI inference on consumer GPUs, achieving up to 39.3 tok/s on an 8GB RTX 4060 laptop.

0 favorites 0 likes
#gpu-performance

Qwen3.8 27b just exceeded my expectations on svg generation :D

Reddit r/LocalLLaMA ↗ · 2026-08-20

A user tested the Qwen3.8 27B model on SVG generation with a complex prompt and found the results exceeded expectations, highlighting the model's capability in creative tasks.

0 favorites 0 likes
#gpu-performance

I pushed Qwen3.8-27B to 99 tps single request and 1150 tps with a batch request on a RTX 3090

Reddit r/LocalLLaMA ↗ · 2026-08-17

The author optimized the Qwen3.8-27B model inference on an RTX 3090 GPU, achieving up to 99 tokens per second for single requests and 1150 tps with batch processing through various quantization and optimization techniques, and released the updated code on GitHub.

0 favorites 0 likes
#gpu-performance

@akshay_pachaar: how do you know whether your GPU is compute-bound or memory-bound? here's a simple explanation: your model weights sit …

X AI KOLs Following ↗ · 2026-08-16 Cached

The article explains how to determine if a GPU workload is compute-bound or memory-bound by analyzing operations per byte fetched from HBM, using NVIDIA's H100 as an example, and discusses how batching and prompt length affect performance.

0 favorites 0 likes
#gpu-performance

I let the agent test its own model upgrade instead of trusting the release notes. It found 3 things throttling itself

Reddit r/AI_Agents ↗ · 2026-08-15

A developer describes letting their AI agent autonomously test its own model upgrade by running controlled probes and measuring performance, revealing issues that throttled itself.

0 favorites 0 likes
#gpu-performance

@NVIDIAAI: As AI models continue to grow in scale and capability, shaping a model matters just as much as its size. We're introduc…

X AI KOLs Timeline ↗ · 2026-07-13 Cached

NVIDIA introduces a series on AI Model Co-Design, explaining how model dimensions affect GPU performance and the trade-offs between throughput and interactivity for LLM deployment. The first post provides a practical primer on designing hardware-friendly LLMs to improve system throughput and user responsiveness.

0 favorites 0 likes
#gpu-performance

@populartourist: Having worked consistently with Qwen3.6 27B NVFP4 on repos - it's clear that this quant is not reliable, at least for c…

X AI KOLs Timeline ↗ · 2026-06-15 Cached

The user reports that the Qwen3.6 27B NVFP4 quantization is unreliable for coding, with inconsistent quality despite high throughput, and suggests that Q4_K_M may be more consistent.

0 favorites 0 likes
#gpu-performance

@TeksEdge: Unsloth released the fastest Qwen3.6-27B MTP GGUF I've tested. Time to upgrade. Compared to the previous GGUF, Q4/Q6 XL…

X AI KOLs Timeline ↗ · 2026-05-12

Unsloth has released an optimized GGUF version of the Qwen3.6-27B MTP model, achieving significantly faster inference speeds (up to 114 tok/s on an RTX 5090) compared to previous quantizations.

0 favorites 0 likes
← Back to home

Submit Feedback