expert-cache

Tag

Cards List
#expert-cache

Qwen3.8-Flash-Next on 2x3090 + DDR4: 17 → 25-29 t/s decode with the expert cache PR

Reddit r/LocalLLaMA · 6d ago

User benchmarks and details a GPU-resident expert cache PR in llama.cpp that boosts decode speed for Qwen3.8-Flash-Next on a dual RTX 3090 system from 17 to 25-29 tokens per second.

0 favorites 0 likes
← Back to home

Submit Feedback