yarn-scaling

Tag

Cards List
#yarn-scaling

(NInfer Fork) I wanted to have a 1M context Qwen-3.8 27B, tp2, dual 5090s

Reddit r/LocalLLaMA · 21h ago

A developer forked NInfer, a C++20/CUDA inference engine, to add tensor-parallelism and YaRN rope scaling, enabling Qwen3.8-27B to run with a 1M token context on dual 5090 GPUs and outperforming vLLM in specific decode scenarios.

0 favorites 0 likes
#yarn-scaling

Qwen 3.6 27B is solid up to 262K context. How high have you guys gone above that using Rope/Yarn scaling?

Reddit r/LocalLLaMA · 2026-07-16

A user shares their experience pushing Qwen 3.6 27B to 262K context with coherent results, and discusses using Rope/Yarn scaling to go higher, along with kv-cache swapping strategies for RTX 3090 Ti.

0 favorites 0 likes
← Back to home

Submit Feedback