@sachindetrax: 262K context. On a 16GB RTX 5070 Ti. Qwen 3.8 27B Q3 hits ~25 tok/s while an adaptive llama.cpp fork streams KV cache b…

X AI KOLs Timeline Tools

Summary

An adaptive KV cache streaming fork of llama.cpp enables running the Qwen 3.8 27B model with 262K context on a 16GB RTX 5070 Ti GPU, achieving ~25 tok/s by efficiently managing memory between RAM and VRAM.

262K context. On a 16GB RTX 5070 Ti. Qwen 3.8 27B Q3 hits ~25 tok/s while an adaptive llama.cpp fork streams KV cache between RAM VRAM. Stock llama.cpp starts thrashing around ~120K context. Same consumer GPU. 2x+ the usable context. This could be huge for local LLMs. Config - .\llama-server.exe ^ --model "Qwen3.8-27B-UD-Q3_K_XL.gguf" ^ --ctx-size 262144 ^ -fa on ^ -ctk q8_0 ^ -ctv q4_0 ^ -ngl all ^ -np 1 ^ -b 256 ^ -ub 256 ^ --kv-stream-stage-mib 2304 I’m pushing the limits further. Follow for the benchmarks, or tell me what you want me to test next. GitHub: https://github.com/sachin-detrax/llama.cpp-adaptive-kv-streaming…
Original Article
View Cached Full Text

Cached at: 08/29/26, 04:06 PM

262K context. On a 16GB RTX 5070 Ti.

Qwen 3.8 27B Q3 hits ~25 tok/s while an adaptive llama.cpp fork streams KV cache between RAM VRAM.

Stock llama.cpp starts thrashing around ~120K context.

Same consumer GPU. 2x+ the usable context.

This could be huge for local LLMs.

Config - .\llama-server.exe ^ –model “Qwen3.8-27B-UD-Q3_K_XL.gguf” ^ –ctx-size 262144 ^ -fa on ^ -ctk q8_0 ^ -ctv q4_0 ^ -ngl all ^ -np 1 ^ -b 256 ^ -ub 256 ^ –kv-stream-stage-mib 2304

I’m pushing the limits further. Follow for the benchmarks, or tell me what you want me to test next.

GitHub: https://github.com/sachin-detrax/llama.cpp-adaptive-kv-streaming…


sachin-detrax/llama.cpp-adaptive-kv-streaming

Source: https://github.com/sachin-detrax/llama.cpp-adaptive-kv-streaming

Adaptive KV Streaming for llama.cpp

This branch adds an experimental, block-granular KV cache streaming path to the CUDA llama-server. It is intended for running long contexts when model weights leave too little VRAM for the complete KV cache.

With --kv-stream-stage-mib N, the authoritative KV tensors are stored in pinned host memory while a bounded CUDA pool is shared by resident KV pages and a transfer ring. The runtime adapts that split as the context grows: it keeps as many pages resident as the budget allows, reclaims resident space for staging when more streaming is required, and prefetches later layers while the current layer computes. This avoids relying on uncontrolled Unified Memory page thrashing and preserves exact attention over the full context.

Detailed project story, design, implementation, and benchmark results are in Running Qwen 27B on 16G VRAM with Full Context Length: Building Adaptive KV Cache Streaming for llama.cpp.

This is research code tailored to our current NVIDIA CUDA configuration: an RTX 5070 Ti with 16 GB VRAM, unsloth/Qwen3.8-27B-GGUF UD-Q3_K_XL, a 262144-token context, Flash Attention, a Q8_0 K cache, a Q4_0 V cache, and one server slot. Other models, KV cache quantization combinations, parallel slots, and non-CUDA backends are not yet supported or validated. Expanding model and KV quantization support is follow-up work.

Build the modified server

Install a C++ compiler, CMake, and the CUDA toolkit, then run this command from the repository root:

cmake -S . -B build -DGGML_CUDA=ON -DGGML_CUDA_FA_ALL_QUANTS=ON -DCMAKE_BUILD_TYPE=Release && cmake --build build --config Release --target llama-server -j

The executable is created at build/bin/llama-server.

Example using the tested cache configuration:

./build/bin/llama-server \
  --model /path/to/model.gguf \
  --ctx-size 262144 \
  -fa on \
  -ctk q8_0 \
  -ctv q4_0 \
  -ngl all \
  -np 1 \
  --kv-stream-stage-mib 2304

The best value for --kv-stream-stage-mib depends on the model, context capacity, GPU, and other VRAM consumers. Start conservatively and increase it while checking startup and peak VRAM use.

Optional Unified Memory for model weights

Adaptive KV streaming works with or without Unified Memory. Leave GGML_CUDA_ENABLE_UNIFIED_MEMORY unset for ordinary CUDA device allocations. To make GPU-offloaded model buffers CUDA managed allocations, launch the same server with the environment variable enabled:

GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 \
./build/bin/llama-server \
  --model /path/to/model.gguf \
  --ctx-size 262144 \
  -fa on \
  -ctk q8_0 \
  -ctv q4_0 \
  -ngl all \
  -np 1 \
  --kv-stream-stage-mib 2304

With this flag, CUDA-backed model buffers, including GPU-offloaded weights, are allocated with cudaMallocManaged and their pages can migrate between VRAM and host memory. The adaptive resident-page and transfer-ring pool is intentionally different: it is still allocated with cudaMalloc, so that fixed-size pool remains physically allocated in VRAM instead of becoming managed memory. UVM is therefore optional for this branch and does not change the KV streaming pool into pageable storage.

Recreate the benchmark graph

The benchmark driver automatically selects the largest practical adaptive KV pool for each configured context capacity, sweeps from 8K through the requested maximum, and generates the CSV, PNG, and SVG results:

python3 -m pip install matplotlib

python3 benchmarks/benchmark_kv_stream.py \
  --model /path/to/model.gguf \
  --max-context 192K

The only required arguments are the model GGUF and maximum context. See benchmarks/README.md for the pool-probing algorithm, generated files, optional settings, and resumable output directories.


Upstream llama.cpp README

llama.cpp

llama

LLM inference in C/C++

License: MIT Release Server Docker Winget

manifesto / ggml / ops / maintainer PRs / compile times / lib llama API / llama-server REST API

Quick start

A few options to get llama.cpp installed on your machine:

Once installed:

# Download and run a model directly from Hugging Face
llama cli -hf ggml-org/Qwen3.5-0.8B-GGUF

# Launch OpenAI-compatible API server
llama serve -hf ggml-org/Qwen3.5-0.8B-GGUF
VLM session with `llama cli` VLM session with llama cli Built-in web UI against `llama serve` running Qwen 3.6 Built-in web UI against llama serve

Description

The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud.

  • Plain C/C++ implementation without any dependencies
  • Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks
  • AVX, AVX2, AVX512 and AMX support for x86 architectures
  • RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures
  • 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use
  • Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA)
  • Vulkan and SYCL backend support
  • CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity

The llama.cpp project is build on top of the ggml library.

Supported backends

BackendTarget devices
BLASAll
BLISAll
CANNAscend NPU
CUDANvidia GPU
HIPAMD GPU
Hexagon [In Progress]Snapdragon
IBM zDNNIBM Z & LinuxONE
MUSAMoore Threads GPU
MetalApple Silicon
OpenCLAdreno GPU
OpenVINO [In Progress]Intel CPUs, GPUs, and NPUs
RPCAll
SYCLIntel GPU
VirtGPUVirtGPU APIR
VulkanGPU
WebGPUAll
ZenDNNAMD CPU

Documentation

Tools

Development

Contributing

  • Contributors can open PRs
  • Collaborators will be invited based on contributions
  • Maintainers can push to branches in the llama.cpp repo and merge PRs into the master branch
  • Any help with managing issues, PRs and projects is very appreciated!
  • Read the CONTRIBUTING.md for more information

Acknowledgements

  • yhirose/cpp-httplib - Single-header HTTP server, used by llama-server - MIT license
  • stb-image - Single-header image format decoder, used by multimodal subsystem - Public domain
  • nlohmann/json - Single-header JSON library, used by various tools/examples - MIT License
  • miniaudio.h - Single-header audio format decoder, used by multimodal subsystem - Public domain
  • subprocess.h - Single-header process launching solution for C and C++ - Public domain

Similar Articles

Running Qwen3.6 35b a3b on 8gb vram and 32gb ram ~190k context

Reddit r/LocalLLaMA

The author shares a high-performance local inference configuration for running Qwen3.6 35B A3B on limited hardware (8GB VRAM, 32GB RAM) using a modified llama.cpp with TurboQuant support, achieving ~37-51 tok/sec with ~190k context.

Getting close to 100K context on 32GB VRAM with Qwen3.6-27 at Q8

Reddit r/LocalLLaMA

A user shares their attempts and configurations to achieve up to 115K context on a Q8-quantized Qwen3.6-27B model using 32GB VRAM on an RTX 5090, with benchmark results and trade-offs between context length and kv-cache quantization.

Qwen3.6-35B-A3B Q4 262k context on 8GB 3070 Ti = +30tps

Reddit r/LocalLLaMA

The author shares detailed tuning tips for running the Qwen3.6-35B-A3B MoE model on an 8GB RTX 3070 Ti with up to 262k context using llama.cpp, achieving 30+ tps, and notes a 25% speed boost when switching from Windows to Ubuntu Server.