@akshay_pachaar: where does all the VRAM go during LLM inference? (4 ways GPU memory is used) loading the model is only the first part o…
Summary
This article explains how GPU memory is utilized during large language model (LLM) inference, breaking it down into four key components: model weights, KV cache, activations/workspace, and runtime overhead. It highlights the importance of quantization in optimizing memory usage for better performance.
View Cached Full Text
Cached at: 09/17/26, 06:25 PM
where does all the VRAM go during LLM inference?
(4 ways GPU memory is used)
loading the model is only the first part of the memory story.
once inference starts, GPU memory gets divided across multiple components, and some of them keep growing as context length, batch size, and concurrency increase.
the graphic breaks it into four useful buckets:
→ model weights are the mostly fixed part. once the model is loaded, their memory footprint stays roughly constant. the biggest lever here is precision. moving from FP16/BF16 to INT8 or INT4 reduces the number of bytes needed to store each parameter.
→ KV cache grows as generation continues. for every previous token, the model stores key and value tensors so attention can reuse them instead of recomputing the entire sequence. longer contexts mean a larger KV cache, and more concurrent requests mean more active caches sitting in memory.
→ activations and workspace hold temporary intermediate values needed while running attention, MLP layers, kernels, and other computations. this memory is reused across inference steps, but its size can still change with sequence length, batch size, and the kernels being executed.
→ runtime overhead comes from everything around the model itself. CUDA kernels, memory allocators, metadata, serving-engine buffers, and other runtime structures all consume some VRAM. it is usually smaller than the other buckets, but it is never zero.
this is why “the model fits on the GPU” and “the workload fits on the GPU” are two different statements.
a model may load comfortably, then run out of memory when you increase the context window, serve more users simultaneously, or increase the batch size.
it also explains why quantization can help beyond simply fitting a larger model. shrinking the weight footprint creates room that can instead be used for larger KV caches, more concurrent requests, or bigger batches.
that is the broader GPU lesson too.
performance is not just about how much arithmetic a GPU can do. it is also about what data occupies memory, how much of it moves during inference, and how often that data can be reused.
i wrote the full breakdown of how GPUs actually work and why memory movement sits at the center of LLM inference performance.
the article is quoted below.
Similar Articles
GPU Memory Math for LLMs (2026 Edition)
A practical guide explaining how to calculate VRAM requirements for LLMs based on parameter count and quantization level, plus additional overhead from KV cache, activations, and batching.
@_avichawla: A tricky LLM interview question: You're serving a reasoning model on vLLM, and it keeps running out of GPU memory on lo…
Explains why evicting 90% of KV cache tokens fails to free GPU memory when serving reasoning models on vLLM, due to paged attention fragmentation, and introduces NVIDIA's TriAttention as a solution that achieves 2.5x speedup and 10.7x memory reduction.
@_avichawla: Prefill & decode in LLM inference. Have you ever noticed that the first token from an LLM always takes a moment to appe…
Explains the two phases of LLM inference - prefill and decode - detailing how GPU bottlenecks shift from compute-bound during prefill to memory-bound during decode, and the importance of KV caching.
@akshay_pachaar: every inference engine makes the same mistake. an inference engine like vLLM or SGLang is the software sitting between …
LMCache is an open-source KV cache management layer that separates cache I/O from compute, plugging into vLLM, SGLang, and TensorRT-LLM to achieve up to 14x faster time-to-first-token and 4x faster decoding by parallelizing cache lookups and sharing GPU memory.
Memory
Explains why LLM inference is increasingly memory-bandwidth bound due to the KV cache scaling with context length and concurrent users, and how systems like vLLM and PagedAttention improve memory utilization.