@akshay_pachaar: where does all the VRAM go during LLM inference? (4 ways GPU memory is used) loading the model is only the first part o…

X AI KOLs Timeline News

Summary

This article explains how GPU memory is utilized during large language model (LLM) inference, breaking it down into four key components: model weights, KV cache, activations/workspace, and runtime overhead. It highlights the importance of quantization in optimizing memory usage for better performance.

where does all the VRAM go during LLM inference? (4 ways GPU memory is used) loading the model is only the first part of the memory story. once inference starts, GPU memory gets divided across multiple components, and some of them keep growing as context length, batch size, and concurrency increase. the graphic breaks it into four useful buckets: → model weights are the mostly fixed part. once the model is loaded, their memory footprint stays roughly constant. the biggest lever here is precision. moving from FP16/BF16 to INT8 or INT4 reduces the number of bytes needed to store each parameter. → KV cache grows as generation continues. for every previous token, the model stores key and value tensors so attention can reuse them instead of recomputing the entire sequence. longer contexts mean a larger KV cache, and more concurrent requests mean more active caches sitting in memory. → activations and workspace hold temporary intermediate values needed while running attention, MLP layers, kernels, and other computations. this memory is reused across inference steps, but its size can still change with sequence length, batch size, and the kernels being executed. → runtime overhead comes from everything around the model itself. CUDA kernels, memory allocators, metadata, serving-engine buffers, and other runtime structures all consume some VRAM. it is usually smaller than the other buckets, but it is never zero. this is why “the model fits on the GPU” and “the workload fits on the GPU” are two different statements. a model may load comfortably, then run out of memory when you increase the context window, serve more users simultaneously, or increase the batch size. it also explains why quantization can help beyond simply fitting a larger model. shrinking the weight footprint creates room that can instead be used for larger KV caches, more concurrent requests, or bigger batches. that is the broader GPU lesson too. performance is not just about how much arithmetic a GPU can do. it is also about what data occupies memory, how much of it moves during inference, and how often that data can be reused. i wrote the full breakdown of how GPUs actually work and why memory movement sits at the center of LLM inference performance. the article is quoted below.
Original Article
View Cached Full Text

Cached at: 09/17/26, 06:25 PM

where does all the VRAM go during LLM inference?

(4 ways GPU memory is used)

loading the model is only the first part of the memory story.

once inference starts, GPU memory gets divided across multiple components, and some of them keep growing as context length, batch size, and concurrency increase.

the graphic breaks it into four useful buckets:

→ model weights are the mostly fixed part. once the model is loaded, their memory footprint stays roughly constant. the biggest lever here is precision. moving from FP16/BF16 to INT8 or INT4 reduces the number of bytes needed to store each parameter.

→ KV cache grows as generation continues. for every previous token, the model stores key and value tensors so attention can reuse them instead of recomputing the entire sequence. longer contexts mean a larger KV cache, and more concurrent requests mean more active caches sitting in memory.

→ activations and workspace hold temporary intermediate values needed while running attention, MLP layers, kernels, and other computations. this memory is reused across inference steps, but its size can still change with sequence length, batch size, and the kernels being executed.

→ runtime overhead comes from everything around the model itself. CUDA kernels, memory allocators, metadata, serving-engine buffers, and other runtime structures all consume some VRAM. it is usually smaller than the other buckets, but it is never zero.

this is why “the model fits on the GPU” and “the workload fits on the GPU” are two different statements.

a model may load comfortably, then run out of memory when you increase the context window, serve more users simultaneously, or increase the batch size.

it also explains why quantization can help beyond simply fitting a larger model. shrinking the weight footprint creates room that can instead be used for larger KV caches, more concurrent requests, or bigger batches.

that is the broader GPU lesson too.

performance is not just about how much arithmetic a GPU can do. it is also about what data occupies memory, how much of it moves during inference, and how often that data can be reused.

i wrote the full breakdown of how GPUs actually work and why memory movement sits at the center of LLM inference performance.

the article is quoted below.

Similar Articles

GPU Memory Math for LLMs (2026 Edition)

Reddit r/LocalLLaMA

A practical guide explaining how to calculate VRAM requirements for LLMs based on parameter count and quantization level, plus additional overhead from KV cache, activations, and batching.

Memory

Reddit r/artificial

Explains why LLM inference is increasingly memory-bandwidth bound due to the KV cache scaling with context length and concurrent users, and how systems like vLLM and PagedAttention improve memory utilization.