@akshay_pachaar: How work is organized inside a GPU. A GPU does not treat a large computation as one job. It keeps dividing that job int…
Summary
This article explains how GPUs organize computational work into kernels, grids, blocks, and warps to achieve massive parallelism and hide memory latency through streaming multiprocessors.
View Cached Full Text
Cached at: 09/28/26, 01:36 PM
How work is organized inside a GPU.
A GPU does not treat a large computation as one job. It keeps dividing that job into smaller pieces until thousands of simple operations can run together.
The easiest way to understand this is to follow the work from the program you launch to the arithmetic the chip performs.
- A kernel creates the full workload.
A kernel is a function meant to run repeatedly across different pieces of data. When you launch one, it creates a grid. The grid represents all the work required for that launch.
- The grid is divided into thread blocks.
Each block owns one portion of the workload. Its threads stay together on the same streaming multiprocessor, or SM, so they can coordinate and exchange values through fast shared memory.
The GPU schedules blocks independently. This allows blocks from a large grid to spread across many SMs without needing to coordinate with one another.
- Each block is divided into warps.
A thread is one lane performing the kernel’s instructions on its own data. The hardware collects threads into fixed groups of 32 called warps.
A block containing 256 threads therefore becomes eight warps. These warps are the units the SM actually chooses between during execution.
All 32 threads in a warp receive the same instruction. They perform it together, but on different values. This is where the GPU gets its width.
It also explains why branching can hurt performance. If threads in one warp choose different paths, the GPU must run each path separately while temporarily disabling the threads that took the other one.
- Warps execute inside an SM.
An SM contains compute units, warp schedulers, registers, and shared memory. It can keep many warps resident at once, even though only some execute during a given clock tick.
When one warp requests data from slower memory and has to wait, the scheduler selects another ready warp. Switching is extremely cheap because every resident warp already has its state stored on the SM.
The GPU does not eliminate memory delays. It hides them by always having another warp ready to run.
This hierarchy also explains why workload size matters. Too few blocks leave SMs unused. Too few resident warps leave the scheduler with nothing to execute during a memory stall.
The complete path is simple.
A kernel creates a grid → The grid contains blocks → Blocks contain warps → Warps contain 32 threads → Blocks are assigned to SMs → Their warps are scheduled onto compute units.
That structure is the foundation behind GPU parallelism, latency hiding, and the need for large batches of similar work.
I wrote the full breakdown on how GPUs work, why they are organized this way, and what that design means for LLM performance.
The article is quoted below.
Similar Articles
@akshay_pachaar: https://x.com/akshay_pachaar/status/2087928032904523980
An educational thread explaining how GPUs work, focusing on the memory-compute asymmetry that dominates LLM serving performance, and demonstrating how techniques like quantization, speculative decoding, and continuous batching follow from that fundamental constraint.
@akshay_pachaar: how do you know whether your GPU is compute-bound or memory-bound? here's a simple explanation: your model weights sit …
The article explains how to determine if a GPU workload is compute-bound or memory-bound by analyzing operations per byte fetched from HBM, using NVIDIA's H100 as an example, and discusses how batching and prompt length affect performance.
@akshay_pachaar: GPU architecture, clearly explained. The usual assumption is that a faster GPU means more compute, so a chip rated for …
The article clarifies that GPU performance in AI inference is limited by memory bandwidth rather than compute power, using the NVIDIA H100 as an example to explain GPU architecture and its effect on token generation rates.
@amitiitbhu: How does a GPU work for Deep Learning? Read here: https://outcomeschool.com/blog/how-does-a-gpu-work-for-deep-learning…
This article explains how GPUs work for deep learning, covering why they are ideal for parallel matrix computations, the difference between CPU and GPU, and key concepts like VRAM and Tensor Cores.
@amitiitbhu: Just published: How do CUDA Kernels work? Read here:
This article explains how CUDA Kernels work on GPUs, covering the basics of CUDA, GPU architecture, threads, blocks, and their role in AI and machine learning.