@vivekgalatage: GPU Programming Fundamentals https://youtu.be/Cl2B_hmg4gA William Brandon, a performance engineer at Anthropic, outline…
Summary
A summary of William Brandon's (performance engineer at Anthropic) GPU programming fundamentals lecture, emphasizing that understanding the streaming multiprocessor (SM) structure of GPU hardware is key to predicting performance, rather than starting solely from the software abstraction of thread blocks/threads.
View Cached Full Text
Cached at: 08/04/26, 10:07 AM
GPU Programming Fundamentals
https://t.co/dig3Ldb73D
William Brandon, a performance engineer at Anthropic, outlines fundamental hardware-first principles for GPU programming - one of the best lectures. https://t.co/lQF7LL2ECl
TL;DR: The key to GPU programming performance is not to think of your program as an arbitrary number of thread blocks and threads, but to understand that GPU hardware is essentially composed of a finite number of streaming multiprocessors (SMs).
Speaker and Course Background
This talk is presented by William Brandon. He is currently a performance engineer at Anthropic, on leave from his PhD program at MIT, with a research focus on improving the efficiency of LLMs through kernel optimization and model architecture improvements. Last fall, he co-created and taught the first accelerated computing course in MIT’s history, and led the design of its CUDA-based programming assignments. Before entering graduate school, he worked on deep learning compilers at NVIDIA.
The content of this talk is primarily adapted from the highlights of the first two weeks of the MIT course 6S4. The course was co-created by William Brandon, his advisor Dr. Kelly, and another PhD student, Nikita Lazrov. All materials are publicly available online; the speaker specifically recommends Labs 1 and 2 as extended versions of this talk’s content.
Target Audience
This talk is intended for:
- Anyone planning to write kernels soon, or sometime in the future;
- Those who are still building their foundational understanding of GPU programming;
- People in machine learning who are curious about writing kernels but don’t yet have hands-on experience.
For more experienced listeners, much of the content may be basic, but the speaker hopes it can help everyone see these concepts from a slightly different angle, or provide additional clarity.
Since the core content is about “how to think about GPU hardware performance,” these principles apply to any language for writing GPU programs: whether writing kernels directly with CUDA, or using higher-level DSLs such as Triton or NVIDIA’s CUTLASS 4. Because they are all principles of the underlying platform.
The Standard View of a CUDA Program
From the surface of the CUDA programming language interface, a CUDA program is:
- Define a function in a C-like language;
- Use a launch command to run many copies of that function in parallel on the GPU.
When launching a CUDA kernel, you typically specify two parameters:
- The number of thread blocks;
- The number of threads per thread block.
Then the GPU runs all threads in all these blocks in parallel to complete the computation.
From a software perspective, this model is roughly: there is a set of blocks, each block contains a set of threads, and each thread is a tiny program that executes work in parallel with other threads.
Puzzle 1: Reducing Threads per Block Doesn’t Change Performance
Suppose you have a kernel that runs fairly quickly, for example, one for vector addition. The current launch configuration is:
- 128 thread blocks;
- 1024 threads per block.
1024 is the maximum number of threads NVIDIA allows a thread block to have.
Now run an experiment: reduce the number of threads per block from 1024 to 256.
A beginner might think: since a thread block can have up to 1024 threads, now using only 256 might only utilize 25% of the hardware’s capability, and the program would run 4 times slower.
But in practice, the result is often: for many kernels, going from 1024 to 256 runs at almost the same speed.
This looks a bit mysterious: the program runs on fewer threads, each thread does more serial work, yet the speed doesn’t change. So what are the other three-quarters of the threads actually doing?
Puzzle 2: Slightly Increasing Thread Blocks Dramatically Reduces Performance
Go back to the original configuration: 128 thread blocks, 1024 threads per block.
This time, don’t change the number of threads; instead, adjust the number of thread blocks: increase from 128 to 133. That’s only about a 4% increase. It seems like you’re just partitioning the problem slightly differently and running on slightly more blocks, so it seems unlikely to have much of an impact on performance.
But if you actually run this experiment, especially on an H100 GPU, for some kernels, a 4% increase in the number of thread blocks can make the program run 2 times slower.
These two puzzles show that thinking of a CUDA program purely in terms of “threads” and “thread blocks” is a very useful abstraction, because they are the concepts NVIDIA provides and that you have to work with; but if you want to predict and reason about program performance, this model is highly incomplete.
The Essence of a GPU: A Bunch of Little Squares
Rather than asking “What is CUDA and how do I use it?”, the better question is:
What is a GPU, and what can it do?
This is an image of a GPU board from NVIDIA. The board has all sorts of electronic components. But if you zoom in on the actual GPU chip, you’ll notice a striking visual feature: the GPU is made up of a bunch of little squares.
The speaker emphasizes that if you could only describe what a GPU is in one sentence, a better mental picture is:
A GPU is a fixed number of little squares.
This picture is far better than imagining the GPU as a “vague blob running countless thread blocks and threads, with their numbers set arbitrarily.”
What are these little squares?
NVIDIA calls them Streaming Multiprocessors, often abbreviated as SMs.
So exactly how many SMs are there? Around a hundred. In this image, there are 144 on the die. In reality, if you buy an H100 and ask it “how many SMs do you have?”, it will tell you about 132, because some of the 144 SMs didn’t pass the manufacturing process and are disabled in software. For most GPUs on the market, the number of SMs you get varies within roughly a factor of two.
Source: @vivekgalatage: GPU Programming Fundamentals (https://www.youtube.com/watch?v=Cl2B_hmg4gA)
Similar Articles
https://www.youtube.com/watch?v=aE0onltJlOo
This lecture introduces the flexible evolution of GPU architecture as a SIMD (vector/array) processor, discusses data parallelism, memory bank grouping, bank conflicts, serial bottlenecks, and the history of SIMD instructions (such as MMX), emphasizing how GPUs leverage data parallelism and deal with serial bottlenecks.
@levidiamode: Day 138/365 of GPU Programming One of my favorite lectures I've watched this year is Stanford's CS336 lecture 7 on GPU …
A learner shares enthusiasm for Stanford CS336 lecture 7 on GPU parallelism, which covers fundamental operations and connects them to multi-GPU setups and parallelism techniques like tensor, data, and pipeline parallelism.
@akshay_pachaar: https://x.com/akshay_pachaar/status/2087928032904523980
An educational thread explaining how GPUs work, focusing on the memory-compute asymmetry that dominates LLM serving performance, and demonstrating how techniques like quantization, speculative decoding, and continuous batching follow from that fundamental constraint.
@vivekgalatage: The portal into the background of GPU architecture https://www2.eecs.berkeley.edu/Pubs/TechRpts/2016/Archive/EECS-2016-…
A technical report from UC Berkeley EECS that provides an in-depth exploration of GPU architecture background and design.
@vivekgalatage: Best structured reference I've found for GPU optimization - 450 papers, 14 years of research. Some techniques will have…
A tweet shares a structured reference of 450 papers on GPU optimization spanning 14 years, noting that while some techniques evolve, the mental models remain useful. It also references a lecture on GPU architectures by Onur Mutlu.