@kmeanskaran: https://x.com/kmeanskaran/status/2105635344385450151

X AI KOLs Timeline News

Summary

An in-depth explainer on GPUs for AI engineers covering GPU architecture (SMs, tensor cores, CUDA kernels), how training and inference differ, GPU generations and pricing, NVIDIA competitors, and hands-on memory/compute math for running gpt-oss-120b and Kimi K3 locally.

https://t.co/KU5iSJRQRr
Original Article
View Cached Full Text

Cached at: 10/02/26, 04:35 AM

Everything an AI Engineer Needs to Know About GPUs

What is inside a GPU, how training and inference use it differently, the generations, what you actually rent, NVIDIA’s competitors, and running models locally.

If you build with AI, you use GPUs every day. But most of us never look at the hardware. We call it “the GPU” and move on.

That works until you have to make a real decision: which GPU to rent, why one model fits and another crashes, or why an H200 beats an H100 that has the same compute on paper.

This article walks through the hardware in the order you meet it, with exact math for two real open-weight models and real prices. No hardware background needed.

The two models:

  • gpt-oss-120b from OpenAI: 117 billion parameters, 5.1 billion active per token. Small enough for one GPU.

  • Kimi K3 from Moonshot AI: 2.8 trillion parameters, 104 billion active. One of the largest open models of 2026.

Both are mixture-of-experts models. Each layer holds many small “expert” networks, and each token only runs through a few of them: 4 of 128 for gpt-oss-120b, 16 of 896 for Kimi K3. So two numbers matter. Total parameters decide how much memory you need. Active parameters decide how much math and data movement each token costs.

1. The kinds of GPUs

GPUs come in three kinds:

  • Datacenter: racked servers with fast links between GPUs. Example: NVIDIA B200.

  • Workstation: one large card for professional desktops. Example: RTX Pro 6000.

  • Personal: gaming PCs. Example: GeForce RTX 5090.

Datacenter GPUs run in one of three places: the cloud (AWS, GCP, or “neoclouds” like CoreWeave and Nebius), on-premise (your company’s own datacenter), or air-gapped (on-premise with no outside network).

GPU landscape

GPU landscape

What this means for you: unless you work at a large enterprise or a government, you will rent datacenter GPUs in the cloud. Most of this article is about those.

2. Inside a GPU

A CPU is great at doing complicated things one after another. A GPU is a throughput machine: it does one simple operation on thousands of numbers at once. AI is mostly matrix multiplication, which is exactly that kind of work.

Here is an H100 from the outside, heatsink removed:

GPU bird view

GPU bird view

The middle square is the chip, with 80 billion transistors per NVIDIA’s Hopper deep dive. Around it sit five memory stacks, 80 GB in total. A GPU is really these two things: compute, and memory beside it.

Inside GPU

Inside GPU

Compute

The chip is made of Streaming Multiprocessors (SMs). An H100 has 132. Each SM has three kinds of compute:

  • CUDA cores: work on single numbers.

  • Tensor Cores: work on small matrices at once. They do the matrix math, so they matter most for AI.

  • Special function units: handle sin, cos and log. Softmax needs these.

Tensor Cores do one operation, Matrix Multiply and Accumulate: multiply A by B, add C, store the result.

The work is done by threads. A CPU has dozens to hundreds. A GPU has tens of thousands, in groups of 32 called warps. When one warp waits for data, the GPU switches to another in a single clock cycle, so the math rarely stops.

The code those threads run is a CUDA kernel: a function you launch once, and thousands of threads run it at the same time on different data. The simplest example, from NVIDIA’s CUDA programming guide:

There is no loop. Each thread adds one pair of numbers. When you call model(x) in PyTorch, every matmul, norm and softmax becomes a kernel like this.

CUDA kernels

CUDA kernels

A real example is attention’s Q, K and V. In vLLM’s gpt-oss code, the three weight matrices are stored side by side and computed in one multiplication. For gpt-oss-120b, each token’s 2,880 numbers go through one 2,880 × 5,120 matrix, which gives Q, K and V together. The GPU splits that into small tiles, one per group of threads.

QKV inside GPU

QKV inside GPU

Compute is measured in FLOPS, floating-point operations per second. Two tips for spec sheets. First, use the dense number, not the “with sparsity” one: the H100 page lists 1,979 teraFLOPS for BF16 with sparsity, so about 990 dense. Second, FLOPS roughly double each time you halve the precision, so always compare GPUs at the same precision.

Now the exact math. A forward pass costs about 2 FLOPs per active parameter per token, one multiply and one add.

  • gpt-oss-120b: 2 × 5.1 billion = 10.2 GFLOPs per token. A 1,000-token prompt is 10.2 TFLOPs: about 10 ms on an H100 running at its full 990 teraFLOPS.

  • Kimi K3: 2 × 104 billion = 208 GFLOPs per token, about 20 times more. The same 1,000-token prompt is 208 TFLOPs: about 210 ms on one H100.

Those are best cases. Real kernels never hit the full spec-sheet number.

Memory

A GPU has two kinds of memory:

  • VRAM (DRAM): the big memory, in gigabytes. The “V” is for video, from the graphics days. Today it is HBM (high-bandwidth memory).

  • SRAM: small, fast memory on the chip, used as caches. There are three levels: L0 for each Tensor Core, L1 inside each SM (256 KB on an H100), and L2 shared by all SMs (50 MB).

VRAM matters in two ways:

  • Capacity decides what fits. It must hold the weights plus the KV cache, the saved keys and values of past tokens. A common rule: weights plus at least 50% extra room for the cache. gpt-oss-120b is about 65 GB, so it fits on an 80 GB H100 with about 15 GB left for the cache. That is less than the rule asks for. The H200’s 141 GB is the comfortable fit.

  • Bandwidth decides how fast tokens come out. To make each new token, the GPU reads every active weight from VRAM.

Here is that math for gpt-oss-120b on one H100, one user. Each token uses 4 experts per layer across 36 layers: about 1.9 GB of 4-bit expert weights. Add attention, the router and the output layer, which stay in BF16: about 3.1 GB. That is about 5 GB read per token.

  • Moving the weights: 5 GB ÷ 3.35 TB/s ≈ 1.5 ms.

  • Doing the math: 10.2 GFLOPs ÷ 990 teraFLOPS ≈ 0.01 ms.

The GPU waits about 145 times longer than it computes. The ceiling is about 670 tokens per second for one user, before the KV cache and real-world overhead take their share.

Decoding step by step

Decoding step by step

What this means for you: the rule of thumb is simple. Compute (FLOPS) is the bottleneck for reading the prompt, called prefill, and for image and video generation. Memory bandwidth is the bottleneck for generating tokens, called decode. That is why the H200 exists: same compute as the H100, but 141 GB at 4.8 TB/s instead of 80 GB at 3.35 TB/s.

3. Training vs inference

The same GPU does two very different jobs. For years training was the main use of GPUs. Now inference is becoming the dominant one.

Training vs Inference on GPU

Training vs Inference on GPU

Training runs three steps for every batch: a forward pass to make predictions, a backward pass to compute gradients, and an update to the weights. The scaling laws paper estimates this at about 6 FLOPs per parameter per token, because the backward pass costs about twice the forward pass. For gpt-oss-120b, that is 6 × 5.1 billion = 30.6 GFLOPs per training token, three times the inference cost.

Training also needs far more memory. Hugging Face’s training memory guide breaks it down for mixed-precision training with the Adam optimizer:

  • Weights: 6 bytes per parameter (a 16-bit copy for the math and a 32-bit copy for stable updates).

  • Optimizer states: 8 bytes per parameter.

  • Gradients: 4 bytes per parameter.

  • Activations: saved from the forward pass for the backward pass, and they grow with batch size and sequence length.

That is about 18 bytes per parameter before activations, and here every parameter counts, not just the active ones. For gpt-oss-120b: 117 billion × 18 bytes ≈ 2.1 TB just to train, against about 65 GB to serve.

Inference only runs the forward pass: about 2 FLOPs per active parameter per token. Memory holds the weights (2 bytes per parameter in BF16, less when quantized) and the KV cache. Nothing else.

The shape of the work differs too. Training pushes huge batches through at once, so each weight fetched from memory is used many times and the Tensor Cores stay busy. Decode makes one token at a time per user, so the GPU mostly waits on memory.

What this means for you: a GPU that is “good for training” is not automatically best for serving. For training, you care about FLOPS, total memory and how fast GPUs talk to each other. For inference, you care about memory bandwidth, room for the KV cache, and cost per token.

4. Quantization

If each weight takes 2 bytes, you can store it in fewer bits instead. That is quantization: 8-bit is half the size, 4-bit is a quarter.

The trick is a shared scale. Hugging Face’s quantization guide writes it as x = S * (x_q - Z). A tiny example: the weights 0.50, -1.27 and 0.03 with a scale of 0.01 are stored as 50, -127 and 3. Multiply by 0.01 to get them back.

Quantization on GPU

Quantization on GPU

Older GPUs store small weights and turn them back into BF16 for the math. Newer ones compute directly in low precision: FP8 on Hopper, FP4 on Blackwell. FP8 is twice as fast on paper, but that does not turn into twice the real speed.

Real examples: gpt-oss-120b ships in a 4-bit format called MXFP4. 117 billion parameters would be 234 GB in BF16, but it ships at about 65 GB, and Hugging Face’s gpt-oss post confirms “the 120B fits in a single 80 GB GPU”. Kimi K3 was trained to run with 4-bit weights, called quantization-aware training. Its 2.8 trillion parameters would be about 5.6 TB in BF16. The checkpoint is about 1.56 TB.

What this means for you: quantization mainly buys memory, so a model fits on fewer or cheaper GPUs. Always test quality on your own tasks.

5. GPU generations

NVIDIA names have two parts: the letter is the generation, the number is the size. The H100 replaced the A100. The H200 is a bigger Hopper. Generations are named after scientists.

You will still see Turing (T4) and Ampere (A10, A100) in older systems. The ones that matter now:

  • Ada Lovelace (L4, L40): cheaper, supports FP8, but no NVLink, the fast GPU-to-GPU link. Good for small models.

  • Hopper (H100, H200): added FP8. New enough to be fast, established enough that every framework supports it. Most inference runs here.

  • Blackwell (B200, B300): added FP4, so 4-bit models like gpt-oss-120b and Kimi K3 run natively. The B200 has 180 GB, the B300 288 GB. The new gold standard for big models.

  • Rubin (2026): HBM4 memory, up to 288 GB at 22 TB/s per GPU, starting to ship as of October 2026. Plus Rubin CPX, a separate chip just for prefill, expected at the end of 2026.

  • Feynman (2028): next on NVIDIA’s roadmap.

NVIDIA’s own Grace and Vera CPUs link to GPU memory at 900 GB/s, fast enough to park old KV cache in the larger CPU memory.

What this means for you: the generation letter tells you which precisions you get, whether there is NVLink, and how mature the software is. My advice on new generations: wait for real benchmarks, because software takes about a year to catch up.

6. What you actually rent

In the cloud you rent an instance, not a GPU: GPUs plus CPUs, RAM, storage, networking and the links between GPUs. Any of these can be your bottleneck.

When a model is too big, GPUs work together. The standard unit is a node: eight GPUs in one server, connected by NVLink (900 GB/s per GPU on Hopper, 1,800 GB/s on Blackwell) and NVSwitch, which links every GPU to every other. Nodes connect over InfiniBand, a network roughly ten times slower than NVLink.

Multi-GPU

Multi-GPU

Because of that speed gap, vLLM’s docs recommend splitting each layer across the GPUs inside a node (tensor parallelism) and splitting groups of layers across nodes (pipeline parallelism). Kimi K3 is a real example. Its weights are about 1.56 TB, more than one 8 × H200 node holds (8 × 141 GB = 1,128 GB). NVIDIA’s own Kimi K3 recipe runs it on 32 H200s: tensor parallelism across the 8 GPUs of each node, pipeline parallelism across 4 nodes. That is 4,512 GB in total, so most of the memory goes to the KV cache. On Blackwell Ultra, one 8 × B300 node (2,304 GB) holds the weights on its own.

When a model is too small, you have the opposite problem. A model with a couple of billion parameters cannot keep an H100 busy. MIG splits one GPU into up to seven smaller, isolated GPUs, each with its own share of memory.

7. How to pick a GPU

Five quick steps:

  • Weights: Hugging Face’s rule of thumb is “roughly 2 * X GB of VRAM in bfloat16” for X billion parameters. Half that for 8-bit, a quarter for 4-bit.

  • Overhead: Hugging Face Accelerate’s docs say to “add up to an additional 20%”.

  • KV cache: leave at least 50% extra room.

  • Speed: weight size ÷ memory bandwidth = the best case per token for one user.

  • Cost: hourly price, and how busy you can keep it. A busy GPU shares each trip to memory across many users.

With Hugging Face Inference Endpoints prices on AWS, as I checked them at the end of September 2026:

  • gpt-oss-120b (~65 GB), cheapest: 1 × A100 80 GB, $2.50/hour. It fits, with about 15 GB left for the KV cache.

  • gpt-oss-120b, many users: 1 × H200, $5/hour. 141 − 65 = 76 GB for the KV cache, and 4.8 TB/s brings the per-token ceiling down to 5 GB ÷ 4.8 TB/s ≈ 1 ms.

  • Kimi K3 (~1.56 TB): too big for the largest instance on the list, 8 × H200 with 1,128 GB at $40/hour. You need a multi-node setup like the one in section 6.

8. Beyond NVIDIA

NVIDIA leads, but it is not alone. The alternatives include AMD’s Instinct GPUs, AWS Inferentia and Trainium, Google’s TPU, Cerebras’s wafer-sized chip, and Groq’s SRAM-based LPU. Two recent changes: AMD has moved to its MI400 series, and NVIDIA licensed Groq’s technology, with the NVIDIA Groq 3 LPX now in full production.

Each bets on one edge: memory bandwidth (Cerebras, Groq), power efficiency (Furiosa, Qualcomm), or cloud integration (Amazon, Google). All share one challenge: without CUDA, they must rebuild the software stack.

What this means for you: an alternative chip can be faster or cheaper for a specific workload, but check software support for your exact model first.

Wrapping up

Personally, I focus on the basics of LLMs and a bit of GPUs, and that has been enough to be useful. The people who tell you to learn CUDA first are describing the last ten percent of the field as if it were the entrance.

If you keep one habit from this article: before you rent anything, do two divisions. Weights ÷ VRAM tells you if it fits. Active weights ÷ bandwidth tells you how fast it can go.

An honest note: I have not benchmarked these GPUs myself, so the speeds here are spec-sheet math. And hardware facts age fast: sources from only a few months ago were already behind on Apple, AMD and Groq when I checked.

Thanks.

Follow @kmeanskaran

Similar Articles

@akshay_pachaar: https://x.com/akshay_pachaar/status/2087928032904523980

X AI KOLs Following

An educational thread explaining how GPUs work, focusing on the memory-compute asymmetry that dominates LLM serving performance, and demonstrating how techniques like quantization, speculative decoding, and continuous batching follow from that fundamental constraint.

@snowboat84: https://x.com/snowboat84/status/2061962883651731602

X AI KOLs Timeline

This article is the first part of the AI Engineering Panorama series. From a historical perspective, it reviews the evolution of GPUs from gaming graphics cards to AI accelerators, the bold bet of CUDA, the independent path of Google's TPU, and why NVIDIA ultimately prevailed. It also provides a detailed analysis of the underlying logic of AI infrastructure such as chips, supply chain, networking, and power.