Tag
An intuitive, visual introduction to information theory covering entropy, mutual information, and channel capacity, assuming only basic probability. The paper explains fundamental limits of compression and transmission.
This article provides a visual guide to the Transformer architecture in Large Language Models, covering self-attention, causal self-attention, masked multi-head attention, and the output layer with step-by-step explanations and examples.
An interactive visual guide that explains how large language models work, from tokenization through attention, transformer blocks, and text generation, built by Roy van Rijn.