@freeman1266: You don't need math to understand most AI papers—just understand this chain: token → embedding → position encoding → attention → FFN → residual stream → next-token prediction. LLMs essentially stack Transf…
Summary
A Chinese science tweet that intuitively explains the core chain of LLMs (Large Language Models): from token, embedding, position encoding, attention, FFN to residual stream and next-token prediction, helping readers without a math background understand AI papers.
View Cached Full Text
Cached at: 06/15/26, 05:07 PM
You don’t need to understand math to get most AI papers—just grasp this pipeline:
token → embedding → positional encoding → attention → FFN → residual stream → next-token prediction
An LLM is essentially a stack of Transformer blocks. Each time it generates, it asks the same question: given all the context so far, what’s the most reasonable next token?
A few key concepts:
Token: LLMs don’t read text—they read integer IDs. How many r’s are in “strawberry”? The model can’t count them. Not because it’s dumb, but because it processes by token, not by letter.
Embedding: Each token is mapped to a 4096-dimensional vector. They aren’t hand-labeled—they are trained “semantic coordinates.”
Attention: Each token asks: whose information in the context should I absorb the most? Query asks the question, Key gets matched, Value delivers the content.
FFN: Attention handles information transport; FFN handles information processing—the model’s “knowledge warehouse” lives largely here.
Residual stream: Each layer doesn’t scrap and rewrite—it just adds a stroke onto the existing understanding. Information accumulates and flows forward.
Next-token prediction: The final output isn’t a sentence—it’s a leaderboard of candidates. Temperature controls how adventurous you are; top-p controls how many candidates are considered.
Once you understand this pipeline, those “spooky phenomena”—like why a clean prompt still fails, or why longer context costs more—all make sense.
Similar Articles
@Potatoloogs: How LLMs Actually Work Inside: From Token to Next-Token – A Complete Overview of Nine Core Mechanisms a) Tokenization: The model doesn't read text, it reads integers · Text is first split into subword pieces, then mapped to integer IDs; modern LLM vocabularies typically have tens of thousands to...
This article systematically outlines the nine core mechanisms inside modern LLMs, from tokenization to next-token prediction, including tokenization, embedding, positional encoding, attention, multi-head attention, feed-forward networks, etc., and compares architectural differences between various models.
@vincemask: Put together, this is the complete AI pipeline: Underlying principles → Model operation → Capability optimization → Product deployment. Breaking it into 4 layers makes it clear: 1. Principle layer: AI's foundation. Neural networks, tokenization, embeddings, attention, Transformer. Addresses: how models understand text, semantics, and context. ...
This post divides the complete AI pipeline into four layers: Principle layer, LLM operation layer, Optimization layer, and System layer, explaining respectively how models understand language, generate answers, optimize performance, and deliver products.
@CamilleRoux: Une explication bien faite du fonctionnement interne des LLMs : tokens, embeddings, positional encoding, attention, fee…
This tweet shares a well-made explanation of the internal workings of LLMs, covering tokens, embeddings, positional encoding, attention, and feed-forward networks, via a blog post by 0xkato.
How LLMs Actually Work (26 minute read)
A detailed walkthrough of how transformer-based LLMs work, covering tokenization, embeddings, attention, and next-token prediction without heavy math.
@tanzhengmc97: https://x.com/tanzhengmc97/status/2066531753762656730
Explained the operating principles of large models in easy-to-understand language, including word vectors, Transformer attention mechanism, next-word prediction training, and emergent abilities, suitable for beginners to understand basic AI concepts.