@tan_maty: Oh my god, the AI Stanford course shared by the awesome @alisawuffles who starts at OpenAI next week — I found it! Must-see for beginners! I've already learned it (and lost my mind), come join me! I feel my English improving too! Stanford CS336: Language Mod…
Summary
Stanford CS336 aims to teach students how to build language models from scratch, with deep understanding of the full-stack design of data, systems, and models. The course videos are publicly available and suitable for AI beginners.
View Cached Full Text
Cached at: 06/24/26, 12:25 PM
Holy cow, the AI Stanford course shared by the fairy sister @alisawuffles, who is going to work at OpenAI next week, I found it — a must-watch for beginners!
I’ve already learned it to death, come join me, I feel my English has improved too!
Stanford CS336: Language Modeling from Scratch
https://t.co/RAxxIU7X12 https://t.co/U3JgLPPcnf
TL;DR: Stanford CS336 aims to have students build language models from scratch, deeply understand the full-stack design of data, systems, and models, and emphasize algorithmic efficiency at scale.
Course Introduction: Building Language Models from Scratch
CS336 is a Stanford University course whose core goal is to let students build a complete language model by hand — from the data system to the modeling. The course is co-taught by Percy and Tatsu, with a TA team including Roit, Neil, and Marcel. As Percy said: “To understand it, you must build it.” All lectures for this course are uploaded to YouTube for learners worldwide.
Why This Course? Breaking Researchers’ “Abstract Black Box”
Percy points out a current crisis: researchers are becoming increasingly disconnected from the underlying technology. Eight years ago, researchers would implement and train AI models themselves; six years ago, they would at least download BERT and fine-tune it. Now, many people accomplish tasks just by prompting proprietary models. While abstraction layers bring convenience, these abstractions are leaky — you don’t truly understand that it’s essentially just “strings in, strings out.” Foundational research requires breaking down the entire tech stack and co-designing data, systems, and models. The purpose of this course is to keep foundational research alive.
Limitations of Small Models: Two Manifestations of Scale Changes
Industrial language models are astonishing in scale: GPT-4 is rumored to have 1.8 trillion parameters with a training cost of up to $100 million; XAI is building a cluster with 200,000 H100s; over $500 billion in investments are planned for the next four years. But frontier models are out of reach for most people, so the course focuses on small language models. However, small models may not be representative for two reasons:
-
Change in FLOPs ratios: In small Transformers, the attention layer and MLP layer have roughly equivalent FLOPs; but when parameters scale to 175 billion, MLP layers completely dominate. If you only optimize attention at small scale, its effect is swamped at large scale.
-
Emergent behaviors: A 2022 paper by Jason Wei shows that as training FLOPs increase, the accuracy of certain tasks (like in-context learning) jumps suddenly at a critical point. If you stay only at small scale, you might incorrectly conclude that “language models don’t work.”
What the Course Can Teach: Mechanisms, Mindsets, and Intuitions
Percy categorizes knowledge into three types:
- Mechanisms (teachable): Transformer implementation, model parallelism, efficient GPU utilization, etc.
- Mindsets (more important): Squeezing performance out of hardware, taking scaling seriously. This mindset was pioneered by OpenAI and is key to leading the next generation of AI models.
- Intuition (only partially teachable): Which data and modeling decisions lead to good models. Because architectures and datasets that work well at small scale may fail at large scale. Still, getting two and a half out of three is a good deal.
On intuition, Percy cites a paper introducing the Swish activation function, whose conclusion was: “We cannot explain it, we can only say it’s God’s grace” — the experiments speak for themselves.
The Real Meaning of the Bitter Lesson: Algorithms Matter at Scale
The “bitter lesson” is often misinterpreted as “scale is everything, algorithms don’t matter.” Percy believes this is completely wrong. The correct interpretation is: Model accuracy = Efficiency × Resources invested. Efficiency is even more important at large scale because when you spend hundreds of millions of dollars, you can’t afford to waste resources like you would on a local cluster. A 2020 OpenAI paper showed that algorithm efficiency for training to a certain accuracy on ImageNet improved by 44x between 2012 and 2019 (exceeding Moore’s Law). Similar results exist for language models.
Therefore, the right framework is: given a certain compute and data budget, what is the best model you can build? As a researcher, the goal is to improve the efficiency of algorithms.
Historical Review: Evolution of Language Models
- Early days: Shannon used language models to estimate English entropy; in 2007, Google trained a five-gram n-gram model based on 2 trillion tokens (more tokens than GPT-3), but it lacked emergent behavior.
- 2010s: Deep learning revolution — Yoshua Bengio’s first neural language model in 2003, sequence-to-sequence models (Ilya et al.), Adam optimizer, attention mechanisms, 2017’s “Attention Is All You Need” (Transformer), expansion of mixture-of-experts models, model parallelism techniques (already capable of training models with hundreds of billions of parameters).
- Foundation model trend: ELMo, BERT, T5 pretrained on large-scale text and adapted to downstream tasks.
- Key inflection point: OpenAI combined components and pushed scaling laws, producing GPT-2, GPT-3. Since then, it has split into closed-source models (accessed via APIs) and open-source models (e.g., EleutherAI, Meta, Bloom, Alibaba, DeepSeek, etc.). Openness levels include: closed-source, open weights (architecture details but no dataset details), open-source (all weights and data available).
Current Landscape and Course Approach
Current frontier models come from OpenAI, Anthropic, xAI, Google, Meta, DeepSeek, Alibaba, Tencent, etc. The course will revisit the technical principles of these components, get as close as possible to best practices of frontier models, while leveraging open-source community information and inferences from closed-source models. The course adopts an executable lecture format, embedding code to explain step by step.
Source: https://www.youtube.com/watch?v=SQ3fZ1sAqXI
Similar Articles
@Michaelzsguo: Alisa Liu mentioned the Stanford course CS336: Language Modeling from Scratch while preparing for an OpenAI interview. If you want to systematically learn LLM now, or if you plan to pursue AI research / MTS / ML e…
Recommends the Stanford open course CS336: Language Modeling from Scratch, which systematically explains the full training pipeline of language models from scratch, suitable for those preparing for AI interviews or wanting to deeply learn LLM.
@WangNextDoor2: Stanford CS146S: A Must-Take AI Programming Introductory Course https://heyuan110.com/zh/posts/ai/2026-02-24-stanford-cs146s-overview/…
Stanford University's new CS146S course systematically teaches AI programming (Vibe Coding), covering LLM principles, Agent architecture, MCP, etc. All resources are free and publicly available, marking AI programming as a formal engineering discipline.
@li9292: How to join OpenAI? Just master the following courses: 1. Stanford's "Language Modeling from Scratch" course: http://cs336.stanford.edu/spring2025/ 2. After gaining breadth, she dives deep into each concept, using blogs, papers, and ChatGPT…
This tweet recommends Stanford's CS336 course and a series of learning resources as a preparation path for joining OpenAI.
@tan_maty: I'm blown away by this course, a must-see for CS majors: CS336, a course that's recently become legendary in the AI community. Building large language models from scratch. This course is offered by Stanford, taught by top NLP experts Percy Liang and Tatsunori Hashim…
A thread promoting Stanford's CS336 course on building language models from scratch, taught by NLP experts Percy Liang and Tatsunori Hashimoto, emphasizing hands-on understanding.
@FinanceYF5: Tonight, skip a TV show and finish this 2-hour 34-minute Stanford course. It covers from Tokenization, BPE to Transformer, pre-training, RLHF, DPO, and token-by-token generation, fully deconstructing how large models like ChatGPT and Claude are built…
A recommended Stanford course on AI that details the principles behind building large language models, covering Tokenization, BPE, Transformer, pre-training, RLHF, and DPO.