@tan_maty: Oh my god, the AI Stanford course shared by the awesome @alisawuffles who starts at OpenAI next week — I found it! Must-see for beginners! I've already learned it (and lost my mind), come join me! I feel my English improving too! Stanford CS336: Language Mod…

X AI KOLs Timeline Events

Summary

Stanford CS336 aims to teach students how to build language models from scratch, with deep understanding of the full-stack design of data, systems, and models. The course videos are publicly available and suitable for AI beginners.

Oh my god, the AI Stanford course shared by the awesome @alisawuffles who starts at OpenAI next week — I found it! Must-see for beginners! I've already learned it (and lost my mind), come join me! I feel my English improving too! Stanford CS336: Language Modeling from Scratch https://t.co/RAxxIU7X12 https://t.co/U3JgLPPcnf
Original Article
View Cached Full Text

Cached at: 06/24/26, 12:25 PM

Holy cow, the AI Stanford course shared by the fairy sister @alisawuffles, who is going to work at OpenAI next week, I found it — a must-watch for beginners!

I’ve already learned it to death, come join me, I feel my English has improved too!

Stanford CS336: Language Modeling from Scratch

https://t.co/RAxxIU7X12 https://t.co/U3JgLPPcnf


TL;DR: Stanford CS336 aims to have students build language models from scratch, deeply understand the full-stack design of data, systems, and models, and emphasize algorithmic efficiency at scale.

Course Introduction: Building Language Models from Scratch

CS336 is a Stanford University course whose core goal is to let students build a complete language model by hand — from the data system to the modeling. The course is co-taught by Percy and Tatsu, with a TA team including Roit, Neil, and Marcel. As Percy said: “To understand it, you must build it.” All lectures for this course are uploaded to YouTube for learners worldwide.

Why This Course? Breaking Researchers’ “Abstract Black Box”

Percy points out a current crisis: researchers are becoming increasingly disconnected from the underlying technology. Eight years ago, researchers would implement and train AI models themselves; six years ago, they would at least download BERT and fine-tune it. Now, many people accomplish tasks just by prompting proprietary models. While abstraction layers bring convenience, these abstractions are leaky — you don’t truly understand that it’s essentially just “strings in, strings out.” Foundational research requires breaking down the entire tech stack and co-designing data, systems, and models. The purpose of this course is to keep foundational research alive.

Limitations of Small Models: Two Manifestations of Scale Changes

Industrial language models are astonishing in scale: GPT-4 is rumored to have 1.8 trillion parameters with a training cost of up to $100 million; XAI is building a cluster with 200,000 H100s; over $500 billion in investments are planned for the next four years. But frontier models are out of reach for most people, so the course focuses on small language models. However, small models may not be representative for two reasons:

  1. Change in FLOPs ratios: In small Transformers, the attention layer and MLP layer have roughly equivalent FLOPs; but when parameters scale to 175 billion, MLP layers completely dominate. If you only optimize attention at small scale, its effect is swamped at large scale.

  2. Emergent behaviors: A 2022 paper by Jason Wei shows that as training FLOPs increase, the accuracy of certain tasks (like in-context learning) jumps suddenly at a critical point. If you stay only at small scale, you might incorrectly conclude that “language models don’t work.”

What the Course Can Teach: Mechanisms, Mindsets, and Intuitions

Percy categorizes knowledge into three types:

  • Mechanisms (teachable): Transformer implementation, model parallelism, efficient GPU utilization, etc.
  • Mindsets (more important): Squeezing performance out of hardware, taking scaling seriously. This mindset was pioneered by OpenAI and is key to leading the next generation of AI models.
  • Intuition (only partially teachable): Which data and modeling decisions lead to good models. Because architectures and datasets that work well at small scale may fail at large scale. Still, getting two and a half out of three is a good deal.

On intuition, Percy cites a paper introducing the Swish activation function, whose conclusion was: “We cannot explain it, we can only say it’s God’s grace” — the experiments speak for themselves.

The Real Meaning of the Bitter Lesson: Algorithms Matter at Scale

The “bitter lesson” is often misinterpreted as “scale is everything, algorithms don’t matter.” Percy believes this is completely wrong. The correct interpretation is: Model accuracy = Efficiency × Resources invested. Efficiency is even more important at large scale because when you spend hundreds of millions of dollars, you can’t afford to waste resources like you would on a local cluster. A 2020 OpenAI paper showed that algorithm efficiency for training to a certain accuracy on ImageNet improved by 44x between 2012 and 2019 (exceeding Moore’s Law). Similar results exist for language models.

Therefore, the right framework is: given a certain compute and data budget, what is the best model you can build? As a researcher, the goal is to improve the efficiency of algorithms.

Historical Review: Evolution of Language Models

  • Early days: Shannon used language models to estimate English entropy; in 2007, Google trained a five-gram n-gram model based on 2 trillion tokens (more tokens than GPT-3), but it lacked emergent behavior.
  • 2010s: Deep learning revolution — Yoshua Bengio’s first neural language model in 2003, sequence-to-sequence models (Ilya et al.), Adam optimizer, attention mechanisms, 2017’s “Attention Is All You Need” (Transformer), expansion of mixture-of-experts models, model parallelism techniques (already capable of training models with hundreds of billions of parameters).
  • Foundation model trend: ELMo, BERT, T5 pretrained on large-scale text and adapted to downstream tasks.
  • Key inflection point: OpenAI combined components and pushed scaling laws, producing GPT-2, GPT-3. Since then, it has split into closed-source models (accessed via APIs) and open-source models (e.g., EleutherAI, Meta, Bloom, Alibaba, DeepSeek, etc.). Openness levels include: closed-source, open weights (architecture details but no dataset details), open-source (all weights and data available).

Current Landscape and Course Approach

Current frontier models come from OpenAI, Anthropic, xAI, Google, Meta, DeepSeek, Alibaba, Tencent, etc. The course will revisit the technical principles of these components, get as close as possible to best practices of frontier models, while leveraging open-source community information and inferences from closed-source models. The course adopts an executable lecture format, embedding code to explain step by step.

Source: https://www.youtube.com/watch?v=SQ3fZ1sAqXI

Similar Articles