@_yusufknl: As someone who's been shipping LLMs since the GPT-2 days, this lecture on cross-entropy from a Stanford math grad is th…

X AI KOLs Timeline News

Summary

A practitioner recommends a free 33-minute lecture on cross-entropy that reframes language models as compression rather than next-word prediction, likening it to a Stanford ML PhD qualifier.

As someone who's been shipping LLMs since the GPT-2 days, this lecture on cross-entropy from a Stanford math grad is the closest thing to an ML PhD qualifying exam I've ever seen released publicly for free. Everyone thinks language models predict the next word. They don't. They compress language. Once you see the math, you can't unsee it. 33 minutes. Bookmark & watch today.
Original Article
View Cached Full Text

Cached at: 07/20/26, 09:29 AM

As someone who’s been shipping LLMs since the GPT-2 days, this lecture on cross-entropy from a Stanford math grad is the closest thing to an ML PhD qualifying exam I’ve ever seen released publicly for free.

Everyone thinks language models predict the next word. They don’t. They compress language. Once you see the math, you can’t unsee it.

33 minutes. Bookmark & watch today.

Similar Articles