cross-entropy

Tag

Cards List
#cross-entropy

@_yusufknl: In 1948, Claude Shannon invented the math behind every LLM you use today. He tested it by making his wife guess the nex…

X AI KOLs Timeline · 4d ago Cached

A detailed walkthrough explains how Claude Shannon's 1948 information theory underlies LLMs and shows that the 'next-token prediction' story is misleading, linking compression and prediction mathematically.

0 favorites 0 likes
#cross-entropy

@_yusufknl: As someone who's been shipping LLMs since the GPT-2 days, this lecture on cross-entropy from a Stanford math grad is th…

X AI KOLs Timeline · 5d ago Cached

A practitioner recommends a free 33-minute lecture on cross-entropy that reframes language models as compression rather than next-word prediction, likening it to a Stanford ML PhD qualifier.

0 favorites 0 likes
#cross-entropy

Conservation Laws for Diffusion Models

arXiv cs.LG · 2026-07-14 Cached

This paper develops conservation laws for diffusion models using generalized extrinsic information transfer (GEXIT) functions, showing that the cross-entropy can be characterized as an integral of local information-theoretic derivatives along the noise path, unifying likelihood characterization for discrete and continuous diffusion.

0 favorites 0 likes
#cross-entropy

What Does the Weight Norm Control in Grokking? Logit-Scale Mediation under Cross-Entropy

arXiv cs.LG · 2026-06-18 Cached

The paper investigates whether weight norm directly controls the grokking delay in neural networks or if its effect is mediated by logit scale and softmax saturation under cross-entropy loss. Experiments show that the delay is almost entirely explained by the effective logit scale, with weight norm contributing negligibly.

0 favorites 0 likes
#cross-entropy

Neural Collapse by Design: Learning Class Prototypes on the Hypersphere

arXiv cs.LG · 2026-05-21 Cached

This paper shows that cross-entropy and supervised contrastive learning are both forms of prototype learning on the hypersphere and proposes normalized losses (NTCE and NONL) that achieve Neural Collapse by design, outperforming standard methods.

0 favorites 0 likes
#cross-entropy

A Closed-Form Upper Bound for Admissible Learning-Rate Steps in Belief-Space Dynamics

arXiv cs.LG · 2026-05-11 Cached

This paper derives a closed-form upper bound for admissible learning-rate steps in belief-space dynamics using KL divergence and Bregman geometry, focusing on cross-entropy classification.

0 favorites 0 likes
← Back to home

Submit Feedback