Tag
The paper provides a unified probabilistic framework for large language models, describing them through probability measures, training via maximum-likelihood estimation, and text generation as stochastic simulation, with insights into phenomena like hallucination and the role of diffusion models.