Tag
This paper presents a Bayesian model of intercomprehension—understanding a related language without training—using a noisy-channel approach. It compares the model's predictions to human behavior across three language pairs, showing better alignment than larger zero-shot models.
The paper proposes a Shannon Scaling Law that models LLM training as information transmission over a noisy channel, explaining non-monotonic performance phenomena like catastrophic overtraining and quantization-induced degradation, and demonstrating superior predictive accuracy over traditional scaling laws.