Induction Heads Interpolate N-Grams

arXiv cs.LG Papers

Summary

This paper studies transformers trained on Markov chains and identifies that induction heads implement soft context-matching and Dirichlet-style smoothing, showing that transformers regularize in-context estimation rather than simply counting n-grams.

arXiv:2607.02800v1 Announce Type: new Abstract: Induction heads are attention circuits believed to underlie in-context learning in transformers, yet a precise characterization of the estimators they implement remains elusive. We study transformers trained on order-$k$ Markov chains and identify two complementary smoothing mechanisms. First, at finite attention-weight scale, the circuit implements a soft context-matching estimator: it aggregates contributions from exact and partial context matches, weighted exponentially by their overlap, and induces a data-dependent interpolation across context orders analogous to Jelinek-Mercer smoothing. Second, a beginning-of-sequence (BOS) token induces additive pseudo-counts, recovering Dirichlet-style smoothing. We construct a disentangled transformer implementing both mechanisms and show that trained transformers recover the predicted attention patterns. Across settings where pseudo-count smoothing is optimal or lower-order contexts provide structured evidence, trained transformers match or outperform classical count-based baselines. Our results bridge mechanistic interpretability of induction heads with classical statistical smoothing, revealing that transformers learn to regularize in-context estimation rather than simply count.
Original Article
View Cached Full Text

Cached at: 07/07/26, 04:39 AM

# Induction Heads Interpolate N-Grams
Source: [https://arxiv.org/abs/2607.02800](https://arxiv.org/abs/2607.02800)
[View PDF](https://arxiv.org/pdf/2607.02800)

> Abstract:Induction heads are attention circuits believed to underlie in\-context learning in transformers, yet a precise characterization of the estimators they implement remains elusive\. We study transformers trained on order\-$k$ Markov chains and identify two complementary smoothing mechanisms\. First, at finite attention\-weight scale, the circuit implements a soft context\-matching estimator: it aggregates contributions from exact and partial context matches, weighted exponentially by their overlap, and induces a data\-dependent interpolation across context orders analogous to Jelinek\-Mercer smoothing\. Second, a beginning\-of\-sequence \(BOS\) token induces additive pseudo\-counts, recovering Dirichlet\-style smoothing\. We construct a disentangled transformer implementing both mechanisms and show that trained transformers recover the predicted attention patterns\. Across settings where pseudo\-count smoothing is optimal or lower\-order contexts provide structured evidence, trained transformers match or outperform classical count\-based baselines\. Our results bridge mechanistic interpretability of induction heads with classical statistical smoothing, revealing that transformers learn to regularize in\-context estimation rather than simply count\.

## Submission history

From: Francesco D'Angelo \[[view email](https://arxiv.org/show-email/bdc7f83e/2607.02800)\] **\[v1\]**Thu, 2 Jul 2026 22:19:26 UTC \(1,477 KB\)

Similar Articles

Comparing Transformers and Hybrid Models at the Token Level

Lobsters Hottest

This paper analyzes token-level prediction differences between transformers and hybrid attention-recurrent models using Olmo 3 and Olmo Hybrid, finding that hybrids improve on semantic state tracking while transformers excel at n-gram copying and syntactic bracket matching.