historical-language-models

Tag

Cards List
#historical-language-models

Pretraining Language Models on Historical Text

arXiv cs.CL · 2026-06-03 Cached

This paper introduces TypewriterLM, a 7.24B parameter language model trained exclusively on English text predating 1913, along with TypewriterCorpus (a 54B-token cleaned historical corpus) and instruction-tuning datasets to avoid temporal leakage and lookahead bias. It also presents a benchmark suite, History-Event, for evaluating temporal grounding and leakage.

0 favorites 0 likes
← Back to home

Submit Feedback