Tag
The paper presents a modular pipeline for extracting structured text and metadata from historical newspaper scans, yielding an open dataset of billions of tokens from millions of scans.