Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers
Summary
The paper presents a modular pipeline for extracting structured text and metadata from historical newspaper scans, yielding an open dataset of billions of tokens from millions of scans.
View Cached Full Text
Cached at: 09/03/26, 07:51 AM
Paper page - Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers
Source: https://huggingface.co/papers/2608.18972
Abstract
A modular pipeline extracts structured text and metadata from historical newspaper scans using small interpretable models, yielding a large open dataset.
Historical newspapers are an abundant record of public life, but their dense, irregular and sometimes noisy layouts make computational access to these materials both challenging and limited. We present the Institutional Newspapers Pipeline, a modular system we jointly designed with Boston Public Library to extract high-quality, structured datasets from historical newspaper scans. It was architected so that each step remains interpretable and customizable, and so that the pipeline as a whole remains computationally frugal enough to run on workstation-level hardware. The pipeline runs each scan through a multi-step process: it segments scans into individual type-agnostic crops and performsOCRon each resulting segment before then performing text analysis,type classification,reading order detection, named entities recognition,subject classification,language detection, and pre-computedembeddings generationon every crop. We ran this pipeline against a portion of Boston Public Library’s holdings and released the results as an open dataset. The optical character recognition (OCR) output represents 16.3 billion o200k_base tokens across 83.1 million individual crops, extracted from 1,473,635 public domain newspaper scans published between 1795 and 1930. This report describes our methods for each processing step, the small models we trained, as well as the evaluation results and dataset-scale measurements we collected in the process. It accompanies the release of the pipeline, models, and dataset. We position this work as a substantial step towards unlocking high-quality data from tens of millions of newspaper scans.
View arXiv pageView PDFProject pageAdd to collection
Models citing this paper3
#### institutional/institutional-newspapers-crop-classifier-image-yolo26m-cls Image Classification• Updated14 days ago • 1
#### institutional/institutional-newspapers-crop-classifier-text-model2vec Text Classification• 32.4M• Updatedabout 19 hours ago • 12 • 1
#### institutional/institutional-newspapers-segmenter-yolo26x Updated14 days ago • 1 • 1
Datasets citing this paper1
#### institutional/institutional-newspapers-bpl Viewer• Updated14 days ago • 1.47M • 1.28k • 2
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.18972 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
Leveraging Morphology for Historical Script Metrological Analysis
This paper presents a transformer-based architecture with prototype learning that enables scalable paleographic measurements from historical documents using only line-level transcriptions, demonstrating effectiveness on a 160-page codex with minimal training data.
Built an AI pipeline that transforms financial news into structured analysis
Built an AI pipeline that converts financial news into structured analysis including sentiment, risks, and opportunities, focusing on consistency through prompt engineering and validation.
N Newsletters to 1 Digest, Built for AI Engineers
A writeup of a multi-stage LLM pipeline that fetches newsletter articles, summarizes, de-duplicates, scores importance and urgency, and publishes a digest, with robust prompt-injection defenses and self-validation.
Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics
This paper introduces Sampled-BPE, a lightweight token-level auditing pipeline for web-scale Chinese corpora, and applies it to reveal widespread but uneven pollution across open Chinese datasets and Common Crawl snapshots. It also releases a hierarchical dataset of over 630k polluted token records with context and explanations.
Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application
This paper benchmarks open-source OCR, LLM, and VLM systems for structured information extraction in a high-risk public sector application, finding that VLMs generally outperform OCR+LLM pipelines but most configurations struggle in zero-shot settings, emphasizing the critical role of input quality.