Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers

Hugging Face Daily Papers Papers

Summary

The paper presents a modular pipeline for extracting structured text and metadata from historical newspaper scans, yielding an open dataset of billions of tokens from millions of scans.

Historical newspapers are an abundant record of public life, but their dense, irregular and sometimes noisy layouts make computational access to these materials both challenging and limited. We present the Institutional Newspapers Pipeline, a modular system we jointly designed with Boston Public Library to extract high-quality, structured datasets from historical newspaper scans. It was architected so that each step remains interpretable and customizable, and so that the pipeline as a whole remains computationally frugal enough to run on workstation-level hardware. The pipeline runs each scan through a multi-step process: it segments scans into individual type-agnostic crops and performs OCR on each resulting segment before then performing text analysis, type classification, reading order detection, named entities recognition, subject classification, language detection, and pre-computed embeddings generation on every crop. We ran this pipeline against a portion of Boston Public Library's holdings and released the results as an open dataset. The optical character recognition (OCR) output represents 16.3 billion o200k_base tokens across 83.1 million individual crops, extracted from 1,473,635 public domain newspaper scans published between 1795 and 1930. This report describes our methods for each processing step, the small models we trained, as well as the evaluation results and dataset-scale measurements we collected in the process. It accompanies the release of the pipeline, models, and dataset. We position this work as a substantial step towards unlocking high-quality data from tens of millions of newspaper scans.
Original Article
View Cached Full Text

Cached at: 09/03/26, 07:51 AM

Paper page - Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers

Source: https://huggingface.co/papers/2608.18972

Abstract

A modular pipeline extracts structured text and metadata from historical newspaper scans using small interpretable models, yielding a large open dataset.

Historical newspapers are an abundant record of public life, but their dense, irregular and sometimes noisy layouts make computational access to these materials both challenging and limited. We present the Institutional Newspapers Pipeline, a modular system we jointly designed with Boston Public Library to extract high-quality, structured datasets from historical newspaper scans. It was architected so that each step remains interpretable and customizable, and so that the pipeline as a whole remains computationally frugal enough to run on workstation-level hardware. The pipeline runs each scan through a multi-step process: it segments scans into individual type-agnostic crops and performsOCRon each resulting segment before then performing text analysis,type classification,reading order detection, named entities recognition,subject classification,language detection, and pre-computedembeddings generationon every crop. We ran this pipeline against a portion of Boston Public Library’s holdings and released the results as an open dataset. The optical character recognition (OCR) output represents 16.3 billion o200k_base tokens across 83.1 million individual crops, extracted from 1,473,635 public domain newspaper scans published between 1795 and 1930. This report describes our methods for each processing step, the small models we trained, as well as the evaluation results and dataset-scale measurements we collected in the process. It accompanies the release of the pipeline, models, and dataset. We position this work as a substantial step towards unlocking high-quality data from tens of millions of newspaper scans.

View arXiv pageView PDFProject pageAdd to collection

Models citing this paper3

#### institutional/institutional-newspapers-crop-classifier-image-yolo26m-cls Image Classification• Updated14 days ago • 1 #### institutional/institutional-newspapers-crop-classifier-text-model2vec Text Classification• 32.4M• Updatedabout 19 hours ago • 12 • 1 #### institutional/institutional-newspapers-segmenter-yolo26x Updated14 days ago • 1 • 1

Datasets citing this paper1

#### institutional/institutional-newspapers-bpl Viewer• Updated14 days ago • 1.47M • 1.28k • 2

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2608.18972 in a Space README.md to link it from this page.

Collections including this paper1

Similar Articles

Leveraging Morphology for Historical Script Metrological Analysis

Hugging Face Daily Papers

This paper presents a transformer-based architecture with prototype learning that enables scalable paleographic measurements from historical documents using only line-level transcriptions, demonstrating effectiveness on a 160-page codex with minimal training data.

N Newsletters to 1 Digest, Built for AI Engineers

Reddit r/AI_Agents

A writeup of a multi-stage LLM pipeline that fetches newsletter articles, summarizes, de-duplicates, scores importance and urgency, and publishes a digest, with robust prompt-injection defenses and self-validation.

Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics

arXiv cs.CL

This paper introduces Sampled-BPE, a lightweight token-level auditing pipeline for web-scale Chinese corpora, and applies it to reveal widespread but uneven pollution across open Chinese datasets and Common Crawl snapshots. It also releases a hierarchical dataset of over 630k polluted token records with context and explanations.