historical-documents

Tag

Cards List
#historical-documents

When the Knowledge Base Becomes the Gold Standard: Measuring Resource-Shared Evaluation Loops in Entity-Level Machine Translation

arXiv cs.CL · 15h ago Cached

This paper measures the self-referential evaluation loop when a knowledge base is used as the gold standard for entity-level machine translation in low-resource historical domains, showing that gains from KB injection are confined to overlapping segments and do not reflect true translation quality.

0 favorites 0 likes
#historical-documents

Leveraging External Knowledge for Historical Document Restoration via Retrieval-Augmented Large Language Models

arXiv cs.CL · 2026-07-27 Cached

This paper introduces ARI, a framework that uses retrieval-augmented large language models to restore illegible portions of historical documents, significantly improving named entity restoration by combining implicit LLM knowledge with explicitly retrieved external historical context.

0 favorites 0 likes
#historical-documents

ICDAR 2026 HIPE-OCRepair Competition on LLM-Assisted OCR Post-Correction for Historical Documents

arXiv cs.CL · 2026-07-10 Cached

This paper presents the HIPE-OCRepair-2026 competition at ICDAR 2026, evaluating LLM-assisted OCR post-correction for historical documents in English, French, and German. Results show that modern LLM systems significantly improve OCR quality, but overcorrection on low-noise inputs remains a challenge.

0 favorites 0 likes
#historical-documents

DistilledGemma: Balanced Efficiency-Accuracy for Person-Place Relation Extraction from Multilingual Historical Articles

arXiv cs.CL · 2026-06-30 Cached

This paper presents DistilledGemma, a system for person-place relation extraction from multilingual historical newspaper articles using a three-stage knowledge distillation pipeline from a 26B Gemma teacher to a 2.3B student, achieving competitive accuracy and efficiency in the HIPE-2026 shared task.

0 favorites 0 likes
#historical-documents

Learning Diachronic Representations of Ancient Greek Letterforms

arXiv cs.LG · 2026-06-25 Cached

This paper introduces three datasets (Hell-Char, PaLit-Char, Med-Char) for diachronic representation learning of ancient Greek letterforms and proposes a similarity-weighted supervised contrastive loss with lacuna-driven augmentation to robustly learn character embeddings across centuries of handwriting variation.

0 favorites 0 likes
#historical-documents

Leveraging Morphology for Historical Script Metrological Analysis

Hugging Face Daily Papers · 2026-06-08 Cached

This paper presents a transformer-based architecture with prototype learning that enables scalable paleographic measurements from historical documents using only line-level transcriptions, demonstrating effectiveness on a 160-page codex with minimal training data.

0 favorites 0 likes
← Back to home

Submit Feedback