Tag
This paper measures the self-referential evaluation loop when a knowledge base is used as the gold standard for entity-level machine translation in low-resource historical domains, showing that gains from KB injection are confined to overlapping segments and do not reflect true translation quality.
This paper introduces ARI, a framework that uses retrieval-augmented large language models to restore illegible portions of historical documents, significantly improving named entity restoration by combining implicit LLM knowledge with explicitly retrieved external historical context.
This paper presents the HIPE-OCRepair-2026 competition at ICDAR 2026, evaluating LLM-assisted OCR post-correction for historical documents in English, French, and German. Results show that modern LLM systems significantly improve OCR quality, but overcorrection on low-noise inputs remains a challenge.
This paper presents DistilledGemma, a system for person-place relation extraction from multilingual historical newspaper articles using a three-stage knowledge distillation pipeline from a 26B Gemma teacher to a 2.3B student, achieving competitive accuracy and efficiency in the HIPE-2026 shared task.
This paper introduces three datasets (Hell-Char, PaLit-Char, Med-Char) for diachronic representation learning of ancient Greek letterforms and proposes a similarity-weighted supervised contrastive loss with lacuna-driven augmentation to robustly learn character embeddings across centuries of handwriting variation.
This paper presents a transformer-based architecture with prototype learning that enables scalable paleographic measurements from historical documents using only line-level transcriptions, demonstrating effectiveness on a 160-page codex with minimal training data.