Tag
AI Historian is an AI agent system that helps historians organize and verify person-centred temporal clues from dispersed historical narratives, reducing the cost of historical research while achieving high accuracy in temporal localization.
This paper investigates whether automatic evaluation metrics for machine translation are reliable for Classical Chinese to English translation, using a diagnostic framework based on minimal pairs. It finds all metrics have blind spots, with MetricX24 performing best overall.
Ancient Library is an online reader offering 1,060 Greek and Latin classics where clicking any word displays its lemma, morphology, and full dictionary entries.
Daniel van Strien uploaded a dataset of 1,080,814 public domain images from 49,455 digitised books (c.1510–1900) from the British Library to Hugging Face Hub, organised into four configs by image type.
This paper presents a computational framework for automatically compiling collected commentaries on classical Chinese texts, preserving contextual dependencies of inline notes via prompt chaining and cross-source clustering.
This paper compares human linguists in training, a trained linguist, and LLMs on annotating evaluative language using Appraisal theory, finding that LLMs achieve strong performance and can assist in complex annotation tasks.
This paper recasts fine-grained intertextuality extraction in Classical Chinese histories as an agentic LLM task, grounding reuse in exact character spans and a five-dimension typology, validated by expert-adjudicated benchmarks and scaled to the Twenty-Four Histories.
The paper introduces TimeCapsule, a 1.2B-parameter LLM trained exclusively on Victorian texts (1800-1875) to generate historically plausible analogies for modern concepts, achieving a 45.4% perplexity reduction over GPT-2 on Victorian prose while exhibiting computational sensemaking through structured ignorance of the future.
This paper explores machine learning approaches, including encoder-based models and fine-tuned LLMs, for automatic thematic indexing of large literary corpora, using Voltaire's complete works as a test case. The best model achieves F1 scores up to 0.67, with implications for structured access to large-scale literary and historical editions.
This paper introduces neural approaches for reversing character-set simplification and abbreviation in medieval text, using one-to-one and banded RNNs trained with self-supervision or parallel corpora, and presents a Python library for letter lemmatization.
This paper presents the HIPE-OCRepair-2026 competition at ICDAR 2026, evaluating LLM-assisted OCR post-correction for historical documents in English, French, and German. Results show that modern LLM systems significantly improve OCR quality, but overcorrection on low-noise inputs remains a challenge.
Presents an open, offline word-level digital reader of the Prasthānatrayī with Śaṅkara's Bhāṣya, featuring clickable word analysis, concordance, and a hybrid pipeline using rule-based and LLM-assisted methods.
This paper presents a two-phase entity resolution pipeline to link narrator names from the Sanadset corpus to two biographical databases, enabling construction of a large transmission graph enriched with cross-source metadata.
This paper presents a use case from the European project LLMs4EU and ALT-EDIC infrastructure, focusing on adapting foundation models to Social Sciences and Humanities (SSH) research practices by integrating knowledge graphs and multilingual scholarly corpora. The approach aims to support tasks like question answering and literature review while ensuring domain sensitivity and regulatory compliance.
This paper presents a reproducible machine-learning pipeline for structural pattern mining in Inka khipus, using unsupervised clustering and supervised classification on the Open Khipu Repository to identify structural groups and validate prior findings, with all code and data openly available.
This paper systematically studies how temporal metadata can be structurally embedded into named entity recognition (NER) models for historical texts. Experiments with absolute and relative temporal representations injected via early or late fusion mechanisms show that late fusion strategies yield more robust performance on French and German historical datasets.
This paper presents a computational analysis of lexical transmission across Buddhist, Shakta, and Vaishnava traditions in Bengal, examining how words and concepts moved between these religious communities.
This paper applies computational stylometry to English translations of the Pali Canon, examining vocabulary differences across the Sutta, Vinaya, and Abhidhamma divisions.
This paper uses TF-IDF and neural models on the Complete Tang Poems to predict poets' regional origins from linguistic features, finding detectable regional fingerprints, distance-decay effects, and temporal modulation of the signal.
This paper introduces Darshana Graph, a parallel commentary corpus for comparative Indian philosophy, and presents stylometric and exploratory graph analyses.