digital-humanities

Tag

Cards List
#digital-humanities

AI Historian: Helping historians organize and verify person-centred temporal clues from dispersed historical narratives

arXiv cs.CL · 2026-09-01 Cached

AI Historian is an AI agent system that helps historians organize and verify person-centred temporal clues from dispersed historical narratives, reducing the cost of historical research while achieving high accuracy in temporal localization.

0 favorites 0 likes
#digital-humanities

Do Evaluation Metrics Detect Errors in Classical Chinese to English Translations?

arXiv cs.CL · 2026-08-11 Cached

This paper investigates whether automatic evaluation metrics for machine translation are reliable for Classical Chinese to English translation, using a diagnostic framework based on minimal pairs. It finds all metrics have blind spots, with MetricX24 performing best overall.

0 favorites 0 likes
#digital-humanities

Ancient Library – 1,060 Greek/Latin texts, click any word to parse it

Hacker News Top · 2026-08-07 Cached

Ancient Library is an online reader offering 1,060 Greek and Latin classics where clicking any word displays its lemma, morphology, and full dictionary entries.

0 favorites 0 likes
#digital-humanities

@vanstriendaniel: Uploaded a dataset of 1,080,814 public domain images, mostly from 19th-century books, to the Hub. https://huggingface.c…

X AI KOLs Following · 2026-08-06 Cached

Daniel van Strien uploaded a dataset of 1,080,814 public domain images from 49,455 digitised books (c.1510–1900) from the British Library to Hugging Face Hub, organised into four configs by image type.

0 favorites 0 likes
#digital-humanities

From Inline Notes to Collected Commentaries: Toward Context-Preserving Organization of Exegetical Knowledge in Classical Chinese Texts

arXiv cs.CL · 2026-08-03 Cached

This paper presents a computational framework for automatically compiling collected commentaries on classical Chinese texts, preserving contextual dependencies of inline notes via prompt chaining and cross-source clustering.

0 favorites 0 likes
#digital-humanities

Challenges in annotations by humans and LLMs: A case study of evaluative language

arXiv cs.CL · 2026-07-31 Cached

This paper compares human linguists in training, a trained linguist, and LLMs on annotating evaluative language using Appraisal theory, finding that LLMs achieve strong performance and can assist in complex annotation tasks.

0 favorites 0 likes
#digital-humanities

Beyond Similarity: Grounded Agentic Extraction and Expert-Adjudicated Evaluation of Intertextuality in Classical Chinese Histories

arXiv cs.CL · 2026-07-31 Cached

This paper recasts fine-grained intertextuality extraction in Classical Chinese histories as an agentic LLM task, grounding reuse in exact character spans and a five-dimension typology, validated by expert-adjudicated benchmarks and scaled to the Twenty-Four Histories.

0 favorites 0 likes
#digital-humanities

TimeCapsule: Generative Hallucination as a Method for Historical Sensemaking

arXiv cs.CL · 2026-07-29 Cached

The paper introduces TimeCapsule, a 1.2B-parameter LLM trained exclusively on Victorian texts (1800-1875) to generate historically plausible analogies for modern concepts, achieving a 45.4% perplexity reduction over GPT-2 on Victorian prose while exhibiting computational sensemaking through structured ignorance of the future.

0 favorites 0 likes
#digital-humanities

Automatic Thematic Indexing of Large Literary Corpora: A Machine Learning Approach to Voltaire's Complete Works

arXiv cs.CL · 2026-07-13 Cached

This paper explores machine learning approaches, including encoder-based models and fine-tuned LLMs, for automatic thematic indexing of large literary corpora, using Voltaire's complete works as a test case. The best model achieves F1 scores up to 0.67, with implications for structured access to large-scale literary and historical editions.

0 favorites 0 likes
#digital-humanities

Letter Lemmatization: One-to-one and Banded RNNs for Reversing Character-Set Simplification and Abbreviation in Medieval Text

arXiv cs.CL · 2026-07-13 Cached

This paper introduces neural approaches for reversing character-set simplification and abbreviation in medieval text, using one-to-one and banded RNNs trained with self-supervision or parallel corpora, and presents a Python library for letter lemmatization.

0 favorites 0 likes
#digital-humanities

ICDAR 2026 HIPE-OCRepair Competition on LLM-Assisted OCR Post-Correction for Historical Documents

arXiv cs.CL · 2026-07-10 Cached

This paper presents the HIPE-OCRepair-2026 competition at ICDAR 2026, evaluating LLM-assisted OCR post-correction for historical documents in English, French, and German. Results show that modern LLM systems significantly improve OCR quality, but overcorrection on low-noise inputs remains a challenge.

0 favorites 0 likes
#digital-humanities

A Word-Level Digital Reader of the Prasthanatrayi with Sankara's Bhasya: Corpus, Method, and an Open, Offline Reading Aid for the Advaita Vedanta Canon

arXiv cs.CL · 2026-07-09 Cached

Presents an open, offline word-level digital reader of the Prasthānatrayī with Śaṅkara's Bhāṣya, featuring clickable word analysis, concordance, and a hybrid pipeline using rule-based and LLM-assisted methods.

0 favorites 0 likes
#digital-humanities

Linking Hadith Narrator Identities Across Heterogeneous Arabic Biographical Databases: A Multi-Signal Entity Resolution Pipeline

arXiv cs.CL · 2026-07-08 Cached

This paper presents a two-phase entity resolution pipeline to link narrator names from the Sanadset corpus to two biographical databases, enabling construction of a large transmission graph enriched with cross-source metadata.

0 favorites 0 likes
#digital-humanities

Integrating knowledge graphs and multilingual scholarly corpora for domain-adaptive LLMs in SSH

arXiv cs.AI · 2026-07-08 Cached

This paper presents a use case from the European project LLMs4EU and ALT-EDIC infrastructure, focusing on adapting foundation models to Social Sciences and Humanities (SSH) research practices by integrating knowledge graphs and multilingual scholarly corpora. The approach aims to support tasks like question answering and literature review while ensuring domain sensitivity and regulatory compliance.

0 favorites 0 likes
#digital-humanities

Structural Pattern Mining in Inka Khipus: Unsupervised Clustering, Provenance Classification, and a Computational Validation of the Santa Valley Match

arXiv cs.CL · 2026-07-02 Cached

This paper presents a reproducible machine-learning pipeline for structural pattern mining in Inka khipus, using unsupervised clustering and supervised classification on the Open Khipu Repository to identify structural groups and validate prior findings, with all code and data openly available.

0 favorites 0 likes
#digital-humanities

A Study of Temporal Fusion Strategies for Named Entity Recognition in Historical Texts

arXiv cs.CL · 2026-06-29 Cached

This paper systematically studies how temporal metadata can be structurally embedded into named entity recognition (NER) models for historical texts. Experiments with absolute and relative temporal representations injected via early or late fusion mechanisms show that late fusion strategies yield more robust performance on French and German historical datasets.

0 favorites 0 likes
#digital-humanities

From Vajrayana Tara to Bengali Baul: A Computational Study of Lexical Transmission Across Buddhist, Shakta, and Vaishnava Traditions in Bengal

arXiv cs.CL · 2026-06-26 Cached

This paper presents a computational analysis of lexical transmission across Buddhist, Shakta, and Vaishnava traditions in Bengal, examining how words and concepts moved between these religious communities.

0 favorites 0 likes
#digital-humanities

Three Buddhist Vocabularies: Computational Stylometry of the English Pali Canon across Sutta, Vinaya, and Abhidhamma

arXiv cs.CL · 2026-06-25 Cached

This paper applies computational stylometry to English translations of the Pali Canon, examining vocabulary differences across the Sutta, Vinaya, and Abhidhamma divisions.

0 favorites 0 likes
#digital-humanities

Predicting Poets' Origins from Verse: A Computational Analysis of Regional Linguistic Fingerprints in the Complete Tang Poems

arXiv cs.CL · 2026-06-24 Cached

This paper uses TF-IDF and neural models on the Complete Tang Poems to predict poets' regional origins from linguistic features, finding detectable regional fingerprints, distance-decay effects, and temporal modulation of the signal.

0 favorites 0 likes
#digital-humanities

Darshana Graph: A Parallel Commentary Corpus for Comparative Indian Philosophy, with Stylometric and Exploratory Graph Analyses

arXiv cs.CL · 2026-06-17 Cached

This paper introduces Darshana Graph, a parallel commentary corpus for comparative Indian philosophy, and presents stylometric and exploratory graph analyses.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback