Tag
This paper presents a human-in-the-loop workflow for creating a corpus of LLM-based simplifications of scientific summaries, using SciSummNet and GPT-4o-mini with non-expert and expert feedback to improve cross-disciplinary accessibility.
A new annotated corpus of persuasion techniques in Bulgarian, Polish, and Russian, covering parliamentary debates and social media, with 25 fine-grained techniques and baseline models for detection and classification.
Introduces CKTN, the first multilingual corpus and benchmark for Cham, Khmer, and Tay-Nung languages, addressing fragmentation in multilingual encoders with a script-aware adaptation recipe.
Presents an open, offline word-level digital reader of the Prasthānatrayī with Śaṅkara's Bhāṣya, featuring clickable word analysis, concordance, and a hybrid pipeline using rule-based and LLM-assisted methods.
Introduces RMISC, a large-scale real-world multivariate time series corpus with around 200 datasets and 142 billion time points, and demonstrates that pretraining time series foundation models on real-world multivariate data improves zero-shot generalization compared to synthetic data.
An 18-year-old Tunisian student introduces an open-source machine translation pipeline and parallel corpus for Tunisian Darija in Arabizi script, built from scratch with a small 15.6M-parameter Transformer and an honest baseline BLEU of 3.89, and calls for contributors to ethically expand the corpus.
This paper introduces MultiSynt/MT, a trillion-token multilingual parallel corpus created by translating English pre-training data into 36 languages. Experiments show that LLMs trained on this translated data achieve performance comparable to native data with fewer tokens, though some evaluation blind spots and cultural gaps remain.
This paper introduces SEFORA, a public corpus of instructor feedback on student essays, and UniMatch, a reference-based evaluation framework for assessing LLM-generated feedback. Experiments show that current LLMs struggle to match instructor feedback, achieving at most 0.4 F1.
Introduces MMEE, a multilingual multi-emotion emphasis corpus of 10,000 utterances across 7 languages and 34 emotions, and benchmarks emphasis detection models under various transfer settings, finding that multilingual training improves robustness while monolingual models show limited zero-shot transfer.
This paper introduces two new Czech corpora, Hlava Cor and Hlava AD, designed to study human label variation in coreference and discourse relations. The corpora feature multiple annotations and annotator explanations, achieving 60-65% inter-annotator agreement and revealing systematic differences in interpretation.
We present the second consolidated version of the Prague Dependency Treebank, a 4-million-token manual multilingual annotation resource covering morphology, syntax, semantics, coreference, and discourse, along with compatible lexicons.
Compiled the public quarterly reports, notes, and interviews of fund manager Zheng Xi into a structured corpus, and built it as a traceable skill across AI platforms for real data-driven investment research Q&A and fund analysis.
A new paper builds and releases a corpus of 2.2 million city and county laws from across the United States, making publicly accessible legal texts that were technically public but hard to find.
Jacob Li introduces 'Machine Studying' as a new problem in continual learning: how AI systems develop expertise in an unfamiliar domain given only a corpus of documents, distinct from avoiding catastrophic forgetting.
This paper introduces Darshana Graph, a parallel commentary corpus for comparative Indian philosophy, and presents stylometric and exploratory graph analyses.
Introduces LOCUS, a comprehensive corpus of U.S. local ordinance codes designed to enable machine-readable legal AI research, covering codes from 9,239 cities and counties with ModernBERT-based classifiers for analysis.
AAbAAC is a manually annotated corpus of 115 PubMed abstracts for autoimmunity information extraction, focusing on entities like autoimmune diseases and autoantibodies. The study demonstrates improved NER performance after fine-tuning on this corpus.
HKJudge is the first sentence-level expert-annotated legal discourse corpus for Hong Kong criminal judgments, featuring a two-tier discourse schema and benchmark evaluations of BERT-based and LLM models.
This paper introduces TypewriterLM, a 7.24B parameter language model trained exclusively on English text predating 1913, along with TypewriterCorpus (a 54B-token cleaned historical corpus) and instruction-tuning datasets to avoid temporal leakage and lookahead bias. It also presents a benchmark suite, History-Event, for evaluating temporal grounding and leakage.
KletterMix is a high-quality German pretraining corpus built by translating a state-of-the-art English pretraining dataset into German while preserving structure and diversity. Controlled experiments show models trained on KletterMix achieve measurable improvements on German-language benchmarks.