corpus

Tag

Cards List
#corpus

A Human-in-the-Loop Corpus for LLM-Based Simplification of Scientific Summaries

arXiv cs.CL · 12h ago Cached

This paper presents a human-in-the-loop workflow for creating a corpus of LLM-based simplifications of scientific summaries, using SciSummNet and GPT-4o-mini with non-expert and expert feedback to improve cross-disciplinary accessibility.

0 favorites 0 likes
#corpus

A Corpus of Persuasion Techniques in Slavic Languages

arXiv cs.CL · 2026-07-14 Cached

A new annotated corpus of persuasion techniques in Bulgarian, Polish, and Russian, covering parliamentary debates and social media, with 25 fine-grained techniques and baseline models for detection and classification.

0 favorites 0 likes
#corpus

Echoes Across Vietnam's Highlands, Delta, and Coast: A Multilingual Corpus for Cham, Khmer, and Tay-Nung

arXiv cs.CL · 2026-07-10 Cached

Introduces CKTN, the first multilingual corpus and benchmark for Cham, Khmer, and Tay-Nung languages, addressing fragmentation in multilingual encoders with a script-aware adaptation recipe.

0 favorites 0 likes
#corpus

A Word-Level Digital Reader of the Prasthanatrayi with Sankara's Bhasya: Corpus, Method, and an Open, Offline Reading Aid for the Advaita Vedanta Canon

arXiv cs.CL · 2026-07-09 Cached

Presents an open, offline word-level digital reader of the Prasthānatrayī with Śaṅkara's Bhāṣya, featuring clickable word analysis, concordance, and a hybrid pipeline using rule-based and LLM-assisted methods.

0 favorites 0 likes
#corpus

RMISC: A Large-scale Real-world Multivariate Corpus for Time Series Foundation Models

arXiv cs.AI · 2026-07-08 Cached

Introduces RMISC, a large-scale real-world multivariate time series corpus with around 200 datasets and 142 billion time points, and demonstrates that pretraining time series foundation models on real-world multivariate data improves zero-shot generalization compared to synthetic data.

0 favorites 0 likes
#corpus

I built an open, from-scratch MT pipeline + parallel corpus for Tunisian Darija (Arabizi) early baseline, and I'm growing it into a curated community corpus [P]

Reddit r/MachineLearning · 2026-07-05

An 18-year-old Tunisian student introduces an open-source machine translation pipeline and parallel corpus for Tunisian Darija in Arabizi script, built from scratch with a small 15.6M-parameter Transformer and an honest baseline BLEU of 3.89, and calls for contributors to ethically expand the corpus.

0 favorites 0 likes
#corpus

MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages

arXiv cs.CL · 2026-07-02 Cached

This paper introduces MultiSynt/MT, a trillion-token multilingual parallel corpus created by translating English pre-training data into 36 languages. Experiments show that LLMs trained on this translated data achieve performance comparable to native data with fewer tokens, though some evaluation blind spots and cultural gaps remain.

0 favorites 0 likes
#corpus

SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework

arXiv cs.CL · 2026-07-02 Cached

This paper introduces SEFORA, a public corpus of instructor feedback on student essays, and UniMatch, a reference-based evaluation framework for assessing LLM-generated feedback. Experiments show that current LLMs struggle to match instructor feedback, achieving at most 0.4 F1.

0 favorites 0 likes
#corpus

Do Speech Emphasis Models Generalize across Languages and Emotions?

arXiv cs.CL · 2026-06-29 Cached

Introduces MMEE, a multilingual multi-emotion emphasis corpus of 10,000 utterances across 7 languages and 34 emotions, and benchmarks emphasis detection models under various transfer settings, finding that multilingual training improves robustness while monolingual models show limited zero-shot transfer.

0 favorites 0 likes
#corpus

Introducing corpora Hlava Cor and Hlava AD: Human Label Variation in Coreference and Discourse Relations

arXiv cs.CL · 2026-06-25 Cached

This paper introduces two new Czech corpora, Hlava Cor and Hlava AD, designed to study human label variation in coreference and discourse relations. The corpora feature multiple annotations and annotator explanations, achieving 60-65% inter-annotator agreement and revealing systematic differences in interpretation.

0 favorites 0 likes
#corpus

Prague Dependency Treebank -- Consolidated 2.0: Enriching a Complex Annotation Scheme

arXiv cs.CL · 2026-06-24 Cached

We present the second consolidated version of the Prague Dependency Treebank, a 4-million-token manual multilingual annotation resource covering morphology, syntax, semantics, coreference, and discourse, along with compatible lexicons.

0 favorites 0 likes
#corpus

@geekbb: Organized the quarterly reports, notes, and interviews of fund manager Zheng Xi from over a decade into a structured corpus, built as a traceable AI skill, enabling AI to conduct investment research Q&A and fund analysis based on real data rather than model hallucinations. https://github.com/lyra81604/zhengxi-views…

X AI KOLs Timeline · 2026-06-22 Cached

Compiled the public quarterly reports, notes, and interviews of fund manager Zheng Xi into a structured corpus, and built it as a traceable skill across AI platforms for real data-driven investment research Q&A and fund analysis.

0 favorites 0 likes
#corpus

@barrowjoseph: New paper: every law in America is technically public. But not really, until now! With @DenisPeskoff at UC Berkeley, we…

X AI KOLs Following · 2026-06-19 Cached

A new paper builds and releases a corpus of 2.2 million city and county laws from across the United States, making publicly accessible legal texts that were technically public but hard to find.

0 favorites 0 likes
#corpus

@jacobli99: To compare procedures for machine studying, we start by defining expertise. The corpus is always available at test time…

X AI KOLs Following · 2026-06-17 Cached

Jacob Li introduces 'Machine Studying' as a new problem in continual learning: how AI systems develop expertise in an unfamiliar domain given only a corpus of documents, distinct from avoiding catastrophic forgetting.

0 favorites 0 likes
#corpus

Darshana Graph: A Parallel Commentary Corpus for Comparative Indian Philosophy, with Stylometric and Exploratory Graph Analyses

arXiv cs.CL · 2026-06-17 Cached

This paper introduces Darshana Graph, a parallel commentary corpus for comparative Indian philosophy, and presents stylometric and exploratory graph analyses.

0 favorites 0 likes
#corpus

Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States

Hugging Face Daily Papers · 2026-06-17 Cached

Introduces LOCUS, a comprehensive corpus of U.S. local ordinance codes designed to enable machine-readable legal AI research, covering codes from 9,239 cities and counties with ModernBERT-based classifiers for analysis.

0 favorites 0 likes
#corpus

AAbAAC: An Annotated Corpus for Autoimmunity Information Extraction

arXiv cs.AI · 2026-06-12 Cached

AAbAAC is a manually annotated corpus of 115 PubMed abstracts for autoimmunity information extraction, focusing on entities like autoimmune diseases and autoantibodies. The study demonstrates improved NER performance after fine-tuning on this corpus.

0 favorites 0 likes
#corpus

HKJudge: A Legal Discourse-Annotated Corpus for Interpreting What Courts Find, How They Reason, and What They Rule

arXiv cs.CL · 2026-06-08 Cached

HKJudge is the first sentence-level expert-annotated legal discourse corpus for Hong Kong criminal judgments, featuring a two-tier discourse schema and benchmark evaluations of BERT-based and LLM models.

0 favorites 0 likes
#corpus

Pretraining Language Models on Historical Text

arXiv cs.CL · 2026-06-03 Cached

This paper introduces TypewriterLM, a 7.24B parameter language model trained exclusively on English text predating 1913, along with TypewriterCorpus (a 54B-token cleaned historical corpus) and instruction-tuning datasets to avoid temporal leakage and lookahead bias. It also presents a benchmark suite, History-Event, for evaluating temporal grounding and leakage.

0 favorites 0 likes
#corpus

KletterMix: Climbing Toward High-Quality German Pretraining Data

Hugging Face Daily Papers · 2026-06-02

KletterMix is a high-quality German pretraining corpus built by translating a state-of-the-art English pretraining dataset into German while preserving structure and diversity. Controlled experiments show models trained on KletterMix achieve measurable improvements on German-language benchmarks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback