nlp-dataset

Tag

Cards List
#nlp-dataset

ThaiTrees: Thai Syntactic Dependency Trees Across Domains

arXiv cs.CL ↗ · 2026-09-24 Cached

ThaiTrees introduces a 342M-token automatically parsed corpus of Thai text across domains, using a reproducible pipeline under Universal Dependencies to enable syntactic research and analysis.

0 favorites 0 likes
#nlp-dataset

FineWeb-CLaR: Culture, Language, and Region Annotations for Benchmark-Aligned Corpus Auditing

arXiv cs.CL ↗ · 2026-09-23 Cached

The paper introduces FineWeb-CLaR, a large-scale annotated dataset that labels web documents with culture, language, and region metadata to enable auditing of cultural representation in language model pretraining data and benchmarks.

0 favorites 0 likes
#nlp-dataset

CVSS-X: A Multilingual Speech-to-Speech Translation Corpus for 28 Languages

arXiv cs.CL ↗ · 2026-09-15 Cached

CVSS-X is a large-scale synthetic speech-to-speech translation corpus extending CVSS to enable translation from English into 28 languages, with over 16,000 hours of parallel speech pairs for bidirectional research.

0 favorites 0 likes
← Back to home

Submit Feedback