Tag
ThaiTrees introduces a 342M-token automatically parsed corpus of Thai text across domains, using a reproducible pipeline under Universal Dependencies to enable syntactic research and analysis.
The paper introduces FineWeb-CLaR, a large-scale annotated dataset that labels web documents with culture, language, and region metadata to enable auditing of cultural representation in language model pretraining data and benchmarks.
CVSS-X is a large-scale synthetic speech-to-speech translation corpus extending CVSS to enable translation from English into 28 languages, with over 16,000 hours of parallel speech pairs for bidirectional research.