@github: New open source dataset for your next build. The GitHub Multilingual Repositories Dataset spans 40M+ repos and 80M+ cla…
Summary
GitHub released the Multilingual Repositories Dataset, covering over 40 million repositories and 80 million classification rows, with insights into non-English READMEs, issues, and PRs. Korean leads in issue text while Portuguese tops READMEs.
View Cached Full Text
Cached at: 07/11/26, 09:25 AM
New open source dataset for your next build. 📊
The GitHub Multilingual Repositories Dataset spans 40M+ repos and 80M+ classification rows, showing where non-English READMEs, issues, and PRs live.
Korean tops issue text; Portuguese tops READMEs (3M+ repos).
Similar Articles
Accelerating researchers and developers building multilingual AI with a new open dataset (7 minute read)
GitHub announces the GitHub Multilingual Repositories Dataset, an open metadata dataset covering over 80 million classification rows across 40 million repositories to help researchers and developers build multilingual AI tools.
@DivyanshT91162: I think GitHub repos are quietly replacing half the SaaS industry. This one turns a prompt into a live dataset. Type: "…
Promotes a GitHub repo that lets users describe a dataset in natural language and have AI agents research the web to build a structured table, exportable to CSV, with automatic refreshes.
RedBench: A Universal Dataset for Comprehensive Red Teaming of Large Language Models
RedBench introduces a universal dataset aggregating 37 benchmark datasets with 29,362 samples across 22 risk categories and 19 domains to enable standardized and comprehensive red teaming evaluation of large language models. The work addresses inconsistencies in existing red teaming datasets and provides baselines, evaluation code, and open-source resources for assessing LLM robustness against adversarial prompts.
@DataScienceDojo: Google's 𝐥𝐚𝐧𝐠𝐞𝐱𝐭𝐫𝐚𝐜𝐭 has crossed 37k stars on GitHub. The core idea: point an LLM at unstructured text and g…
Google's open-source tool 'langextract' uses LLMs to extract structured data from unstructured text with grounded character positions, crossing 37k GitHub stars.
@Fluyeporlaweb: The 10 repos that have grown the fastest this June on GitHub: 1. pewdiepie-archdaemon/odysseus PewDiePie (111M subscrib…
A Twitter thread lists the 10 fastest-growing repositories on GitHub in June 2025, covering AI workspaces, token compression, agent prompt optimization, video generation, voice cloning, stock analysis, research agents, and more.