Tag
The article discusses the process of utilizing open training data from Hugging Face to train Marin’s 535B model, which involved 25T tokens from 152 datasets with permissible licenses.
The article presents a tool that displays the release age and training cutoff dates for 20 current AI models, highlighting how stale the data might be in each model.
This paper investigates membership inference in language models by using exact duplication counts from open pretraining corpora of OLMo-2 and Pythia, revealing that typical duplication levels show minimal exposure traces and that apparent membership signals are often confounded with factors like sentence fame.
Planet Labs has announced an open satellite feed, providing public access to its satellite imagery data.
This paper introduces a large open corpus of soil compaction tests and establishes machine-learning baselines for predicting compaction parameters, emphasizing physics-constrained models for practical screening.
Reactor Atlas is an interactive map and database that catalogs all nuclear reactors on Earth since 1954, featuring technical records, animated explanations, and real-time monitoring of outages and events.
SolarWM introduces an open framework and unified training recipe for building interactive video world models with scalable training across diverse data sources, enabling long-horizon real-time rollouts.
Gregory Kurtzer, founder of Rocky Linux, announces OpenWALDO, a project aimed at opening up AI training data to foster true open-source AI development.
World Train Map is an interactive atlas of 1,247 notable train routes across 120 countries, recently renamed from TrainRouter. It offers route facts, statistics, and an open dataset.
Daniel van Strien uploaded a dataset of 1,080,814 public domain images from 49,455 digitised books (c.1510–1900) from the British Library to Hugging Face Hub, organised into four configs by image type.
This paper argues that Latin America lacks the datasets and benchmarks layers needed to build its own AI, and proposes DataHub, an open, incentive-driven platform with a task-first ontology to index and contribute regional datasets.
This freeCodeCamp handbook teaches readers how to build a production-grade ETL pipeline in Python using real flood data, covering incremental loading, type coercion, deduplication, and idempotency.
An MIT-developed medical database evolved into PhysioNet, a global biomedical data repository used by thousands of researchers worldwide, as detailed in a recent Nature Health paper.
Folding Globes offers free printable paper globes built from openly licensed data, with themes and a globe builder, as a MapScaping project.
PhononBench-MP40 is a benchmark dataset of phonon stability labels and spectra for over 46,000 Materials Project-derived crystals, designed to evaluate workflow-defined phonon stability and support materials screening.
The Book Prize Index is a searchable index of award-winning books, designed to make literary prize records open and useful to readers, researchers, librarians, and publishers.
A 3D indoor map of Shinjuku Station built with Three.js using open government data from Japan's Ministry of Land.
The New York City Office of Technology and Innovation publishes a weekly updated dataset of building footprints for all 1.08 million buildings in NYC, providing a detailed and current geospatial foundation for city data.
OpenTK voegt internetconsultaties en notificaties toe, waarmee gebruikers vroegtijdig op de hoogte worden gebracht van nieuwe wetgevingstrajecten.
A data-driven visualization project that reconstructs every match of FIFA World Cup 2026 from recorded events, offering an impression built entirely from data.