datasets

Tag

Cards List
#datasets

Manufacturing a gold standard eval dataset before launch

Reddit r/AI_Agents · 2026-07-23

A developer shares the challenge of creating a gold standard evaluation dataset for an AI product with no users, considering synthetic data generation and adversarial testing to avoid post-launch restructuring.

0 favorites 0 likes
#datasets

When Does Machine Learning Beat Value Sorting? A Three-Dataset Diagnostic of Exposure-Weighted Shipment Prioritization

arXiv cs.AI · 2026-07-22 Cached

This paper investigates when machine learning outperforms traditional value sorting for exposure-weighted shipment prioritization using three datasets.

0 favorites 0 likes
#datasets

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges

Hugging Face Daily Papers · 2026-07-21 Cached

This survey examines computational humor understanding in multimodal LLMs, covering methods, datasets, evaluation protocols, and challenges such as shortcut-prone evaluation and weak evidence grounding.

0 favorites 0 likes
#datasets

Position: Every Ground Truth is a Human Construction, not an Objective Truth

arXiv cs.LG · 2026-07-14 Cached

This position paper argues that ground truth datasets in machine learning are not objective truths but human constructions shaped by choices, and advocates for articulating these choices to improve reliability, transparency, and accountability.

0 favorites 0 likes
#datasets

@seclink: This project is interesting. It open-sources some datasets that can be used for algorithm training and research.

X AI KOLs Timeline · 2026-07-11 Cached

This project open-sources datasets that can be used for algorithm training and research.

0 favorites 0 likes
#datasets

@HowToPrompt__: someone open-sourced 15TB of physics simulations that would take a national lab and millions in supercomputer time to r…

X AI KOLs Timeline · 2026-07-10 Cached

The Well is an open-source collection of 15TB of physics simulation datasets spanning 16 domains, created by Flatiron Institute and 11 universities, designed to train PDE surrogate models, enabling researchers to replace expensive supercomputer simulations with neural networks.

0 favorites 0 likes
#datasets

Mental Health Disorder Detection Beyond Social Media: A Systematic Review of Available Datasets

arXiv cs.CL · 2026-07-07 Cached

A systematic review of non-social media free-text datasets for mental health disorder detection, identifying biases and gaps in current resources.

0 favorites 0 likes
#datasets

LeRobot v0.6.0: Imagine, Evaluate, Improve

Hugging Face Blog · 2026-07-07 Cached

LeRobot v0.6.0 is a major release of Hugging Face's robot learning library, adding world model policies (VLA-JEPA, FastWAM, LingBot-VA), new VLAs, reward models, six simulation benchmarks, depth sensing, VLM-powered dataset annotation, custom video encoding, cloud training, and a deployment CLI for human-in-the-loop corrections.

0 favorites 0 likes
#datasets

@ClementDelangue: As America turns 250, we put together 250 open AI milestones from the US: open models, datasets, demos, papers, and too…

X AI KOLs Following · 2026-07-04 Cached

HuggingFace CEO Clement Delangue compiled 250 open AI milestones from the US, highlighting contributions like Transformers, PyTorch, BERT, GPT-2, and Llama, with a call to maintain openness in AI development.

0 favorites 0 likes
#datasets

@0x0SojalSec: Awesome cybersecurity datasets for machine learning/Model training, - Network traffic - MaI/ware - Web attacks - Phishi…

X AI KOLs Timeline · 2026-07-02 Cached

A collection of cybersecurity datasets for machine learning and model training, covering network traffic, malware, web attacks, phishing, and more, including notable public research datasets like LANL, CTU-13, and UNSW-NB15.

0 favorites 0 likes
#datasets

@tom_doerr: Curated GNN papers, datasets, and implementation tools https://github.com/dair-ai/GNNs-Recipe…

X AI KOLs Timeline · 2026-06-26 Cached

A curated collection of GNN papers, datasets, and implementation tools, hosted on GitHub.

0 favorites 0 likes
#datasets

AI datasets by their very nature are backward-looking. Creativity by its very nature is forward-looking.

Reddit r/artificial · 2026-06-18

Strauss Zelnick explains that AI is limited by backward-looking data and can reproduce the known but not create breakthroughs, placing value on human decisions about what to build.

0 favorites 0 likes
#datasets

How did China develop AI so quickly recently if most work was done in USA ?

Reddit r/ArtificialInteligence · 2026-06-14

This article discusses how China has rapidly advanced in AI despite being a latecomer, questioning the sources of datasets, computing power, and algorithms that enabled companies like DeepSeek to catch up with US leaders like OpenAI and Google.

0 favorites 0 likes
#datasets

@gui_penedo: Today we’re announcing Macrodata Labs. Over the last few years, @HKydlicek and I have been turning a large part of the …

X AI KOLs Following · 2026-06-11 Cached

Today, Macrodata Labs announced its launch, along with Refiner, an open-source framework for processing robotics datasets. The framework aims to help teams extract more signal from demonstrations and sensor data.

0 favorites 0 likes
#datasets

Before we spend months processing open-source robotics datasets, tell us why this is a bad idea [D]

Reddit r/MachineLearning · 2026-05-30

The author, an ML student, questions the robotics community about data interoperability issues and proposes an experiment to normalize and enrich public robotics datasets for better reuse.

0 favorites 0 likes
#datasets

Open Multimodal Datasets and Open-Source Software for Data-Driven Modeling of Multiphase Transport and Thermal Systems

arXiv cs.LG · 2026-05-25 Cached

This paper presents open multimodal datasets and open-source software packages for reproducible AI-enabled thermal-fluid research, introducing a spatial-temporal dimensionality framework and tools like SeqReg for sequence regression.

0 favorites 0 likes
#datasets

@MaziyarPanahi: Open-source clinical AI is a niche that most days feels like shouting into a PubMed-shaped void. And yet. 5,000 of you …

X AI KOLs Following · 2026-05-22 Cached

A developer thanks 5,000 followers for supporting OpenMed, an open-source clinical AI project, and hints at upcoming model and dataset releases.

0 favorites 0 likes
#datasets

HuggingFace benchmark datasets now let you filter by model size

Reddit r/LocalLLaMA · 2026-05-20

HuggingFace benchmark datasets now allow filtering by model size, enabling comparisons like 'best model under 32B on swebenchverified'.

0 favorites 0 likes
#datasets

What if i really wanna train an AI from scratch?

Reddit r/artificial · 2026-05-19

A personal reflection on the challenges and allure of training an AI model from scratch, highlighting the difficulties with data, hardware, and scaling, while noting that surprisingly good small models can be trained on modest hardware.

0 favorites 0 likes
#datasets

Why are realistic datasets for agent workflows still so hard to find?

Reddit r/AI_Agents · 2026-05-15

A discussion on the scarcity of realistic datasets for AI agent workflows, noting that existing benchmarks fail to capture messy production scenarios like tool failures, ambiguous requests, and long conversational drift, and seeking recommendations for better datasets.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback