dataset

Tag

Cards List
#dataset

I open sourced 12M agentic decision examples for training Jev style models

Reddit r/artificial · yesterday

The author has open-sourced Jev Decisions v1, a dataset of 12 million examples for training AI models on agentic decisions like tool selection and routing, to address data gaps in agent decision-making.

0 favorites 0 likes
#dataset

SambaGraph: Action-Reaction Spatio-Temporal Graphs for Soccer Tactical Response Modeling

arXiv cs.LG · yesterday Cached

SambaGraph introduces a spatio-temporal graph dataset and benchmark for modeling soccer tactical responses, curated from 2022 FIFA World Cup data to classify actions and retrieve defensive examples.

0 favorites 0 likes
#dataset

Extending FunctionGemma for Practical On-Device Mobile Function Calling

arXiv cs.LG · yesterday Cached

This research paper extends FunctionGemma 270M for practical on-device Android workflows by introducing a synthetic dataset and fine-tuning the model, achieving improved accuracy for function calling while balancing performance and coverage.

0 favorites 0 likes
#dataset

ARAFA: An LLM-Generated Arabic Fact-Checking Dataset

arXiv cs.CL · yesterday Cached

This paper introduces Arafa, a large-scale Arabic fact-checking dataset generated using LLMs, aimed at addressing the scarcity of resources for automatic fact-checking in Arabic.

0 favorites 0 likes
#dataset

CTSpinoPelvic1K: spine, pelvis, ribs and femora in one coordinate frame, annotated for lumbosacral transitional anatomy

arXiv cs.AI · yesterday Cached

The CTSpinoPelvic1K dataset introduces annotated CT images of spine, pelvis, ribs, and femora in one coordinate frame, focusing on lumbosacral transitional anatomy to aid in vertebra classification and clinical assessment.

0 favorites 0 likes
#dataset

Mining Legal Arguments in U.S. Corporate Case Law

arXiv cs.CL · yesterday Cached

This paper introduces an expert-annotated dataset of 42 U.S. federal tax opinions for legal argument mining, featuring functional labels and tree-structured arguments, and demonstrates classification and retrieval experiments.

0 favorites 0 likes
#dataset

IntLawNER: A Named Entity Recognition Dataset and Benchmark in International Law

arXiv cs.AI · yesterday Cached

IntLawNER is a new named entity recognition dataset and benchmark for international law, covering gold-annotated sentences from legal texts and evaluating model performance with few-shot improvements.

0 favorites 0 likes
#dataset

QontoFAQ: A better Information Retrieval Benchmark [R]

Reddit r/MachineLearning · 2d ago

QontoFAQ introduces a new benchmark and metric for information retrieval, focusing on improving relevance measurement for answering product questions, with associated code and dataset released.

0 favorites 0 likes
#dataset

A Synthetic Multivariate Refrigerator Time-Series Dataset for Predictive Maintenance

arXiv cs.LG · 2d ago Cached

This paper introduces a synthetic multivariate time-series dataset for refrigerator predictive maintenance, generated using a physics-inspired simulator to support failure prediction and degradation analysis.

0 favorites 0 likes
#dataset

@billxbf: Today we give Superintelligence back to its owners. Introducing Skill2Env , the most aligned and diverse dataset to fue…

X AI KOLs Timeline · 3d ago Cached

Introducing Skill2Env, an aligned and diverse dataset designed to fuel modern Agentic Reinforcement Learning research.

0 favorites 0 likes
#dataset

Chinese Competitive Debating Dataset and Benchmark

arXiv cs.CL · 3d ago Cached

This paper introduces a dataset and benchmark for evaluating large language models' understanding of competitive Chinese-language debate, featuring 148 matches with professional adjudication and tasks at match, stage, and speaker levels.

0 favorites 0 likes
#dataset

GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay

Hugging Face Daily Papers · 3d ago Cached

The paper introduces GameHorizon Suite, a unified data and evaluation framework for assessing AI models' capabilities in gameplay across multiple temporal horizons, featuring an annotation pipeline, large-scale dataset, and reproducible benchmark.

0 favorites 0 likes
#dataset

DeepSWE-mini a 16 instance subset of DeepSWE that replicates the rankings of the leaderboard

Reddit r/LocalLLaMA · 6d ago

Released a new dataset called DeepSWE-mini, a 16-instance subset of DeepSWE designed for efficient benchmarking of local AI models by replicating leaderboard rankings.

0 favorites 0 likes
#dataset

HuRo: Robotizing Human Videos for Scalable VLA Pretraining

Hugging Face Daily Papers · 6d ago Cached

This paper presents HuRo, a pipeline for robotizing human videos to create scalable VLA pretraining data, showing significant improvements in task completion and robustness on real-world manipulation tasks.

0 favorites 0 likes
#dataset

OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation

Hugging Face Daily Papers · 6d ago Cached

The paper introduces OmniVBench, a comprehensive benchmark for omni reference-to-video generation, and the Omni-R2VDataset, a large-scale training dataset, to evaluate and improve R2V models.

0 favorites 0 likes
#dataset

Emotion Experience, Expression, and Perception: Emotion Analysis on Multimodal Social Media Posts

arXiv cs.CL · 2026-09-17 Cached

This paper introduces the Mult2EMo dataset for studying emotion expression and perception in multimodal social media posts, finding that reconstruction is challenging, particularly when posts rely heavily on images.

0 favorites 0 likes
#dataset

Paint-Anything: Unified Any-Color Control for Image Generation and Editing

Hugging Face Daily Papers · 2026-09-17 Cached

Paint-Anything introduces a unified hex-prompt interface for color control in image generation and editing, trained on a custom dataset and evaluated on a new benchmark showing significant improvements in color fidelity.

0 favorites 0 likes
#dataset

ParsHate: A Benchmark Dataset for Hate and Target Detection in Persian

arXiv cs.CL · 2026-09-16 Cached

Introduces ParsHate, a manually annotated dataset of 10,000 Persian tweets for hate speech and target detection, showing moderate performance of current models and emphasizing the need for more advanced methods.

0 favorites 0 likes
#dataset

How Humans and LLMs Read Gender into Gender-Neutral Physical Descriptions

arXiv cs.CL · 2026-09-16 Cached

This study introduces the GAPA dataset to demonstrate that physical descriptions often carry gender associations, and LLMs show systematic misalignment with human ratings, challenging assumptions about gender-neutral communication.

0 favorites 0 likes
#dataset

Not all Negation Cues are Equal: Affixal Negations Yield Better Negation Understanding

arXiv cs.CL · 2026-09-15 Cached

The paper introduces NegCue, a large-scale negation dataset, and shows that pre-training with affixal negation yields better improvements for negation understanding in language models compared to single-word negation.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback