data-synthesis

Tag

Cards List
#data-synthesis

Infinity-Parser2 Technical Report

arXiv cs.AI · 2026-07-10 Cached

The Infinity-Parser2 technical report presents a large multimodal model for end-to-end document parsing, featuring a scalable data synthesis pipeline and multi-task reinforcement learning. It achieves state-of-the-art results on multiple benchmarks while releasing open-source model variants and a 5-million-sample bilingual corpus.

0 favorites 0 likes
#data-synthesis

Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing

Hugging Face Daily Papers · 2026-06-30 Cached

This paper introduces Goku, a million-scale dataset and benchmark for instruction-based video editing, supporting multi-task and structural manipulations. The accompanying model, Goku-Edit, achieves up to +8% improvement on instruction following over open-source models.

0 favorites 0 likes
#data-synthesis

RODS: Reward-Driven Online Data Synthesis for Multi-Turn Tool-Use Agents

arXiv cs.AI · 2026-06-18 Cached

This paper introduces RODS, a reward-driven online data synthesis method that addresses the depletion of informative samples in static datasets for multi-turn tool-use agent training. It achieves comparable performance to larger offline pipelines with significantly fewer trajectories.

0 favorites 0 likes
#data-synthesis

Edu-Theater: A Data-Efficient Agent Framework for Scalable Learner Behavior Simulation through Staging Roll-Call

arXiv cs.LG · 2026-06-16 Cached

Edu-Theater is a data-efficient agent framework that uses LLM-powered generative agents to simulate learner behavior in educational settings. It employs a cohort-aware roll-call paradigm to infer learner states with fewer data and computational resources, achieving higher simulation accuracy.

0 favorites 0 likes
#data-synthesis

Geometry-Aware Tabular Diffusion

arXiv cs.LG · 2026-06-03 Cached

Introduces Geometry-Aware Tabular Diffusion (GATD), which augments tabular diffusion denoisers with explicit pairwise geometric features. Achieves state-of-the-art performance on ten benchmarks while using significantly fewer parameters.

0 favorites 0 likes
#data-synthesis

Domain-Specific Data Synthesis for LLMs via Minimal Sufficient Representation Learning

Hugging Face Daily Papers · 2026-05-29 Cached

DOMINO is a novel framework that learns minimal sufficient domain representations from reference examples to synthesize domain-specific data for LLMs, improving code benchmark performance without requiring explicit domain descriptions.

0 favorites 0 likes
#data-synthesis

Knowledge Distillation for Low-Resource Open-source Text-to-SQL Model

arXiv cs.CL · 2026-05-25 Cached

This paper proposes a knowledge-aware Text-to-SQL framework that uses knowledge distillation to improve performance in low-resource settings by constructing task-specific knowledge bases and generating synthetic training data. Experiments on seven benchmarks show substantial improvements, especially for open-source models.

0 favorites 0 likes
#data-synthesis

Terminal-World: Scaling Terminal-Agent Environments via Agent Skills

arXiv cs.CL · 2026-05-21 Cached

Terminal-World introduces a fully automated pipeline that uses agent skills to synthesize high-quality training data for terminal agents, enabling models to outperform baselines with only 1.2% of the training data. The method co-derives task instructions, environments, and teacher trajectories from skill primitives.

0 favorites 0 likes
#data-synthesis

Are Rationales Necessary and Sufficient? Tuning LLMs for Explainable Misinformation Detection

arXiv cs.CL · 2026-05-20 Cached

This paper proposes a pipeline for fine-tuning LLMs specifically for explainable misinformation detection and introduces LonsRex, a data synthesis method to generate necessary and sufficient rationales, addressing limitations of naive filtering based solely on label correctness.

0 favorites 0 likes
#data-synthesis

Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning

Hugging Face Daily Papers · 2026-05-20 Cached

Uni-Edit proposes using intelligent image editing as a single general task to simultaneously improve unified multimodal models' understanding, generation, and editing capabilities, with an automated data synthesis pipeline creating complex editing instructions.

0 favorites 0 likes
#data-synthesis

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale

Hugging Face Daily Papers · 2026-05-14 Cached

FrontierSmith automatically generates diverse open-ended coding problems from closed-ended tasks, improving LLM coding performance on benchmarks through enhanced agent interactions and training data synthesis.

0 favorites 0 likes
#data-synthesis

Covering Human Action Space for Computer Use: Data Synthesis and Benchmark

Hugging Face Daily Papers · 2026-05-12 Cached

This paper introduces CUActSpot, a multimodal benchmark for evaluating computer-use agents, and a renderer-based data synthesis pipeline. The proposed Phi-Ground-Any-4B model outperforms open-source models under 32B parameters.

0 favorites 0 likes
#data-synthesis

CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution

arXiv cs.CL · 2026-04-20 Cached

CoEvolve proposes an agent-data mutual evolution framework for training LLM agents through closed-loop, interaction-driven learning that adapts both the agent and its training data distribution. The method extracts feedback signals from rollout trajectories to guide LLM-based task synthesis, demonstrating significant improvements (15-19% absolute gains) across multiple Qwen models on AppWorld and BFCL benchmarks.

0 favorites 0 likes
#data-synthesis

Dual-View Training for Instruction-Following Information Retrieval

Hugging Face Daily Papers · 2026-04-20 Cached

A dual-view data synthesis method using polarity reversal boosts instruction-following retrieval performance by 45% on the FollowIR benchmark.

0 favorites 0 likes
#data-synthesis

WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization

Papers with Code Trending · 2025-07-20 Cached

WebShaper is a formalization-driven framework for synthesizing information-seeking datasets using set theory and Knowledge Projections, achieving state-of-the-art performance on GAIA and WebWalkerQA benchmarks among open-source agents.

0 favorites 0 likes
← Back to home

Submit Feedback