Tag
FINESSE is an agent-based simulation framework and benchmark dataset for generating synthetic multimodal financial event sequences, addressing data scarcity and supporting tasks like fraud detection and balance forecasting.
ToolGrad introduces an efficient method for generating tool-use datasets using textual gradients, enabling better LLM training with lower cost and improved performance on out-of-distribution tasks.
The paper introduces a generator that creates consistent fictional enterprise data without real datasets, using reference-free evaluation methods to ensure realism, and includes a hosted service for building relational databases from business questions.
This paper investigates model collapse in multi-model ecosystems where AI-generated text is recycled into training data. It finds that market concentration has minimal effect on the speed or destination of collapse, which is primarily influenced by the sources of the training pool.
An AI developer shares that incorporating reinforcement learning and synthetic data improved their voice model's accuracy and inference speed, with plans to open-source the project.
The developer announces plans to accelerate the LFM2.5-Audio-1.5B model using reinforcement learning methods and synthetic data, with a focus on English, and shares optimizations for the Mimi codec model to improve speed and quality for open-source release.
This paper examines how different AI model checkpoints exhibit varying susceptibility to collapse when trained recursively on shared model-generated data, revealing that fragility is an intrinsic property that can be inferred from short iterations.
Microsoft Research presents FrogNano, a 4B coding agent trained exclusively via reinforcement learning with online task synthesis, achieving competitive performance without distillation from larger models.
MedDeID is an on-premises framework for de-identifying clinical text using real or synthetic training data, achieving high accuracy in detecting personally identifiable information with minimal over-redaction.
This paper introduces SpatialBlock-15k, a synthetic dataset for block-stacking problems, to enhance 3D spatial reasoning in large vision-language models, demonstrating improved performance and generalization to real-world tasks.
This paper demonstrates that causal fairness mechanisms, specifically edge cuts on causal graphs, are portable across various synthetic data generator families including GANs and diffusion models, with minimal impact on data fidelity and utility.
Introduces Xiaomi-TabLDM, a tabular foundation model that leverages synthetic data and in-context learning for superior prediction accuracy without task-specific fine-tuning, achieving top rankings on multiple benchmarks.
This paper empirically studies contrastive pretraining with synthetic semantic supervision for code embeddings in small transformers, showing significant gains over baselines and competitiveness with larger models.
This paper investigates the use of synthetic data for improving Thai OCR, leading to the development of Wayu-Paxa-OCR-Zero, which achieves competitive performance without real OCR labels.
This paper presents a method to build a compact fixed-voice Thai TTS system using synthetic speech from a larger model, evaluating its performance and introducing an 82M-parameter model for on-device deployment.
This paper introduces a tri-agent framework for evaluating and aligning the question clarification capabilities of large language models, using three LLM-based agents to simulate and assess clarification dialogues.
This paper evaluates relational foundation models on high-cardinality data, revealing that context window limitations lead to performance drops, which can be mitigated by simple pre-aggregation steps.
The paper proposes a synthetic simulation framework and benchmark for evaluating and updating knowledge in large language models, demonstrating a 14.23% improvement over existing methods.
Atlas is a tool that reconstructs spaces from a few photos and generates photorealistic RGB and depth data for robotic simulation, enabling robots to be trained and tested in more diverse environments.
The paper introduces ATOM, a framework that distinguishes between operand and operator perturbations in synthetic data, showing that operator perturbations are critical for LLM performance and proposing improved filtering strategies.