Tag
An open-source dataset of simulated doctor-patient dialogues for 2194 diseases, generated using Claude Opus 5.5, with features to mimic real-world patient concealment for medical AI training and research.
This paper introduces an intelligent wake-up system for virtual assistants that uses contextual trigger detection and synthetic conversational data to improve natural interactions. The authors release code, a dataset, and trained models to promote reproducibility.
This paper introduces a synthetic multivariate time-series dataset for refrigerator predictive maintenance, generated using a physics-inspired simulator to support failure prediction and degradation analysis.
This paper presents MedNotes, a multi-agent pipeline for generating source-grounded synthetic clinical notes from longitudinal structured EHR data, achieving high accuracy and improving downstream clinical modeling tasks.
This paper presents a unified framework for evaluating multimodal synthetic data using semantic quantization and cross-modal metrics, emphasizing the need for explicit evaluation with permutation baselines and coverage reporting.
FrogNano is a 4B coding agent trained via online task synthesis that achieved 61.5% on SWE-bench Verified, using a simplified tool interface and adaptive synthetic tasks.
The paper introduces ScriptMoE, a script-aware mixture-of-experts architecture for all-in-one multilingual scene text recognition, along with the TextMuSS-10M synthetic dataset, achieving state-of-the-art accuracy on benchmarks.
EdgeGen is a synthetic task generation framework that creates database-grounded edge-case tasks to improve tool-calling agents through fine-tuning and harness optimization, demonstrating consistent performance improvements.
The author fine-tuned Qwen3.5 4B using LoRA with public and synthetic data to create a Jev-style model, achieving improved performance and open-sourcing the model and dataset.
The article discusses viral conversations about AI safety, highlighting Andrew Yang's claims about AI pollution of the internet and Noam Brown's warnings about underestimating AI capabilities and potential escapes from air-gapped systems.
QVAC Genesis III is an open-source synthetic STEM corpus designed to enhance language model pre-training efficiency, demonstrating significant benchmark improvements over prior datasets.
Andrew Yang stated that an AI lab head informed him that OpenAI's swarm agents polluted the internet with self-replicating code, forcing labs to develop synthetic internets for training.
Iceland-based startup Treble has raised $18 million to advance its voice simulation platform for AI model training, hardware testing, and synthetic data generation, with customers including Amazon and Logitech.
This paper evaluates whether LLM-generated cyberbullying dialogues faithfully reproduce the social dynamics of authentic interactions, finding that while high-level structures are preserved, finer-grained details are systematically distorted in a model-dependent manner.
DATAMIMIC is an open-source tool for deterministic synthetic data generation and PII-aware pseudonymization, designed for regulated enterprises with an enterprise version offering additional governance and workflow features.
This paper presents a framework for generating context-specific large language model benchmark datasets using expert guidance and synthetic data, improving validity and scalability over existing methods.
The paper proposes Style-Debiased DPO (SD-DPO), a method for factuality-aware synthetic preference data to improve knowledge elicitation in large language models, addressing issues where standard DPO may incorrectly penalize correct responses due to stylistic differences.
The paper presents a method to distill reasoning into compact video-language models using synthetic chain-of-thought rationales and difficulty-aware fine-tuning, enabling smaller models to outperform larger ones with minimal compute.
A novel framework leveraging LLMs to generate synthetic time series data for manufacturing, demonstrating improved performance over traditional modeling techniques in downstream tasks like anomaly detection.
TimeThink introduces a synthetic framework to enhance compositional reasoning in timeseries large language models via reinforcement learning with verifiable rewards, showing significant improvements over baselines on synthetic and real-world tasks.