synthetic-data

Tag

Cards List
#synthetic-data

@frxiaobei: Friends working on medical AI, take a look at this open-source dataset. Simulated doctor-patient dialogues for 2194 dis…

X AI KOLs Timeline ↗ · 15h ago Cached

An open-source dataset of simulated doctor-patient dialogues for 2194 diseases, generated using Claude Opus 5.5, with features to mimic real-world patient concealment for medical AI training and research.

0 favorites 0 likes
#synthetic-data

Training Intelligent Voice Assistant Wakeup with Controllable Synthetic Conversations

arXiv cs.AI ↗ · 3d ago Cached

This paper introduces an intelligent wake-up system for virtual assistants that uses contextual trigger detection and synthetic conversational data to improve natural interactions. The authors release code, a dataset, and trained models to promote reproducibility.

0 favorites 0 likes
#synthetic-data

A Synthetic Multivariate Refrigerator Time-Series Dataset for Predictive Maintenance

arXiv cs.LG ↗ · 5d ago Cached

This paper introduces a synthetic multivariate time-series dataset for refrigerator predictive maintenance, generated using a physics-inspired simulator to support failure prediction and degradation analysis.

0 favorites 0 likes
#synthetic-data

A Multi-Agent Pipeline for Source-Grounded Synthetic Note Generation from Longitudinal Structured EHR

arXiv cs.CL ↗ · 5d ago Cached

This paper presents MedNotes, a multi-agent pipeline for generating source-grounded synthetic clinical notes from longitudinal structured EHR data, achieving high accuracy and improving downstream clinical modeling tasks.

0 favorites 0 likes
#synthetic-data

Beyond the Stitching Assumption: A Unified Framework for Multimodal Synthetic Data Evaluation via Semantic Quantization

arXiv cs.CL ↗ · 5d ago Cached

This paper presents a unified framework for evaluating multimodal synthetic data using semantic quantization and cross-modal metrics, emphasizing the need for explicit evaluation with permutation baselines and coverage reporting.

0 favorites 0 likes
#synthetic-data

@rohanpaul_ai: A 4B coding agent reached 61.5% on SWE-bench Verified without frontier-model distillation by combining a simpler tool i…

X AI KOLs Timeline ↗ · 6d ago Cached

FrogNano is a 4B coding agent trained via online task synthesis that achieved 61.5% on SWE-bench Verified, using a simplified tool interface and adaptive synthetic tasks.

0 favorites 0 likes
#synthetic-data

All-in-One Multilingual Scene Text Recognition with Script-aware Mixture-of-Experts

Hugging Face Daily Papers ↗ · 6d ago Cached

The paper introduces ScriptMoE, a script-aware mixture-of-experts architecture for all-in-one multilingual scene text recognition, along with the TextMuSS-10M synthetic dataset, achieving state-of-the-art accuracy on benchmarks.

0 favorites 0 likes
#synthetic-data

EDGEGEN: Improving Tool-Calling Agents Beyond Happy Paths with Synthetic Edge Case Generation

Hugging Face Daily Papers ↗ · 6d ago Cached

EdgeGen is a synthetic task generation framework that creates database-grounded edge-case tasks to improve tool-calling agents through fine-tuning and harness optimization, demonstrating consistent performance improvements.

0 favorites 0 likes
#synthetic-data

A Jev-style model fine-tuned on Qwen3.5 4B

Reddit r/LocalLLaMA ↗ · 2026-09-20

The author fine-tuned Qwen3.5 4B using LoRA with public and synthetic data to create a Jev-style model, achieving improved performance and open-sourcing the model and dataset.

0 favorites 0 likes
#synthetic-data

AI safety conversations have gotten unbelievable

TechCrunch AI ↗ · 2026-09-19 Cached

The article discusses viral conversations about AI safety, highlighting Andrew Yang's claims about AI pollution of the internet and Noam Brown's warnings about underestimating AI capabilities and potential escapes from air-gapped systems.

0 favorites 0 likes
#synthetic-data

QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training

arXiv cs.AI ↗ · 2026-09-18 Cached

QVAC Genesis III is an open-source synthetic STEM corpus designed to enhance language model pre-training efficiency, demonstrating significant benchmark improvements over prior datasets.

0 favorites 0 likes
#synthetic-data

Andrew Yang says an AI lab head told him yesterday the OpenAI swarm agents "polluted the internet" with "code to self-replicate and create bot swarms," and that the labs now "have to create synthetic internets to train their bots."

Reddit r/ArtificialInteligence ↗ · 2026-09-17

Andrew Yang stated that an AI lab head informed him that OpenAI's swarm agents polluted the internet with self-replicating code, forcing labs to develop synthetic internets for training.

0 favorites 0 likes
#synthetic-data

Iceland-based Treble raises $18 million for its voice simulation platform

TechCrunch AI ↗ · 2026-09-17 Cached

Iceland-based startup Treble has raised $18 million to advance its voice simulation platform for AI model training, hardware testing, and synthetic data generation, with customers including Amazon and Logitech.

0 favorites 0 likes
#synthetic-data

Do Social Patterns Hold in Synthetic Data? Analyzing Cyberbullying Dynamics in LLM-Generated and Authentic Dialogues

arXiv cs.CL ↗ · 2026-09-17 Cached

This paper evaluates whether LLM-generated cyberbullying dialogues faithfully reproduce the social dynamics of authentic interactions, finding that while high-level structures are preserved, finer-grained details are systematically distorted in a model-dependent manner.

0 favorites 0 likes
#synthetic-data

Datamimic – don't let your coding agent invent its own test world

Hacker News Top ↗ · 2026-09-16 Cached

DATAMIMIC is an open-source tool for deterministic synthetic data generation and PII-aware pseudonymization, designed for regulated enterprises with an enterprise version offering additional governance and workflow features.

0 favorites 0 likes
#synthetic-data

A Framework for Generating Valid Context-Specific Benchmarks through Expert Guidance

arXiv cs.AI ↗ · 2026-09-16 Cached

This paper presents a framework for generating context-specific large language model benchmark datasets using expert guidance and synthetic data, improving validity and scalability over existing methods.

0 favorites 0 likes
#synthetic-data

Style-Debiased DPO: Updating LLM Knowledge with Factuality-Aware Synthetic Preference Data

arXiv cs.CL ↗ · 2026-09-16 Cached

The paper proposes Style-Debiased DPO (SD-DPO), a method for factuality-aware synthetic preference data to improve knowledge elicitation in large language models, addressing issues where standard DPO may incorrectly penalize correct responses due to stylistic differences.

0 favorites 0 likes
#synthetic-data

Efficient Reasoning Distillation: Small Video-Language Models via Synthetic CoT and Difficulty-Aware Fine-Tuning

arXiv cs.LG ↗ · 2026-09-16 Cached

The paper presents a method to distill reasoning into compact video-language models using synthetic chain-of-thought rationales and difficulty-aware fine-tuning, enabling smaller models to outperform larger ones with minimal compute.

0 favorites 0 likes
#synthetic-data

LLMs as Master Forgers: Generating Synthetic Time Series Data for Manufacturing

arXiv cs.LG ↗ · 2026-09-16 Cached

A novel framework leveraging LLMs to generate synthetic time series data for manufacturing, demonstrating improved performance over traditional modeling techniques in downstream tasks like anomaly detection.

0 favorites 0 likes
#synthetic-data

TimeThink: Eliciting Compositional Reasoning in Timeseries Large Language Models

arXiv cs.AI ↗ · 2026-09-15 Cached

TimeThink introduces a synthetic framework to enhance compositional reasoning in timeseries large language models via reinforcement learning with verifiable rewards, showing significant improvements over baselines on synthetic and real-world tasks.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback