synthetic-data

Tag

Cards List
#synthetic-data

FINESSE: An Agent-Based Simulator and Benchmark Dataset for Multimodal Financial Event Sequences

arXiv cs.LG ↗ · 2026-09-14 Cached

FINESSE is an agent-based simulation framework and benchmark dataset for generating synthetic multimodal financial event sequences, addressing data scarcity and supporting tasks like fraud detection and balance forecasting.

0 favorites 0 likes
#synthetic-data

ToolGrad: Efficient tool-use dataset generation with textual “gradients” (3 minute read)

TLDR AI ↗ · 2026-09-14 Cached

ToolGrad introduces an efficient method for generating tool-use datasets using textual gradients, enabling better LLM training with lower cost and improved performance on out-of-distribution tasks.

0 favorites 0 likes
#synthetic-data

Generating a Consistent Enterprise: Synthesis and Reference-Free Evaluation of Multi-System Business Data

arXiv cs.AI ↗ · 2026-09-12 Cached

The paper introduces a generator that creates consistent fictional enterprise data without real datasets, using reference-free evaluation methods to ensure realism, and includes a hosted service for building relational databases from business questions.

0 favorites 0 likes
#synthetic-data

The Oligarch Barely Steers Model Collapse in Multi-Model Ecosystems

arXiv cs.AI ↗ · 2026-09-12 Cached

This paper investigates model collapse in multi-model ecosystems where AI-generated text is recycled into training data. It finds that market concentration has minimal effect on the speed or destination of collapse, which is primarily influenced by the sources of the training pool.

0 favorites 0 likes
#synthetic-data

@kadirnardev: I had been very focused on continuous pretraining and collecting real-world voice data, but after talking with Maxime L…

X AI KOLs Following ↗ · 2026-09-11 Cached

An AI developer shares that incorporating reinforcement learning and synthetic data improved their voice model's accuracy and inference speed, with plans to open-source the project.

0 favorites 0 likes
#synthetic-data

@kadirnardev: From now on, I will work on accelerating the LFM2.5-Audio-1.5B model. I will also work on RL methods and synthetic data…

X AI KOLs Following ↗ · 2026-09-11 Cached

The developer announces plans to accelerate the LFM2.5-Audio-1.5B model using reinforcement learning methods and synthetic data, with a focus on English, and shares optimizations for the Mimi codec model to improve speed and quality for open-source release.

0 favorites 0 likes
#synthetic-data

A Fragility Spectrum for Recursive Language-Model Training

arXiv cs.CL ↗ · 2026-09-11 Cached

This paper examines how different AI model checkpoints exhibit varying susceptibility to collapse when trained recursively on shared model-generated data, revealing that fragility is an intrinsic property that can be inferred from short iterations.

0 favorites 0 likes
#synthetic-data

Microsoft trained a 4B coding agent almost entirely with Reinforcement Learning, without a bigger teacher

Reddit r/ArtificialInteligence ↗ · 2026-09-10 Cached

Microsoft Research presents FrogNano, a 4B coding agent trained exclusively via reinforcement learning with online task synthesis, achieving competitive performance without distillation from larger models.

0 favorites 0 likes
#synthetic-data

MedDeID enables locally governed clinical-text de-identification from real or synthetic training data

arXiv cs.CL ↗ · 2026-09-10 Cached

MedDeID is an on-premises framework for de-identifying clinical text using real or synthetic training data, achieving high accuracy in detecting personally identifiable information with minimal over-redaction.

0 favorites 0 likes
#synthetic-data

SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem

Hugging Face Daily Papers ↗ · 2026-09-07 Cached

This paper introduces SpatialBlock-15k, a synthetic dataset for block-stacking problems, to enhance 3D spatial reasoning in large vision-language models, demonstrating improved performance and generalization to real-world tasks.

0 favorites 0 likes
#synthetic-data

Portable Causal Fairness Across Synthetic Data Generator Families

arXiv cs.LG ↗ · 2026-09-04 Cached

This paper demonstrates that causal fairness mechanisms, specifically edge cuts on causal graphs, are portable across various synthetic data generator families including GANs and diffusion models, with minimal impact on data fidelity and utility.

0 favorites 0 likes
#synthetic-data

Xiaomi-TabLDM: A Tabular Foundation Model Technical Report

arXiv cs.AI ↗ · 2026-09-04 Cached

Introduces Xiaomi-TabLDM, a tabular foundation model that leverages synthetic data and in-context learning for superior prediction accuracy without task-specific fine-tuning, achieving top rankings on multiple benchmarks.

0 favorites 0 likes
#synthetic-data

Synthetic Semantic Supervision for Contrastive Code Representation Learning in Small Transformers: An Empirical Study

arXiv cs.AI ↗ · 2026-09-04 Cached

This paper empirically studies contrastive pretraining with synthetic semantic supervision for code embeddings in small transformers, showing significant gains over baselines and competitiveness with larger models.

0 favorites 0 likes
#synthetic-data

How Far Can Synthetic Data Take Thai OCR?

arXiv cs.CL ↗ · 2026-09-04 Cached

This paper investigates the use of synthetic data for improving Thai OCR, leading to the development of Wayu-Paxa-OCR-Zero, which achieves competitive performance without real OCR labels.

0 favorites 0 likes
#synthetic-data

Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

arXiv cs.CL ↗ · 2026-09-04 Cached

This paper presents a method to build a compact fixed-voice Thai TTS system using synthetic speech from a larger model, evaluating its performance and introducing an 82M-parameter model for on-device deployment.

0 favorites 0 likes
#synthetic-data

A Tri-Agent Framework for Evaluating and Aligning Question Clarification Capabilities of Large Language Models

arXiv cs.CL ↗ · 2026-09-03 Cached

This paper introduces a tri-agent framework for evaluating and aligning the question clarification capabilities of large language models, using three LLM-based agents to simulate and assess clarification dialogues.

0 favorites 0 likes
#synthetic-data

Context Window Failures in Relational Foundation Models

arXiv cs.LG ↗ · 2026-09-02 Cached

This paper evaluates relational foundation models on high-cardinality data, revealing that context window limitations lead to performance drops, which can be mitigated by simple pre-aggregation steps.

0 favorites 0 likes
#synthetic-data

Synthetic Worlds for Temporal Evaluation and Knowledge Updating in LLMs

arXiv cs.CL ↗ · 2026-09-02 Cached

The paper proposes a synthetic simulation framework and benchmark for evaluating and updating knowledge in large language models, demonstrating a 14.23% improvement over existing methods.

0 favorites 0 likes
#synthetic-data

@theworldlabs: For robotic simulation, Atlas reconstructs a space from just a few photos and generates the photorealistic RGB and dept…

X AI KOLs Following ↗ · 2026-09-01 Cached

Atlas is a tool that reconstructs spaces from a few photos and generates photorealistic RGB and depth data for robotic simulation, enabling robots to be trained and tested in more diverse environments.

0 favorites 0 likes
#synthetic-data

Quantifying Error Tolerance in Synthetic Data: An Atomic-level Operand vs. Operator Perturbation Study

arXiv cs.CL ↗ · 2026-09-01 Cached

The paper introduces ATOM, a framework that distinguishes between operand and operator perturbations in synthetic data, showing that operator perturbations are critical for LLM performance and proposing improved filtering strategies.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback