Tag
RelightFormer introduces a feed-forward generative Transformer for direct single- and multi-view image relighting, using cross-attention for illumination injection and permutation-invariant encodings for unordered views, trained on a massive synthetic dataset to achieve state-of-the-art visual quality.
This paper introduces a synthetic Bengali speech dataset of 10,000 audio-text pairs for telecom customer care scenarios, generated using OmniVoice voice-cloning, and evaluates it with an ASR model, achieving low word error rates.
QuixiAI releases SYN-1B, a synthetic pretraining dataset designed to teach models rule tracking, belief updating, and evidence preservation over long contexts, available on Hugging Face.
AnyBokeh is a physics-guided framework for any-to-any bokeh editing that estimates source blur states and transfers optical characteristics between different focus and aperture settings without requiring all-in-focus reconstruction.
GORGO introduces a proxy architecture for LLM inference that jointly optimizes network latency, prefill cost, and queueing delay using evolutionary strategy tuning on a new synthetic dataset, improving p95 TTFT by 6.9-15.5% and end-to-end latency by 14.3-30.9%.
SupraLabs released SupraWeather-Nano-Preview, a small FT-Transformer model for classifying weather phenomena from tabular meteorological features, trained on synthetic data as an architecture experiment.
This paper introduces a large-scale synthetic dataset (WATER-S) and a specialized model (WATERec) to advance WordArt-oriented scene text recognition, achieving state-of-the-art accuracy on irregular artistic text benchmarks.
This paper presents COVA-X, an expanded synthetic multi-turn conversation dataset for smishing detection, and shows that Longformer now outperforms XGBoost, confirming that transformer models benefit from larger training corpora.
This paper introduces MagBridge-Battery, a synthetic dataset of 6,760 magnetic-field signatures for Li-ion battery state-of-health diagnostics, combining real magnetic morphology with real degradation labels to bridge the gap in public magnetic-sensing battery data.
This paper introduces a framework for synthesizing long-term medical dialogue datasets using LLMs, and creates MediLongChat with three benchmark tasks to evaluate healthcare agents' memory and reasoning capabilities. Experiments show that even state-of-the-art LLMs struggle with these tasks.