DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation
Summary
DataEvolver is a self-evolving multi-agent framework that leverages feedback from rejected samples to iteratively enhance data quality for text-rich image generation, achieving 85.3% OCR-F1 improvement on TextScenesHQ and 35.3% on LongTextBench using PixArt-alpha.
View Cached Full Text
Cached at: 07/01/26, 11:42 AM
Paper page - DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation
Source: https://huggingface.co/papers/2606.31537
Abstract
DataEvolver is a self-evolving multi-agent framework that improves text-rich image generation by leveraging feedback from rejected samples to iteratively enhance data quality.
Text-rich image generationis one of the most challenging settings in image generation, since models must simultaneously produce visually realistic images and render legible, semantically aligned, and layout-consistent text. Existing data pipelines usually follow a static crawl-filter-freeze paradigm. They collect candidate samples, filter them once, and freeze the accepted data for training. However, rejected samples are usually discarded, although they often contain useful failure signals such as OCR errors and semantic mismatches. As a result, later construction rounds may repeat the same failure modes. To address these limitations, we propose DataEvolver, a self-evolvingmulti-agent frameworkfor text-rich imagedata construction. DataEvolver treatsdata constructionas feedback-driven construction policy evolution. A Retriever collects candidate samples, a Verifier assigns quality scores and rejection causes, a Critic summarizes round-level feedback intosemantic feedback, and a Generator completes under-covered regions through targeted synthesis. The updated feedback memory then guides the next construction round. Experiments ontext-rich image generationbenchmarks show that DataEvolver produces more useful training data than fixed-dataset baselines under matcheddata budgets. At the 0.75M scale on PixArt-alpha, DataEvolver improvesOCR-F1over the strongest baseline by 85.3 percent on TextScenesHQ and 35.3 percent on LongTextBench. The improvements are consistent across both evaluated benchmarks and also transfer to Show-o2, indicating that the benefit of DataEvolver is not tied to a singledownstream generator. These results suggest that rejected samples can provide actionable feedback for improving text-rich imagedata construction.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2606\.31537
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2606.31537 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2606.31537 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2606.31537 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation
GenEvolve is a self-evolving image generation framework that uses tool-orchestrated trajectories and visual experience distillation to iteratively improve generative capabilities, achieving state-of-the-art performance.
EvoMap/evolver
Evolver is a GEP-powered self-evolution engine for AI agents that automates prompt optimization and creates auditable, reusable evolution assets. The project is transitioning from fully open source to source-available while maintaining backward compatibility with existing MIT and GPL-3.0 releases.
EvoDS: Self-Evolving Autonomous Data Science Agent with Skill Learning and Context Management
EvoDS is a self-evolving autonomous data science agent that improves via reinforcement learning-driven skill acquisition and adaptive context compression, outperforming open-source agents by 28.9% on benchmarks.
CoEvolve: Training LLM Agents via Agent-Data Mutual Evolution
CoEvolve proposes an agent-data mutual evolution framework for training LLM agents through closed-loop, interaction-driven learning that adapts both the agent and its training data distribution. The method extracts feedback signals from rollout trajectories to guide LLM-based task synthesis, demonstrating significant improvements (15-19% absolute gains) across multiple Qwen models on AppWorld and BFCL benchmarks.
DALL·E: Creating images from text
OpenAI introduces DALL·E, a 12-billion parameter transformer model that generates images from text descriptions by treating text and images as a single token stream. The model demonstrates diverse capabilities including creating anthropomorphized objects, combining disparate concepts, rendering text, and performing image inpainting tasks.