Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis
Summary
Introduces Poplar, a scalable Specify-Render-Inspect pipeline for synthesizing human-centric image datasets, and releases Poplar-9K, a curated dataset of 9,401 image-text pairs with auditable inspection records.
View Cached Full Text
Cached at: 08/04/26, 05:37 AM
Paper page - Poplar: A Scalable Pipeline for Human-Centric Image Dataset Synthesis
Source: https://huggingface.co/papers/2608.00440
Abstract
Recentimagegeneratorscansynthesizeconvincinghuman-centricimages,yetproducingausefulcollectionremainsdifferentfromproducingasinglesuccessfulimage.Ahuman-centricdatasetmustcovervariedpeopleandcontexts,avoidimplausibleattributecombinations,preserveaneverydayphotographiccharacter,andexposequality-controldecisionsatscale.WepresentPoplar,areproducibleSpecify--Render--Inspectpipelineforhuman-centricimagedatasetsynthesis.Specifysamplesstructuredattributesundercommonsenseconstraintsandverbalizesthemasphotography-orientedprompts.Renderusesarealism-adaptedimagegeneratoracrosscomposition-awareaspectratiosandretriesobvioustechnicalfailures.Inspectappliesasinglestructuredvision--languagereviewtoeachcandidate,preservingtheoriginalpromptwhilerejectingintrinsicimagedefectsormaterialpromptmismatches.UsingPoplar,weconstructPoplar-9K:9,401curatedhuman-centricimage--textpairsretainedfrom11,765reviewedcandidates(79.9\%acceptance).Wereleasethedatasettogetherwiththepipeline,configurations,immutablegenerationprompts,andauditableinspectionrecordsasacompactresourceforbuildingcustomizablehuman-centriccollections.
View arXiv pageView PDFProject pageGitHub5Add to collection
Get this paper in your agent:
hf papers read 2608\.00440
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2608.00440 in a model README.md to link it from this page.
Datasets citing this paper1
#### choucsan/Poplar-9K Viewer• Updatedabout 2 hours ago • 9.4k • 204 • 3
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2608.00440 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
@drfeifei: I’m very excited by this new benchmark dataset for visual generation that is suitable for the modern era of large scale…
Introducing GPIC (Giant Permissive Image Corpus), a large-scale dataset of 100M VLM-captioned image-text pairs for training and 1M pairs for benchmarking, fully permissive for research and commercial use.
PixVerve: Advancing Native UHR Image Generation to 100MP with a Large-Scale High-Quality Dataset
This paper introduces PixVerve-95K, a large-scale open-source dataset of 95K ultra-high-resolution (100MP) images with annotations, and PixVerve-Bench, a benchmark for evaluating native 100MP text-to-image generation, extending existing T2I models to unprecedented resolutions.
Urban-ImageNet: A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception
Urban-ImageNet is a large-scale multi-modal dataset and evaluation benchmark for urban space perception from social media imagery, supporting scene classification, cross-modal retrieval, and instance segmentation tasks across 61 urban sites in 24 Chinese cities.
SGOCR: A Spatially-Grounded OCR-focused Pipeline & V1 Dataset [P]
SGOCR is an open-source dataset pipeline for generating spatially-grounded, OCR-focused visual question answering (VQA) tuples with rich metadata to support diverse VLM training. The pipeline uses a multi-stage approach combining models like Nvidia's nemotron-ocr-v2, Gemma4, Qwen3-VL, and Gemini-2.5-Flash, along with an agentic optimization loop.
DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis
Introduces DF3DV-1K, a large-scale real-world dataset with 1,048 scenes and 89,924 images for distractor-free novel view synthesis, along with a benchmark of nine methods and an application improving radiance field methods via fine-tuning a diffusion-based 2D enhancer.