TableVerse: A Large-scale Tabletop Dataset with Real-world Grounded Layouts for Generalizable Manipulation
Summary
TableVerse introduces a fully automated Real2Sim pipeline that converts unstructured, in-the-wild images into high-fidelity, simulation-ready tabletop environments with accurate metrics and physical stability, along with a large-scale dataset (TableVerse-100K) for generalizable robotic manipulation.
View Cached Full Text
Cached at: 07/24/26, 05:06 AM
Paper page - TableVerse: A Large-scale Tabletop Dataset with Real-world Grounded Layouts for Generalizable Manipulation
Source: https://huggingface.co/papers/2607.21017
Abstract
Thedevelopmentofgeneralizableroboticmanipulationpoliciesisinherentlyboundedbytheavailabilityoflarge-scale,high-fidelityscenedata.Whilerecentautomatedsynthesismethodsattempttobridgethisgapviatext-to-layouthallucinationorsimplifiedproceduralgeneration,theyfrequentlysufferfromphysicalimplausibilityandfailtocapturethecomplex,denseclutterofactualhumanenvironments.Inthispaper,weintroduceTableVerse,afullyautomatedReal2Simpipelinethatshiftstheparadigmfromimaginativelayoutgenerationtodeterministicreconstructionfromunstructured,in-the-wildimagedata.Ourframeworkseamlesslyprocessesunscriptedinternetmediaintohigh-fidelity,simulation-readytabletopenvironmentswithaccuratemetricscales,authentictopologies,andverifiedmechanicalstability.Furthermore,anautomatedtask-conditionedtrajectorygenerationframeworkisintegratedtosynthesizehigh-quality,collision-freepick-and-placedemonstrations.Leveragingthiscompletepipeline,weconstructtheTableVerse-100KDataset,alarge-scalecorpuscomprising100,000unique,physicallyconsistentenvironmentspairedwithinteractivemanipulationtrajectories.Bycapturingdiverseassetcompositions,realisticspatialdistributions,andhigh-qualitydemonstrations,TableVerse-100Kestablishesahighlyscalableandhigh-fidelitydatafoundation,providingsignificantvaluetofacilitatefutureresearchingeneralizableroboticmanipulationtasks.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2607\.21017
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.21017 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.21017 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.21017 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models
Deform360 is a large-scale visuotactile dataset with 198 objects and over 215 hours of observations for studying deformable object dynamics, enabling comparison between 2D video and 3D particle world models for robotic manipulation.
TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity
Introduces TableVista, a comprehensive benchmark for evaluating foundation models on multimodal table reasoning under visual and structural complexity, comprising 3,000 problems expanded into 30,000 multimodal samples. Evaluation of 29 models reveals performance degradation on complex layouts and vision-only settings.
TT4D: A Pipeline and Dataset for Table Tennis 4D Reconstruction From Monocular Videos
This paper introduces TT4D, a novel pipeline and large-scale dataset for reconstructing table tennis gameplay in 4D from monocular videos. It features a unique lift-first approach that estimates 3D ball trajectories and spin before time segmentation, enabling robust reconstruction even with occlusions.
Image2Sim: Scaling Embodied Navigation via Generative Neural Simulator
Image2Sim is a neural simulation framework that creates high-fidelity interactive environments from RGB-D images, enabling scalable training for embodied navigation agents. It generates nearly 20K scenes and over 10 million training samples, showing strong benchmark improvements and effective real-world zero-shot transfer.
VCG-Bench: Towards A Unified Visual-Centric Benchmark for Structured Generation and Editing
VCG-Bench is a unified benchmark for evaluating vision-language models on structured diagram generation and editing tasks, introducing a 'Diagram-as-Code' paradigm using symbolic mxGraph XML and a taxonomized dataset of 1,449 diagrams across 6 domains.