WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence
Summary
WildCity introduces a large-scale multimodal dataset for city-scale urban navigation and spatial representation, collected by autonomous fleets. It provides 18 long trajectories and establishes baselines for reconstruction and closed-loop simulation to advance AI systems that can perceive and reason about city-scale environments.
View Cached Full Text
Cached at: 07/09/26, 07:51 AM
Paper page - WildCity: A Real-World City-Scale Testbed for Rendering, Simulation, and Spatial Intelligence
Source: https://huggingface.co/papers/2607.06838 Authors:
,
,
,
,
,
,
,
,
,
,
Abstract
WildCity presents a large-scale multimodal dataset for urban navigation and spatial representation, enabling research into AI systems that can perceive and reason about city-scale environments similar to human cognitive capabilities.
Humans can navigate an unfamiliar city and gradually form a coherent spatial mental map spanning tens of square kilometers. Can AI buildspatial representations at a comparable scale? Although recent foundation models have advancedscene reconstructionandembodied intelligence, scaling to entire cities remains an open challenge, primarily due to the lack ofcity-scale data. To bridge the gap, we introduce WildCity, a real-worldmultimodal datasetcollected byautonomous fleetstraversing complexurban environments. Our dataset includes 18 trajectories, each averaging 83.7 kilometers in length, and preserves the core challenges of in-the-wildperception, e.g., dynamic objects, lighting variations, and imperfect camera poses. We further establish an urban-tailored reconstruction baseline and convert the reconstructed environments into aclosed-loop simulator. Beyond the dataset and baseline, we systematically analyze the key challenges on the path to simulation-readyurban digital twins: scalability, extrapolation, and uncertainty. Ultimately, WildCity aims to catalyze progress not only in city-scale rendering, but more broadly in the pursuit of AI that can perceive, remember, and reason across space at a scale comparable to human cognition. Project page: https://han-xiangyu.github.io/Wild-City/
View arXiv pageView PDFProject pageGitHub2Add to collection
Get this paper in your agent:
hf papers read 2607\.06838
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.06838 in a model README.md to link it from this page.
Datasets citing this paper1
#### Neptune615/Wild-City Preview• Updatedabout 2 hours ago • 24 • 1
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.06838 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
CityRAG: Stepping Into a City via Spatially-Grounded Video Generation
CityRAG introduces a video generative model that produces long, physically grounded, 3D-consistent videos of real-world cities using geo-registered data, enabling realistic navigation and simulation for robotics and autonomous driving.
CityBehavEx: A Scalable and Empirically Validated LLM-Assisted Urban Simulation Platform
CityBehavEx is a scalable LLM-assisted urban simulation platform that combines established human mobility models with fine-tuned cross-encoders to generate realistic, empirically validated mobility patterns for city-sized populations, demonstrating 100,000 agents over 75 days in under one hour on a single consumer GPU.
SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks
SpatialWorld is a unified benchmark for evaluating interactive spatial reasoning in multimodal agents across diverse real-world tasks, revealing that even the strongest models achieve low task success rates.
Urban-ImageNet: A Large-Scale Multi-Modal Dataset and Evaluation Framework for Urban Space Perception
Urban-ImageNet is a large-scale multi-modal dataset and evaluation benchmark for urban space perception from social media imagery, supporting scene classification, cross-modal retrieval, and instance segmentation tasks across 61 urban sites in 24 Chinese cities.
@thesupermanmx: Japanese researchers created a system that simulates an entire city by generating up to 1 million virtual residents who…
Japanese researchers developed CitySim, an urban simulator powered by LLMs that populates a digital twin of Tokyo with up to 1 million autonomous AI agents, accurately predicting real-world patterns like commuting and shopping behavior.