Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene
Summary
Mira-Scene introduces a compositional 3D scene reconstruction framework using pixel-aligned canonical coordinate maps for accurate object layouts, achieving significant improvements in layout accuracy over existing methods.
View Cached Full Text
Cached at: 09/22/26, 07:25 AM
Paper page - Mira-Scene: Pixel-Aligned Layouts for Generative 3D Scene
Source: https://huggingface.co/papers/2609.23796
Abstract
Single-image3Dobjectgenerationcannowproducehigh-fidelityassets,yetaccuratelyplacingthemintoacoherentscenelayoutremainsanopenchallenge.Acentraldifficultyliesinhowobjectlayoutisrepresented.Holisticmethodsabsorbplacementintoascene-levelgenerationprocess,sacrificingobject-leveldetail.Compositionalmethodspreserveobjectfidelitybydecouplinggeometryfromlayout,buttypicallyparameterizelayoutassparse,unboundedposevariablesthataredifficulttolearnandgeneralizepoorlyunderscarcescene-levelsupervision.WepresentMira-Scene,acompositional3Dscenereconstructionframeworkthatreplacessparseposeregressionwithdense,boundedcorrespondencerecovery.AtitscoreistheCanonicalCoordinateMap(CCM),apixel-alignedfieldthatmapseachvisibleobjectpixeltoasurfacecoordinateintheobject’sboundedcanonicalspace.Whenpairedwithascene-spacePointCloudMap(PCM)frommonoculargeometryestimation,CCMinducesdensecanonical-to-scenecorrespondencesfromwhichobjecttransformationsarerecoveredthroughrobustgeometricalignment.BecauseCCMoperatesinboundedcanonicalspace,itprovidesastablepredictiontargetthatcanbetrainedfromscalableobject-level3Ddatawithoutrequiringscene-levellayoutannotations.Mira-ScenefurtherintroducesamultimodaldiffusiontransformerthatjointlygeneratesobjectgeometryandCCMs,usingmodality-specificexpertstreamswithsharedattentionandpositionalencodingtopromotegeometry-layoutconsistency.Experimentsonindoor,outdoor,synthetic,andin-the-wildscenesshowthatMira-Scenesubstantiallyoutperformsstrongbaselinesinlayoutaccuracy,achievingrelativegainsof39.8%in3D-IoUand16.5%in2D-IoUoverSAM3D,usinglimitedopen-sourcetrainingdata.
View arXiv pageView PDFProject pageGitHub8Add to collection
Get this paper in your agent:
hf papers read 2609\.23796
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.23796 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.23796 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.23796 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution
SceneMosaic combines learned image priors with vision-language agents to efficiently generate diverse, physically valid indoor scenes, achieving a 24x speedup and improved physical validity over existing methods.
Pixal3D: Pixel-Aligned 3D Generation from Images
Pixal3D introduces a pixel-aligned 3D generation approach that improves fidelity by establishing direct pixel-to-3D correspondences through back-projection conditioning, addressing issues in canonical space generation.
PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space
PixWorld presents a unified pixel-space diffusion approach for 3D scene reconstruction and generation, overcoming limitations of latent-space methods by using direct image-level supervision and geometry-aware feature alignment. The method outperforms prior generation methods and matches state-of-the-art reconstruction methods.
Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling
Lucida proposes a method for composable indoor scene reconstruction that distributes requirements across parsing, generation, and placement using a VLM policy to create high-fidelity editable replicas from cluttered captures.
GenRecon: Bridging Generative Priors for Multi-View 3D Scene Reconstruction
GenRecon introduces a method for 3D scene reconstruction that integrates generative 3D priors with multi-view image conditioning, achieving high-fidelity, editable mesh reconstructions of indoor environments and outperforming existing methods by 16%.