Tag
TRACE introduces a novel approach to active 3D reconstruction by optimizing full sensor trajectories for ergodic coverage of scene information, outperforming next-best-view baselines with a 1.5 dB PSNR improvement.
Introduces UniWorld-View, a unified framework for large-baseline novel view synthesis from monocular inputs, integrating occlusion-aware point cloud rendering with video diffusion models for precise camera control and geometric consistency.
This paper introduces a rest-state framework that reconstructs articulated objects from a single closed configuration, using explicit meshes, vision-language outputs, and video diffusion models to generate and validate articulation hypotheses without observed motion.
MAGiSt3R is a multi-agent feed-forward framework that achieves real-time 3D reconstruction from monocular RGB videos at 10 FPS, using a merging model (MAGMA) to combine local point maps and pose graph optimization to reduce drift.
A Chinese open-source model reconstructs 3D scenes in real-time from a single camera video, no LiDAR required, achieves 20fps on a single GPU with performance superior to traditional optimization methods. Fully open-source.
This paper proposes a 4D digital viewing system for archiving the creation process of Ikebana using synchronized cameras, eye tracking, and 3D Gaussian Splatting, enabling interactive viewing from arbitrary angles.
The user used the Kimi K3 model to reconstruct a digital version of the Yingxian Wooden Pagoda, containing 10,611 components and four lighting scenes (morning, noon, dusk, night), demonstrating the model's 3D reconstruction capability.
China open-sourced a real-time 3D scene reconstruction model that works from a single regular video without LiDAR, achieving 20 FPS on a single GPU and maintaining stability over 10,000+ frames.
Introduces PhiCalNet, a neural network for single-shot fringe projection profilometry that avoids shape-prior shortcuts by outputting wrapped phase and using a fixed calibration layer, achieving 3.3x lower error than baseline UNet on a synthetic benchmark.
A model is proposed to fill in missing depth data from depth cameras when encountering transparent surfaces like glass walls, addressing a common sensor limitation.
PixWorld presents a unified pixel-space diffusion approach for 3D scene reconstruction and generation, overcoming limitations of latent-space methods by using direct image-level supervision and geometry-aware feature alignment. The method outperforms prior generation methods and matches state-of-the-art reconstruction methods.
Someone connected Blender to Fable 5, and in just 20 minutes, it reconstructed the entire New York City building complex based on public building data, with the model to scale. The poster believes this 'check data before acting' approach is smarter than Opus 4.8, reflecting that AI is beginning to have the ability to 'do its homework'.
PointDiT presents a minimalist pixel-space diffusion transformer using a plain ViT architecture for monocular geometry estimation, outperforming complex latent-based models while maintaining simplicity and robustness in ambiguous regions.
This paper proposes a feed-forward framework that decomposes 3D scenes into instance-structured token groups from unposed multi-view images, enabling direct object-level reconstruction, segmentation, and manipulation without 3D annotations.
Lift4D is a test-time optimization framework that reconstructs complete 4D geometry, appearance, and deformation of dynamic objects from a single monocular in-the-wild video, improving over prior methods on challenging sequences with occlusions and non-rigid motion.
The article describes a project that uses Fourier series and epicycles to reconstruct a human face on the three faces of a cube, demonstrating how sinusoids can generate complex shapes.
GeneralVLA-2 introduces GeoFuse-MV3D for improved 3D reconstruction and a governed KnowledgeBank for better memory management in robotic manipulation tasks, achieving performance gains on several benchmarks.
SpatialAvatar-0 introduces a multi-stage reconstruction method for high-quality 4D head avatars using a shared FLAME-mesh-bound Gaussian representation, achieving superior performance across benchmarks with reduced iterations.
World Tracing introduces a generative pixel-aligned geometry representation that predicts 3D points aligned with observed pixels while completing occluded surfaces. It uses a diffusion transformer trained with pixel-space flow matching, achieving strong performance on visible-surface reconstruction and complete geometry generation across object, scene, and dynamic benchmarks.
Surflo is a feed-forward 3D reconstruction model that compresses unposed RGB views into latent tokens and decodes consistent 3D surface points via flow matching, enabling variable-resolution output and outperforming existing methods in speed.