@theworldlabs: Today we are sharing three new research papers, each exploring a new way to generate 3D content by leveraging large-sca…
Summary
World Labs announces three new research papers focused on generating 3D content using large-scale generative models and 2D priors, led by interns Hao Zhang, Bert Duisterhof, and Ben Tunnels.
View Cached Full Text
Cached at: 06/12/26, 05:01 PM
Today we are sharing three new research papers, each exploring a new way to generate 3D content by leveraging large-scale generative models and 2D priors.
These projects were led by our incredible interns @HaoZhang623 @BDuisterhof @DrTunnels
[1/4]
World Tracing predicts full 3D from a single image. It outputs a stack of depth values for each input pixel, peeling the world into layers and predicting them with a diffusion model. This predicts full 3D (even occluded surfaces) while remaining faithful to the image.
[2/4]
Modality Forcing adapts text-to-image models to reason jointly about text, images, and depth. It shows that text-to-image is a scalable pretraining objective for 3D reasoning, and how text-to-RGBD, depth estimation, and depth-to-image can be unified in a single model.
[3/4]
Flex4DHuman lifts monocular video into dynamic 4D Gaussians. A video diffusion model is finetuned to generate synchronized multiview videos which are distilled into 4D Gaussians. With this method, a video of a person dancing can be lifted to 4D and composted into a 3D world.
Similar Articles
@theworldlabs: OpenArt now lets you turn a single image into a persistent 3D world creators can direct with precise control. Wide shot…
OpenArt now allows users to transform a single image into a persistent 3D world with precise camera control, powered by the World Labs API.
@zhiwen_fan_: paper from dust3r’s team
A paper from the DUSt3R team proposes sparse auto-regressive modeling for 3D scene generation from multi-view images, using a voxel-aligned 3D latent space and an occupancy-aware masked autoregressive transformer.
@vincieye: Generating full 3D scenes from just a few photos usually means hallucinating blind spots or blowing up memory. SPAR3S s…
SPAR3S is a sparse auto-regressive generative model that efficiently generates complete 3D scenes from sparse multi-view images without requiring 3D ground truth data, using a voxel-aligned latent space and differentiable 3D Gaussian Splatting.
@svpino: This world generation model is ridiculous. I uploaded a single image (the one on the left), and it created a complete 3…
A user showcases Hyper3D WorldGen, an AI model that generates a complete, editable 3D scene from a single image, highlighting its impressive capabilities.
@FinanceYF5: World Labs has released Atlas: This is the world's first multimodal world model, capable of generating images and video…
World Labs has released Atlas, the world's first multimodal world model capable of generating images and videos with pixel-level camera control and reconstructing them into 3D, enabling world modeling and spatial-temporal simulation.