@theworldlabs: Today we are sharing three new research papers, each exploring a new way to generate 3D content by leveraging large-sca…

X AI KOLs Following Papers

Summary

World Labs announces three new research papers focused on generating 3D content using large-scale generative models and 2D priors, led by interns Hao Zhang, Bert Duisterhof, and Ben Tunnels.

Today we are sharing three new research papers, each exploring a new way to generate 3D content by leveraging large-scale generative models and 2D priors. These projects were led by our incredible interns @HaoZhang623 @BDuisterhof @DrTunnels [1/4] https://t.co/5KBSG1AuBv
Original Article
View Cached Full Text

Cached at: 06/12/26, 05:01 PM

Today we are sharing three new research papers, each exploring a new way to generate 3D content by leveraging large-scale generative models and 2D priors.

These projects were led by our incredible interns @HaoZhang623 @BDuisterhof @DrTunnels

[1/4]

World Tracing predicts full 3D from a single image. It outputs a stack of depth values for each input pixel, peeling the world into layers and predicting them with a diffusion model. This predicts full 3D (even occluded surfaces) while remaining faithful to the image.

[2/4]

Modality Forcing adapts text-to-image models to reason jointly about text, images, and depth. It shows that text-to-image is a scalable pretraining objective for 3D reasoning, and how text-to-RGBD, depth estimation, and depth-to-image can be unified in a single model.

[3/4]

Flex4DHuman lifts monocular video into dynamic 4D Gaussians. A video diffusion model is finetuned to generate synchronized multiview videos which are distilled into 4D Gaussians. With this method, a video of a person dancing can be lifted to 4D and composted into a 3D world.

Similar Articles

@zhiwen_fan_: paper from dust3r’s team

X AI KOLs Timeline

A paper from the DUSt3R team proposes sparse auto-regressive modeling for 3D scene generation from multi-view images, using a voxel-aligned 3D latent space and an occupancy-aware masked autoregressive transformer.