@NielsRogge: Very cool work!! Modality Forcing gets SOTA on 4 out of 5 monocular depth estimation benchmarks. Explore the paper and …

X AI KOLs Following Papers

Summary

Bardienus Duisterhof introduces Modality Forcing, a recipe for post-training text-to-image (T2I) models that achieves state-of-the-art results on 4 out of 5 monocular depth estimation benchmarks.

Very cool work!! Modality Forcing gets SOTA on 4 out of 5 monocular depth estimation benchmarks. 🏆 Explore the paper and evals here: https://t.co/i9WcxlpIdY https://t.co/eKNlbOUqWu
Original Article
View Cached Full Text

Cached at: 06/15/26, 05:05 PM

Very cool work!!

Modality Forcing gets SOTA on 4 out of 5 monocular depth estimation benchmarks. 🏆

Explore the paper and evals here: https://t.co/i9WcxlpIdY https://t.co/eKNlbOUqWu

Bardienus Duisterhof (@BDuisterhof): Introducing Modality Forcing, a recipe for post-training T2I models for SOTA RGB-Depth generation!

Text-to-image (T2I) models learn rich representations of the spatial world.

How do we build on this prior for high-quality depth generation?

https://t.co/uJjGHNiDBu

🧵 [1/6]

Similar Articles

How Modalities Learn Together (49 minute read)

TLDR AI

A systematic study from Meta FAIR, Reality Labs, and Oxford on multimodal pretraining, revealing asymmetric knowledge flow between modalities, synergy vs. competition dynamics, the benefits of early unification, and efficient training recipes validated with 13.5B MoE models.