@NielsRogge: Very cool work!! Modality Forcing gets SOTA on 4 out of 5 monocular depth estimation benchmarks. Explore the paper and …
Summary
Bardienus Duisterhof introduces Modality Forcing, a recipe for post-training text-to-image (T2I) models that achieves state-of-the-art results on 4 out of 5 monocular depth estimation benchmarks.
View Cached Full Text
Cached at: 06/15/26, 05:05 PM
Very cool work!!
Modality Forcing gets SOTA on 4 out of 5 monocular depth estimation benchmarks. 🏆
Explore the paper and evals here: https://t.co/i9WcxlpIdY https://t.co/eKNlbOUqWu
Bardienus Duisterhof (@BDuisterhof): Introducing Modality Forcing, a recipe for post-training T2I models for SOTA RGB-Depth generation!
Text-to-image (T2I) models learn rich representations of the spatial world.
How do we build on this prior for high-quality depth generation?
https://t.co/uJjGHNiDBu
🧵 [1/6]
Similar Articles
One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models
Introduces MultiDepth-3k, a benchmark to evaluate depth-layer preferences in monocular depth foundation models, and shows Laplacian Visual Prompting can alter reported depth layers, suggesting complementary geometric hypotheses exist across models.
@NielsRogge: Great paper, made it available here: https://paperswithcode.co/paper/98589 Check how it compares to other text-to-image…
A paper on text-to-image generation is released with open-sourced code, models, and full training recipe, comparing performance against other models.
Masked depth modeling with sensor-validity masking: reports best RMSE on 7 of 8 masked/sparse depth benchmarks, plus a controlled encoder-init study[R]
This paper proposes masked depth modeling with sensor-validity masking, achieving best RMSE on 7 out of 8 masked/sparse depth benchmarks, with a controlled encoder-init study.
How Modalities Learn Together (49 minute read)
A systematic study from Meta FAIR, Reality Labs, and Oxford on multimodal pretraining, revealing asymmetric knowledge flow between modalities, synergy vs. competition dynamics, the benefits of early unification, and efficient training recipes validated with 13.5B MoE models.
Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models
This paper investigates modality preference in omni-modal large language models (OLLMs), revealing a paradigm shift from text-dominance to visual preference. The authors introduce a conflict-based benchmark and layer-wise probing to diagnose cross-modal hallucinations using internal model signals.