Masked depth modeling with sensor-validity masking: reports best RMSE on 7 of 8 masked/sparse depth benchmarks, plus a controlled encoder-init study[R]

Reddit r/MachineLearning Papers

Summary

This paper proposes masked depth modeling with sensor-validity masking, achieving best RMSE on 7 out of 8 masked/sparse depth benchmarks, with a controlled encoder-init study.

No content available
Original Article

Similar Articles

Vision Pretraining for Dense Spatial Perception

Hugging Face Daily Papers

This paper introduces masked boundary modeling, a self-supervised paradigm for vision pretraining that learns sub-pixel boundary representations to improve dense spatial perception. The resulting model, LingBot-Vision, demonstrates significant improvements in depth estimation and other downstream tasks, showing that boundary modeling is a scalable pretraining principle for spatially structured visual representations.

chenxwh/depth-anything-v2

Replicate Explore

Depth Anything V2 is a monocular depth estimation model that significantly outperforms V1 in fine-grained details and robustness, offering faster inference and higher accuracy than SD-based models. It is available on Replicate under varying licenses.