Tag
LingBot-Vision is a family of self-supervised vision transformer backbones for dense spatial perception, using masked boundary modeling to capture semantic and geometric structures for tasks like depth estimation and segmentation.
LingBot Vision, a self-supervised vision backbone family from Ant Group, uses masked boundary modeling to achieve state-of-the-art performance on dense spatial perception tasks, beating the larger DINOv3 model on NYU-Depth v2.