adirik/grounding-dino
Summary
Grounding DINO is an open-vocabulary object detection model that can detect arbitrary objects based on text descriptions, now available on Replicate.
View Cached Full Text
Cached at: 05/08/26, 06:25 AM
Similar Articles
idea-research/ram-grounded-sam
Recognize Anything Model (RAM) is a strong image tagging model with zero-shot generalization, now combined with Grounded-Segment-Anything for open-set object detection and segmentation, significantly outperforming CLIP and BLIP.
Ant Group released LingBot-Vision: DINO-family vision backbones in 4 sizes, and the 0.3B ViT-L matches DINOv3-7B on NYUv2 depth with ~23x fewer params
Ant Group released LingBot-Vision, a family of DINO-style vision backbones in 4 sizes; the 0.3B ViT-L matches DINOv3-7B on NYUv2 depth with ~23x fewer parameters, showcasing significant efficiency gains.
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
LocateAnything proposes Parallel Box Decoding for unified visual grounding and object detection, decoding geometric elements as atomic units to improve throughput and localization accuracy, supported by a large-scale dataset of 138M samples.
Playing with Vision Embeddings
This post explores DINOv3 vision embeddings by generating images that correspond to specific embedding directions, using gradient optimization and augmentation strategies to invert the model.
@AdinaYakup: LingBot Vision A self-supervised vision backbone family for dense spatial perception from Ant Group @robbyant_brain - A…
LingBot Vision, a self-supervised vision backbone family from Ant Group, uses masked boundary modeling to achieve state-of-the-art performance on dense spatial perception tasks, beating the larger DINOv3 model on NYU-Depth v2.