Tag
Recognize Anything Model (RAM) is a strong image tagging model with zero-shot generalization, now combined with Grounded-Segment-Anything for open-set object detection and segmentation, significantly outperforming CLIP and BLIP.