Tag
A CLIP-based embedding model hosted on Replicate that generates 768-dimensional embeddings for both images and text using the clip-vit-large-patch14 architecture, costing ~$0.00022 per run.
A model on Replicate that outputs CLIP ViT-L/14 features for text and images, allowing similarity computation between inputs.