@berryxia: Bro! Jina just dropped a huge one today! Jina-embeddings-v5-omni is here! It's their first unified Embedding model that truly supports text + image + audio + video! (Multimodal EMB~!) Two si...
Summary
Jina has released Jina-embeddings-v5-omni, the first unified multimodal embedding model supporting text, images, audio, and video. The model is available in Small and Nano versions, is backward compatible with existing indexes, and boasts strong performance. It is now available on Hugging Face and via the Jina API.
Similar Articles
@JinaAI_: jina-embeddings-v5-omni is here! Our first universal embedding model for text, images, audio, and video. Available in t…
Jina AI has released jina-embeddings-v5-omni, a universal embedding model supporting text, images, audio, and video with back-compatible indexing capabilities.
jina-embeddings-v5-omni: Text-Geometry-Preserving Multimodal Embeddings via Frozen-Tower Composition
This paper introduces jina-embeddings-v5-omni, a suite of multimodal embedding models that extend text embeddings to image, audio, and video using frozen-tower composition. The method trains only 0.35% of the total weights, maintaining text geometry while achieving competitive state-of-the-art performance with significantly lower computational cost.
zsxkib/jina-clip-v2
Jina CLIP v2 is an improved multimodal embedding model supporting 89 languages, high-resolution images, and flexible embedding dimensions, with 3% better performance than v1 and state-of-the-art results on multilingual benchmarks.
@_philschmid: Gemini Embedding 2 now GA! One embedding model that understand text, images, video, audio, and PDFs! 5 modalities in a …
Google releases Gemini Embedding 2 for general availability, offering a single model that embeds text, images, video, audio, and PDFs into one unified space across 100+ languages without needing audio transcription.
tencent/WeMM-Embedding 9B/4B/2B
Tencent introduces WeMM-Embedding, a series of universal multimodal embedding models in 9B, 4B, and 2B sizes, built on Qwen3.5, supporting text, images, videos, and visual documents for embedding generation.