VultronRetriever family of models released on HuggingFace![R]

Reddit r/MachineLearning Models

Summary

Vultr released the VultronRetriever family of models, including Prime-8B, Core-4.5B, and Flash-0.8B, which achieve top rankings on the MTEB leaderboard with high efficiency and offline capabilities.

Thrilled to announce the VultronRetriever family of models, which were announced during Raise Summit Paris and demonstrated running Q&A and embedding documents on the iPhone, fully offline! πŸ“± Some highlights from the VultronRetriever model family: πŸ₯‡ Each model ranks #1 in its respective class on the MTEB Leaderboard, with VultronRetrieverPrime-8B as the global #1 πŸ“¦ VultronRetrieverPrime-8B has up to 16x smaller index storage footprint and 12x higher throughput versus previous 9B-class leaders 🎯 VultronRetrieverCore-4.5B ranks second only to Prime on the leaderboard, outperforming models twice its size ⚑ VultronRetrieverFlash-0.8B outperforms models up to 5x its size, runs cool on edge devices, and indexes up to 60 images per minute, fully offline! 🐍 Deploying the VultronRetriever models with the Hydra Architecture gives you late interaction retrieval at unparalleled precision, plus generation at up to half the memory of comparable models πŸ§ͺ All models were trained on datasets with 0% cross-dataset duplication and 0% eval contamination, and show no overfitting on privately run MTEB evals Grab them, break them, make them your own πŸ”§ πŸ† Prime: https://huggingface.co/vultr/VultronRetrieverPrime-Qwen3.5-8B βš™οΈ Core: https://huggingface.co/vultr/VultronRetrieverCore-Qwen3.5-4.5B ⚑ Flash: https://huggingface.co/vultr/VultronRetrieverFlash-Qwen3.5-0.8B πŸ“Š MTEB Leaderboard: https://mteb-leaderboard.hf.space/benchmark/ViDoRe(v3)) 🐍 Hydra Architecture: https://arxiv.org/abs/2603.28554
Original Article

Similar Articles

@swyx: roundup of links:

X AI KOLs Following

NVIDIA releases Cosmos 3 (Mixture-of-Transformers models up to 64B), Nemotron 3 Ultra (550B-A55B LLM), and previews RTX Spark personal superchip at Computex 2026, achieving SOTA on multiple open model leaderboards.