VultronRetriever family of models released on HuggingFace![R]
Summary
Vultr released the VultronRetriever family of models, including Prime-8B, Core-4.5B, and Flash-0.8B, which achieve top rankings on the MTEB leaderboard with high efficiency and offline capabilities.
Similar Articles
@svpino: Tiny, specialized, and open models are the future! The TwiL-LM family of models is now available on HuggingFace for Traβ¦
TwiL-LM, a family of tiny specialized open models, is released on HuggingFace. The 3B version outperforms OpenAI's 120B gpt-oss on formal reasoning benchmarks and runs efficiently on consumer hardware.
NVIDIA Nemotron 3 Embed Ranks #1 Overall on RTEB, Advancing Agentic Retrieval
NVIDIA releases Nemotron 3 Embed, a collection of open embedding models that top the RTEB leaderboard, featuring an 8B flagship model and efficient 1B variants for production-scale retrieval.
Qwen3.6-35B-A3B and 9B are officially on the public Terminal-Bench 2.0 leaderboard!
Qwen3.6-35B-A3B and Qwen3.5-9B models are officially on the Terminal-Bench 2.0 leaderboard, with little-coder achieving 24.6% on the 35B variant, surpassing Gemini 2.5 Pro and Qwen3-Coder-480B, while the 9B model shows that sub-10B local models can compete on hard agentic benchmarks.
@swyx: roundup of links:
NVIDIA releases Cosmos 3 (Mixture-of-Transformers models up to 64B), Nemotron 3 Ultra (550B-A55B LLM), and previews RTX Spark personal superchip at Computex 2026, achieving SOTA on multiple open model leaderboards.
@TraffAlex: AI MODELS FOR 32GB VRAM β TOP 17 CHEAT SHEET Hit the HuggingFace API, grabbed real .gguf Q4 sizes. Every link = direct β¦
A cheat sheet listing top AI models optimized for 32GB VRAM using GGUF Q4 quantization, with direct download links from HuggingFace. Includes models from Qwen, DeepSeek, Llama, and Mistral families, with tips on quantization and context settings.