Tag
Hulu-Med is a transparent medical vision-language model that unifies understanding across text, 2D/3D images, and video, achieving state-of-the-art performance on 30 benchmarks while being fully open-source.
MinerU2.5 is a 1.2B-parameter vision-language model that achieves state-of-the-art document parsing accuracy with high computational efficiency using a coarse-to-fine parsing strategy.
SmolDocling is a compact 256M parameter vision-language model designed for end-to-end multi-modal document conversion. It introduces a new universal markup format called DocTags to capture page elements with location, competing with models 27 times larger.
olmOCR is an open-source toolkit using a fine-tuned vision language model to extract clean text from PDFs while preserving structure, optimized for large-scale batch processing.
moondream2 is a compact vision language model designed for efficient edge device inference, with benchmark results and usage instructions provided.
NVIDIA releases a reference blueprint for building vision agents and AI-powered video analytics applications, including real-time intelligence, downstream analytics, and agentic workflows for search, summarization, and Q&A.