Tag
ColNanoVDR introduces the first document-free query distillation framework for multi-vector visual document retrieval, using optimal transport to align student and teacher query embeddings, enabling faster encoding while retaining high performance.
Introduces Generative Late-Interaction Embeddings (GLIE) for compressing visual document retrieval vectors, improving accuracy under storage constraints by regenerating full embeddings on demand.
EVIE-Preview-4.5B is a state-of-the-art multilingual Visual Document Retrieval model from Tencent that uses ColBERT-style late interaction and achieves leading performance on ViDoRe benchmarks with 128-dimensional token embeddings.
ConceptFormer learns continuous latent concept representations for visual document retrieval, bridging visual evidence and semantic relevance without text intermediates, achieving significant improvements over baselines.
DistilVDR is a compact 524M visual document retriever distilled from an 8B teacher via cosine alignment, achieving near-teacher accuracy on ViDoRe with 15.6x smaller indexes and faster indexing.
Argus-Retriever is a new late-interaction visual document retriever that adapts document representation to the query, achieving SOTA performance on ViDoRe benchmarks with a smaller index.
This article provides a technical guide on training and fine-tuning multimodal embedding and reranker models using the Sentence Transformers library, demonstrating performance improvements on Visual Document Retrieval tasks with Qwen3-VL.