Tag
The paper introduces a bias-variance theory for dense retrieval, comparing shared and dual projections, and proposes the CARS method to select optimal geometry based on directional signal estimation from training data.
The paper proposes a task-adapted retrieval approach for mapping free-form location phrases to geographic entities in people search, using a prompt-asymmetric bi-encoder that improves relevance, especially for non-canonical queries, as demonstrated in production and benchmark evaluations.
This paper introduces an open benchmark and a specialized bi-encoder model for natural language code retrieval in the 1C:Enterprise ecosystem, addressing the lack of domain-specific resources for Russian-language code search.
This paper evaluates retrieval quality in RAG systems for Bengali agricultural advisory, finding performance varies by query type and language conditions, and introduces a benchmark dataset to highlight the need for disaggregated evaluation in low-resource settings.
Proposes Retrieval Grounding Latent Reasoning (RGLR), a latent reasoning framework for dense retrieval that explicitly connects intermediate latent transitions with retrieval improvements, outperforming baselines on reasoning-intensive tasks.
UEmbed is a decoder-only multimodal embedding model that produces both sparse and dense representations in a single forward pass, released at 2B, 4B, and 9B scales. It outperforms existing public-data-trained multimodal embedding models on MMEB-v2 and remains competitive on BEIR.
This paper presents fully open DenseOn and LateOn retrieval models, trained on curated English data and extended to multilingual settings via translate-train, achieving state-of-the-art BEIR results for their parameter size.
A controlled scaling study of retrieval-augmented generation paradigms finds that BM25 lexical retrieval outperforms agentic and graph-based retrieval at scale, while agentic search only leads on small corpora.
This paper describes LLM-INSTRUCT, the winning system for the UZH Shared Task at ArgMining 2026 on paragraph-level argument mining. The system uses constraint-aware retrieval and selective debate to improve accuracy and schema compliance.
SkillSight is a training-free retrieval framework that calibrates shared background in skill descriptions to improve skill retrieval accuracy for LLM agents, achieving up to 20.21 percentage point improvement in Recall@10 over dense retrievers.
This paper systematically studies in-context retrieval at million-token scale, introducing BlockSearch, a 0.6B LM retriever, and analyzing attention dilution. The model matches or outperforms dense retrieval on benchmarks like MS MARCO and NQ, and significantly outperforms on tasks requiring different similarity notions, highlighting the potential of in-context retrieval while emphasizing attention control under extreme context growth.
DREAM trains dense retrieval embeddings by using autoregressive language model attention to supervise query-document similarity, eliminating the need for labeled data. It consistently outperforms baselines on BEIR and RTEB benchmarks across model scales.
This paper identifies document-side early compression as a failure mode in long-document dense retrieval and introduces the Evidence Dilution Index (EDI) to measure it. The authors propose DICE, a training-free method that splits documents into chunks, encodes them independently, and aggregates them into a single vector, significantly improving retrieval on long documents.
MCompassRAG enhances retrieval-augmented generation by enriching chunk representations with topic metadata and using LLM-teacher distillation, achieving 8.24% average improvement in information efficiency with over 5x lower latency compared to strong baselines.
ECI_sem is a training-free method for ranking hard negative sources in dense retrieval using frozen embeddings, achieving strong performance on MS MARCO and BEIR benchmarks.
The LateOn model with 140M parameters achieves strong results, and the community is excited about advances in multi-vector models including new CPU indexes and multilingual support.
The paper proposes Latent Terms, a method using Sparse Autoencoders to extract BM25-ready sparse features from frozen dense retrievers, achieving competitive performance without retrieval-specific training.
CoHyDE introduces an iterative co-training procedure for an LLM rewriter and a dense encoder to improve tool retrieval from large API catalogs. It outperforms single-component baselines, especially on vague queries, by training both components together using InfoNCE and DPO.
Xetrieval is a mechanistic framework that explains dense retrieval by enhancing sentence embeddings with reasoning information and decomposing them into interpretable sparse features, providing feature-level explanations for retrieval decisions without expensive autoregressive generation.
This paper benchmarks Google Embeddings 2 against five open-source models for multilingual dense retrieval and RAG, finding GE2 top in accuracy but slower, with mE5-L as a competitive low-latency alternative.