dense-retrieval

Tag

Cards List
#dense-retrieval

When Does Dense Retrieval Need Asymmetric Geometry? A Bias-Variance Theory of Shared and Dual Projections

Hugging Face Daily Papers ↗ · 4d ago Cached

The paper introduces a bias-variance theory for dense retrieval, comparing shared and dual projections, and proposes the CARS method to select optimal geometry based on directional signal estimation from training data.

0 favorites 0 likes
#dense-retrieval

From Location Phrases to Geographic Entities: Task-Adapted Retrieval for People Search

arXiv cs.AI ↗ · 2026-09-01 Cached

The paper proposes a task-adapted retrieval approach for mapping free-form location phrases to geographic entities in people search, using a prompt-asymmetric bi-encoder that improves relevance, especially for non-canonical queries, as demonstrated in production and benchmark evaluations.

0 favorites 0 likes
#dense-retrieval

Natural Language Code Retrieval for 1C:Enterprise: An Open Benchmark and Efficient Bi-Encoder

arXiv cs.CL ↗ · 2026-08-21 Cached

This paper introduces an open benchmark and a specialized bi-encoder model for natural language code retrieval in the 1C:Enterprise ecosystem, addressing the lack of domain-specific resources for Russian-language code search.

0 favorites 0 likes
#dense-retrieval

Where Does Retrieval Fail? Evaluating RAG Architectures for Agricultural Advisory

arXiv cs.CL ↗ · 2026-08-18 Cached

This paper evaluates retrieval quality in RAG systems for Bengali agricultural advisory, finding performance varies by query type and language conditions, and introduces a benchmark dataset to highlight the need for disaggregated evaluation in low-resource settings.

0 favorites 0 likes
#dense-retrieval

Retrieval Grounding Latent Reasoning for Dense Retrieval

arXiv cs.AI ↗ · 2026-08-17 Cached

Proposes Retrieval Grounding Latent Reasoning (RGLR), a latent reasoning framework for dense retrieval that explicitly connects intermediate latent transitions with retrieval improvements, outperforming baselines on reasoning-intensive tasks.

0 favorites 0 likes
#dense-retrieval

UEmbed: Unified Sparse and Dense Multimodal Embeddings

Hugging Face Daily Papers ↗ · 2026-08-03 Cached

UEmbed is a decoder-only multimodal embedding model that produces both sparse and dense representations in a single forward pass, released at 2B, 4B, and 9B scales. It outperforms existing public-data-trained multimodal embedding models on MMEB-v2 and remains competitive on BEIR.

0 favorites 0 likes
#dense-retrieval

DenseOn with the LateOn: Fully Open Dense and Late-Interaction Models for Multilingual, Long-Context, and Code Search

arXiv cs.CL ↗ · 2026-07-30 Cached

This paper presents fully open DenseOn and LateOn retrieval models, trained on curated English data and extended to multilingual settings via translate-train, achieving state-of-the-art BEIR results for their parameter size.

0 favorites 0 likes
#dense-retrieval

BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms

Hugging Face Daily Papers ↗ · 2026-07-30 Cached

A controlled scaling study of retrieval-augmented generation paradigms finds that BM25 lexical retrieval outperforms agentic and graph-based retrieval at scale, while agentic search only leads on small corpora.

0 favorites 0 likes
#dense-retrieval

LLM-INSTRUCT at UZH Shared Task 2026: Constraint-Aware Retrieval and Selective Debate for Paragraph-Level Argument Mining

arXiv cs.CL ↗ · 2026-07-24 Cached

This paper describes LLM-INSTRUCT, the winning system for the UZH Shared Task at ArgMining 2026 on paragraph-level argument mining. The system uses constraint-aware retrieval and selective debate to improve accuracy and schema compliance.

0 favorites 0 likes
#dense-retrieval

SkillSight: Seeing Through Shared Descriptions for Accurate Skill Retrieval

arXiv cs.AI ↗ · 2026-07-22 Cached

SkillSight is a training-free retrieval framework that calibrates shared background in skill descriptions to improve skill retrieval accuracy for LLM agents, achieving up to 20.21 percentage point improvement in Recall@10 over dense retrievers.

0 favorites 0 likes
#dense-retrieval

Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale

arXiv cs.CL ↗ · 2026-07-03 Cached

This paper systematically studies in-context retrieval at million-token scale, introducing BlockSearch, a 0.6B LM retriever, and analyzing attention dilution. The model matches or outperforms dense retrieval on benchmarks like MS MARCO and NQ, and significantly outperforms on tasks requiring different similarity notions, highlighting the potential of in-context retrieval while emphasizing attention control under extreme context growth.

0 favorites 0 likes
#dense-retrieval

DREAM: Dense Retrieval Embeddings via Autoregressive Modeling

Hugging Face Daily Papers ↗ · 2026-06-23 Cached

DREAM trains dense retrieval embeddings by using autoregressive language model attention to supervise query-document similarity, eliminating the need for labeled data. It consistently outperforms baselines on BEIR and RTEB benchmarks across model scales.

0 favorites 0 likes
#dense-retrieval

Lost in a Single Vector: Improving Long-Document Retrieval with Chunk Evidence Aggregation

arXiv cs.CL ↗ · 2026-06-18 Cached

This paper identifies document-side early compression as a failure mode in long-document dense retrieval and introduces the Evidence Dilution Index (EDI) to measure it. The authors propose DICE, a training-free method that splits documents into chunks, encodes them independently, and aggregates them into a single vector, significantly improving retrieval on long documents.

0 favorites 0 likes
#dense-retrieval

MCompassRAG: Topic Metadata as a Semantic Compass for Paragraph-Level Retrieval

arXiv cs.CL ↗ · 2026-06-18 Cached

MCompassRAG enhances retrieval-augmented generation by enriching chunk representations with topic metadata and using LLM-teacher distillation, achieving 8.24% average improvement in information efficiency with over 5x lower latency compared to strong baselines.

0 favorites 0 likes
#dense-retrieval

ECI_{sem}: Semantic Residual Effective Contrastive Information for Evaluating Hard Negatives

Hugging Face Daily Papers ↗ · 2026-06-05 Cached

ECI_sem is a training-free method for ranking hard negative sources in dense retrieval using frozen embeddings, achieving strong performance on MS MARCO and BEIR benchmarks.

0 favorites 0 likes
#dense-retrieval

@raphaelsrty: At 140 million parameters, our LateOn model yield strong results Unrelated to LateOn, I'm really excited by what's happ…

X AI KOLs Following ↗ · 2026-05-30 Cached

The LateOn model with 140M parameters achieves strong results, and the community is excited about advances in multi-vector models including new CPU indexes and multilingual support.

0 favorites 0 likes
#dense-retrieval

@_reachsumit: Latent Terms: Dense Retrievers Contain Trivially Extractable BM25-ready Zipfian Vocabularies @bclavie et al. extract in…

X AI KOLs Following ↗ · 2026-05-29 Cached

The paper proposes Latent Terms, a method using Sparse Autoencoders to extract BM25-ready sparse features from frozen dense retrievers, achieving competitive performance without retrieval-specific training.

0 favorites 0 likes
#dense-retrieval

CoHyDE: Iterative Co-Training of LLM Rewriter & Dense Encoder for Tool Retrieval

arXiv cs.AI ↗ · 2026-05-29 Cached

CoHyDE introduces an iterative co-training procedure for an LLM rewriter and a dense encoder to improve tool retrieval from large API catalogs. It outperforms single-component baselines, especially on vague queries, by training both components together using InfoNCE and DPO.

0 favorites 0 likes
#dense-retrieval

Xetrieval: Mechanistically Explaining Dense Retrieval

Hugging Face Daily Papers ↗ · 2026-05-28 Cached

Xetrieval is a mechanistic framework that explains dense retrieval by enhancing sentence embeddings with reasoning information and decomposing them into interpretable sparse features, providing feature-level explanations for retrieval decisions without expensive autoregressive generation.

0 favorites 0 likes
#dense-retrieval

Benchmarking Google Embeddings 2 against Open-Source Models for Multilingual Dense Retrieval and RAG Systems

arXiv cs.CL ↗ · 2026-05-25 Cached

This paper benchmarks Google Embeddings 2 against five open-source models for multilingual dense retrieval and RAG, finding GE2 top in accuracy but slower, with mE5-L as a competitive low-latency alternative.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback