Tag
REPAIR is a self-evolving data augmentation framework for scientific dense retrievers that resolves long-tail confusion via fact-verified iterative refinement, demonstrating significant performance improvements on materials science and biomedical benchmarks.
This paper presents a three-stage training pipeline for developing compact dense retrievers, introducing PolDense for Polish and EuroDense for European languages, which achieve strong performance with significantly reduced parameters compared to larger models.
Researchers extract indexable, BM25-ready sparse features from frozen dense retrievers using reconstruction-trained sparse autoencoders.
This paper investigates whether positional bias in dense retrievers originates from architecture or training data, finding that training data distribution strongly influences bias and that balanced training can reduce sensitivity by up to 87% while maintaining retrieval performance.