@_reachsumit: Latent Terms: Dense Retrievers Contain Trivially Extractable BM25-ready Zipfian Vocabularies @bclavie et al. extract in…
Summary
The paper proposes Latent Terms, a method using Sparse Autoencoders to extract BM25-ready sparse features from frozen dense retrievers, achieving competitive performance without retrieval-specific training.
View Cached Full Text
Cached at: 05/29/26, 11:58 PM
Latent Terms: Dense Retrievers Contain Trivially Extractable BM25-ready Zipfian Vocabularies
@bclavie et al. extract indexable, BM25-ready sparse features from frozen dense retrievers using reconstruction-trained Sparse Autoencoders.
📝 https://t.co/WRIaCu2xIm
Latent Terms: Dense Retrievers Contain Trivially Extractable BM25-ready Zipfian Vocabularies
Source: https://arxiv.org/html/2605.29384 Benjamin Clavié1,2,Sean Lee1, Aamir Shakir1,Makoto P. Kato2,3 1Mixedbread AI, 2National Institute of Informatics (NII),3University of Tsukuba Correspondence:[email protected]
Abstract
We proposeLatent Terms, a method revealing that models trained for dense retrieval, whether single- or multi-vector, learn representations that can trivially be decomposed into retrieval-ready sparse features. When trained on frozen retrievers, Sparse Autoencoders without any retrieval-specific adjustments extract a latent vocabulary with approximately Zipfian collection statistics, directly suitable for classical sparse retrieval scoring via BM25. This approach enables sparse retrieval while requiring no learned expansion objective or sparse retrieval supervision whatsoever, and can be readily applied to any dense retriever.Latent Termsis able to match or outperform single-vector scoring methods from its own base model as well as comparable SPLADE variants. In addition, it substantially outperforms its base model on LIMIT, a task specifically designed to highlight the failures of single-vector retrieval. Overall, our results highlight that neural retrievers contain more expressive and indexable structure than their default scoring functions expose, but that other methods can nonetheless be leveraged.
Latent Terms: Dense Retrievers Contain Trivially Extractable BM25-ready Zipfian Vocabularies
Benjamin Clavié1,2, Sean Lee1,Aamir Shakir1,Makoto P. Kato2,31Mixedbread AI,2National Institute of Informatics (NII),3University of TsukubaCorrespondence:[email protected]
1Introduction
Neural information retrieval is deeply tied to representation learning. A retrieval model, typically built on a pre-trained language model backbone, is trained to produce representations that can be searched through a particular scoring interfaceMiutra and Craswell (2018). In practice, neural retrievers are often categorized by the representations they expose at inference time and by the operators used to score them. Dense single-vector retrievers encode queries and documents into one vector each and score them with dot product or cosine similarityYateset al.(2021). Late-interaction, or dense multi-vector, retrievers expose sets of token-level vectors and score them with operations such as MaxSimKhattab and Zaharia (2020). Learned sparse retrievers such as SPLADE expose sparse vocabulary weights that can be indexed and searched efficientlyFormalet al.(2021b).
Meanwhile, Sparse Autoencoders (SAEs), models trained to map a model activation into a higher-dimensional sparse code and then reconstruct the original activation from that code, have become a widely used tool for analyzing the internal representations of neural networksCunninghamet al.(2023); Gaoet al.(2024).
In this work, we ask whether the interface exposed by a retriever captures all of the retrieval-relevant structure learned by the model. Recent work has shown that single-vector retrievers can sometimes be adapted into strong multi-vector retrieversClavié (2024); Chaffin (2025), suggesting that trained retrievers may encode useful retrieval structure beyond what their default scoring interface exposes. We study a complementary question: do dense retrievers also contain sparse, indexable structure, even when they are not trained to produce sparse representations?
Specifically, we hypothesize that SAEs could recover such structure from dense retrievers, by converting a model’s representations into a “quasi-lexical” latent vocabulary. To test this hypothesis, we introduceLatent Terms. Given a frozen retriever,Latent Termsencodes queries and documents, projects their final-layer token representations through an SAE, and applies BM25Robertsonet al.(1995)directly over the resulting sparse activations, treating activated feature indices as vocabulary terms and transformed activation magnitudes as term weights.
Importantly,Latent Termsdoes not train a sparse retriever with retrieval supervision, nor use any normally-required sparse training methods such as learned expansion objectivesFormalet al.(2021b), hard negativesXionget al.(2021), or sparsity regularization, with FLOPs regularization being the most common and alternatives remaining an active area of researchPorcoet al.(2025). Instead, the SAE is trained only with the standard SAE reconstruction objective over web text extracted from FineWeb-EduPenedoet al.(2024). All sparsity comes from the SAE itself, while ranking is performed by a classical BM25 scorer over the resulting features.
We applyLatent Termsto multiple dense retrievers with varying original retrieval performance. Despite its simplicity,Latent Termsconsistently extracts strong sparse retrieval performance from all evaluated frozen backbones: it matches comparable SPLADE variants111Comparable models defined as competitive models developed around the same time period.and outperforms the base model’s single-vector cosine similarity approach on both single-vector backbones tested. The gains are especially pronounced on benchmarks designed to expose limitations of single-vector models, further suggesting thatLatent Termscan leverage relevant structure that is present in the model but inaccessible through its single-vector scoring interface.
We then show that SAE features learned from retrieval models form a latent vocabulary whose collection statistics resemble those of natural-language terms, providing BM25 with meaningful document-frequency statistics. Qualitative analysis supports the view that SAEs extract a meaningful vocabulary, which contains a mixture of lexical as well as both narrow and broad semantic units, combining sparse indexability with a vocabulary induced from the neural retriever’s internal representation.
Overall, our results suggest a different view of neural retrieval models. A model’s default scoring function is not necessarily the only useful way to access its retrieval knowledge. Dense retrievers can contain sparse, expressive, and indexable structure that their inference interface does not expose, and this structure can be recovered with a reconstruction-trained SAE and classical sparse IR methods such as BM25.
ContributionsIn summary, our contributions are: We (i) introduceLatent Terms, a simple method for converting frozen retriever activations into BM25-searchable sparse representations using reconstruction-trained SAEs; (ii) show that these latent vocabularies support strong sparse retrieval without sparse retrieval supervision; and (iii) propose an analysis of why the method works by showing that the generated vocabulary has term-like collection statistics and a mix of meaningful semantic and lexical units.
2Background
2.1Sparse Autoencoders
Sparse Autoencoders (SAEs) are shallow neural networks trained to represent a dense activationh∈ℝdh\in\mathbb{R}^{d}using a higher-dimensional sparse codez∈ℝ≥0mz\in\mathbb{R}_{\geq 0}^{m}, where typicallym≫dm\gg d. They are built on an encoder-decoder architecture, made up of an encoderfencf_{\mathrm{enc}}and decoderfdecf_{\mathrm{dec}}:
z=fenc(h),h^=fdec(z),z=f_{\mathrm{enc}}(h),\qquad\hat{h}=f_{\mathrm{dec}}(z),(1)Trained jointly with an objective comprising a reconstruction term, encouraging information preservation, and a sparsity penalty so each input activates only a small subset of latent features:
ℒSAE(h)=‖h−h^‖22+λ‖z‖1.\mathcal{L}_{\mathrm{SAE}}(h)=\|h-\hat{h}\|_{2}^{2}+\lambda\|z\|_{1}.(2) SAEs have become a common tool in neural network, and specifically language-model interpretability. In the latter, the activations on which the SAE is trained are individual token-level activations. This approach is used to decompose dense neural activations into features that are more localized and interpretable than individual coordinates of the original representationCunninghamet al.(2023), thus facilitating the process of interpreting otherwise “black-boxed” neural activationsLieberumet al.(2024); Templetonet al.(2025). Indeed, rather than interpreting activations through potentially polysemantic individual dimensions, SAEs aim to learn a basis in which different latent dimensions can map to a specific pattern or conceptBrickenet al.(2023).
2.2Okapi BM25
Okapi Best Match 25, more frequently referred to as just BM25Robertsonet al.(1995), is a ubiquitous method in classical information retrieval which remains surprisingly competitive against modern neural methods, especially with proper per-dataset parameter tuningKamphuiset al.(2020). Given a queryQQand a documentDD, BM25 scoresDDby summing the contributions of query terms that occur in the document:
BM25(Q,D)\displaystyle\operatorname{BM25}(Q,D)=∑t∈QIDF(t)f(t,D)(k1+1)f(t,D)+k1KD,\displaystyle=\sum_{t\in Q}\operatorname{IDF}(t)\frac{f(t,D)(k_{1}+1)}{f(t,D)+k_{1}K_{D}},(3)KD\displaystyle K_{D}=1−b+b|D|avgdl.\displaystyle=1-b+b\frac{|D|}{\operatorname{avgdl}}.Here,f(t,D)f(t,D)is the frequency of termttin documentDD,|D||D|is the length ofDD,avgdl\operatorname{avgdl}is the average document length in the collection,k1k_{1}controls term-frequency saturation andbbcontrols document-length normalization. The inverse document frequency term is commonly defined as
IDF(t)=logN−n(t)+0.5n(t)+0.5,\operatorname{IDF}(t)=\log\frac{N-n(t)+0.5}{n(t)+0.5},(4)whereNNis the total number of documents andn(t)n(t)is the number of documents containing termtt.
Traditionally, BM25 is used directly on textual inputs, with various levels of pre-processing, functioning as a bag-of-words method defined over lexical terms. Its scoring combines inverse document frequency, term-frequency saturation, and document-length normalization. However, its underlying assumptions are not inherently lexical: BM25 can in principle be applied to any set of sparse features with meaningful collection frequencies, magnitudes, and lengths.
2.3Learned Sparse Retrieval
Learned sparse retrieval encompasses a family of neural methods that preserve the efficiency of classical lexical retrieval while leveraging language models to improve its representation. Broadly, the general recipe is to keep sparse, vocabulary-indexed representations, but have them be learned or enhanced by a model rather than directly derived from the surface form of the input text.
This taken many forms throughout the years, with early instantiations addressing queries and documents separately. Initially, work focused on the document-side: DeepCTDai and Callan (2019)and DeepImpactMalliaet al.(2021)developed contextualized methods to learn sparse document representations, while Doc2Query approachesGospodinovet al.(2023); Nogueiraet al.(2019)focused on vocabulary expansion to mitigate the document/query vocabulary mismatch problem. uniCOILGaoet al.(2021)further expanded on these methods by attempting to reconcile them, learning scalar weights over lexical terms with optional vocabulary expansion.
Following this early work, SPLADEFormalet al.(2021b)proposed handling weighting and expansion jointly within a single end-to-end model, leveraging the language modeling abilities of a pre-trained model such as BERTDevlinet al.(2019). Given an input sequencexx, SPLADE reuses the language modeling head of its base encoder to project each contextualized token representationhih_{i}onto the encoder vocabularyVV, before aggregating these projections into a single sparse vectorw(x)∈ℝ≥0|V|w(x)\in\mathbb{R}^{|V|}_{\geq 0}:
wj(x)=maxi∈1..|x|log(1+ReLU(MLM(hi)j)).w_{j}(x)=\max_{i\in 1..|x|}\log\!\left(1+\operatorname{ReLU}\!\left(\operatorname{MLM}(h_{i})_{j}\right)\right).(5)with the log-ReLU transformation ensuring non-negative vocabulary weights and a final pooling operation allowing each vocabulary item to be activated by the most relevant input position. The resulting sparse vector can therefore contain both observed terms and expansion terms predicted by the language model. Scoring is then defined as sparse vector matches, commonly expressed with an inner product:
s(q,d)=⟨w(q),w(d)⟩=∑j∈Vwj(q)wj(d).s(q,d)=\langle w(q),w(d)\rangle=\sum_{j\in V}w_{j}(q)w_{j}(d).(6)As the majority of coordinates are zero, these representations can be indexed and searched with efficient indexing methods such as inverted indexes, while still benefiting from contextualized neural term weighting and expansion.
Achieving competitive retrieval with SPLADE-style models, however, requires considerably more than simply applying a masked language-modeling head. In addition to the common complexities of retrieval training, such as mined hard negatives or knowledge distillation techniquesLassanceet al.(2024), SPLADE’s performance can be sensitive to factors other model families are robust to, such as the tokenization methodHu (2026), and requires explicit sparsity regularization during trainingFormalet al.(2021b,2024).
2.4Other SAE-Based Retrieval Work
CL-SRParket al.(2025)proposed the use of SAEs on the task of reconstructing the final, single-vector representations of a dense retriever. In doing so, they demonstrated that not only do the extracted features provide a degree of interpretability, and showed that the resulting latent features could be scored in a SPLADE-like manner, using an inner product to perform retrieval, which they dubbed a form ofConcept-LevelSparseRetrieval. However, while a promising avenue, its retrieval performance is substantially degraded compared to that of its original single-vector retriever and is built on a fully in-domain setting, with in-domain queries and the same corpus used for training the SAE and evaluating retrieval downstream. Furthermore, CL-SR does not explore the use of SAEs on token-level representations, and instead instead focusing on final, single-vector representations.
Concurrent work such as SPLAREFormalet al.(2026)and BM25-VHanet al.(2026)have also both proposed different ways of leveraging SAEs as vocabularies. BM25-VHanet al.(2026)proposes the BM25 over SAE-generated features from a CLIPRadfordet al.(2021)-like vision encoder, but restricts its exploration to the use of such method as a high-recall, low-ranking-quality first stage retriever which requires a second stage using the model’s normal scoring function instead. Meanwhile, SPLARE uses an SAE-generated vocabulary over a frozen LLM. This vocabulary is then used as the basis for full retrieval training, employing a SPLADE-like training pipeline and achieving moderate but consistent improvements over the same training pipeline using the original model vocabulary insteadFormalet al.(2026).
3Latent Terms: BM25 over SAE Features Extracted from Dense Retrievers
At a high level,Latent Termstakes a frozen dense retrieverRR, trains an SAE on the activationsRRproduces over unlabeled text, and uses the resulting sparse latent code as a vocabulary on which BM25 is applied. This approach relies on three broad steps: training the SAE, constructing latent representation, and BM25 scoring over the latent vocabulary.
3.1Training SAEs on Frozen Retrievers
LetRRbe a frozen dense retriever that maps an input sequencex=(x1,…,x|x|)x=(x_{1},\ldots,x_{|x|})to a set of contextualized token representations
R(x)=(h1,…,h|x|),hi∈ℝd,R(x)=(h_{1},\ldots,h_{|x|}),\qquad h_{i}\in\mathbb{R}^{d},(7)whereddis the retriever’s final hidden dimension. Importantly, we make no assumptions about howRRis normally scored at inference time: the process is the same whetherRRis a single-vector retriever, where individual tokens would be pooled into one document-level vector, or a multi-vector model. In all cases,Latent Termsreads activations from the final hidden states of the backbone model.
We train an SAE on token-level activations drawn fromRRrun over unlabeled web text. Specifically, we sample passages from FineWeb-EduPenedoet al.(2024). Every resulting token activationhih_{i}is treated as an independent training example for the SAE. Training minimizes the standard reconstruction objective discussed in Section2.1.
During training, we set the SAE’s total latent vocabulary dimension to 32,768 terms, within the same order of magnitude as common monolingual tokenizersDevlinet al.(2019), and fix the top-k sparsity to 16 active features per token to ensure that the representations we will use for retrieval will naturally remain sparse. We found that increasing the latent vocabulary size, training data volume, or training-time top-k did not meaningfully improve downstream retrieval, which was robust to these hyperparameter choices overall. We provide further details on hyperparameters in AppendixA. Importantly, the SAE never sees data which is directly in-domain for retrieval tasks.
Following this training, the trained SAE encoderfencf_{\mathrm{enc}}is frozen. It is used to project individual token into the 32,768latent vocabularyVSAE={1,…,m}V_{\mathrm{SAE}}=\{1,\ldots,m\}which will serve as our retrieval vocabulary in subsequent steps. The decoderfdecf_{\mathrm{dec}}plays no further role and is discarded.
3.2Constructing Latent Sparse Representations
At indexing and query time, we usefencf_{\mathrm{enc}}as a token-level sparse projector. Given an inputxx, which can be either a document to be indexed or a query, we first obtain its token-level dense activations fromRR, then apply the SAE encoder independently at each position, relying on the backbone model’s own contextualization of each token:
zi=fenc(hi)∈ℝ≥0m.z_{i}=f_{\mathrm{enc}}(h_{i})\in\mathbb{R}^{m}_{\geq 0}.(8)Eachziz_{i}is sparse by construction, and themmcoordinates ofziz_{i}are entries in the fixedlatent vocabularyVSAE={1,…,m}V_{\mathrm{SAE}}=\{1,\ldots,m\}learned by the SAE.
To produce a single representation per input, we aggregate the per-token codes by sum-pooling. Our experiments confirmed that max-pooling resulted in consistent minor performance degradation compared to sum-pooling. We believe this to be due to sum-pooling preserving the cumulative evidence contributed by repeated feature activations across the input, while max-pooling retains only the single strongest activation of each feature, discarding repeated weaker activations that contribute useful evidence. After pooling, we finally apply an element-wise activation transformϕ:ℝ≥0→ℝ≥0\phi:\mathbb{R}_{\geq 0}\to\mathbb{R}_{\geq 0}:
w~(x)=∑i=1|x|zi,w(x)=ϕ(w~(x))∈ℝ≥0m,\tilde{w}(x)=\sum_{i=1}^{|x|}z_{i},\qquad w(x)=\phi\!\left(\tilde{w}(x)\right)\in\mathbb{R}^{m}_{\geq 0},(9) Forϕ\phiwe consider sublinear transforms such asϕ(u)=uα\phi(u)=u^{\alpha}forα∈(0,1)\alpha\in(0,1). This activation transform is beneficial because the summed SAE activationsw~j(d)=∑izi,j\tilde{w}_{j}(d)=\sum_{i}z_{i,j}fundamentally differ from the term frequencies in natural language, where eachtfj(d)\operatorname{tf}_{j}(d)is a count. On the other hand, each per-token activationzi,jz_{i,j}produced by the SAE encoder is a non-negative real-valued magnitude carrying signal rather than a simple indicator of token presence. For all experiments, we use the square-root transformϕ(u)=u\phi(u)=\sqrt{u}as the default parameter, and otherwise includeϕ\phiin our hyperparameter tuning process as described in Section4.2.1.
The resulting representation is a sparse, non-negative vectorw(x)∈ℝ≥0mw(x)\in\mathbb{R}^{m}_{\geq 0}per input. Its support
supp(w(x))={j∈VSAE:wj(x)>0}\operatorname{supp}(w(x))=\{j\in V_{\mathrm{SAE}}:w_{j}(x)>0\}(10)identifies the features activated inxx, while the magnitudeswj(x)w_{j}(x)capture theϕ\phi-transformed activation strength of each feature across all tokens ofxx.
3.3BM25 Scoring over Latent Features
Finally, at retrieval time, given the sparse representationsw(q)w(q)andw(d)w(d)defined above,Latent Termsscores query-document pairs by applying the Okapi BM25 formula over the latent vocabularyVSAEV_{\mathrm{SAE}}. One adjustment to the lexical formulation is needed: sincewj(q)w_{j}(q)is a real-valued activation rather than a binary indicator of term presence, we retain it as an explicit per-feature weight on each summand of the BM25 sum. Effectively, in the lexical case,jjcorresponds to a term andwj(D)w_{j}(D)is its term frequency inDD; in ourLatent Termssetting, we interpret the BM25 frequency term asf(j,D)=wj(D)f(j,D)=w_{j}(D). We otherwise perform scoring and the inverse document frequency (IDF) calculation over the term activations as in Section2.2.
In this context, all of BM25’s structural mechanisms, saturating term-frequency contributions, document-length normalization, and IDF-based downweighting of pervasive features, transfer largely unchanged from natural language toLatent Terms. Thus, indexing and retrieval with this method is “plug-and-play” with existing infrastructure and can be carried out with any standard BM25 implementation that accepts a custom vocabulary: the latent feature indices serve directly as vocabulary entries in an inverted index.
4Experimental Setup
4.1SAE Training
Training SettingThe SAEs are all trained using the same method mimicking standard best practices, following the Top-K SAE architecture introduced byGaoet al.(2024). We did not find any significant downstream improvements with SAE variants such as JumpReLURajamanoharanet al.(2024)or BatchTopKBussmannet al.(2024). The decoder is initialized with Kaiming initialization and the encoder is initialized with the transposed weights of the decoder, followingBrickenet al.(2023). We use the AdamW optimizer with a maximum learning rate of0.0010.001with 5% linear warmup followed by a cosine decay to 0. All trainings are performed on a single A100 GPU, taking under two hours per run. All SAEs are trained 5 times with different random seeds to minimise variance, with results reported as the average of these 5 runs.
Dense Encoder Backbones.We apply our method to three models to showcase its applicability to encoders with different training methods and scoring functions. Specifically, we use ContrieverIzacardet al.(2022)222Specifically, the nthakur/contriever-base-msmarco available on the HuggingFace hub, as there appears to be multiple variants with conflicting reported results., a widely studied single-vector model, contemporary with SPLADE-v2Formalet al.(2021a). We also evaluate the single-vector retrieval model nomic-embed-text-v1.5Nussbaumet al.(2025)(Nomic), a model using more modern training methods, contemporary to SPLADE-v3Lassanceet al.(2024)and with much stronger downstream performance than Contriever, to ensure that our method’s gains are not restricted to weaker models. Finally, we also use GTE-ModernColBERTChaffin (2025)(GTE-MC), a strong multi-vector model following ColBERTKhattab and Zaharia (2020).
4.2Retrieval Evaluation Setup
Baselines.
We report the results of various baselines to contextualise our method’s performance. Our sparse baselines are lexical BM25 as well as multiple generations of SPLADE models: SPLADE-v2, SPLADE-v2-Distill, and SPLADE-v3. We also report the results of all three of our chosen backbone models evaluated in their normal scoring setting, and ofLatent Termsover a non-finetuned BERTDevlinet al.(2019)to confirm that our approach requires structures learned during retrieval training. Main Evaluation Data.We report our main results over standard information retrieval benchmarks. Specifically, our main results are obtained by evaluating all methods on the widely usedBEIRThakuret al.(2021)evaluation suite, containing 15 datasets across a variety of domains and which is currently the de facto standardised way of evaluating English information retrieval models. LIMIT.To further understand the information extracted by ourLatent Termsmethod, we also report results on LIMITWelleret al.(2026), a benchmark specifically designed to test the theoretical limitations of single-vector retrieval while being trivial for lexical models: while the strongest single-vector models score under 10% on its main metric, Recall@20, BM25 reaches a score of above 95%.
4.2.1BM25 Tuning
BM25 introduces two main tunable parameters, controlling term frequency penalties and length regularization, which are known to have a potentially substantial impact on retrieval performanceHsuet al.(2026); Kamphuiset al.(2020), andLatent Termsadditionally introduces the tunable knob of theϕ\phitransform applied to raw model activations. For best results, it is common practice to tune BM25 to individual datasets to best match their idiosyncrasiesHsuet al.(2026); He and Ounis (2005). We find thatLatent Termsis very resilient to BM25 hyperparameter choices, with AppendixDpresenting a full comparison of results with and without tuning.
5Results
5.1Main Retrieval Results
Table 1:Main retrieval results. Italicized values inLatent Termsrows indicate that theLatent Termsvariant improves over its base retriever. BM25-based methods reported with tuned parameters. All results reported as [email protected] present our main results on the full BEIR collection in Table1. Overall, we observe thatLatent Terms, while fully out-of-domain, is a capable retriever across all evaluated settings. While it lags behind MaxSim scoring used by GTE-ModernColBERT, it outperforms the native cosine similarity scoring for both single-vector models evaluated. As expected, performance is particularly weak when used with an unfinetuned BERT, indicating that the necessary information is learned during retrieval fine-tuning. Comparison with SPLADE.When paired with a backbone from the same era as a given SPLADE variant,Latent Termsconsistently outperforms it. Combined with Nomic, it outperforms SPLADE-v3, while it outperforms the no-knowledge distillation variant of SPLADE-v2 when paired with Contriever. Interesting dataset-level differences are noticeable: on domain-specific tasks such as FiQA or TREC-Covid, comparableLatent Termsresults strongly outperform SPLADE variants. However, on ArguAna, an argument mining task where lexical overlap is particularly important, SPLADE outperforms it, and the gap in performance is also significantly narrower on NQ, a large-scale QA dataset, again characterized by strong lexical overlap between questions and answers. Comparison with Dense Models.The comparison ofLatent Termsapproaches with their dense backbones in their default scoring setting reveals some interesting patterns. As commonly thought, it appears that ColBERT’s MaxSim scoring allows more of the model’s learned knowledge to be expressed via its scoring mechanism, and thus remains noticeably stronger than itsLatent Termsvariant. On the other hand,Latent Termsis strong against both single-vector backbones, on both a weaker model such as Contriever or a competitive, near-state-of-the-art model like Nomic, but the magnitude of the performance differences is notable: whileLatent Terms+Nomic only very slightly outperforms its backbone, essentially just matching its overall performance with different strengths and weaknesses,Latent Terms+Contriever is vastly superior to its native scoring setting. We hypothesize that Contriever’s comparatively lighter training regimen produces good latent representations but does not fully saturate the model’s final scoring pathway, leaving room for sparse extraction to recover additional signal. We hypothesize that Contriever’s comparatively lighter training regimen produces good latent representations but does not fully exploit the model’s final scoring pathway, leaving room for sparse extraction to recover additional signal. We intend to further explore this in future work.
5.2LIMIT
Table 2:Recall@k of selected models on LIMIT.Next, Table2presents the results of selected methods on LIMIT. On common retrieval methods, our results reproduce those of the original paper: BM25 reaches extremely strong performance, closely followed by multi-vector retrieval models, with SPLADE models reaching weaker results and single-vector models completely collapsing on this task. This is line with what the paper introducing LIMIT proposes: a deliberately simple task with noisy lexical attributes designed to crowd out signal in single-vector representations.
Latent Termsappears to significantly recover performance, reaching a score over 25 times higher when applied to Contriever compared to its single-vector setting. This does offer strong insight further supporting the idea that the inherent limits of single-vector retrievers lie in their scoring mechanism,butthat this scoring mechanism also offers sufficient training signal for the underlying model to learn to generate representations that are able to at least partially capture this signal. With the GTE-MC backbone, MaxSim once again allows it to be the strongest neural model evaluated by directly leveraging token-level signal. However, it appears that while its dense representation mechanism has better ranking abilities thanLatent Termsapplied over it, it also reaches earlier recall saturation: its Recall@1000 caps out at 87.95%, almost identical to its Recall@100, whereasLatent Termsreaches 97.75%, suggesting better long-tail performance and virtually matching purely lexical approaches.
5.3Overall
Figure 1:Frequency distribution of activated features.Overall, these results appear to suggest thatLatent Termsis able to extract a retrieval-native vocabulary from dense retrievers, and that this vocabulary can be used as a capable retriever when combined with traditional IR methods for handling sparse representations. We believe that the overall strength of MaxSim on all tasks, and the results ofLatent Termsvariants on both classical tasks such as BEIR, and LIMIT, where it recovers performance that is otherwise collapsed in single-vector scoring methods both reinforce our hypothesis: Dense retrievers learn expressive representations, but there exist cases where they cannot be expressed in a way that is captured by their scoring mechanism.
6The Anatomy ofLatent Terms
Table 3:Representative features sampled from each of the three qualitatively identified categories.### 6.1SAE Features Have Term-Like Collection Statistics
We now explore why BM25 works out-of-the-box on the generated sparse features while dot product scoring does not, in complete contrast with SPLADE models, for which BM25 scoring yields significantly degraded results (AppendixC).
We attribute this to the distributional properties of SAE features. Indeed, BM25 was developed specifically to match the distributional properties of natural language, which is understood to follow a quasi-Zipfian distributionYuet al.(2018), meaning that ther−thr-thmost common word appears roughly1/r1/rtimes as often as the most common oneZipf (1932). This shape is key to enabling the weighting parameters of BM25 to act as an effective way to discriminate between documents.
Figure1presents the term distribution over the full MS MARCONguyenet al.(2016)collection for SPLADE-v3,Latent Terms+Nomic and natural language. The features generated by SPLADE diverge from Zipf’s law via a lack of dominant, quasi-stopword features, common in natural language. On the other hand, SAE-generated features, while not a perfect fit, are Zipfian in nature. Notably, they generated very pronounced saturated terms at the top of the distribution before adopting a curve that remains less steep than that of natural terms’, which is seemingly sufficient to leverage BM25’s weighting mechanisms. AppendixBpresents a rapid exploration of saturated term pruning.
Figure 2:Distribution of features by feature types.
6.2What Do Latent Retrieval Terms Capture?
Next, we qualitatively identify three categories of features, presented with examples in Table3, and design a simple automated annotation process during which we use Gemini 3 ProGoogle DeepMind (2025)to annotate all vocabulary terms. We then randomly sample 500 of its annotations and manually review them, finding perfect human-LLM agreement. The results of this annotation process are presented in Figure2. The majority of features fall within thebroad topicalcategory, with a third of the features being lexical and just 10% being narrow semantic ones. This distribution suggests that the information captured byLatent Termstend towards a form of “hybrid” semantic-lexical representation, with around two-thirds of its features being primarily semantic and the remaining third focusing on purely lexical matches. Interestingly, this falls in line with the existing literature, which has long argued that purely semantic matching misses informationYateset al.(2021)that can be recovered through hybrids of semantic and lexical methodsCormacket al.(2009).
7Conclusion
In this paper, we introduceLatent Terms, which demonstrates that dense retrievers contain more information than just the than what is exposed through their cosine similarity-based scoring mechanism. These features can be extracted by Sparse Autoencoders without any retrieval-specific modifications, yielding extracted features that are Zipfian in nature and approach the distribution of natural language. We further show that these features are suitable for BM25 scoring, originally designed for lexical terms, and can reach retrieval performance that surpasses that of the same model used in its native single-vector similarity setting. Finally, qualitative analysis reveals that these features capture multiple categories of information, creating a “hybrid” mix of lexical and semantic features. These results are obtained without any retrieval-specific data during the training of the SAE, highlighting that dense retrievers naturally learn meaningful, indexable sparse representations. We believe that these results encourage future research into decoupling the study of scoring operators from that of retrieval representation learning to better understand what truly limits the expressivity of retrievers.
Limitations
We identify six main limitations to our study, which we plan to address in future work.
Language.Our study largely focuses on the English language, as it is the highest resource language to demonstrate the mechanisms studied. We believe future work should extend this approach to both non-English monolingual and multilingual models, as it is possible that a shared, cross-lingual vocabulary could be surfaced by the SAE encoder.
Lexical/Semantic Hybridification.While our analysis reveals that the features hybridize to an extent, our results do not appear to fully match the strength of a true Dense + Sparse hybrid method as it suffers from some drawbacks on datasets where one or the other is typically strong. However, we believe that our results are encouraging and point towardsLatent Termspotentially paving the way to better hybridification of feature types which warrants further exploration.
SAE Variants and SAE Limitations.We have two limitations related to SAEs: The first is that while we did not find BatchTopKBussmannet al.(2024)or JumpReLURajamanoharanet al.(2024)to outperform our Top-K SAEGaoet al.(2024), the SAE literature is growing rapidly. Whether some other SAE variant is better suited to extracting retrieval-adapted features remains an open question. Additionally, while SAEs are popular models, they are known to have inherent limitations that have not yet been overcomeSharkeyet al.(2025), notably around feature completenessLeasket al.(2025)and potential dependency on the training datasetKissaneet al.(2024).
Sparsity.This study does not deeply explore how the sparsity generated by our method actually manifests, and whether there are techniques to make it more efficient, such as by eliminating saturated terms from the index as proposed byLassance and Clinchant (2022)to increase the efficiency of SPLADE, especially as Figure1appears to indicate there exists many such terms.
Alternate Scoring Approaches.Our approach demonstrates that dense models contain extractable sparse features, using BM25 as the scoring mechanism. However, BM25 is just one of many scoring methods, and although it is empirically strong for lexical terms, future work should explore scoring methods that could be better suited toLatent Terms.
Latent Terms+ColBERT as a separate class.In this study, we explore the use of our method applied to one ColBERT model, but otherwise treat the late interaction family of retrieval models as a special class of dense retrievers. While this is taxonomically reasonable, ColBERT’s strong token-level signal may warrant dedicated extraction methods that exploit late-interaction structure more directly.
Ethical Considerations
All retrieval models are currently understood to contain poorly-understood biases, and can potentially result in downstream issues should they surface such biased results which are not understood be biased. While we believe this to be an issue that warrants further work to be alleviated, our work, focusing on extracting representations within existing models, does not meaningfully carry considerably greater risk than existing retrieval methods. Due to its non-generative nature, we believe that our work is unlikely to be able to result in significant harm.
References
- T. Bricken, A. Templeton, J. Batson, B. Chen, A. Jermyn, T. Conerly, N. Turner, C. Anil, C. Denison, A. Askell,et al.(2023)Towards monosemanticity: decomposing language models with dictionary learning.Transformer Circuits.Cited by:Table 4,§2.1,§4.1.
- BatchTopK sparse autoencoders.InNeurIPS 2024 Workshop on Scientific Methods for Understanding Deep Learning,External Links:LinkCited by:§4.1,Limitations.
- A. Chaffin (2025)GTE-ModernColBERT.Note:HuggingFace HubExternal Links:LinkCited by:§1,§4.1.
- B. Clavié (2024)JaColBERTv2.5: optimising multi-vector retrievers to create state-of-the-art japanese retrievers with constrained resources.External Links:2407.20750,LinkCited by:§1.
- G. V. Cormack, C. L. A. Clarke, and S. Buettcher (2009)Reciprocal rank fusion outperforms condorcet and individual rank learning methods.InProceedings of the 32nd International ACM SIGIR Conference on Research and Development in Information Retrieval,SIGIR ’09,New York, NY, USA,pp. 758–759.External Links:ISBN 9781605584836,Link,DocumentCited by:§6.2.
- H. Cunningham, A. Ewart, L. Riggs, R. Huben, and L. Sharkey (2023)Sparse autoencoders find highly interpretable features in language models.External Links:2309.08600,LinkCited by:§1,§2.1.
- Z. Dai and J. Callan (2019)Context-aware sentence/passage term importance estimation for first stage retrieval.arXiv preprint arXiv:1910.10687.Cited by:§2.3.
- J. Devlin, M. Chang, K. Lee, and K. Toutanova (2019)BERT: pre-training of deep bidirectional transformers for language understanding.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers),pp. 4171–4186.Cited by:§2.3,§3.1,§4.2.
- T. Formal, C. Lassance, B. Piwowarski, and S. Clinchant (2021a)SPLADE v2: sparse lexical and expansion model for information retrieval.arXiv preprint arXiv:2109.10086.Cited by:§4.1.
- T. Formal, C. Lassance, B. Piwowarski, and S. Clinchant (2024)Towards effective and efficient sparse neural information retrieval.ACM Trans. Inf. Syst.42(5).External Links:ISSN 1046-8188,Link,DocumentCited by:§2.3.
- T. Formal, M. Louis, H. Déjean, and S. Clinchant (2026)Learning retrieval models with sparse autoencoders.InThe Fourteenth International Conference on Learning Representations,External Links:LinkCited by:§2.4.
- T. Formal, B. Piwowarski, and S. Clinchant (2021b)SPLADE: sparse lexical and expansion model for first stage ranking.InProceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp. 2288–2292.Cited by:§1,§1,§2.3,§2.3.
- L. Gao, T. D. la Tour, H. Tillman, G. Goh, R. Troll, A. Radford, I. Sutskever, J. Leike, and J. Wu (2024)Scaling and evaluating sparse autoencoders.External Links:2406.04093,LinkCited by:Table 4,§1,§4.1,Limitations.
- L. Gao, Z. Dai, and J. Callan (2021)COIL: revisit exact lexical match in information retrieval with contextualized inverted list.InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty, and Y. Zhou (Eds.),Online,pp. 3030–3042.External Links:Link,DocumentCited by:§2.3.
- Google DeepMind (2025)Gemini 3 pro model card.Model cardGoogle DeepMind.Note:Model release: November 2025. Accessed: 2026-05-22External Links:LinkCited by:§6.2.
- M. Gospodinov, S. MacAvaney, and C. Macdonald (2023)Doc2Query–: when less is more.InAdvances in Information Retrieval: 45th European Conference on Information Retrieval, ECIR 2023, Dublin, Ireland, April 2–6, 2023, Proceedings, Part II,Berlin, Heidelberg,pp. 414–422.External Links:ISBN 978-3-031-28237-9,Link,DocumentCited by:§2.3.
- D. Han, E. Park, and S. Seo (2026)Visual words meet bm25: sparse auto-encoder visual word scoring for image retrieval.External Links:2603.05781,LinkCited by:§2.4.
- B. He and I. Ounis (2005)Term frequency normalisation tuning for bm25 and dfr models.InAdvances in Information Retrieval,D. E. Losada and J. M. Fernández-Luna (Eds.),Berlin, Heidelberg,pp. 200–214.External Links:ISBN 978-3-540-31865-1Cited by:§4.2.1.
- T. Hsu, J. Yang, and J. Lin (2026)Rethinking agentic search with pi-serini: is lexical retrieval sufficient?.External Links:2605.10848,LinkCited by:§4.2.1.
- X. Hu (2026)Beyond bm25 and dense embeddings: how we built smart and interpretable retrieval at faire.Note:The Craft, Medium PostExternal Links:LinkCited by:§2.3.
- G. Izacard, M. Caron, L. Hosseini, S. Riedel, P. Bojanowski, A. Joulin, and E. Grave (2022)Unsupervised dense information retrieval with contrastive learning.Transactions on Machine Learning Research.Cited by:§4.1.
- C. Kamphuis, A. P. de Vries, L. Boytsov, and J. Lin (2020)Which bm25 do you mean? a large-scale reproducibility study of scoring variants.InEuropean Conference on Information Retrieval,pp. 28–34.External Links:DocumentCited by:§2.2,§4.2.1.
- O. Khattab and M. Zaharia (2020)Colbert: efficient and effective passage search via contextualized late interaction over bert.InProceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval,pp. 39–48.Cited by:§1,§4.1.
- C. Kissane, R. Krzyzanowski, N. Nanda, and A. Conmy (2024)Saes are highly dataset dependent: a case study on the refusal direction.InAlignment Forum,Cited by:Limitations.
- C. Lassance and S. Clinchant (2022)An efficiency study for splade models.InProceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval,SIGIR ’22,New York, NY, USA,pp. 2220–2226.External Links:ISBN 9781450387323,Link,DocumentCited by:Limitations.
- C. Lassance, H. Déjean, T. Formal, and S. Clinchant (2024)SPLADE-v3: new baselines for splade.arXiv preprint arXiv:2403.06789.Cited by:§2.3,§4.1.
- P. Leask, B. Bussmann, M. Pearce, J. Bloom, C. Tigges, N. Al Moubayed, L. Sharkey, and N. Nanda (2025)Sparse autoencoders do not find canonical units of analysis.InInternational Conference on Learning Representations,Vol.2025,pp. 53617–53642.Cited by:Limitations.
- T. Lieberum, S. Rajamanoharan, A. Conmy, L. Smith, N. Sonnerat, V. Varma, J. Kramár, A. Dragan, R. Shah, and N. Nanda (2024)Gemma scope: open sparse autoencoders everywhere all at once on gemma 2.External Links:2408.05147,LinkCited by:§2.1.
- A. Mallia, O. Khattab, T. Suel, and N. Tonellotto (2021)Learning passage impacts for inverted indexes.InProceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval,SIGIR ’21,New York, NY, USA,pp. 1723–1727.External Links:ISBN 9781450380379,Link,DocumentCited by:§2.3.
- B. Miutra and N. Craswell (2018)An introduction to neural information retrieval.Found. Trends Inf. Retr.13(1),pp. 1–126.External Links:ISSN 1554-0669,Link,DocumentCited by:§1.
- T. Nguyen, M. Rosenberg, X. Song, J. Gao, S. Tiwary, R. Majumder, and L. Deng (2016)MS MARCO: a human generated machine reading comprehension dataset.choice2640,pp. 660.Cited by:Appendix B,§6.1.
- R. Nogueira, W. Yang, J. Lin, and K. Cho (2019)Document expansion by query prediction.arXiv preprint arXiv:1904.08375.Cited by:§2.3.
- Z. Nussbaum, J. X. Morris, A. Mulyar, and B. Duderstadt (2025)Nomic embed: training a reproducible long context text embedder.Transactions on Machine Learning Research.Note:Reproducibility CertificationExternal Links:ISSN 2835-8856,LinkCited by:§4.1.
- S. Park, T. Kim, and Y. Ko (2025)Decoding dense embeddings: sparse autoencoders for interpreting and discretizing dense retrieval.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.),Suzhou, China,pp. 26468–26485.External Links:Link,Document,ISBN 979-8-89176-332-6Cited by:§2.4.
- G. Penedo, H. Kydlíček, L. B. Allal, A. Lozhkov, M. Mitchell, C. Raffel, L. Von Werra, and T. Wolf (2024)The fineweb datasets: decanting the web for the finest text data at scale.InProceedings of the 38th International Conference on Neural Information Processing Systems,NIPS ’24,Red Hook, NY, USA.External Links:ISBN 9798331314385Cited by:§1,§3.1.
- A. Porco, D. Mehra, I. Malioutov, K. Radhakrishnan, M. Keymanesh, D. Preoţiuc-Pietro, S. MacAvaney, and P. Cheng (2025)An alternative to flops regularization to effectively productionize splade-doc.InProceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval,SIGIR ’25,New York, NY, USA,pp. 2789–2793.External Links:ISBN 9798400715921,Link,DocumentCited by:§1.
- A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark,et al.(2021)Learning transferable visual models from natural language supervision.InInternational conference on machine learning,pp. 8748–8763.Cited by:§2.4.
- S. Rajamanoharan, T. Lieberum, N. Sonnerat, A. Conmy, V. Varma, J. Kramár, and N. Nanda (2024)Jumping ahead: improving reconstruction fidelity with jumprelu sparse autoencoders.arXiv preprint arXiv:2407.14435.Cited by:§4.1,Limitations.
- S. E. Robertson, S. Walker, S. Jones, M. M. Hancock-Beaulieu, M. Gatford,et al.(1995)Okapi at trec-3.Nist Special Publication Sp109,pp. 109.Cited by:§1,§2.2.
- L. Sharkey, B. Chughtai, J. Batson, J. Lindsey, J. Wu, L. Bushnaq, N. Goldowsky-Dill, S. Heimersheim, A. Ortega, J. I. Bloom, S. Biderman, A. Garriga-Alonso, A. Conmy, N. Nanda, J. M. Rumbelow, M. Wattenberg, N. Schoots, J. Miller, W. Saunders, E. J. Michaud, S. Casper, M. Tegmark, D. Bau, E. Todd, A. Geiger, M. Geva, J. Hoogland, D. Murfet, and T. McGrath (2025)Open problems in mechanistic interpretability.Transactions on Machine Learning Research.Note:Survey CertificationExternal Links:ISSN 2835-8856,LinkCited by:Limitations.
- A. Templeton, T. Conerly, J. Marcus, J. Lindsey, T. Bricken, B. Chen, A. Pearce, C. Citro, E. Ameisen, A. Jones,et al.(2025)Scaling monosemanticity: extracting interpretable features from Claude 3 Sonnet.Transformers Circuits.Cited by:§2.1.
- N. Thakur, N. Reimers, A. Rücklé, A. Srivastava, and I. Gurevych (2021)Beir: a heterogenous benchmark for zero-shot evaluation of information retrieval models.arXiv preprint arXiv:2104.08663.Cited by:§4.2.
- O. Weller, M. Boratko, I. Naim, and J. Lee (2026)On the theoretical limitations of embedding-based retrieval.InThe Fourteenth International Conference on Learning Representations,External Links:LinkCited by:§4.2.
- L. Xiong, C. Xiong, Y. Li, K. Tang, J. Liu, P. N. Bennett, J. Ahmed, and A. Overwijk (2021)Approximate nearest neighbor negative contrastive learning for dense text retrieval.InInternational Conference on Learning Representations,External Links:LinkCited by:§1.
- A. Yates, R. Nogueira, and J. Lin (2021)Pretrained transformers for text ranking: bert and beyond.InProceedings of the 14th ACM International Conference on web search and data mining,pp. 1154–1156.Cited by:§1,§6.2.
- S. Yu, C. Xu, and H. Liu (2018)Zipf’s law in 50 languages: its structural pattern, linguistic interpretation, and cognitive motivation.arXiv preprint arXiv:1807.01855.Cited by:§6.1.
- G. K. Zipf (1932)Selected studies of the principle of relative frequency in language.Cited by:§6.1.
Appendix ASAE Parameters
Table 4:SAE training hyperparameters used for allLatent Termsruns reported in the main results.The full parameters we used for the final model are presented in Table4.
Appendix BTerm Pruning
Table 5:Effect of pruning the most-activated latent features on retrieval quality. Percentage changes relative to no pruning shown in parentheses.Section6.1showed thatLatent Termsfeatures have a heavy head, with many features appearing to be saturated. This prompts investigation into whether or not pruning such terms would hurt. In Table5, we present the performance impact of pruning the most frequent terms from the vocabulary prior to indexing on performance, on MS MARCONguyenet al.(2016), the dataset used in Section6.1.
We notice that performance appears to be robust to pruning the top 1% of features, but that it otherwise rapidly degrades with more aggressive pruning. This suggests that while the most saturated terms appear to have little discriminative values, the rest of the heavy head of the distribution does play a discriminative role.
Appendix CDot Product vs BM25 Scoring
Table 6:Comparison of dot-product and BM25 scoring on select datasets. All values report nDCG@10. Best values per dataset are bolded.Latent Termsuses Nomic in all cases.In Table6, we present the results of dot product scoring used with Nomic+Latent Termsas well as BM25 scoring used with SPLADE-v3, on selected BEIR datasets in the interest of computational efficiency.
These results appear to show that dot product scoring is unsuitable for the features extracted from an SAE as presented in our study, with strong degradation on all evaluated datasets. Interestingly, the opposite hold true for SPLADE: BM25 scoring significantly degrades performance, while dot product scoring preserves it.
We believe these results to be expected: SPLADE is trained with an explicit regularization objective which encourages it to move away from the distributional attributes expected by BM25, while our earlier results reveal that the features extracted by our method have a term-like distribution that is particularly suitable for it, but are not shaped for inner-product scoring.
Appendix DImpact of BM25 Tuning
Table 7:Comparison of tuned and untunedLatent Termsretrieval. All values report [email protected]7shows a comparison between tuned and un-tuned results, usingLatent Termsover Nomic. We show that tuning hyperparameters does result in a performance improvement on most datasets, but that the effect appears to be moderate, with default parameters remaining strong. Default parameters are defined as ak1k1of 8,bblength penalty set to 0.7, and using square root transforms asϕ\phifor both documents and queries.
Similar Articles
@mixedbreadai: By now, everyone knows that single-vector embedding models are hugely limiting for modern workflows. But they contain t…
Single-vector embedding models can be used to extract sparse latent terms, and BM25 can turn this vocabulary into a strong retriever.
@bclavie: Very excited to finally share this one after sitting on it for far too long! It's very topical now. Blog post coming ve…
Researchers extract indexable, BM25-ready sparse features from frozen dense retrievers using reconstruction-trained sparse autoencoders.
Why Advanced Encoders Lag on Sparse Retrieval? The Answer and an Approach to Bridging Vocabulary Gaps
This paper identifies a vocabulary gap as the root cause why advanced encoders like ModernBERT underperform in learned sparse retrieval, and proposes Vocabulary Transfer (VT), a model-agnostic framework that migrates encoders to sparse-friendly vocabularies, achieving state-of-the-art on the BEIR benchmark.
Training-Free Lexical-Dense Fusion for Conversational-Memory Retrieval
This paper proposes a training-free, CPU-only retrieval method that fuses BM25 lexical scores with late-interaction dense scores for conversational memory retrieval, achieving up to +17.2 points improvement on LoCoMo Hit@1 over late interaction alone across six encoders. The study provides controlled ablations on pooling operators, reranker effects, and benchmark robustness, framing the gain as a division of labor between dense and lexical signals.
@_reachsumit: No More K-means:Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval @Veritas2026 et al. replace vector clus…
This paper proposes Single-stage Sparse Retrieval (SSR), which replaces K-means clustering with sparse autoencoders and inverted indexing, achieving 15x faster indexing and halved retrieval latency while improving accuracy on the BEIR benchmark.