Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings
Summary
The paper identifies that LLM text embeddings overly express high-frequency uninformative tokens and proposes EmbedFilter, a linear transformation that filters out this subspace to improve semantic representations and enable dimensionality reduction.
View Cached Full Text
Cached at: 06/08/26, 07:14 AM
Paper page - Your UnEmbedding Matrix is Secretly a Feature Lens for Text Embeddings
Source: https://huggingface.co/papers/2606.07502
Abstract
Text embeddings from large language models are enhanced by EmbedFilter, a linear transformation that reduces the influence of high-frequency tokens and improves semantic representations while enabling dimensionality reduction.
Large language modelsexhibit impressive zero-shot capabilities across a wide range of downstream tasks. However, they struggle to function as off-the-shelf embedding models, leading to suboptimal performance on massive text embedding benchmarks. In this paper, we identify a potential cause underlying this deficiency. Our motivation stems from an unexpected observation:text embeddingstend to align with frequent but uninformative tokens when projected onto the vocabulary space. We argue that this excessive expression ofhigh-frequency tokenssuppresses the model’s ability to capture nuanced semantics. To address this, we introduce EmbedFilter, a simplelinear transformationdesigned to refinetext embeddingsderived from LLMs directly. Specifically, we uncover that theunembedding matrixwithin LLMs encodes a latent space that is actively writing these frequent tokens into embedding space. By filtering out this subspace, EmbedFilter suppress the influence ofhigh-frequency tokens, thereby enhancingsemantic representations. As a compelling byproduct, this enables an inherentdimensionality reduction, lowering index storage and speedup retrieval while fully preserving the refined embedding quality. Our experiments across multiple LLM backbones demonstrate that LLMs equipped with EmbedFilter achieve superior zero-shot downstream performance even with significantly reduced embedding dimensions. We hope our findings provide deeper insights into the mechanisms of LLM-based representations and inspire more principled designs to improvetext embeddingstraining. Our code is available at https://github.com/CentreChen/EmbFilter.
View arXiv pageView PDFGitHub2Add to collection
Get this paper in your agent:
hf papers read 2606\.07502
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2606.07502 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2606.07502 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2606.07502 in a Space README.md to link it from this page.
Collections including this paper1
Similar Articles
@vintcessun: Turns out LLM text embeddings are hijacked by high-frequency tokens (periods, articles)! The unembedding matrix implicitly defines a low-rank subspace dominated by these uninformative expressions. This is the root cause of LLMs' poor performance as universal embeddings, and the contamination is subtle. EmbedFilter…
This study reveals that LLM text embeddings are hijacked by high-frequency tokens (e.g., periods, articles) and proposes EmbedFilter, which performs SVD on the unembedding matrix and subtracts the projection component to release true semantics, achieving zero-training-cost dimensionality reduction and retrieval efficiency gains.
@v0xium: LLM Inference Engineering: Embedding Models Explained 1. An Embedding Model (EM) converts a chunk of text, or any other…
This article explains embedding models and their role in LLM inference, covering architectures, traffic profiles, and optimization techniques like quantization.
The Embedder's Dilemma: LLMs Are Better, but at What Cost?
The paper compares large language models and embedding models across 37 tasks, finding that while aggregate performance is similar, embedding models are far cheaper and faster, supporting a division of labor for cost-efficiency.
Geometric Filtering of LLM-Generated Samples for Few-Shot Text Classification
This paper proposes a geometric filtering framework that selects high-quality LLM-generated samples by evaluating their Euclidean distance to real class examples in an embedding space, improving few-shot text classification performance by +2.61 percentage points over SMOTE.
Query Lens: Interpreting Sparse Key-Value Features with Indirect Effects
Query Lens extends Logit Lens to interpret sparse autoencoder features by jointly considering encoder-side key features and decoder-side value features, and accounting for indirect effects from downstream modules. The paper also introduces the Subspace Channel Hypothesis, suggesting downstream modules read features through layer-specific subspaces.