@bo_wangbo: okay maybe it's a good time? We have a small colbert model trained at pplx, it is a continue-training of pplx-embed-0.6…
Summary
Perplexity AI releases pplx-embed-v1-late-0.6b, a small ColBERT late-interaction embedding model for retrieval, fine-tuned from their existing embedding model and optimized for MaxSim scoring, now open-source on HuggingFace.
View Cached Full Text
Cached at: 05/19/26, 12:39 AM
okay maybe it’s a good time? We have a small colbert model trained at pplx, it is a continue-training of pplx-embed-0.6b, so native multilingual, just made it open and added a section how to use MaxSim kernel: https://huggingface.co/perplexity-ai/pplx-embed-v1-late-0.6b…
perplexity-ai/pplx-embed-v1-late-0.6b · Hugging Face
Source: https://huggingface.co/perplexity-ai/pplx-embed-v1-late-0.6b
pplx-embed-v1-late-0.6b: Late-Interaction Embeddings
pplx\-embed\-v1\-late\-0\.6bis a token-level late-interaction embedding model for retrieval with MaxSim scoring. It is continued training ofpplx\-embed\-v1\-0\.6busingContrastiveLossto optimize token-level MaxSim.
Token-level embedding dim is128, which hits the fast path of the optionalerikkaum/maxsimMaxSim kernel.
https://huggingface.co/perplexity-ai/pplx-embed-v1-late-0.6b#usageUsage
Using PyLate (indexing + retrieval)``` from pylate import indexes, models, retrieve
model = models.ColBERT( model_name_or_path=“perplexity-ai/pplx-embed-v1-late-0.6b”, trust_remote_code=True, )
documents_ids = [“1”, “2”, “3”] documents = [ “Scientists explore the universe driven by curiosity.”, “Children learn through curious exploration.”, “Historical discoveries began with curious questions.”, ]
index = indexes.PLAID( index_folder=“pylate-index”, index_name=“pplx-embed-v1-late-0.6b”, override=True, ) documents_embeddings = model.encode(documents, is_query=False) index.add_documents(documents_ids=documents_ids, documents_embeddings=documents_embeddings)
retriever = retrieve.ColBERT(index=index) queries_embeddings = model.encode([“What motivates scientific discovery?”], is_query=True) scores = retriever.retrieve(queries_embeddings=queries_embeddings, k=3) print(scores)
Using the`erikkaum/maxsim`kernel \(fast MaxSim scoring\)Fused MaxSim for reranking, pair scoring, or evaluation\. Supports CUDA \(sm\_80/86/89\) and Metal \(Apple Silicon\); fp32/fp16/bf16 in, fp32 out; forward\-only\.
import torch from kernels import get_kernel from pylate import models
device = “cuda” if torch.cuda.is_available() else “mps” model = models.ColBERT( model_name_or_path=“perplexity-ai/pplx-embed-v1-late-0.6b”, trust_remote_code=True, device=device, ) maxsim = get_kernel(“erikkaum/maxsim”, version=1, trust_remote_code=True)
q_emb = model.encode([“What motivates scientific discovery?”], is_query=True, convert_to_tensor=True) d_emb = model.encode([ “Scientists explore the universe driven by curiosity.”, “Children learn through curious exploration.”, “Historical discoveries began with curious questions.”, ], is_query=False, convert_to_tensor=True)
Pad to [B=1, n_candidates, Ld_max, dim] for score_candidates_padded.
Lq, dim = q_emb[0].shape n, Ld_max = len(d_emb), max(d.shape[0] for d in d_emb) queries_pad = q_emb[0].unsqueeze(0).to(device, torch.float16) documents_pad = torch.zeros(1, n, Ld_max, dim, device=device, dtype=torch.float16) for i, d in enumerate(d_emb): documents_pad[0, i, : d.shape[0]] = d.to(device, torch.float16) query_lengths = torch.tensor([Lq], dtype=torch.int32, device=device) doc_lengths = torch.tensor([[d.shape[0] for d in d_emb]], dtype=torch.int32, device=device)
scores = maxsim.score_candidates_padded(queries_pad, documents_pad, query_lengths, doc_lengths) print(scores[0].tolist()) # fp32 scores per candidate
For ragged variable\-length pair scoring \(eval, distillation, hard\-negative mining\), use`maxsim\.score\_pairs\_packed\(\.\.\.\)`instead — see the[kernel card](https://huggingface.co/kernels/erikkaum/maxsim)for the packed API\.
## [https://huggingface.co/perplexity-ai/pplx-embed-v1-late-0.6b#performance](https://huggingface.co/perplexity-ai/pplx-embed-v1-late-0.6b#performance)Performance
We evaluate`pplx\-embed\-v1\-late\-0\.6b`on two standard late\-interaction retrieval suites and report the average nDCG@10:
- **BEIR**— average over 15 English retrieval tasks\.
- **MIRACL**— average over 18 languages\.
Benchmark`pplx\-embed\-v1\-late\-0\.6b`ReferenceBEIR \(15 tasks\)56\.61colbert\-zero: 55\.43MIRACL \(18 langs\)66\.62jina\-colbert\-v2: 62\.28
## [https://huggingface.co/perplexity-ai/pplx-embed-v1-late-0.6b#technical-details](https://huggingface.co/perplexity-ai/pplx-embed-v1-late-0.6b#technical-details)Technical Details
This model uses late interaction: queries and documents are encoded as token\-level vectors and scored with MaxSim rather than pooled into a single vector\.
For background on the base embedding family, see the[`pplx\-embed\-v1\-0\.6b`](https://huggingface.co/perplexity-ai/pplx-embed-v1-0.6b)model card and the technical report:[https://arxiv\.org/abs/2602\.11151](https://arxiv.org/abs/2602.11151)\.
> **Erik Kaunismäki (@ErikKaum):**
> Releasing my first kernel on @huggingface:
> MaxSim
>
> Late-interaction retrieval (ColBERT / PyLate) bottlenecks on materializing the full similarity matrix. This kernel avoids it by using tiled scoring with simdgroup_matrix (Metal) and WMMA.
>
> Result is 3–5× speedup compared to naive
Similar Articles
@AmelieTabatta: ColBERT models continue to embarrass models 54× their sizes , this is why we trust late interaction @LightOnIO . A 1-ye…
The article highlights how ColBERT models, despite being smaller and older, outperform larger models like Qwen3-embed-8B when coupled with late interaction techniques and minimal fine-tuning.
@bo_wangbo: We causally trained a lot of SOTA search models internally, shall we make some small release from time to time
暗示即将以低调方式发布一个强大的开源多语言ColBERT搜索模型。
@LightOnIO: Reason-ModernColBERT topped BrowseComp-Plus with just 149M parameters. Now, Agent-ModernColBERT adds ~10% on top. Reach…
LightOn released Agent-ModernColBERT, a 149M parameter open-source retrieval model that achieves performance comparable to GPT-5 combined with Qwen3-Embed-8B by integrating agent reasoning traces into queries.
@liquidai: Introducing LFM2.5-Embedding-350M and LFM2.5-ColBERT-350M: two multilingual retrieval models built for ultra-fast and a…
Liquid AI introduces LFM2.5-Embedding-350M and LFM2.5-ColBERT-350M, two multilingual retrieval models optimized for fast and accurate search across 11 languages, with latency as low as 1.5ms.
@lateinteraction: it can never be too late for some late interaction - so cool @sirupsen @turbopuffer !
Turbopuffer announces beta support for late interaction, enabling models like ColBERT to represent text as token-level vectors, combining a fast single-vector ANN first pass with exact late interaction reranking to improve recall.