How we built a SOTA search engine using PostgreSQL, pgvector, and Qwen3 embeddings [P]

Reddit r/MachineLearning News

Summary

A technical breakdown of how Papers with Code built a state-of-the-art hybrid search engine combining keyword and semantic search using PostgreSQL, pgvector, and Qwen3 embeddings, powered by Hugging Face infrastructure.

I wrote a technical breakdown of how search works on Papers with Code. The system combines keyword and semantic search, which produced better results than either approach alone. The stack includes: PostgreSQL with pgvector Qwen3-Embedding-0.6B for text embeddings Hugging Face Jobs with an NVIDIA L4 for batch embedding generation Hugging Face Buckets for storing artifacts A live embedding model served through Hugging Face Inference Endpoints The same infrastructure also powers the “related papers” recommendations shown on individual paper pages. Full write-up: How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code I’d be interested to hear how others are implementing hybrid search for research papers or similarly technical content. Disclosure: I work at Hugging Face and on Papers with Code.
Original Article

Similar Articles

Efficient GPU Retrieval for Semantic Search

arXiv cs.AI

This paper introduces a policy-aligned retrieval framework for semantic search on LinkedIn, leveraging embeddings partitioned into category-supervised segments and a two-stage GPU architecture to improve recall and precision, with significant gains validated in A/B testing.