@burkov: The revolutionary doc2vec paper is now on @ChapterPal: https://chapterpal.com/s/2c5qp4s2/distributed-representations-of…
Summary
The revolutionary doc2vec paper is now on @ChapterPal: https://chapterpal.com/s/2c5qp4s2/distributed-representations-of-sentences-and-documents… This paper made the notion of a document/sentence embedding mainstream. And even if today doc2vec is no longer the SOTA algorithm, it has influenced everything we are taking for granted today.
Similar Articles
@shubh6200: Spent some time reading this over the weekends and honestly I wish it existed a few years ago. every AI tutorial we wat…
A tweet recommends an arXiv paper that explains the mathematical foundations of Transformers, covering tokenization, embeddings, multi-headed attention, and KV caching for applied mathematicians.
@jerryjliu0: Our team is at CVPR 2026 if you want to come say hi :)
Jerry Liu's team is presenting ParseBench, a comprehensive document understanding benchmark for VLMs, at CVPR 2026. The benchmark includes 2,000 pages of real-world enterprise documents with evaluation metrics for tables, charts, and visual grounding.
@0x0SojalSec: Document Index for Vectorless, Reasoning-based RAG, build RAG without Vector DBs. this open-source library that uses do…
PageIndex is an open-source library for vectorless, reasoning-based RAG that uses hierarchical document trees instead of embeddings, achieving 98.7% on FinanceBench. It enables context-aware retrieval without vector databases or chunking.
@omarsar0: NEW AI paper worth bookmarking. This is something I called early, and this paper confirms it: verification has emerged …
This paper from Stanford, NVIDIA, and UC Berkeley introduces LLM-as-a-Verifier, a training-free verification framework that uses continuous scoring from LLM logits to improve accuracy across coding, robotics, and medical domains, achieving state-of-the-art results on multiple benchmarks.
@mixedbreadai: By now, everyone knows that single-vector embedding models are hugely limiting for modern workflows. But they contain t…
Single-vector embedding models can be used to extract sparse latent terms, and BM25 can turn this vocabulary into a strong retriever.