Tag
This article critiques Retrieval-Augmented Generation (RAG) systems, demonstrating through experiments that embeddings fail to capture contextual details like contradictions, leading to hallucinations, and emphasizes the essential role of strong underlying models for accurate AI responses.
This post discusses the limitations of traditional databases for modern applications and promotes converged databases as a unified solution to handle diverse data types like relational, JSON, and vector data in a single system.
Manticore Search introduces automatic document chunking for vector search, improving recall for long documents by splitting them into chunks and embedding each chunk.
The article explains how agent memory systems can leak data due to post-filter tenant scoping in vector search and recommends scoping at write time to prevent cross-tenant data exposure.
Antfly has rewritten their search-and-inference database from Go to pure Zig for improved performance and zero dependencies. The article details their technical decisions, first principles, and implementations of vector search algorithms.
Alpine Agentic Search is an AI community meetup in Zurich on October 7, featuring technical talks on vector search, reranking, and agents for search workloads, with speakers from Nebius, NVIDIA, and Tavily.
An individual built a local long-term memory system for AI agents using markdown files and a local index, enabling persistent memory across sessions and including tests for false memories, with plans to potentially productize it.
The article discusses the challenge of memory staleness in long-running AI agents, where context becomes outdated and contradictory, and seeks practical solutions for maintaining reliable memory over time.
This article presents vector-bench, a tool designed to benchmark vector indexes across different databases with consistent conditions to accurately measure approximate nearest neighbor search performance.
This article describes a reference RAG agent implementation using Mastra and Elasticsearch vector store, demonstrated on a corpus of 500 sci-fi movies, with the agent built in approximately 60 lines of TypeScript.
Bonsai discovered a pervasive issue where embedding vectors are unnecessarily cast from float32 to float64, causing 'Float Bloat' that wastes storage and bandwidth, with an estimated global impact of over 20 Petabytes of overhead.
The author built a retrieval engine from scratch to benchmark HNSW against FAISS, finding that brute force search is faster for small document sets and that encoder latency dominates retrieval time, while RRF fusion of BM25 and dense retrieval improves quality significantly.
The article introduces Polign DB, a typed, stateless database designed for agent memory, optimized for edge deployment and cost efficiency. It demonstrates a demo showing how it can manage structured memory for AI agents with minimal resource usage.
A 30-question breakdown explaining how embeddings, vector search, and retrieval systems work, covering similarity metrics, training methods, indexing techniques, and evaluation.
This article provides a deep-dive into Retrieval-Augmented Generation (RAG) and vector search, with measured benchmarks on 100,000 documents showing the trade-offs between exact search and IVF index for speed and recall.
LatticeDB is an embedded property-graph database that integrates vector and full-text indexing into a single-file format, enabling graph traversal, similarity search, and BM25 search in one query layer for local applications.
This article explains embedding models and their role in LLM inference, covering architectures, traffic profiles, and optimization techniques like quantization.
This blog post explains how Hugging Face's Inference Endpoints, Jobs, and Buckets are used to power a hybrid search system for Papers with Code, combining keyword and vector search to improve AI research accessibility.
A user shares their experience of switching from multiple independent systems to Elasticsearch, which can handle logging, search, and monitoring tasks simultaneously, and introduces its distributed features based on Apache Lucene and its application in AI.
This article compares three common shapes of agent memory systems—file-based, structured store, and experience-based—and evaluates their effectiveness through a benchmark using a common agent loop and open-weight model.