Tag
The article critiques the 'remember everything' approach for AI memory in companies, advocating for separate current-state and historical views to prevent agents from using outdated information.
The article introduces GCB and KRE as two layers to optimize token usage and context management in persistent multi-agent AI systems, reducing costs while maintaining capability.
DI-Bench is a pipeline for systematically generating benchmarks for enterprise data intelligence tasks, combining knowledge retrieval and analytical computation. It evaluates models and finds they achieve only 32% accuracy on tasks where business rules modify computations.
This paper investigates the role of conditional memory in language models for scientific reasoning, introducing a Knowledge Boundary-Aware Router to dynamically activate memory based on task-specific inputs to enhance performance and avoid distractions.
SSE-Bio proposes a structured self-evolving agent with a trainable retrieval policy to improve multi-hop biomedical reasoning, demonstrating significant performance gains over baselines.
Introduces Semantic Compression Trees (SCT) for hierarchical knowledge retrieval in RAG, demonstrating efficiency gains in token usage but mixed retrieval performance compared to flat methods.
MathForm introduces a framework for mathematical autoformalization using knowledge retrieval and verification-guided refinement, yielding the FormalVerse dataset and an 8B model that outperforms specialized baselines.
This paper proposes VaseMuseum, a lightweight multimodal agent framework that combines 3D digitization with vision-language models to create an interactive digital museum for ancient Greek pottery, addressing challenges of evidence grounding and hallucination through source-level and response-level reliability control.
Introduces IsoSci, a benchmark of isomorphic cross-domain science problem pairs that separates reasoning ability from domain knowledge retrieval in LLM evaluation. The study finds that 91.3% of reasoning-mode gains are knowledge-dependent, challenging common assumptions about chain-of-thought reasoning.
This paper explores cross-lingual prompting strategies to improve access to parametric knowledge in large language models, demonstrating significant gains in knowledge transfer and factual recall across 17 languages on multilingual benchmarks.
Perplexity Brain is a memory system that builds a persistent context graph across tasks, projects, decisions, files, and sources, enabling agents to start with relevant context instead of from scratch, improving answer correctness and reducing task costs.
A developer building a multi-agent operations system for a logistics company discusses the challenge of giving agents institutional knowledge without fine-tuning, opting for a retrieval layer with human-in-the-loop approval.
Kapa announces the launch of Kapa for Agents, a platform that provides a single knowledge search tool for AI agents to access product documentation, code, and tickets, reducing dead ends and improving agent planning.
A discussion about the most useful AI agents actually deployed in production, highlighting simple, single-problem solutions like lead qualification and support triage.
The author shares their experience trying to reduce repetitive support emails with AI, finding that most automated solutions fail, and that a combination of knowledge retrieval, OCR, reply drafting, confidence scoring, and human review is more effective.
Introducing a fully local, free AI knowledge base tool called qmd, supporting Chinese note retrieval, recommended by Shopify founder and Karpathy, saving token costs.
The author discusses using Bluedot's AI meeting data as long-term memory for agents via Claude MCP integration, enabling querying of historical meeting transcripts and action items.
The article explores the possibility that free open-source LLM releases may cease, questioning whether existing models could remain useful through advanced retrieval tooling despite stale knowledge.
NGM is a training-free, plug-and-play memory module for LLMs that enhances performance by using pretrained token embeddings for N-gram knowledge retrieval without additional training or retrieval pipelines, achieving gains of up to 3 points on code generation and knowledge tasks.
Introduces MultiSearch, an RL-based framework that generates multiple queries at each reasoning step and explicitly merges retrieved information to improve signal-to-noise ratio and reasoning accuracy in question-answering tasks.