evidence-retrieval

Tag

Cards List
#evidence-retrieval

Memory Control Signals Emerge Before Action in Long Horizon Agents

arXiv cs.AI ↗ · 3d ago Cached

This paper studies hidden states in long horizon language model agents, revealing that memory compression and recall needs are encoded before actions. It proposes the PaMER framework to reduce context consumption while maintaining task performance through state-guided compression and evidence retrieval.

0 favorites 0 likes
#evidence-retrieval

What Changes When Fact-Verification Scores Improve? Evidence and Answer Accounting Across Trained Verifiers and LLMs

arXiv cs.CL ↗ · 3d ago Cached

The paper quantifies how improvements in fact-verification scores are partitioned between answer accuracy and evidence quality, using trained DeBERTa checkpoints and LLMs across multiple benchmarks.

0 favorites 0 likes
#evidence-retrieval

PACE: Towards Surfacing Hidden Conflicts in User Requests

arXiv cs.CL ↗ · 2026-09-04 Cached

The paper introduces PACE, a dataset for evaluating whether AI models can identify hidden conflicts in user requests by retrieving implicit knowledge base facts, and proposes PaceMaker, a multi-agent framework to enhance conflict-aware decision-making.

0 favorites 0 likes
#evidence-retrieval

ECHO: A Cognitively Inspired, Auditable Memory Plane for Long-Horizon Agents

arXiv cs.AI ↗ · 2026-08-25 Cached

ECHO is a cognitively inspired, auditable memory architecture for long-horizon agents, evaluated on benchmarks like LoCoMo and LongMemEval-S with high retrieval performance.

0 favorites 0 likes
#evidence-retrieval

FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems

arXiv cs.AI ↗ · 2026-08-20 Cached

FinRCA-Bench is a benchmark designed to evaluate evidence retrieval and reasoning capabilities in financial AI systems, providing a standardized approach for assessment and improvement.

0 favorites 0 likes
#evidence-retrieval

Don't Scroll Back: Missing-Evidence Memory for Streaming Dialogue Summarization

Hugging Face Daily Papers ↗ · 2026-08-10 Cached

This paper introduces ReMEMBER, a missing-evidence memory framework for streaming dialogue summarization that retrieves and refines evidence from long histories to resolve gaps in current windows under fixed memory budgets, along with a benchmark for evaluation.

0 favorites 0 likes
#evidence-retrieval

ColGraphRAG: Late-Interaction Evidence Retrieval for Multimodal GraphRAG

arXiv cs.AI ↗ · 2026-07-21 Cached

ColGraphRAG replaces single-vector bi-encoder similarity with late-interaction MaxSim scoring for ranking graph-linked image candidates in multimodal GraphRAG, improving retrieval and QA accuracy on MultimodalQA.

0 favorites 0 likes
#evidence-retrieval

ToE: A Hierarchical and Explainable Claim Verification Framework with Dynamic Multi-source Evidence Retrieval and Aggregation

arXiv cs.AI ↗ · 2026-06-29 Cached

This paper introduces Tree of Evidence (ToE), a hierarchical and explainable claim verification framework that dynamically retrieves and aggregates multi-source evidence using reinforcement learning. Experiments show 4-24 percentage point improvements over baselines, especially against adversarially poisoned inputs from Generative Engine Optimization.

0 favorites 0 likes
#evidence-retrieval

From Scenes to Elements: Multi-Granularity Evidence Retrieval for Verifiable Multimodal RAG

arXiv cs.CL ↗ · 2026-05-15 Cached

This paper introduces GranuVistaVQA, a multimodal benchmark with element-level annotations, and GranuRAG, a framework that treats visual elements as first-class retrieval units for verifiable multimodal RAG, achieving up to 29.2% improvement over baselines.

0 favorites 0 likes
← Back to home

Submit Feedback