FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems
Summary
FinRCA-Bench is a benchmark designed to evaluate evidence retrieval and reasoning capabilities in financial AI systems, providing a standardized approach for assessment and improvement.
View Cached Full Text
Cached at: 08/20/26, 10:14 AM
# FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI Systems Source: [https://arxiv.org/abs/2608.18534](https://arxiv.org/abs/2608.18534) Bibliographic Tools ## Bibliographic and Citation Tools Bibliographic Explorer Toggle Code, Data, Media ## Code, Data and Media Associated with this Article Demos ## Demos Related Papers ## Recommenders and Search Tools About arXivLabs ## arXivLabs: experimental projects with community collaborators arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website\. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy\. arXiv is committed to these values and only works with partners that adhere to them\. Have an idea for a project that will add value for arXiv's community?[**Learn more about arXivLabs**](https://info.arxiv.org/labs/index.html)\.
Similar Articles
FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional Deliverables
Introduces FinProBench, a benchmark for evaluating financial AI agents using role-grounded rubrics derived from real professional deliverables, and proposes an RGRC pipeline that improves evaluation for role-specialized tasks.
FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents
This paper introduces FinanceComplexQA, a comprehensive benchmark for evaluating agentic reasoning on industrial-grade financial documents, featuring bilingual support, expert-level questions, and complex layouts across six scenarios and seven tasks.
FinSkillBench: Evaluating AI Agents and Domain Skills for Investment Management
FinSkillBench is an evaluation suite for AI agents in investment management, showing that curated procedural skills improve performance compared to self-generated skills.
RusFinChain: A Russian Benchmark for Verifiable Chain-of-Thought Reasoning in Finance with Fuzzy-Aligned Evaluation
Introduces RusFinChain, the first Russian-language symbolic benchmark for verifiable chain-of-thought reasoning in finance, spanning 17 domains with 5,280 parameterized examples and enhanced evaluation metrics including fuzzy numeric alignment.
FINESSE-Bench: A Hierarchical Benchmark Suite for Financial Domain Knowledge and Technical Analysis in Large Language Models
This paper introduces FINESSE-Bench, a suite of eight specialized benchmarks with 3,993 questions for hierarchical evaluation of financial competencies in large language models, covering professional certification topics and applied trading tasks.