Tag
Tavily promotes its web retrieval API for financial AI agents, highlighting use cases like real-time risk research, due diligence automation, and AML case work, with security features like zero data retention and prompt injection protection.
Introduces FinProBench, a benchmark for evaluating financial AI agents using role-grounded rubrics derived from real professional deliverables, and proposes an RGRC pipeline that improves evaluation for role-specialized tasks.
This paper introduces DebtBench, the first persona-enriched benchmark for debt collection negotiation, and DebtGPT, a debt collection agent that jointly optimizes financial recovery and interaction experience. Experiments show most LLMs struggle in this realistic scenario, while DebtGPT matches GPT-4o performance.
This paper investigates whether AI agents can predict stock movements based on financial judgment, exploring the frontier of agent-based financial analysis.
This paper evaluates the ability of large language models to perform technical market analysis for trading, using a benchmark approach.
This paper presents an explainable AI approach for detecting anomalies in banking transactions from an internal audit perspective, addressing interpretability and trust in financial security systems.
XAlpha introduces a memory-driven AI quant researcher that integrates financial knowledge and discovery feedback to automate the full hypothesis-to-code alpha discovery loop, achieving stronger performance on CSI300.
This paper systematically evaluates reject inference methods in credit scoring and identifies a failure mode where accuracy improves while recall collapses, creating an illusion of improvement while rejection quality deteriorates. It proposes a controlled exploration strategy that breaks the feedback loop and shows that even minimal exploration rates are sufficient to diagnose the problem.
This paper presents a unified multi-modal framework integrating reinforcement learning, high-frequency trading, game-theoretic approaches, and cross-modal sentiment analysis for intelligent financial systems, claiming significant improvements over single-domain systems.
Leni is a newly launched AI tool for investors, claiming to be the most accurate AI for investment decisions.
This paper investigates the behavioral alignment and representation dynamics of LLM agents in financial trading, introducing the TradeArena testbed and finding measurable pre-failure signatures in planning embeddings that can predict drawdowns with high accuracy across multiple frontier models and stress conditions.
This survey examines computational nondeterminism in financial AI systems, covering tabular models, graph networks, and LLM-based workflows, and proposes a layered evaluation framework for auditability.
Y Combinator is hosting a fintech happy hour on Thursday in New York City, inviting startups working on stablecoins, tokenization, financial AI, agentic commerce, and prediction markets.
Kronos is the world's first open-source foundational large model for financial markets, trained from scratch on 12 billion real candlestick data points, supporting price prediction and volatility forecasting, far outperforming general models, and completely free and open-source.
Google is expanding its new AI-powered Google Finance service to Europe, featuring enhanced AI research, advanced charting visualizations, and live earnings insights with local language support.
A roundup of the fastest-growing GitHub repositories this week, dominated by autonomous financial and coding agent frameworks, with highlights including TradingAgents, a Claude orchestration platform, and OpenAI's Symphony. The overarching theme is multi-agent orchestration and autonomous AI workflows.
AlphaCrafter is a full-stack multi-agent framework for cross-sectional quantitative trading that uses specialized agents for factor mining, screening, and trading to adapt to evolving market conditions.
FinRAG-12B is a 12B-parameter LLM optimized for retrieval-augmented generation in banking, featuring a unified training framework that improves answer quality, citation grounding, and calibrated refusal. The model outperforms GPT-4.1 in citation grounding and is deployed across over 40 financial institutions with significant cost and latency advantages.
Researchers release SAHM, the first Arabic financial benchmark with 14,380 expert-verified instances covering Shari’ah-compliant reasoning, showing large performance gaps for 20 evaluated LLMs.
Kronos is a new foundation model for financial K-line data that uses a specialized tokenizer and autoregressive pre-training to outperform existing models in forecasting and synthetic data generation.