Tag
This paper argues that financial LLM applications require system-level validation beyond benchmark scores, covering data, model design, retrieval, agent behavior, governance, and implementation. It advocates for ongoing validation discipline and a research agenda for system-aware evaluation.