Can Agentic Trading Systems Pay for Their Own Intelligence?
Summary
This paper introduces TradeLens, a trace-grounded diagnostic toolkit for evaluating whether LLM-based agentic trading systems convert their reasoning and tool-use costs into measurable incremental profit, analyzing failure patterns across models like DeepSeek-V3.2 and GLM-4.7.
View Cached Full Text
Cached at: 07/14/26, 04:20 AM
# Can Agentic Trading Systems Pay for Their Own Intelligence?
Source: [https://arxiv.org/html/2607.10286](https://arxiv.org/html/2607.10286)
Qiqi Duan1∗\\ast,Changlun Li2∗\\ast,Chen Wang1∗\\ast,Fan Zhang4,5,Mengxiang Wang1,Dayi Miao1,Peixian Ma2, Jiangpeng Yan3,Liyuan Chen3,Shuoling Liu3,Preslav Nakov4,Yuyu Luo1,2,Nan Tang1,2†\\dagger 1HKUST\(GZ\)2Paradoox AI3E Fund Management Co\., Ltd4MBZUAI5The University of Tokyo ∗\\astEqual contribution\.†\\daggerCorresponding author:nantang@hkust\-gz\.edu\.cn
###### Abstract
Large language model \(LLM\) agents are increasingly used in trading systems, where model reasoning, tool use, and continual decisions incur costs that are expected to produce trading value\. Existing evaluations typically report performance metrics, but rarely examine*agentic viability*: whether dynamic LLM\-mediated decisions convert their induced costs into measurable incremental profit\. To apply this criterion, we introduceTradeLens, a trace\-grounded diagnostic toolkit for evaluating agentic trading systems from their trading records, runtime traces, and deployment configurations\. It reconstructs trading trajectories, attributes profit and cost to interpretable evidence, and diagnoses whether and why an agent pays for its own intelligence\. We conduct extensive analysis across backbone models, capital scales, trading frequencies, and system architectures, together with deployment discussion\. Our results show that viability hinges on intelligence\-to\-profit conversion: models exhibit different failure patterns, such as poor asset selection in DeepSeek\-V3\.2 and negative timing in GLM\-4\.7, while capital scale, trading frequency, and architecture matter only by amplifying or degrading decision\-attributed timing value\. These findings reframe the evaluation of LLM\-based trading agents from capability\-centric performance ranking to trace\-grounded diagnosis of intelligence\-to\-profit conversion\. Our code is available at[https://anonymous\.4open\.science/r/TradeLens](https://anonymous.4open.science/r/TradeLens)\.
Can Agentic Trading Systems Pay for Their Own Intelligence?
Qiqi Duan1∗\\ast, Changlun Li2∗\\ast, Chen Wang1∗\\ast, Fan Zhang4,5, Mengxiang Wang1, Dayi Miao1, Peixian Ma2,Jiangpeng Yan3,Liyuan Chen3,Shuoling Liu3,Preslav Nakov4,Yuyu Luo1,2,Nan Tang1,2†\\dagger1HKUST\(GZ\)2Paradoox AI3E Fund Management Co\., Ltd4MBZUAI5The University of Tokyo∗\\astEqual contribution\.†\\daggerCorresponding author:nantang@hkust\-gz\.edu\.cn
## 1Introduction
Figure 1:Motivation\. A profitable agentic trading system may still fail to create useful trading value, and diagnosing “Why” is challenging because both returns and costs arise from intertwined deployment drivers\.With strong general reasoning abilitiesYuet al\.\([2023a](https://arxiv.org/html/2607.10286#bib.bib54)\); Liuet al\.\([2023a](https://arxiv.org/html/2607.10286#bib.bib1)\); Suma and Dauncey \([2025](https://arxiv.org/html/2607.10286#bib.bib16)\)and increasing adaptation to financial tasksWuet al\.\([2023](https://arxiv.org/html/2607.10286#bib.bib47)\); Chenet al\.\([2025a](https://arxiv.org/html/2607.10286#bib.bib5)\); Liuet al\.\([2023b](https://arxiv.org/html/2607.10286#bib.bib29)\), LLM agents are moving from financial question answering to direct participation in trading workflows\. Recent systems use LLMs as trading copilotsFanet al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib11)\); Yanget al\.\([2024](https://arxiv.org/html/2607.10286#bib.bib51)\), portfolio managersZhaoet al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib60)\); Ko and Lee \([2024](https://arxiv.org/html/2607.10286#bib.bib21)\), and market\-analysis agents equipped with external toolsHanet al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib18)\)\.
Observation\.Despite promising results, a critical gap remains in how agentic trading systems are evaluated\. Most existing studies report*gross profit*or metrics such as Sharpe ratioLiet al\.\([2025b](https://arxiv.org/html/2607.10286#bib.bib27)\); Chenet al\.\([2025b](https://arxiv.org/html/2607.10286#bib.bib6)\); Xieet al\.\([2023](https://arxiv.org/html/2607.10286#bib.bib49)\); Sharpe \([1994](https://arxiv.org/html/2607.10286#bib.bib64)\)\. While useful for measuring trading performance, these metrics are incomplete for deployment: developers must assess not only how much profit a system makes, but also where the profit comes from and how deployment costs affect the realized margin\. This concern echoes recent cost\-aware LLM evaluation, which accounts for monetary output cost and inefficient reasoning computationErolet al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib9)\); Zellinger and Thomson \([2025](https://arxiv.org/html/2607.10286#bib.bib67)\); Wanget al\.\([2025b](https://arxiv.org/html/2607.10286#bib.bib43)\); Zhanget al\.\([2024](https://arxiv.org/html/2607.10286#bib.bib57)\)\. In agentic trading, however, cost has a stricter meaning: it is the price paid for LLM\-mediated intelligence, including model reasoning, tool use, memory retrieval, and continual decision\-making\. Beyond operating within a budgetWanget al\.\([2024](https://arxiv.org/html/2607.10286#bib.bib80)\), the system should also meet the*token\-economy requirement*Chenet al\.\([2026](https://arxiv.org/html/2607.10286#bib.bib81)\): generating enough incremental profit to justify the intelligence consumed by its trading decisions\.
The blind spot\.This requirement creates a profit\-side and cost\-side double blind in evaluating agentic trading systems\. As shown in Figure[1](https://arxiv.org/html/2607.10286#S1.F1), an agentic system may earn a gross profit of 630, but after 250 in deployment costs, its net profit drops to 380; meanwhile, a simple buy\-and\-hold baseline without LLM inference or active trading costs may earn 450 over the same period\. Thus, gross profit can overstate decision value when the agent does not outperform passive market exposure, and can overstate deployability when end\-to\-end costs absorb the realized margin\. We therefore ask:Can agentic trading systems pay for their own intelligence, and how can we diagnose whether LLM\-mediated decisions convert their induced costs into incremental trading value?
Call for rigorous assessment\.Answering this question requires identifying both the value added by LLM intelligence and the costs it incurs\. The decomposition in the lower panel of Figure[1](https://arxiv.org/html/2607.10286#S1.F1)shows that both profit and cost are shaped by the deployment process\. A profitable trajectory may mainly reflect market movement or favorable initial asset selection, while the agent’s dynamic decisions contribute little or even reduce value\. Conversely, higher reasoning, tool\-use, or execution costs are not necessarily wasteful if they produce sufficient active profit\. This makes rigorous assessment necessary for diagnosing whether LLM\-mediated decisions are economically justified\.
Challenge of trace\-grounded diagnosis\.However, performing this diagnosis is difficult because the evidence needed to evaluate LLM\-mediated decisions is fragmented across different records\. Portfolio trajectories show what the system earned, but not whether the gain came from passive exposure, initial allocation, or dynamic agent intervention\. Runtime traces show what the agent did, including model calls, context processing, tool use, retries, and trading actions, but not whether these activities improved asset selection, timing, or portfolio value\. Deployment costs further vary with model choice, context length, tool\-use frequency, trading cadence, and execution behavior, making a single cost proxy insufficient\. Therefore, the challenge is to ground the diagnosis of intelligence\-to\-profit conversion in heterogeneous evidence from trading outcomes, runtime traces, and deployment costs\.
Solution\.To address these challenges, we proposeTradeLens, a trace\-grounded diagnosis toolkit for evaluating whether agentic trading systems can pay for their own intelligence\. This toolkit first reconstructs the investment process from portfolio records, runtime traces, and deployment configurations, and attributes profit and cost to interpretable components\. Based on this evidence, we define two viability notions \(Section[3](https://arxiv.org/html/2607.10286#S3)\):*system viability*, which asks whether the deployed system pays for itself after end\-to\-end costs, and*agentic viability*, which asks whether dynamic LLM\-mediated decisions create enough active value to justify their decision\-induced costs\. Then, an LLM\-based diagnosis agent generates evidence\-grounded explanations and strategy\-level suggestions from the observed trading trajectory, and converts these diagnoses into visual and textual reports \(Section[4](https://arxiv.org/html/2607.10286#S4)\)\. We evaluateTradeLensthrough controlled experiments and practitioner feedback, showing that it reveals failure modes hidden by system’s profitability and helps users assess deployability and identify revision directions \(Section[5](https://arxiv.org/html/2607.10286#S5)\)\.
Collectively, we reframe the evaluation of agentic trading systems from capability\-centric to economically grounded\. To summarize, we make three contributions:
\(i\) We formulate*profit–cost viability*as a trace\-grounded evaluation problem for agentic trading systems, distinguishing whether the overall deployed system pays for itself and whether LLM\-mediated decisions justify their induced costs\.
\(ii\) We buildTradeLens, an open\-source audit toolkit that operationalizes the methodology using trading records, runtime traces, and deployment configurations, and produces evidence\-grounded diagnostic reports\.
\(iii\) We conduct empirical viability studies across backbone model, capital scale, trading frequency, and architecture, together with practitioner feedback, showing how profit sources and intelligence cost jointly determine deployability\.
## 2Related Work
### 2\.1LLM\-based trading agents
LLM\-based trading agents have expanded from financial analysis and signal generationWanget al\.\([2025a](https://arxiv.org/html/2607.10286#bib.bib44)\); Xing \([2024](https://arxiv.org/html/2607.10286#bib.bib50)\)to decision pipelines that use memory, tools, and multi\-agent coordinationYuet al\.\([2024](https://arxiv.org/html/2607.10286#bib.bib55),[2023b](https://arxiv.org/html/2607.10286#bib.bib56)\); Xiaoet al\.\([2024](https://arxiv.org/html/2607.10286#bib.bib48)\); Liet al\.\([2025c](https://arxiv.org/html/2607.10286#bib.bib26)\)\. Recent work also moves evaluation closer to practical trading settings, including cryptocurrency marketsLiet al\.\([2024](https://arxiv.org/html/2607.10286#bib.bib24)\); Luoet al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib32)\), portfolio constructionGuoet al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib17)\); Zhaoet al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib60)\), streaming interactionLiet al\.\([2025a](https://arxiv.org/html/2607.10286#bib.bib28)\); Fanet al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib11)\), and end\-to\-end research platformsSunet al\.\([2023](https://arxiv.org/html/2607.10286#bib.bib40)\); Zhanget al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib58)\)\. Across these works, evaluation mainly focuses on capability or gross trading performance, leaving profit sources and economic viability underexplored\.TradeLensaddresses this gap as a diagnosis toolkit for existing trading agents, checking whether these agents can pay for their intelligence\.
### 2\.2Cost\-aware and trace\-grounded AI evaluation
Recent work increasingly incorporates computational cost into model evaluation\. Cost\-of\-pass estimates the monetary cost of obtaining a successful answerErolet al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib9)\); Zellinger and Thomson \([2025](https://arxiv.org/html/2607.10286#bib.bib67)\); efficient\-reasoning studies examine thinking style and redundancy in agentic systemsWanget al\.\([2025c](https://arxiv.org/html/2607.10286#bib.bib45)\); Linet al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib68)\); Wanget al\.\([2025b](https://arxiv.org/html/2607.10286#bib.bib43)\); Zhanget al\.\([2024](https://arxiv.org/html/2607.10286#bib.bib57)\); and other works advances resource\-constrained evaluationChenet al\.\([2023](https://arxiv.org/html/2607.10286#bib.bib4)\); Huanget al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib19)\); Yanget al\.\([2025b](https://arxiv.org/html/2607.10286#bib.bib52)\)\. These studies mainly treat cost as a resource\-efficiency variable\. For LLM agents, however, costs are also induced by the decision process itself, including reasoning, tool use, memory access, and multi\-step interactionCassanoet al\.\([2023](https://arxiv.org/html/2607.10286#bib.bib82)\)\. Token economics further frames these activities as paid token and compute consumption before downstream actions are producedWanget al\.\([2024](https://arxiv.org/html/2607.10286#bib.bib80)\); Chenet al\.\([2026](https://arxiv.org/html/2607.10286#bib.bib81)\)\. Prior work on agent evaluation and diagnosis has shown that such trajectories provide useful evidence beyond final task outcomesHeet al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib84)\); Ouet al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib83)\)\. In trading workflows, we use this trace\-grounded view in a different setting: agent activities can be aligned with resource use, trading actions, and portfolio outcomes, enabling diagnosis of whether decision\-induced costs are converted into trading value\.
### 2\.3Performance attribution in financial evaluation
Financial evaluation has long separated raw return from interpretable sources of return\. Risk\-adjusted metrics such as the Sharpe ratioSharpe \([1994](https://arxiv.org/html/2607.10286#bib.bib64)\)and regime\-aware analysesAng and Timmermann \([2012](https://arxiv.org/html/2607.10286#bib.bib2)\)show that portfolio outcomes can reflect volatility or market conditions rather than decision skill\. Classical performance attribution decomposes benchmark\-relative returns into allocation, selection, and interaction effectsBrinsonet al\.\([1986](https://arxiv.org/html/2607.10286#bib.bib69),[1991](https://arxiv.org/html/2607.10286#bib.bib70)\), with later work using attribution to study active management skill and dynamic allocation behaviorGrinold and Kahn \([2000](https://arxiv.org/html/2607.10286#bib.bib71)\); Hsuet al\.\([2010](https://arxiv.org/html/2607.10286#bib.bib72)\); Al\-Aradi and Jaimungal \([2018](https://arxiv.org/html/2607.10286#bib.bib74)\)\.
Recent financial AI benchmarks and trading\-agent studies often report predictive accuracy, cumulative return, or portfolio\-level profitKo and Lee \([2024](https://arxiv.org/html/2607.10286#bib.bib21)\); Liet al\.\([2025b](https://arxiv.org/html/2607.10286#bib.bib27)\); Chenet al\.\([2025b](https://arxiv.org/html/2607.10286#bib.bib6)\); Fanet al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib11)\)\. These metrics are useful, but they do not explain whether profit comes from market exposure, asset selection, or dynamic reallocation\.TradeLensadopts a lightweight attribution view for this purpose and joins it with system\-cost accounting, so viability is assessed from decision\-attributed margin rather than gross profit alone\.
## 3Methodology: Profit–Cost Viability
This section defines how we evaluate whether an agentic trading system can pay for the intelligence it consumes\. We focus on two questions: whether the system is economically viable, and whether its agentic decisions generate sufficient incremental value to justify their induced costs\.
### 3\.1Profit and Cost Attribution
Given an evaluation window\[1,T\]\[1,T\], letV0V\_\{0\}be the initial portfolio value andVTdynV\_\{T\}^\{\\mathrm\{dyn\}\}be the realized terminal value of the deployed agentic system\. The cumulative gross profit is
P1:T=VTdyn−V0\.P\_\{1:T\}=V\_\{T\}^\{\\mathrm\{dyn\}\}\-V\_\{0\}\.\(1\)
Profit attribution\.Classical performance attribution decomposes portfolio returns into interpretable sources before evaluating investment decisionsBrinsonet al\.\([1986](https://arxiv.org/html/2607.10286#bib.bib69),[1991](https://arxiv.org/html/2607.10286#bib.bib70)\)\. We adopt this logic for agentic trading systems, but reformulate it at the portfolio\-value level\. For diagnosis, we separate gross profit using two nested counterfactual baselines over the same asset universe and evaluation window\. LetVTsysV\_\{T\}^\{\\mathrm\{sys\}\}denote the terminal value of a passive market benchmark, and letVTbaseV\_\{T\}^\{\\mathrm\{base\}\}denote the terminal value obtained by holding the system’s initial allocation unchanged\. The comparison between these baselines separates aggregate market exposure from initial asset\-selection value, while the remaining difference between the realized dynamic path and the static\-allocation path is attributed to timing\. We define
Psys\\displaystyle P^\{\\mathrm\{sys\}\}=VTsys−V0,\\displaystyle=V\_\{T\}^\{\\mathrm\{sys\}\}\-V\_\{0\},\(2\)Passet\\displaystyle P^\{\\mathrm\{asset\}\}=VTbase−VTsys,\\displaystyle=V\_\{T\}^\{\\mathrm\{base\}\}\-V\_\{T\}^\{\\mathrm\{sys\}\},Ptiming\\displaystyle P^\{\\mathrm\{timing\}\}=VTdyn−VTbase\.\\displaystyle=V\_\{T\}^\{\\mathrm\{dyn\}\}\-V\_\{T\}^\{\\mathrm\{base\}\}\.
Thus,
P1:T=Psys\+Passet\+Ptiming\.P\_\{1:T\}=P^\{\\mathrm\{sys\}\}\+P^\{\\mathrm\{asset\}\}\+P^\{\\mathrm\{timing\}\}\.\(3\)
Here,PsysP^\{\\mathrm\{sys\}\},PassetP^\{\\mathrm\{asset\}\}, andPtimingP^\{\\mathrm\{timing\}\}correspond to market exposure, asset selection, and timing decisions, respectively\. In our analysis,PtimingP^\{\\mathrm\{timing\}\}is treated as the profit component most directly attributable to agentic intervention\.
Cost Attribution\.Transaction cost economics treats market participation as non\-frictionless: information processing, decision\-making, coordination, and execution all consume resourcesDow \([2018](https://arxiv.org/html/2607.10286#bib.bib46)\)\. For agentic trading systems, “intelligence” is produced through paid inference and then realized through execution and deployment pipelines\. For costs, we decompose the deployment cost in each periodttas
Ct=Ctllm\+Cttrd\+Ctinf\+Ctsto,C\_\{t\}=C\_\{t\}^\{\\mathrm\{llm\}\}\+C\_\{t\}^\{\\mathrm\{trd\}\}\+C\_\{t\}^\{\\mathrm\{inf\}\}\+C\_\{t\}^\{\\mathrm\{sto\}\},\(4\)
whereCtllmC\_\{t\}^\{\\mathrm\{llm\}\}is LLM inference cost,CttrdC\_\{t\}^\{\\mathrm\{trd\}\}is trading execution cost,CtinfC\_\{t\}^\{\\mathrm\{inf\}\}is infrastructure and data\-access cost, andCtstoC\_\{t\}^\{\\mathrm\{sto\}\}captures residual hard\-to\-enumerate costs\. The cumulative cost is
C1:T=∑t=1TCt\.C\_\{1:T\}=\\sum\_\{t=1\}^\{T\}C\_\{t\}\.\(5\)
Detailed profit and cost itemization is provided in Appendix[B](https://arxiv.org/html/2607.10286#A2)and Appendix[C](https://arxiv.org/html/2607.10286#A3)\.
### 3\.2Viability Criteria
Viability is evaluated over a common decision window\[1,T\]\[1,T\]\. We then distinguish two levels of viability\. System viability asks whether the whole deployed pipeline pays for itself, whereas agentic viability asks a more specific agent\-evaluation question: whether dynamic LLM intervention itself creates enough incremental value to justify the costs it induces\.
System viability\.System viability asks whether the fully deployed pipeline pays for itself\. The system\-level net profit is
R1:T=P1:T−C1:T\.R\_\{1:T\}=P\_\{1:T\}\-C\_\{1:T\}\.\(6\)
The system is viable if
SystemViable=𝕀\[R1:T≥0\]\.\\mathrm\{SystemViable\}=\\mathbb\{I\}\\\!\\left\[R\_\{1:T\}\\geq 0\\right\]\.\(7\)
Agentic viability\.Agentic viability asks whether dynamic agent intervention itself is economically justified\. It removes passive market exposure \(PsysP^\{\\mathrm\{sys\}\}\) and initial asset\-selection effects \(PassetP^\{\\mathrm\{asset\}\}\), and compares timing profit against costs that are directly induced by dynamic decisions\. We classify costs that scale with token usage, retries, and tool calls as dynamic costs, and treat fixed hosting as static costs\. The dynamic cost is
C1:Tdyn=∑t=1T\(Ctllm\+Cttrd\+Ctsto\),C\_\{1:T\}^\{\\mathrm\{dyn\}\}=\\sum\_\{t=1\}^\{T\}\\left\(C\_\{t\}^\{\\mathrm\{llm\}\}\+C\_\{t\}^\{\\mathrm\{trd\}\}\+C\_\{t\}^\{\\mathrm\{sto\}\}\\right\),\(8\)
excluding static infrastructure cost that does not vary with agent decisions\. The net agentic value is
R1:Tagent=Ptiming−C1:Tdyn\.R\_\{1:T\}^\{\\mathrm\{agent\}\}=P^\{\\mathrm\{timing\}\}\-C\_\{1:T\}^\{\\mathrm\{dyn\}\}\.\(9\)
The agent is viable if
AgenticViable=𝕀\[R1:Tagent≥0\]\.\\mathrm\{AgenticViable\}=\\mathbb\{I\}\\\!\\left\[R\_\{1:T\}^\{\\mathrm\{agent\}\}\\geq 0\\right\]\.\(10\)
## 4TradeLens
Figure 2:The pipeline overview\.Retail traders can utilize the toolkit to evaluate the profit–cost viability of their own agentic systems under realistic assumptions by providing trading results, runtime traces, and basic configurations\. Consequently, they can gain iterative improvements and achieve promising results\.In this section, we describeTradeLensthat operationalizes this methodology by reconstructing trading trajectories, aggregating runtime costs, and generating structured diagnostic reports\.
### 4\.1Overview
The design goal ofTradeLensis to make agentic trading systems auditable without assuming a specific agent architecture\. Different systems may use different LLM backbones, prompting strategies, tool chains, memory modules, or multi\-agent designs\. However, once deployed, they tend to leave similar observable traces\.TradeLenstherefore defines a trace interface around three inputs:*trading results*,*runtime traces*, and*system configurations*\.
Input\.Trading results describe what happened in the market\-facing process, including portfolio values, position weights, executed orders, and transaction logs, used to reconstruct the realized portfolio trajectory\. Runtime traces describe how the system produced its decisions, including model usage, tool calls, retries, latency, and execution events, used to recover cost drivers and system activities\. System configurations describe the deployment setting, including the LLM backbone, capital scale, trading frequency, market window, cost assumptions, and agent architecture\. They provide the context needed to interpret both profit and cost\.
Outputs\.TradeLensproduces two user\-friendly reports \(see Appendix[A](https://arxiv.org/html/2607.10286#A1)\)\. The financial report summarizes portfolio performance, profit attribution, cost attribution, and viability status\. The agentic system diagnosis report explains the failure mode and provides evidence\-grounded revision suggestions for the trader\.
### 4\.2Accounting Layer
As shown in Figure[2](https://arxiv.org/html/2607.10286#S4.F2), the accounting layer reconstructs the trading trajectory and applies the attribution method defined in Section[3](https://arxiv.org/html/2607.10286#S3)to both profit and cost\. The profit module replays trading actions to recover portfolio values and decomposes gross profit into market, selection, and timing effects\. The cost module accounts for commissions, LLM usage, infrastructure, data subscriptions, and uncertainty costs\. The layer outputs numerical audit results, including the action ledger, attribution tables, and diagnostic figures, providing fixed evidence for the diagnosis agent\.
### 4\.3Diagnosis Layer
The diagnosis layer converts accounting results into interpretable system\-level diagnoses\. It checks viability, identifies dominant failure modes, and generates diagnostic reports\. It first extracts key metrics from the audit results and summarizes runtime records into a compact execution profile\. These structured inputs are then combined under a constrained output contract, requiring the final report to follow a fixed diagnostic structure and to ground conclusions in supplied evidence\.
The diagnosis produces two types of analysis\. First, it interprets viability by combining trading outcomes and runtime measurements, identifying whether weak performance stems from insufficient trading profit, high model\-side cost, excessive decisions, or execution frictions\. Second, it maps trace\-supported failure modes to actionable revisions\. For example, frequent decisions with limited incremental profit indicate a cadence–return mismatch\. The agent then prioritizes operational fixes, such as reducing redundant model calls or decision frequency, and structural revisions, such as improving execution logic, portfolio construction, trading architecture, or the LLM’s role in the decision loop\.
## 5Experiments
In this section, we empirically evaluate the economic viability of agentic trading systems through cost\-profit breakdowns\. We conduct backtesting experiments along four deployment\-relevant dimensions: the LLM backbone model, capital scale, trading frequency, and system architecture\. We address the following questions:
- •RQ1 \(Backbone model\): How do LLM backbones differ in profit sources, intelligence cost, and net margin under the same trading pipeline?
- •RQ2 \(Capital scale\): Does capital scaling improve agentic viability by diluting fixed costs, or does it amplify the value and risk of LLM\-mediated decisions?
- •RQ3 \(Trading frequency\): Does higher decision frequency create enough marginal profit to cover additional inference and execution costs?
- •RQ4 \(System architecture\): Does architectural complexity translate into decision\-attributed gains, or merely increase coordination and execution costs?
### 5\.1Setup
Deployment setting\.We instantiateTradeLenson top of a representative agentic trading system, AI\-TraderFanet al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib11)\), to evaluate whether the toolkit can audit realistic decision pipelines\. The system trades a fixed universe of liquid U\.S\. equities over a two\-month market window from December 1, 2025 to January 30, 2026, starting with an initial capital of $100,000\. Market data are accessed through licensed provider APIs/subscriptions, and raw provider data are not redistributed\.
Profit–cost setting\.We adopt a simplified but end\-to\-end cost formulation that covers four deployment\-relevant components: LLM\-side decision cost, trading\-side execution cost, infrastructure cost, and stochastic runtime overhead\. On the profit side, we construct two attribution baselines: a systematic exposure baseline based on the S&P 500 return over the same window to the strategy’s initial capital, and a static allocation baseline that holds the initial portfolio weights unchanged\. Detailed cost assumptions and provider\-specific pricing are reported in Appendix[D\.1](https://arxiv.org/html/2607.10286#A4.SS1)\.
Experiment scenarios\.For backbone model \(RQ1\), we compare flagship LLMs from leading providers, including GPT\-5\.2OpenAI \([2025](https://arxiv.org/html/2607.10286#bib.bib37)\), Gemini 3 FlashGoogle \([2025](https://arxiv.org/html/2607.10286#bib.bib14)\), Claude Sonnet 4\.5Anthropic \([2025](https://arxiv.org/html/2607.10286#bib.bib3)\), Qwen3\-MaxYanget al\.\([2025a](https://arxiv.org/html/2607.10286#bib.bib53)\), DeepSeek\-V3\.2DeepSeek\-AIet al\.\([2024](https://arxiv.org/html/2607.10286#bib.bib30)\), GLM\-4\.7Zenget al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib61)\), Kimi\-k2Baiet al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib42)\), Minimax\-m2\.1MiniMax \([2025](https://arxiv.org/html/2607.10286#bib.bib34)\), Llama\-4\-scoutMeta AI \([2025](https://arxiv.org/html/2607.10286#bib.bib33)\), and Mistral\-large\-3Liuet al\.\([2026](https://arxiv.org/html/2607.10286#bib.bib35)\)\. For RQ2–RQ4, we use GPT\-5\.2 and DeepSeek\-V3\.2 as representative backbones\. RQ2 varies the initial endowment from $10,000 to $500,000 and reports net profit, cost\-to\-capital ratio, and break\-even status\. RQ3 compares daily and hourly decision frequencies using marginal profit minus marginal cost in one month horizon\. RQ4 explores the architecture reasoning complexity by placing AI\-Trader as a moderate design\. We utilize the typical Chain\-of\-Thought prompting \(See details in Appendix[D\.4](https://arxiv.org/html/2607.10286#A4.SS4)\) and DeepFundLiet al\.\([2025a](https://arxiv.org/html/2607.10286#bib.bib28)\)as simple and complex designs, respectively\. All variants use the same trading logic and execution assumptions\.
### 5\.2Results Analysis \(RQ1–RQ4\)
Figure 3:Viability across backbone models\.\(a\) System viability and \(b\) agentic viability: profit \(y\-axis\) versus cost \(x\-axis\) for 10 backbone LLMs under identical trading configurations\.Table 1:Viability across backbone models\. Net profit measures system viability, while agentic profit measures whether timing value covers decision\-induced costs\.RQ1 \(Backbone model\)\.The key difference across LLM backbones lies less in their operating cost than in how they generate, or fail to generate, active profit\. Figure[3](https://arxiv.org/html/2607.10286#S5.F3)and Table[1](https://arxiv.org/html/2607.10286#S5.T1)reveal models viability levels to 3 groups\. First, Mistral\-Large\-3 represents a fully viable backbone: it achieves positive net profit and positive agentic profit, despite relatively high total cost\. Its viability is mainly supported by a large positive timing effect, indicating that strong active reallocation can offset operating burden\. Second, Claude Sonnet 4\.5 represents a system\-viable but agentically weak backbone\. It achieves positive net profit, but its agentic profit is negative, suggesting that the system can cover costs while the dynamic agentic component fails to add value\. This case shows why system viability and agentic viability should be evaluated separately\. Third, most remaining backbones are non\-viable, but for different reasons\. Some models, such as DeepSeek\-V3\.2, suffer mainly from poor asset selection; others, such as GLM\-4\.7 and MiniMax\-M2\.1, are dominated by negative timing effects\. Qwen3\-Max is a borderline case with positive agentic profit but negative net profit, indicating that active improvement is insufficient to overcome total system cost\. Overall, backbone choice changes not only the magnitude of profit and cost, but also the source of economic value and failure\.
Takeaway: Backbones differ mainly in how similar intelligence costs are converted into timing value and net margin\.
Figure 4:Viability across capital scale\.*Net Profit*and*Agentic Profit*under varying initial cash investment scale for GPT\-5\.2 and DeepSeek\-V3\.2\.RQ2 \(Capital scale\)\.Figure[4](https://arxiv.org/html/2607.10286#S5.F4)shows that capital scaling does not uniformly improve viability\. DeepSeek\-V3\.2 exhibits a negative pattern\. Both system and agentic viability remain negative across all scales, with losses expanding sharply at 500k\. These losses are mainly associated with unfavorable timing effects rather than operating costs\.TradeLensidentifies a gap between positive asset selection and poor execution, where heavy LLM usage and timing slippage prevent the system from capturing market upside\. For GPT\-5\.2, system viability remains positive and generally increases with capital, whereas agentic viability is non\-linear\. However, with 500k initial cash, it only gains a modest return compared to the setting with 50k, and this result is mainly due to passive market exposure, while the timing effect is hazardous\.TradeLenstherefore recommends streamlining the inference and order pipeline to reduce latency and slippage, and recover timing value\. Overall, capital scale acts more as an amplifier of backbone\-specific trading behavior than as a simple cost\-dilution mechanism: it can magnify profitable timing decisions, but it can also enlarge strategy\-level losses\. Detailed diagnostic samples are provided in Appendix[D\.2](https://arxiv.org/html/2607.10286#A4.SS2)\.
Takeaway\.Capital scaling does not simply dilute fixed costs; it amplifies the value and risk of model’s timing behavior\.
Figure 5:Viability across trading frequency\.Cumulative net profit and agentic profit over time under hourly \(dashed\) and daily \(solid\) trading frequencies for GPT\-5\.2 and DeepSeek\-V3\.2\. The horizontal line at zero indicates the break\-even line\.RQ3 \(Trading frequency\)\.Figure[5](https://arxiv.org/html/2607.10286#S5.F5)compares cumulative net profit and agentic profit under hourly and daily decision frequencies\. The results show that higher decision frequency does not improve viability\. For both backbones, daily trading outperforms hourly trading in terms of both net profit and agentic profit\. DeepSeek\-V3\.2 improves from a net profit of \-1222\.24 under hourly trading to 29\.84 under daily trading, while GPT\-5\.2 improves from \-585\.46 to \-198\.28\. This pattern is not only caused by higher operating costs\. The diagnosis shows that the hourly setting changes the quality of active decisions\. Hourly trading increases total cost for both models, especially for DeepSeek\-V3\.2, but it also produces worse gross profit\(see Appendix Table[6](https://arxiv.org/html/2607.10286#A4.T6)\)\. This suggests that more frequent decisions introduce additional trading noise and timing errors rather than reliably capturing opportunities\.
Takeaway\.Higher frequency fails when extra decisions add noisy trades and timing errors\.
RQ4 \(System architecture\)\.Table[2](https://arxiv.org/html/2607.10286#S5.T2)shows that architectural design affects viability through active decision quality rather than cost alone\. CoT has the lowest total cost for both backbones, but it does not produce the best outcomes, suggesting that low\-cost reasoning alone is insufficient for trading\. The appendix diagnosis further shows that this failure is not driven by cost accumulation and the key bottleneck is when the agent reallocates capital\. By contrast, DeepFund achieves the strongest performance: it is the only architecture with positive net profit and agentic profit for DeepSeek\-V3\.2, and it also yields the highest values for GPT\-5\.2\. Despite its higher cost, DeepFund’s advantage comes mainly from timing\. It produces positive timing effects for both backbones, showing that its additional reasoning and coordination are converted into better dynamic reallocation\. Thus, the value of an agentic architecture depends not on being cheaper or more complex, but on whether it can convert reasoning cost into better asset selection and timing decisions\.
Takeaway\.Architectural complexity helps only when its added reasoning and coordination become decision\-attributed gains, especially timing value\.
Table 2:Viability across system architecture with $100,000 initial cash\.
### 5\.3Discussion
Sensitivity in Market Regime\.Because trading performance is highly regime\-dependent, we further examine viability across market regimes\. The result shows that market regimes reshape profit sources more strongly than system costs, making attribution necessary for interpreting deployability across environments\. Detailed results are provided in Appendix[D\.5](https://arxiv.org/html/2607.10286#A4.SS5)\.
Forward Diagnosis\.We test whetherTradeLensprovides useful signals for later deployment decisions\. Using the December diagnostic window, it flags hourly trading as cost\- and friction\-intensive, motivating a January 2026 daily\-versus\-hourly comparison\. Although daily trading performs better in early January, both settings remain net negative, showing that cost control alone cannot ensure viability when strategy\-level profit generation is weak\. Detailed results are in Appendix[E\.1](https://arxiv.org/html/2607.10286#A5.SS1)\.
Practitioner Feedback\.We also assess whether the diagnosis is interpretable and useful in private deployments\. We collect feedback from 13 retail traders deployingTradeLens\. Participants find the joint profit–cost diagnosis useful for distinguishing weak decision value, trading friction, and LLM usage\. They also emphasize that actionable deployment support should go further by providing stronger strategy\-level optimization guidance\. Additional details are in Appendix[E\.2](https://arxiv.org/html/2607.10286#A5.SS2)\.
## 6Conclusion and Future Work
We introducedTradeLens, a trace\-grounded profit–cost viability diagnosis toolkit for evaluating agentic trading systems\. It connects trading outcomes, runtime traces, and deployment costs under a shared accounting window, allowing users to assess whether a system is profitable as a whole and whether its LLM\-mediated decisions generate enough incremental value to justify their induced costs\. By distinguishing system viability from agentic viability,TradeLensdiagnoses whether intelligence is converted into economic value\.
Future work can improve friction estimation, extend the toolkit to additional asset classes and execution venues, and study how tool use, multi\-agent coordination, and long\-context reasoning affect the viability frontier\. We also plan to evaluate across broader bull–bear regimes, incorporate finer\-grained latency\-to\-fill logs and slippage models, and compare against stronger non\-LLM baselines\. Viability signals may further be integrated with reinforcement learningJinet al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib20)\); Li and Oliva \([2024](https://arxiv.org/html/2607.10286#bib.bib25)\); Qianet al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib38)\)to optimize agentic systems toward better economic outcomes\.
## Limitations
This work uses controlled experimental settings to demonstrate howTradeLensreveals profit–cost behaviors in agentic trading systems\. Thus, the reported patterns should be viewed as diagnostic observations under specific settings, not general claims about the intrinsic trading ability of any backbone model\. Their generalizability to other asset classes, broader investment universes, longer horizons, or different market regimes remains to be validated\. We use the S&P 500 as a broad\-market benchmark proxy in our experiments\. This choice may introduce slight attribution bias in measuring systematic risk exposure, and users can customize the benchmark to better match their own trading universe and deployment needs\. The practitioner evaluation involves a modest number of participants \(13 retail traders and 5 domain experts\) and collects qualitative feedback rather than controlled experimental evidence of downstream decision improvement\. Additionally, cost inputs such as infrastructure fees and stochastic terms are user\-configured rather than automatically measured; accuracy, therefore, depends on the fidelity of the values provided by the user\.
## Ethical Considerations
All market data used in our experiments is publicly available third\-party data accessed through licensed provider APIs, and we do not redistribute any raw data\. The practitioner feedback study involved voluntary participants, and we collected only qualitative feedback without personally identifiable information, sensitive information, private trading strategies, or portfolio details\. The diagnostic reports generated byTradeLensare intended as informational analysis tools and should not be interpreted as financial advice\. Users deploying agentic trading systems in real markets should exercise independent judgment and comply with applicable regulations\. We acknowledge that cost\-aware evaluation, while promoting transparency, does not eliminate the risks inherent in algorithmic trading, including potential market impact, amplification of biases in LLM reasoning, and the possibility that cost\-optimized configurations may inadvertently encourage excessive risk\-taking\.
## References
- Outperformance and tracking: dynamic asset allocation for active and passive portfolio management\.Applied Mathematical Finance25\(3\),pp\. 268–294\.External Links:[Document](https://dx.doi.org/10.1080/1350486X.2018.1507751)Cited by:[§2\.3](https://arxiv.org/html/2607.10286#S2.SS3.p1.1)\.
- A\. Ang and A\. Timmermann \(2012\)Regime changes and financial markets\.Annu\. Rev\. Financ\. Econ\.4\(1\),pp\. 313–337\.External Links:[Document](https://dx.doi.org/10.1146/ANNUREV-FINANCIAL-110311-101808)Cited by:[§D\.5](https://arxiv.org/html/2607.10286#A4.SS5.p1.1),[§2\.3](https://arxiv.org/html/2607.10286#S2.SS3.p1.1)\.
- Anthropic \(2025\)Cited by:[§5\.1](https://arxiv.org/html/2607.10286#S5.SS1.p3.1)\.
- K\. T\. Y\. Bai, Y\. Bao, G\. Chen, J\. Chen, N\. Chen, R\. Chen, Y\. Chen, Y\. Chen, Y\. Chen, Z\. Chen,et al\.\(2025\)Kimi k2: open agentic intelligence\.arXiv preprint arXiv:2507\.20534\.Cited by:[§5\.1](https://arxiv.org/html/2607.10286#S5.SS1.p3.1)\.
- G\. P\. Brinson, L\. R\. Hood, and G\. L\. Beebower \(1986\)Determinants of portfolio performance\.Financial Analysts Journal42\(4\),pp\. 39–44\.External Links:[Document](https://dx.doi.org/10.2469/FAJ.V42.N4.39)Cited by:[Appendix B](https://arxiv.org/html/2607.10286#A2.p7.1),[§2\.3](https://arxiv.org/html/2607.10286#S2.SS3.p1.1),[§3\.1](https://arxiv.org/html/2607.10286#S3.SS1.p3.2)\.
- G\. P\. Brinson, B\. D\. Singer, and G\. L\. Beebower \(1991\)Determinants of portfolio performance ii: an update\.Financial Analysts Journal47\(3\),pp\. 40–48\.External Links:[Document](https://dx.doi.org/10.2469/FAJ.V47.N3.40)Cited by:[Appendix B](https://arxiv.org/html/2607.10286#A2.p7.1),[§2\.3](https://arxiv.org/html/2607.10286#S2.SS3.p1.1),[§3\.1](https://arxiv.org/html/2607.10286#S3.SS1.p3.2)\.
- F\. Cassano, A\. Gopinath, K\. Narasimhan, N\. Shinn, and S\. Yao \(2023\)Reflexion: language agents with verbal reinforcement learning, 2023\.Advances in Neural Information Processing Systems 368,pp\. 8634–8652\.External Links:[Document](https://dx.doi.org/10.52202/075280-0377)Cited by:[§2\.2](https://arxiv.org/html/2607.10286#S2.SS2.p1.1)\.
- L\. Chen, M\. Zaharia, and J\. Zou \(2023\)Frugalgpt: how to use large language models while reducing cost and improving performance\.Trans\. Mach\. Learn\. Res\.\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2305.05176)Cited by:[§2\.2](https://arxiv.org/html/2607.10286#S2.SS2.p1.1)\.
- L\. Chen, S\. Liu, J\. Yan, X\. Wang, H\. Liu, C\. Li, K\. Jiao, J\. Ying, Y\. Liu, Q\. Yang,et al\.\(2025a\)Advancing financial engineering with foundation models: progress, applications, and challenges\.Engineering\.External Links:[Document](https://dx.doi.org/10.1016/j.eng.2025.11.029)Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p1.1)\.
- Y\. Chen, Z\. Yao, Y\. Liu, J\. Ye, J\. Yu, L\. Hou, and J\. Li \(2025b\)StockBench: Can LLM Agents Trade Stocks Profitably In Real\-world Markets?\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2510.02209)Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p2.1),[§2\.3](https://arxiv.org/html/2607.10286#S2.SS3.p2.1)\.
- Y\. Chen, J\. Chen, C\. He, Y\. Li, Y\. Ji, Y\. Wu, D\. Yang, L\. Diao, L\. Shou, H\. Zhang,et al\.\(2026\)Token economics for llm agents: a dual\-view study from computing and economics\.arXiv preprint arXiv:2605\.09104\.Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p2.1),[§2\.2](https://arxiv.org/html/2607.10286#S2.SS2.p1.1)\.
- DeepSeek\-AI, A\. Liu, B\. Feng, B\. Xue, B\. Wang, B\. Wu, C\. Lu, C\. Zhao, C\. Deng, C\. Zhang,et al\.\(2024\)Deepseek\-v3 technical report\.arXiv\.org\.Cited by:[§5\.1](https://arxiv.org/html/2607.10286#S5.SS1.p3.1)\.
- Y\. Dong, F\. Wu, K\. Zhang, Y\. Dai, S\. Zhang, W\. Ye, S\. Chen, and Z\. Cheng \(2025\)Large language model agents in finance: a survey bridging research, practice, and real\-world deployment\.Conference on Empirical Methods in Natural Language Processing2025,pp\. 17889–17907\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.972)Cited by:[§D\.5](https://arxiv.org/html/2607.10286#A4.SS5.p1.1)\.
- G\. K\. Dow \(2018\)Transaction cost economics\.Handbook of industrial organization1,pp\. 135–182\.External Links:[Document](https://dx.doi.org/10.1017/9781316459423.015)Cited by:[Appendix C](https://arxiv.org/html/2607.10286#A3.p1.1),[§3\.1](https://arxiv.org/html/2607.10286#S3.SS1.p8.1)\.
- M\. H\. Erol, B\. El, M\. Suzgun, M\. Yuksekgonul, and J\. Zou \(2025\)Cost\-of\-pass: an economic framework for evaluating language models\.arXiv\.org\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2504.13359)Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p2.1),[§2\.2](https://arxiv.org/html/2607.10286#S2.SS2.p1.1)\.
- T\. Fan, Y\. Yang, Y\. Jiang, Y\. Zhang, Y\. Chen, and C\. Huang \(2025\)AI\-Trader: Benchmarking Autonomous Agents in Real\-Time Financial Markets\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2512.10971)Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p1.1),[§2\.1](https://arxiv.org/html/2607.10286#S2.SS1.p1.1),[§2\.3](https://arxiv.org/html/2607.10286#S2.SS3.p2.1),[§5\.1](https://arxiv.org/html/2607.10286#S5.SS1.p1.1)\.
- J\. Gatheral \(2010\)No\-dynamic\-arbitrage and market impact\.Quantitative finance10\(7\),pp\. 749–759\.External Links:[Document](https://dx.doi.org/10.1080/14697680903373692)Cited by:[4th item](https://arxiv.org/html/2607.10286#A3.I1.i4.p1.1)\.
- Google \(2025\)Cited by:[§5\.1](https://arxiv.org/html/2607.10286#S5.SS1.p3.1)\.
- R\. C\. Grinold and R\. N\. Kahn \(2000\)Active portfolio management\.The Journal of Financial Data Science\.External Links:[Document](https://dx.doi.org/10.3905/jfds.2021.1.071)Cited by:[§2\.3](https://arxiv.org/html/2607.10286#S2.SS3.p1.1)\.
- T\. Guo, H\. Shen, J\. Huang, Z\. Mao, J\. Luo, Z\. Chen, X\. Liu, B\. Xia, L\. Liu, Y\. Ma,et al\.\(2025\)MASS: Multi\-Agent Simulation Scaling for Portfolio Construction\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2505.10278)Cited by:[§2\.1](https://arxiv.org/html/2607.10286#S2.SS1.p1.1)\.
- S\. Han, J\. Zhang, Y\. Shen, K\. Yan, and H\. Li \(2025\)FinSphere: a real\-time stock analysis agent with instruction\-tuned large language models and domain\-specific tool integration\.Frontiers of Information Technology & Electronic Engineering26\(10\),pp\. 1822–1831\.External Links:[Document](https://dx.doi.org/10.1631/FITEE.2500414)Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p1.1)\.
- P\. He, Z\. Dai, B\. He, H\. Liu, X\. Tang, H\. Lu, J\. Li, J\. Ding, S\. Mukherjee, S\. Wang,et al\.\(2025\)TRAJECT\-bench: a trajectory\-aware benchmark for evaluating agentic tool use\.arXiv preprint arXiv:2510\.04550\.Cited by:[§2\.2](https://arxiv.org/html/2607.10286#S2.SS2.p1.1)\.
- J\. C\. Hsu, V\. Kalesnik, and B\. W\. Myers \(2010\)Performance attribution: measuring dynamic allocation skill\.Financial Analysts Journal66\(6\),pp\. 17–26\.External Links:[Document](https://dx.doi.org/10.2469/faj.v66.n6.3)Cited by:[§2\.3](https://arxiv.org/html/2607.10286#S2.SS3.p1.1)\.
- K\. Huang, Y\. Shi, D\. Ding, Y\. Li, Y\. Fei, L\. Lakshmanan, and X\. Xiao \(2025\)Thriftllm: on cost\-effective selection of large language models for classification queries\.Proceedings of the VLDB Endowment18\(11\),pp\. 4410–4423\.External Links:[Document](https://dx.doi.org/10.14778/3749646.3749702)Cited by:[§2\.2](https://arxiv.org/html/2607.10286#S2.SS2.p1.1)\.
- B\. Jin, T\. Collins, D\. Yu, M\. Cemri, S\. Zhang, M\. Li, J\. Tang, T\. Qin, Z\. Xu, J\. Lu,et al\.\(2025\)Controlling performance and budget of a centralized multi\-agent llm system with reinforcement learning\.arXiv\.org\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2511.02755)Cited by:[§6](https://arxiv.org/html/2607.10286#S6.p2.1)\.
- H\. Ko and J\. Lee \(2024\)Can chatgpt improve investment decisions? from a portfolio management perspective\.Finance Research Letters64,pp\. 105433\.External Links:[Document](https://dx.doi.org/10.1016/j.frl.2024.105433)Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p1.1),[§2\.3](https://arxiv.org/html/2607.10286#S2.SS3.p2.1)\.
- C\. Li, Y\. SHI, C\. Wang, Q\. Duan, R\. RUAN, W\. Huang, H\. Long, L\. Huang, N\. Tang, and Y\. Luo \(2025a\)Time travel is cheating: going live with deepfund for real\-time fund investment benchmarking\.InarXiv\.org,External Links:[Link](https://openreview.net/forum?id=SXADEhZ0sl),[Document](https://dx.doi.org/10.48550/arXiv.2505.11065)Cited by:[§2\.1](https://arxiv.org/html/2607.10286#S2.SS1.p1.1),[§5\.1](https://arxiv.org/html/2607.10286#S5.SS1.p3.1)\.
- H\. Li, Y\. Cao, Y\. Yu, S\. R\. Javaji, Z\. Deng, Y\. He, Y\. Jiang, Z\. Zhu, K\. Subbalakshmi, J\. Huang,et al\.\(2025b\)Investorbench: a benchmark for financial decision\-making tasks with llm\-based agent\.InAnnual Meeting of the Association for Computational Linguistics,pp\. 2509–2525\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.126)Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p2.1),[§2\.3](https://arxiv.org/html/2607.10286#S2.SS3.p2.1)\.
- X\. Li, Y\. Zeng, X\. Xing, J\. Xu, and X\. Xu \(2025c\)HedgeAgents: a balanced\-aware multi\-agent financial trading system\.The Web Conferenceabs/2502\.13165,pp\. 296–305\.External Links:[Document](https://dx.doi.org/10.1145/3701716.3715232)Cited by:[§2\.1](https://arxiv.org/html/2607.10286#S2.SS1.p1.1)\.
- Y\. Li and J\. Oliva \(2024\)Towards cost sensitive decision making\.International Conference on Artificial Intelligence and Statistics\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2410.03892)Cited by:[§6](https://arxiv.org/html/2607.10286#S6.p2.1)\.
- Y\. Li, B\. Luo, Q\. Wang, N\. Chen, X\. Liu, and B\. He \(2024\)Cryptotrade: A reflective llm\-based agent to guide zero\-shot cryptocurrency trading\.InConference on Empirical Methods in Natural Language Processing,pp\. 1094–1106\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-main.63)Cited by:[§2\.1](https://arxiv.org/html/2607.10286#S2.SS1.p1.1)\.
- J\. Lin, X\. Zeng, J\. Zhu, S\. Wang, J\. Shun, J\. Wu, and D\. Zhou \(2025\)Plan and budget: effective and efficient test\-time scaling on large language model reasoning\.arXiv\.org\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2505.16122)Cited by:[§2\.2](https://arxiv.org/html/2607.10286#S2.SS2.p1.1)\.
- A\. Liu, K\. Khandelwal, S\. Subramanian, V\. Jouault, A\. Rastogi, A\. Sad’e, A\. Jeffares, A\. Q\. Jiang, A\. Cahill, A\. Gavaudan,et al\.\(2026\)Ministral 3\.arXiv\.org\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2601.08584)Cited by:[§5\.1](https://arxiv.org/html/2607.10286#S5.SS1.p3.1)\.
- J\. Liu, C\. S\. Xia, Y\. Wang, and L\. Zhang \(2023a\)Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation\.InNeural Information Processing Systems,NIPS ’23,Red Hook, NY, USA,pp\. 21558–21572\.External Links:[Document](https://dx.doi.org/10.52202/075280-0943)Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p1.1)\.
- X\. Liu, G\. Wang, and D\. Zha \(2023b\)Fingpt: democratizing internet\-scale data for financial large language models\.arXiv\.org\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2307.10485)Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p1.1)\.
- Y\. Luo, Y\. Feng, J\. Xu, P\. Tasca, and Y\. Liu \(2025\)LLM\-Powered Multi\-Agent System for Automated Crypto Portfolio Management\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2501.00826)Cited by:[§2\.1](https://arxiv.org/html/2607.10286#S2.SS1.p1.1)\.
- Meta AI \(2025\)External Links:[Link](https://ai.meta.com/blog/llama-4-multimodal-intelligence/)Cited by:[§5\.1](https://arxiv.org/html/2607.10286#S5.SS1.p3.1)\.
- MiniMax \(2025\)External Links:[Link](https://www.minimax.io/news/minimax-m21)Cited by:[§5\.1](https://arxiv.org/html/2607.10286#S5.SS1.p3.1)\.
- Y\. Nevmyvaka, Y\. Feng, and M\. Kearns \(2006\)Reinforcement learning for optimized trade execution\.InInternational Conference on Machine Learning,pp\. 673–680\.External Links:[Document](https://dx.doi.org/10.1145/1143844.1143929)Cited by:[4th item](https://arxiv.org/html/2607.10286#A3.I1.i4.p1.1)\.
- OpenAI \(2025\)Cited by:[§5\.1](https://arxiv.org/html/2607.10286#S5.SS1.p3.1)\.
- T\. Ou, W\. Guo, A\. Gandhi, G\. Neubig, and X\. Yue \(2025\)AgentDiagnose: an open toolkit for diagnosing llm agent trajectories\.InConference on Empirical Methods in Natural Language Processing,pp\. 207–215\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-demos.15)Cited by:[§2\.2](https://arxiv.org/html/2607.10286#S2.SS2.p1.1)\.
- C\. Qian, Z\. Liu, S\. Kokane, A\. Prabhakar, J\. Qiu, H\. Chen, Z\. Liu, H\. Ji, W\. Yao, S\. Heinecke,et al\.\(2025\)XRouter: training cost\-aware llms orchestration system via reinforcement learning\.arXiv\.org\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2510.08439)Cited by:[§6](https://arxiv.org/html/2607.10286#S6.p2.1)\.
- W\. F\. Sharpe \(1994\)The sharpe ratio\.QFINANCE Calculation Toolkit3\(3\),pp\. 169–85\.External Links:[Document](https://dx.doi.org/10.3905/JPM.1994.409501)Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p2.1),[§2\.3](https://arxiv.org/html/2607.10286#S2.SS3.p1.1)\.
- A\. Suma and S\. Dauncey \(2025\)Deepseek\-r1: incentivizing reasoning capability in llms via reinforcement learning\.arXiv\.org\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2501.12948)Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p1.1)\.
- S\. Sun, M\. Qin, W\. Zhang, H\. Xia, C\. Zong, J\. Ying, Y\. Xie, L\. Zhao, X\. Wang, and B\. An \(2023\)Trademaster: a holistic quantitative trading platform empowered by reinforcement learning\.Neural Information Processing Systems36,pp\. 59047–59061\.External Links:[Document](https://dx.doi.org/10.52202/075280-2576)Cited by:[§2\.1](https://arxiv.org/html/2607.10286#S2.SS1.p1.1)\.
- J\. Wang, W\. Ding, and X\. Zhu \(2025a\)Financial analysis: intelligent financial data analysis system based on llm\-rag\.Applied and Computational Engineering145\(1\),pp\. 182–189\.External Links:[Document](https://dx.doi.org/10.54254/2755-2721/2025.22221)Cited by:[§2\.1](https://arxiv.org/html/2607.10286#S2.SS1.p1.1)\.
- J\. Wang, S\. Jain, D\. Zhang, B\. Ray, V\. Kumar, and B\. Athiwaratkun \(2024\)Reasoning in token economies: budget\-aware evaluation of llm reasoning strategies\.InConference on Empirical Methods in Natural Language Processing,pp\. 19916–19939\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2406.06461)Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p2.1),[§2\.2](https://arxiv.org/html/2607.10286#S2.SS2.p1.1)\.
- N\. Wang, X\. Hu, P\. Liu, H\. Zhu, Y\. Hou, H\. Huang, S\. Zhang, J\. Yang, J\. Liu, G\. Zhang,et al\.\(2025b\)Efficient Agents: Building Effective Agents While Reducing Cost\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2508.02694)Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p2.1),[§2\.2](https://arxiv.org/html/2607.10286#S2.SS2.p1.1)\.
- R\. Wang, H\. Wang, B\. Xue, J\. Pang, S\. Liu, Y\. Chen, J\. Qiu, D\. F\. Wong, H\. Ji, and K\. Wong \(2025c\)Harnessing the Reasoning Economy: A Survey of Efficient Reasoning for Large Language Models\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2503.24377)Cited by:[§2\.2](https://arxiv.org/html/2607.10286#S2.SS2.p1.1)\.
- S\. Wu, O\. Irsoy, S\. Lu, V\. Dabravolski, M\. Dredze, S\. Gehrmann, P\. Kambadur, D\. Rosenberg, and G\. Mann \(2023\)Bloomberggpt: a large language model for finance\.arXiv\.org\.Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p1.1)\.
- Y\. Xiao, E\. Sun, D\. Luo, and W\. Wang \(2024\)TradingAgents: Multi\-Agents LLM Financial Trading Framework\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2412.20138)Cited by:[§2\.1](https://arxiv.org/html/2607.10286#S2.SS1.p1.1)\.
- Q\. Xie, W\. Han, X\. Zhang, Y\. Lai, M\. Peng, A\. Lopez\-Lira, and J\. Huang \(2023\)Pixiu: a comprehensive benchmark, instruction dataset and large language model for finance\.Neural Information Processing Systems36,pp\. 33469–33484\.External Links:[Document](https://dx.doi.org/10.52202/075280-1454)Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p2.1)\.
- F\. Xing \(2024\)Designing heterogeneous llm agents for financial sentiment analysis\.ACM Transactions on Management Information Systems16\(1\),pp\. 1–24\.External Links:[Document](https://dx.doi.org/10.1145/3688399)Cited by:[§2\.1](https://arxiv.org/html/2607.10286#S2.SS1.p1.1)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv,et al\.\(2025a\)Qwen3 technical report\.arXiv preprint arXiv:2505\.09388\.Cited by:[§5\.1](https://arxiv.org/html/2607.10286#S5.SS1.p3.1)\.
- H\. Yang, B\. Zhang, N\. Wang, C\. Guo, X\. Zhang, L\. Lin, J\. Wang, T\. Zhou, M\. Guan, R\. Zhang,et al\.\(2024\)FinRobot: an open\-source ai agent platform for financial applications using large language models\.arXiv\.org\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2405.14767)Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p1.1)\.
- L\. Yang, J\. Luo, X\. Liu, Y\. Lou, and Z\. Chen \(2025b\)BAMAS: structuring budget\-aware multi\-agent systems\.AAAI Conference on Artificial Intelligence\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2511.21572)Cited by:[§2\.2](https://arxiv.org/html/2607.10286#S2.SS2.p1.1)\.
- L\. Yu, W\. Jiang, H\. Shi, J\. Yu, Z\. Liu, Y\. Zhang, J\. T\. Kwok, Z\. Li, A\. Weller, and W\. Liu \(2023a\)Metamath: bootstrap your own mathematical questions for large language models\.International Conference on Learning Representations\.Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p1.1)\.
- Y\. Yu, H\. Li, Z\. Chen, Y\. Jiang, Y\. Li, J\. W\. Suchow, D\. Zhang, and K\. Khashanah \(2023b\)Finmem: a performance\-enhanced llm trading agent with layered memory and character design\.IEEE Transactions on Big Data\.External Links:[Document](https://dx.doi.org/10.1109/TBDATA.2025.3593370)Cited by:[§2\.1](https://arxiv.org/html/2607.10286#S2.SS1.p1.1)\.
- Y\. Yu, Z\. Yao, H\. Li, Z\. Deng, Y\. Cao, Z\. Chen, J\. W\. Suchow, R\. Liu, Z\. Cui, D\. Zhang,et al\.\(2024\)Fincon: a synthesized llm multi\-agent system with conceptual verbal reinforcement for enhanced financial decision making\.Neural Information Processing Systems37,pp\. 137010–137045\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2407.06567)Cited by:[§2\.1](https://arxiv.org/html/2607.10286#S2.SS1.p1.1)\.
- M\. J\. Zellinger and M\. Thomson \(2025\)Economic evaluation of llms\.arXiv\.org\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2507.03834)Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p2.1),[§2\.2](https://arxiv.org/html/2607.10286#S2.SS2.p1.1)\.
- G\. T\. A\. Zeng, X\. Lv, Q\. Zheng, Z\. Hou, B\. Chen, C\. Xie, C\. Wang, D\. Yin, H\. Zeng, J\. Zhang,et al\.\(2025\)GLM\-4\.5: agentic, reasoning, and coding \(arc\) foundation models\.External Links:2508\.06471,[Link](https://arxiv.org/abs/2508.06471),[Document](https://dx.doi.org/10.48550/arXiv.2508.06471)Cited by:[§5\.1](https://arxiv.org/html/2607.10286#S5.SS1.p3.1)\.
- G\. Zhang, Y\. Yue, Z\. Li, S\. Yun, G\. Wan, K\. Wang, D\. Cheng, J\. X\. Yu, and T\. Chen \(2024\)Cut the crap: an economical communication pipeline for llm\-based multi\-agent systems\.International Conference on Learning Representations\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2410.02506)Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p2.1),[§2\.2](https://arxiv.org/html/2607.10286#S2.SS2.p1.1)\.
- W\. Zhang, Y\. Zhao, C\. Zong, X\. Wang, and B\. An \(2025\)FinWorld: an all\-in\-one open\-source platform for end\-to\-end financial ai research and deployment\.arXiv\.org\.External Links:[Document](https://dx.doi.org/10.1145/3770854.3785689)Cited by:[§2\.1](https://arxiv.org/html/2607.10286#S2.SS1.p1.1)\.
- T\. Zhao, J\. Lyu, S\. Jones, H\. Garber, S\. Pasquali, and D\. Mehta \(2025\)AlphaAgents: Large Language Model based Multi\-Agents for Equity Portfolio Constructions\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2508.11152)Cited by:[§1](https://arxiv.org/html/2607.10286#S1.p1.1),[§2\.1](https://arxiv.org/html/2607.10286#S2.SS1.p1.1)\.
## Appendix ADiagnosis Report Format
In this part, we provide two formats to show whatTradeLensoutputs in practice; a complete output example is provided in our open source repository\. First, we present an example of a financial report, which reports the exact profit and loss and other trading performance metrics of a trial run\. Based on a detailed financial analysis report,TradeLensgenerates the diagnosis report with ranked suggestions from multiple dimensions, including capital budget, backbone model choice, trading frequency, and architectural intelligence such as modules, agents, and tools\.
Financial Report Format1\.Trading Configuration•Trading Period:•Trading Model:•Assets:2\.Asset & Portfolio State•Cash condition:•Positions:3\.Performance•Profit Attribution and figure•Cost Breakdown and figure•Financial Statement4\.Execution Quality•Opportunity Cost and latency•Execution layer
Diagnosis Report FormatExecutive SummaryOverview of the overall report and top suggestions1\.Trading Performance Analysis•Outcome Snapshot\.•Profit Attribution Analysis\.•Cost Structure Analysis\.•Portfolio Timeline and Risk Signals\.•Execution Quality and LLM Efficiency\.2\.Trading System Diagnosis•System Diagnosis\. Define core restrictive factors and auxiliary limiting conditions of strategy performance; assess the impact of fee expenditure on actual revenue\.•Root Causes\. Summarize fundamental issues behind unsatisfactory performance based on analytical data\.•Top Recommended Actions\. Provide recommendations targeting model parameters, capital allocation, and system framework; predict improvement effects and outline specific implementation approaches\.•Quick Wins vs\. Structural Changes\. Short\-term adjustable optimization measures vs\. long\-term systematic reconstruction and upgrade plans\.
## Appendix BProfit Attribution Details
As presented in Table[3](https://arxiv.org/html/2607.10286#A2.T3), we attribute cumulative gross profit to three portfolio\-level components\. The decomposition is defined by two counterfactual baselines constructed on the same asset universe and evaluation window\.
Table 3:Profit attribution categories for agentic trading systems\.Notation\.Lett∈\{1,…,T\}t\\in\\\{1,\\ldots,T\\\}index decision periods,NNthe number of tradable assets,wi\(t\)w\_\{i\}^\{\(t\)\}the portfolio weight of assetiiat periodttwith∑iwi\(t\)=1\\sum\_\{i\}w\_\{i\}^\{\(t\)\}=1, andri\(t\)r\_\{i\}^\{\(t\)\}its realized return\. LetV0V\_\{0\}denote the initial portfolio value\.
Realized portfolio value\.The dynamic portfolio value under the agent’s actual weight path is
VTdyn=V0∏t=1T\(1\+∑i=1Nwi\(t\)ri\(t\)\)\.V\_\{T\}^\{\\mathrm\{dyn\}\}=V\_\{0\}\\prod\_\{t=1\}^\{T\}\\left\(1\+\\sum\_\{i=1\}^\{N\}w\_\{i\}^\{\(t\)\}r\_\{i\}^\{\(t\)\}\\right\)\.
Systematic exposure baseline\.Letrtmktr\_\{t\}^\{\\mathrm\{mkt\}\}denote the realized market return in periodtt, reflecting aggregate market movement in the investable universe\. InTradeLens,rtmktr\_\{t\}^\{\\mathrm\{mkt\}\}is computed from the benchmark market prices supplied in the audit inputs\. The counterfactual value under passive market exposure is
VTsys=V0∏t=1T\(1\+rtmkt\)\.V\_\{T\}^\{\\mathrm\{sys\}\}=V\_\{0\}\\prod\_\{t=1\}^\{T\}\\left\(1\+r\_\{t\}^\{\\mathrm\{mkt\}\}\\right\)\.The attributed profit isPsys=VTsys−V0P^\{\\mathrm\{sys\}\}=V\_\{T\}^\{\\mathrm\{sys\}\}\-V\_\{0\},i\.e\.,the portion of gross profit that would have been obtained by participating in market movement alone\.
Static allocation baseline\.Holding the initial allocation𝐰\(0\)\\mathbf\{w\}^\{\(0\)\}fixed yields
VTbase=V0∏t=1T\(1\+∑i=1Nwi\(0\)ri\(t\)\)\.V\_\{T\}^\{\\mathrm\{base\}\}=V\_\{0\}\\prod\_\{t=1\}^\{T\}\\left\(1\+\\sum\_\{i=1\}^\{N\}w\_\{i\}^\{\(0\)\}r\_\{i\}^\{\(t\)\}\\right\)\.The attributed profit isPasset=VTbase−VTsysP^\{\\mathrm\{asset\}\}=V\_\{T\}^\{\\mathrm\{base\}\}\-V\_\{T\}^\{\\mathrm\{sys\}\}\.
Timing component\.The residual dynamic component is
Ptiming=VTdyn−VTbase\.P^\{\\mathrm\{timing\}\}=V\_\{T\}^\{\\mathrm\{dyn\}\}\-V\_\{T\}^\{\\mathrm\{base\}\}\.By construction, cumulative gross profit satisfies the identity
P1:T=VTdyn−V0=Psys\+Passet\+Ptiming\.P\_\{1:T\}=V\_\{T\}^\{\\mathrm\{dyn\}\}\-V\_\{0\}=P^\{\\mathrm\{sys\}\}\+P^\{\\mathrm\{asset\}\}\+P^\{\\mathrm\{timing\}\}\.
Relation to classical attribution\.Classical Brinson attribution compares portfolio and benchmark weights at the security levelBrinsonet al\.\([1986](https://arxiv.org/html/2607.10286#bib.bib69),[1991](https://arxiv.org/html/2607.10286#bib.bib70)\)\. Our formulation operates at the portfolio\-value level and uses two nested counterfactual paths rather than a single external benchmark\. This design is tailored to agentic trading systems, where decisions are naturally described as market exposure, initial allocation, and subsequent reallocation, and where the evaluation goal is deployability diagnosing rather than performance ranking\.
## Appendix CCost Taxonomy
Following Williamson’s transaction cost economicsDow \([2018](https://arxiv.org/html/2607.10286#bib.bib46)\), markets are not frictionless: every exchange requires resources for information processing, decision\-making, and execution, beyond the nominal price movement itself\. Such transaction costs arise from bounded rationality and the need to coordinate actions under uncertainty, and they directly shape whether a strategy is viable once implemented\. This perspective is especially relevant for agentic trading systems, where “intelligence” is produced via paid inference and deployed through fee and impact sensitive execution pipelines\. Motivated by this view, we construct an explicit cost model that accounts for end\-to\-end frictions across decision\-making, trading execution, and infrastructure\.
Lett∈\{1,…,T\}t\\in\\\{1,\\dots,T\\\}index decision periods within a time windowTT\(e\.g\., minutes, hours, or days\)\. We decompose the total cost incurred in periodttinto LLM\-side, trading\-side, infra\-side, and stochastic components:
Ct=Ctllm\+Cttrd\+Ctinf\+CtstoC\_\{t\}\\;=\\;C\_\{t\}^\{\\text\{llm\}\}\\;\+\\;C\_\{t\}^\{\\text\{trd\}\}\\;\+\\;C\_\{t\}^\{\\text\{inf\}\}\\;\+\\;C\_\{t\}^\{\\text\{sto\}\}\(11\)
Specifically, the components are computed as follows\. The trading\-side costCttrdC\_\{t\}^\{\\text\{trd\}\}combines fixed charges and activity\-dependent expenses,
Cttrd=Cstattrd\+∑i=1Mtrdγt,itrd⋅activityt,itrd\.C\_\{t\}^\{\\text\{trd\}\}=C\_\{\\text\{stat\}\}^\{\\text\{trd\}\}\+\\sum\_\{i=1\}^\{M\_\{\\text\{trd\}\}\}\\gamma\_\{t,i\}^\{\\text\{trd\}\}\\cdot\\text\{activity\}\_\{t,i\}^\{\\text\{trd\}\}\.
The LLM\-side costCtllmC\_\{t\}^\{\\text\{llm\}\}is determined by token usage and adjusted by the success rate,
ctllm=∑s∈𝒮ps⋅xrsρ,𝒮=\{in,out,cache\},c\_\{t\}^\{\\text\{llm\}\}=\\frac\{\\sum\_\{s\\in\\mathcal\{S\}\}p^\{s\}\\cdot x\_\{r\}^\{s\}\}\{\\rho\},\\qquad\\mathcal\{S\}=\\\{\\text\{in\},\\,\\text\{out\},\\,\\text\{cache\}\\\},whereρ\\rhoaccounts for successful completion rate due to retries\.
The infra\-side cost is modeled as a constant per period,
Ctinf=κ,C\_\{t\}^\{\\text\{inf\}\}=\\kappa,and the stochastic cost captures unanticipated expenses,
Ctsto∼Uniform\(0,n\)\.C\_\{t\}^\{\\text\{sto\}\}\\sim\\text\{Uniform\}\(0,n\)\.
Given the cost taxonomy defined above,TradeLensleverages a*system\-agnostic*cost accounting module that acts as a lightweight middleware between an agent’s execution pipeline and the downstream accounting logic\. Instead of assuming a specific prompting strategy or agent architecture, the taxonomy defines a small set of standardized events \(with timestamps and identifiers\) so that agentic workflows can all be logged and attributed consistently to a decision period and step\. The dynamic portion of each category \(e\.g\.,token\-dependent inference fees, per\-trade frictions such as commissions or slippage\) is traced during execution through event logging\. In contrast, the static portion \(e\.g\.,data subscription or deployment baseline fees\) is configured by end users as part of the evaluation setting, enablingTradeLensto reflect different billing plans and deployment assumptions without modifying the agent\. Concretely, it considers the following aspects:
- •LLM\-side accountingCtllmC\_\{t\}^\{\\text\{llm\}\}: converts traced model usage \(prompt/completion tokens, call counts, retries\) into monetary inference cost under provider pricing \(or user\-specified prices\)\.
- •Trading\-side accountingCttrdC\_\{t\}^\{\\text\{trd\}\}: aggregates execution\-layer costs from orders and fills, including explicit fees \(commissions, exchange/clearing fees, taxes\) and modeled frictions \(spread/slippage\) when measurable\.
- •Infra\-side accountingCtinfC\_\{t\}^\{\\text\{inf\}\}: aggregates deployment costs such as elastic compute/bandwidth/I/O and monitoring/logging overhead; static baselines \(e\.g\.,hosting plans\) are incorporated from user configuration, while dynamic usage is traced during execution\.
- •Stochastic termCtstoC\_\{t\}^\{\\text\{sto\}\}: represents unanticipated or hard\-to\-measure expenses as an additive uncertainty term; it can be set to zero \(best\-case\) or calibrated from real billing data \(conservative\-case\) as discussed in previous workNevmyvakaet al\.\([2006](https://arxiv.org/html/2607.10286#bib.bib36)\); Gatheral \([2010](https://arxiv.org/html/2607.10286#bib.bib12)\)\.
As presented in Table[4](https://arxiv.org/html/2607.10286#A3.T4), we categorize the cost items to four types according to \(i\) where the cost arises, and \(ii\) whether the cost is static \(fixed for a period or billing cycle\) or dynamic\.
Table 4:Cost categories for end\-to\-end accounting\.
## Appendix DExperiment Detail
Across all experiments, we gathered data through an agentic trading system and recorded complete decision trajectories, including portfolio states and executed trades\. We then applyTradeLensto perform profit–cost accounting and diagnostic analysis, allowing us to explain not only whether an agent is economically viable, but also why it succeeds or fails\.
### D\.1Detailed experimental setup
Cost assumptions\.Specifically, we account for fourcost components: \(1\) LLM\-side: LLM token cost is computed from each model’s usage and pricing from the official provider endpoints, and is adjusted by the task success rate \(whereρ\\rho= 98% as described in LLM provider metrics\) to reflect the effective cost per successful completion\. \(2\) Trading\-side: Commission fee is charged per execution, consistent with typical retail brokerage fees111[https://www\.interactivebrokers\.com/en/pricing/commissions\-stocks\.php](https://www.interactivebrokers.com/en/pricing/commissions-stocks.php)\. Thus, we calculate the relevant cost based on that documentation\. \(3\) Infra\-side: Variable infrastructure cost is set to $0\.20 per run, with a fixed monthly data subscription fee \(e\.g\.,$100/month222[https://www\.alphavantage\.co/premium/](https://www.alphavantage.co/premium/)\)\. \(4\) Stochastic: A stochastic cost term\(e\.g\.,max\. $0\.5 per run\) is added\.
Profit attribution baselines\.To complement the cost\-side accounting, we construct twoprofit\-side baselinesfor attribution analysis\. The first is a systematic exposure baseline, which uses the return of the S&P 500 index over the same market window to capture broad market movement\. The second is a static allocation baseline, which follows a buy\-and\-hold strategy by keeping the initial portfolio allocation unchanged throughout the trading period\.
Trading setup\. The trading window is set to a two\-month sideways market period from December 1, 2025, to January 30, 2026, starting with an initial cash of $100,000 and trading a fixed universe of the most liquid 10 liquid U\.S\. equities\. We accessed market data via licensed provider APIs/subscriptions and do not redistribute raw provider data\.
### D\.2Viability across Capital scale
Figure[5](https://arxiv.org/html/2607.10286#A4.T5)shows that increasing capital scale does not lead to a uniform improvement in economic viability\. Although larger capital may dilute fixed operating costs, the main variation in viability comes from profit\-side components, especially asset selection and timing effects\.
For GPT\-5\.2, system viability remains positive across all tested capital scales, although it does not increase monotonically with capital size\. Net profit rises from 75\.02 at 10k to 1041\.22 at 50k, decreases to 246\.97 at 100k, and then increases again to 1780\.03 at 500k\. However, agentic viability exhibits a different pattern\. It is negative at 10k and 50k, becomes slightly positive at 100k, and then drops sharply to \-5697\.07 at 500k\. Table[5](https://arxiv.org/html/2607.10286#A4.T5)suggests that this deterioration at the largest tested scale is mainly driven by a strongly negative timing effect rather than by operating cost\. Although GPT\-5\.2 remains system viable at 500k, its dynamic intervention fails to generate positive active profit after accounting for decision\-induced costs\.
Diagnosis sample: GPT\-5\.2 at 500kTop Recommended Actions:Reduce execution latency and slippage \(System/agent/tool architecture\)\.\- Expected impact: rapidly recapture a large portion of the 3\.98k opportunity cost and reduce timing erosion; could convert current fragile net profit into a robust positive\.\- Implementation idea: move to a trimmed inference pipeline \(pruned prompt/context, async batching, prioritized decisions\), colocate order gateway or use faster broker API paths, and implement pre‑validated order templates to eliminate per‑trade orchestration delays\. Measure by reducing the average latency per trade below 5s and the slippage by half\.
DeepSeek\-V3\.2 shows a consistently weaker pattern\. Both system viability and agentic viability remain negative across all capital scales\. The losses are moderate at 10k–100k but expand substantially at 500k, where system viability decreases to \-5679\.66, and agentic viability decreases to \-15873\.79\. Table[5](https://arxiv.org/html/2607.10286#A4.T5)shows that this deterioration is not primarily caused by system cost, which remains small relative to profit variation\. Instead, the largest losses come from unfavorable profit components, especially the strongly negative timing effect at 500k\. The diagnosis, therefore, points to strategy\-level failure rather than fee burden: DeepSeek\-V3\.2’s market judgment and reallocation decisions introduce substantial capital losses as scale increases\.
Diagnosis sample: DeepSeek\-V3\.2 at 500k\- Token intensity is large \(239k input tokens/day\) — this implies expensive and heavy LLM usage and long prompt/context windows, increasing compute latency and opportunity cost\.\- Distinction: selection skill is present \(asset picks positive\), but execution/system design \(latency, LLM workflow\) is producing missed fills and timing slippage\. Market exposure \(holding fast\-rising assets\) helped, but the system could not capture that upside because of when/how it traded\.
Overall, these results suggest that capital scaling mainly acts as an amplifier of backbone\-specific trading behavior\. For GPT\-5\.2, larger capital can amplify profitable timing decisions and offset losses that would have occurred under the initial buy\-and\-hold allocation\. For DeepSeek\-V3\.2, larger capital instead magnifies strategy\-level losses caused by poor asset selection and ineffective timing\. Therefore, larger capital can dilute fixed costs, but cost dilution alone is insufficient to ensure economic viability; the effect of scaling depends on whether the agentic strategy can generate positive active profit after costs\.
Table 5:Viability across Capital scale, ranging from 10k USD to 500k USD, comparing DeepSeek\-V3\.2 and GPT\-5\.2\. All values are rounded to two decimal places from raw data\.
### D\.3Viability across trading frequency
Figure[5](https://arxiv.org/html/2607.10286#S5.F5)shows that higher decision frequency does not improve viability\. Compared with daily trading, hourly trading reduces gross profit by 1104\.09 for DeepSeek\-V3\.2 and 334\.99 for GPT\-5\.2, while the additional total cost accounts for only 11\.8% and 13\.5% of the corresponding net\-profit deterioration\. This indicates that the main failure mode is not fee burden, but degradation in trading decisions\.
The attribution results \(Table[6](https://arxiv.org/html/2607.10286#A4.T6)\) further reveal that the degradation is concentrated in active components\. For both backbones, hourly trading produces more negative timing effects than daily trading, suggesting that more frequent decision points introduce additional timing errors rather than stable short\-term opportunities\. For DeepSeek\-V3\.2, the failure is more severe: hourly trading worsens both asset selection and timing effects, indicating that frequent reallocation affects not only when the agent trades but also what it holds\. Beyond the aggregate profit–cost accounting, we applyTradeLensto explain why higher decision frequency fails to improve viability\. The diagnosis shows that the hourly setting does not merely increase operating cost; more importantly, it changes the quality of active decisions\. Thus, the hourly setting exposes a strategy\-level problem: when the predictive signal is weak, increasing decision frequency amplifies noise, turnover, and mistimed reallocations instead of improving economic viability\.
Diagnosis sample: DeepSeek\-V3\.2 under hourly tradingTrading cadence is likely too aggressive relative to predictive signal quality: 190 trades/month across 8 assets combined with negative timing indicates frequency is not justified by signal edge\. Portfolio construction also contributed: initial holdings and sizing choices \(see initial holdings versus final positions\) show nontrivial exposures that did not capture market moves\.
Table 6:Viability across trading frequency\.
### D\.4Viability across system architecture
Chain\-of\-thought baseline\.The CoT baseline uses the same market data and portfolio context as the agentic system, but removes tool interaction and multi\-step external execution\. For each trading date, the system first retrieves the current portfolio, daily open prices, previous open and close prices, and recent news for the watchlist\. It then constructs a single prompt that asks the LLM to analyze all tickers, provide a short portfolio\-level synthesis, and output a final trading decision in a structured JSON format\. Only the final JSON block is used for trade execution, while the preceding step\-by\-step analysis is logged as the model’s reasoning trace\.
Table[2](https://arxiv.org/html/2607.10286#S5.T2)compares economic viability across different system architectures under the same initial capital of $100,000\. The results show that the economic value of an architecture does not depend simply on whether it is cheaper or more complex, but on whether additional reasoning and coordination can be converted into better asset selection and timing decisions\.
CoT has the lowest total cost for both backbones, but it does not achieve the best economic outcome\. This suggests that low operating cost alone is insufficient for viability\. When the reasoning structure fails to support effective investment decisions, lower cost may still be accompanied by poor asset selection and ineffective reallocation\. This pattern is especially clear for DeepSeek\-V3\.2, where CoT produces a strongly negative timing effect and the worst agentic profit\.
In contrast, DeepFund shows more stable gains across backbones\. For DeepSeek\-V3\.2, it is the only architecture that achieves both positive net profit and positive agentic profit\. For GPT\-5\.2, it also obtains the highest net profit and agentic profit\. Importantly, DeepFund does not have the lowest total cost\. Its advantage, therefore, comes from the profit side rather than from cost reduction: the additional architectural overhead is offset by better trading decisions\.
The attribution results further show that DeepFund’s advantage is mainly associated with timing\. It produces positive timing effects for both backbones, indicating that the architecture does not merely benefit from market exposure but creates active value through dynamic reallocation\. For retail investors developing agentic trading systems, this suggests that the key challenge is not only to reduce inference cost or add more reasoning steps, but to monitor complete trading trajectories and verify whether reasoning cost is actually converted into better investment outcomes\.
Diagnosis sample: DeepSeek\-V3\.2 under CoTDynamic cost accumulation is tiny relative to portfolio swings, so cumulative cost lines are almost flat while portfolio value moves dominate\. The run exhibits path dependency driven by timing decisions: the timing effect \(\-3,214\.75\) shows losses are concentrated in when trades were executed, not in cost build\-up\.\- Fragility of profitability:\- Given small cost base and positive selection when capital is invested, profitability is fragile but fixable — i\.e\., the strategy can produce positive selection returns if timing is corrected or exposure windows are altered\.\- However, until timing decisions improve, increasing capital or leaving the static subscription in place risks repeated net losses\.
### D\.5Sensitivity in Market Regime
Table 7:Viability across market conditions\. Cumulative profit and cost breakdown with an initial cash of $100,000\.Market regime is typically discussed for its impact on returnsAng and Timmermann \([2012](https://arxiv.org/html/2607.10286#bib.bib2)\); Donget al\.\([2025](https://arxiv.org/html/2607.10286#bib.bib8)\)\. For agentic trading systems, regime shifts may also affect viability by changing both profit opportunities and deployment frictions\. For example, liquidity and volatility can alter execution costs such as spread and slippage, while different market conditions may induce different trading intensity and runtime workload\. Thus, before attributing net\-profit changes to model or system design alone, we test whether profit–cost viability is sensitive to market regimes\.
To isolate temporal and regime effects, we conduct additional experiments under bearish \(2025\-02\-10 to 2025\-04\-07\), sideways \(2025\-12\-01 to 2026\-01\-30\), and bullish \(2025\-06\-02 to 2025\-08\-08\) conditions\. Table[7](https://arxiv.org/html/2607.10286#A4.T7)shows that profit varies much more sharply than cost across regimes\. For both DeepSeek\-V3\.2 and GPT\-5\.2, total costs remain within a relatively narrow range, whereas gross profit, net profit, and agentic profit change substantially with market conditions\.
In the bearish regime, both models are system\-non\-viable, but the failure mechanisms differ\. DeepSeek\-V3\.2 suffers from a large negative timing effect, indicating that dynamic reallocation amplifies downside exposure\. GPT\-5\.2, by contrast, is mainly hurt by negative market exposure and asset selection, while its timing loss is relatively small\. In the sideways regime, GPT\-5\.2 achieves modest positive net profit and agentic profit, whereas DeepSeek\-V3\.2 remains negative, suggesting that the same market condition can still expose backbone\-specific differences in timing and execution quality\. In the bullish regime, both models become strongly profitable, but through different profit sources: DeepSeek\-V3\.2 is dominated by timing gains, while GPT\-5\.2 benefits from both asset selection and timing\.
Overall, the market regime mainly determines whether LLM\-mediated decisions can be converted into profit\. This result reinforces the need for profit–cost diagnosis: similar cost levels can correspond to very different viability outcomes because market exposure, asset selection, and timing respond differently across regimes\.
## Appendix EDeployment User Study
This part evaluates deployability from both technical and user\-facing perspectives: whetherTradeLenscan provide forward\-looking deployment diagnostics, and whether practitioners find such diagnostics useful for interpreting and revising agentic trading systems\.
### E\.1Forward diagnostic analysis
Figure 6:Forward diagnostic analysis of trading frequency\. Cumulative net profit of DeepSeek\-V3\.2 under hourly \(dashed line\) and daily \(solid line\) trading frequencies\. The vertical gray line marks the transition date\.Diagnostic finding in December 2025\.Under the hourly configuration, the testing system starts with $100,000 and ends in December with a gross loss of approximately $1,500, while incurring about $300 in total system cost\. Fixed data subscription fees account for roughly $100, and the main dynamic costs come from trading\-side and LLM\-side usage\. The diagnostic report further identifies excessive token usage per decision and concentrated capital utilization\. Based on these accounting signals,TradeLensflags trading frequency as a major cost driver and recommends reducing decision frequency as a high\-priority adjustment\.
Forward evaluation in January 2026\.We then evaluate the frequency\-reduction suggestion in a later market window by comparing the daily configuration with the hourly baseline\. In January, the daily configuration reduces total system cost from approximately $610 under the hourly baseline to about $124, and achieves higher net profit in early January \(Figure[6](https://arxiv.org/html/2607.10286#A5.F6)\)\. However, net profit remains negative under both configurations\. A reversal occurs on January 28, when the daily configuration abstains from trading under unfavorable market conditions, whereas the hourly configuration continues active trading and partially mitigates losses\. By the end of January, the hourly baseline records a smaller net loss than the reduced\-frequency configuration\.
These results show that the December diagnosis correctly identifies trading frequency as an important cost driver, but cost reduction alone does not guarantee later\-window viability\. When gross returns are constrained by market conditions and strategy\-level edge, lowering decision frequency can reduce cost while still failing to improve final net profit\. Thus, the forward analysis supports the role ofTradeLensas a diagnostic tool: it explains why a system fails to become viable, rather than assuming that a cost\-saving suggestion will always improve performance\.
### E\.2Practitioner Feedback
Table 8:Participant demographics for the deployment\-oriented usability assessment\.Participants\.To complement the forward diagnostic analysis, we collected practitioner feedback on whetherTradeLenshelps users interpret the deployability of private agentic trading systems\. We recruited 13 retail traders who usedTradeLensin their own agentic trading workflows while keeping their strategies, asset universes, trading logs, and portfolio details private\. Participants were recruited through social networks, and a startup company specializing in financial AI agreed to support\. Employees of the company voluntarily participated in the experiment and were allowed to withdraw at any time\. To preserve anonymity, we do not disclose the company name\. In this collaboration, we collected only participants’ relevant professional experience and feedback, and did not collect any personally identifiable information\. Participant demographics are summarized in Table[8](https://arxiv.org/html/2607.10286#A5.T8)\.
Procedure\.To ensure valid participation, we explained the purpose of data collection and how the collected data would be used, and obtained informed consent from all participants\. Participants then received a brief usage guide explaining the required inputs, profit–cost metrics, and report structure\. They then usedTradeLensto audit profit sources, cost drivers, net profit, system viability, and agentic viability for their own runs\. After using the toolkit, they completed a 5\-point Likert questionnaire and provided open\-ended comments\. The questionnaire covered three dimensions: \(i\)*profit–cost diagnosis*, including whether the report helped users understand viability, distinguish market exposure, asset selection, and timing, identify dominant cost drivers, and separate weak decision value from high cost; \(ii\)*diagnostic actionability*, including whether the report supported revision decisions such as model choice, trading frequency, architecture, or execution adjustment; and \(iii\)*usability and integration*, including whether the outputs were understandable and whether the integration effort was acceptable\.
Main feedback\.Overall, participants foundTradeLensuseful for diagnosing why a private trading agent is or is not deployable, instead of only reporting its final profit\. They emphasized that profit and cost should be interpreted jointly, since poor net performance may result from different mechanisms, including weak asset selection, mistimed reallocation, excessive trading friction, high LLM usage, or fixed data and infrastructure fees\. Participants also found the distinction between system viability and agentic viability helpful, because it separates systems that merely benefit from market exposure from those whose dynamic decisions add incremental value\.
From a developer perspective, participants highlighted three practical benefits\. First, the reports made debugging more traceable by linking failure modes to trading records, runtime traces, and cost components\. Second, the configuration\-level recommendations were actionable, such as reducing trading frequency, limiting redundant model calls, or improving execution design\. Third, the reports helped prioritize whether a failure should be addressed through operational tuning or deeper architectural revision\.
Participants also noted several limitations\. They requested stronger strategy\-level guidance after failure modes are identified, especially for asset selection, timing adjustment, and position control\. Some participants suggested making benchmark choice, execution\-cost assumptions, and fixed infrastructure fees more explicit, so that the diagnosis can be better adapted to different deployment settings\. These results support the role ofTradeLensas a practical diagnostic layer for evaluating whether LLM\-mediated decisions convert induced costs into deployable trading value under realistic private deployment constraints\.
Ethics statement\.This feedback study was reviewed and approved by the institutional ethics committee of our university\. All participants were informed of the purpose of the study before providing feedback\. Participation was voluntary, and participants could withdraw at any time\. The collected feedback was used only for research purposes, and all demographic information was anonymized and reported in aggregate form\.Similar Articles
Agentic Trading: When LLM Agents Meet Financial Markets
This paper presents a systematic survey and evidence map of 77 studies on LLM-based trading agents, finding that architectural experimentation is expanding rapidly but evaluation protocols, execution semantics, and reproducibility remain critical bottlenecks.
TradingAgents: Multi-Agents LLM Financial Trading Framework
This paper introduces TradingAgents, a multi-agent LLM framework that simulates real-world trading firms to improve stock trading performance. It utilizes specialized agents for analysis and risk management, demonstrating superior results in cumulative returns and Sharpe ratio compared to baselines.
The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?
This study evaluates whether test-time reasoning in large language models (like DeepSeek, GPT, and Gemini) improves net portfolio returns in trading, finding that additional reasoning does not reliably enhance economic outcomes across various conditions.
AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets
This paper introduces AI-Trader, the first fully automated live benchmark for evaluating LLMs in financial decision-making across US stocks, A-shares, and cryptocurrencies. It highlights that general intelligence does not guarantee trading success and emphasizes the importance of risk control in autonomous agents.
Beyond Agent Architecture: Execution Assumptions and Reproducibility in LLM-Based Trading Systems
This paper reviews and audits execution realism in LLM-based trading research, proposing clearer reporting standards for reproducibility and evaluation comparability.