AQuA: Recursively Self-Improving Quantitative Trading Research Agents

arXiv cs.CL Papers

Summary

AQuA is a research system with two independent language-model-driven agents that recursively self-improve in quantitative trading research, achieving strong information coefficients on crypto and US equities while using sealed sandboxes to prevent data leakage.

arXiv:2608.12841v1 Announce Type: new Abstract: We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from earlier experiments to improve the hypotheses and candidates proposed in later iterations. We present AQuA, which comprises two separate language-model-driven research systems: one for symbolic factor discovery and one for trainable model development. The two systems do not share agents, memories, candidate spaces, or research state. Instead, each independently closes its own research loop by retaining validated evidence and using it to guide subsequent proposals. In this bounded sense, both systems implement recursive self-improvement at the level of the research process. Each system also uses its own sealed sandbox, which fixes the data splits, feature and label definitions, and evaluator while allowing the model to act only through constrained factor expressions or configuration diffs. The factor system, a manager-mediated multi-agent pipeline, discovers and combines factors into a signal that reaches a combined information coefficient of about $0.190$ on a crypto universe. The model system, a config-driven loop over a hybrid time-series architecture, reaches a per-stock information coefficient of $+0.0843$ on US equities and converts it into a threshold long/short strategy with a held-out Sharpe of up to $+2.50$ at a two-leg cost. The strategy is positive in every year from 2021 to 2025.
Original Article
View Cached Full Text

Cached at: 08/14/26, 09:27 AM

# Recursively Self-Improving Quantitative Trading Research Agents
Source: [https://arxiv.org/html/2608.12841](https://arxiv.org/html/2608.12841)
Jiacheng Guo1\*, Suozhi Huang1\*, Yunlong Gao2\*, Zihao Li1, Jian Ge, Xu Kuang3, Mengdi Wang1 1Princeton University2Ant Group3Stanford University

###### Abstract

We study recursive self\-improvement at the level of quantitative\-investment research: whether an autonomous system can use evidence from earlier experiments to improve the hypotheses and candidates proposed in later iterations\. We presentAQuA, which comprises two separate language\-model\-driven research systems: one for symbolic factor discovery and one for trainable model development\. The two systems do not share agents, memories, candidate spaces, or research state\. Instead, each independently closes its own research loop by retaining validated evidence and using it to guide subsequent proposals\. In this bounded sense, both systems implement recursive self\-improvement at the level of the research process\. Each system also uses its own sealed sandbox, which fixes the data splits, feature and label definitions, and evaluator while allowing the model to act only through constrained factor expressions or configuration diffs\. The factor system, a manager\-mediated multi\-agent pipeline, discovers and combines factors into a signal that reaches a combined information coefficient of about0\.1900\.190on a crypto universe\. The model system, a config\-driven loop over a hybrid time\-series architecture, reaches a per\-stock information coefficient of\+0\.0843\+0\.0843on US equities and converts it into a threshold long/short strategy with a held\-out Sharpe of up to\+2\.50\+2\.50at a two\-leg cost\. The strategy is positive in every year from 2021 to 2025\.

## 1Introduction

Quantitative\-investment research searches over a large space of factors and models, and small methodological errors can turn into convincing but non\-reproducible backtests\. A feature that reads future information, a strategy selected on the test set, or a result confined to one favorable regime may all fail out of sample[6](https://arxiv.org/html/2608.12841#bib.bib8);[5](https://arxiv.org/html/2608.12841#bib.bib9)\. Quantitative research therefore relies on frozen data splits, held\-out evaluation, and skepticism toward results that look unusually strong[20](https://arxiv.org/html/2608.12841#bib.bib23)\.

Large language models can now propose hypotheses, write experiments, and revise their search from empirical feedback\. Existing quantitative agents, however, generally focus on either factor discovery[34](https://arxiv.org/html/2608.12841#bib.bib28);[11](https://arxiv.org/html/2608.12841#bib.bib34)or model development[23](https://arxiv.org/html/2608.12841#bib.bib40);[36](https://arxiv.org/html/2608.12841#bib.bib42)\. More importantly, an unconstrained agent can corrupt the evidence on which its later iterations depend\.

A code\-generating agent may inadvertently introduce a temporal\-alignment or preprocessing error that uses information unavailable at prediction time\. A reviewing agent may miss the bug because the code appears semantically plausible\. If the resulting leakage produces a high score, the experiment may be stored as a successful precedent and propagated through later iterations\. Recursive improvement can therefore amplify an undetected error as readily as a genuine discovery\.

Prompt\-level instructions and model\-based review do not provide a reliable integrity boundary: repeated access to a fixed holdout can cause adaptive overfitting[9](https://arxiv.org/html/2608.12841#bib.bib48), and language\-model agents have been observed exploiting misspecified objectives, tests, evaluators, or reward mechanisms[14](https://arxiv.org/html/2608.12841#bib.bib49);[7](https://arxiv.org/html/2608.12841#bib.bib50);[3](https://arxiv.org/html/2608.12841#bib.bib51)\.

AQuA instead makes leakage\-inducing actions unavailable to the agent\. Each system fixes its data pipeline, splits, labels, and evaluator before autonomous iteration begins\. The agent cannot write arbitrary experimental code; it can only emit a program in a restricted domain\-specific language \(DSL\), which contains no operation for modifying or bypassing the sealed data path or evaluator\.

We buildAQuA, comprising two separate recursively self\-improving research systems: one for factor discovery and one for model development\. They do not share agents, memories, candidate spaces, or outputs\. In each system, validated experiments update a local research state that guides later proposals, while the experimental contract used to judge those proposals remains fixed\. Thus, what improves is the research process, not the definition of success\.

![Refer to caption](https://arxiv.org/html/2608.12841v1/figures/figure1_overview.png)Figure 1:Overview of the two AQuA systems\. Factor discovery \(Part I\) and model development \(Part II\) use separate agents, memories, search spaces, and outputs\. Each closes its own loop through hypothesis, construct/train, evaluate, validate, select/combine, and a persistent update that guides the next iteration\.The key design is asymmetric freedom: the agent remains free to explore within its DSL, but the evaluator is outside the adaptive surface\. Part I implements this principle with a manager\-mediated multi\-agent pipeline that proposes falsifiable economic hypotheses, evaluates factors, combines surviving signals, and carries beliefs across runs\. Part II implements it with a config\-driven loop over a hybrid time\-series model; each configuration diff defines one comparable model variant, and its outcome updates the knowledge used to propose the next variant\.

Both systems produce out\-of\-sample signal\. On a crypto universe, Part I reaches a combined signal information coefficient of about0\.1900\.190\. On US equities, Part II reaches a per\-stock information coefficient of\+0\.0843\+0\.0843, versus\+0\.0613\+0\.0613for the strongest baseline, a GRU, under the same evaluator—an absolute improvement of\+0\.0230\+0\.0230and a relative improvement of37\.5%37\.5\\%\. The resulting threshold long/short strategy reaches a held\-out Sharpe of\+2\.50\+2\.50at a two\-leg cost of22bps and is positive in every year from 2021 to 2025\. A stricter walk\-forward evaluation, in which every model and strategy parameter is fixed using only data available before the next test segment, retains a Sharpe of about\+2\.0\+2\.0\.

Our main contributions are:

- •We instantiate recursive self\-improvement in two separate components of quantitative research\.In both factor discovery and model development, validated experiments improve later research decisions without coupling the two systems\.
- •We build an autonomous factor\-discovery system\.A manager\-mediated multi\-agent pipeline proposes, evaluates, and combines factors while carrying empirical beliefs across runs\.
- •We design a hybrid time\-series model and an autonomous development loop\.A config\-driven agent trains directly comparable variants under a sealed evaluation sandbox\.
- •We demonstrate out\-of\-sample predictive and trading performance\.The two systems produce positive signals in crypto and US equities, with the equity strategy remaining profitable under a fully causal walk\-forward evaluation\.

Section[3](https://arxiv.org/html/2608.12841#S3)describes the shared high\-level pattern and the two sealed sandboxes; Sections[4](https://arxiv.org/html/2608.12841#S4)and[5](https://arxiv.org/html/2608.12841#S5)present factor discovery and model development\.

## 2Related Work

#### LLM\-driven alpha mining\.

A fast\-growing line of work uses large language models to mine formulaic alpha factors\. It builds on operator\-based factor search by genetic programming[51](https://arxiv.org/html/2608.12841#bib.bib5);[13](https://arxiv.org/html/2608.12841#bib.bib6)and reinforcement learning[48](https://arxiv.org/html/2608.12841#bib.bib7);[50](https://arxiv.org/html/2608.12841#bib.bib21);[53](https://arxiv.org/html/2608.12841#bib.bib29);[54](https://arxiv.org/html/2608.12841#bib.bib33), and now spans evolutionary and agentic search[19](https://arxiv.org/html/2608.12841#bib.bib2);[38](https://arxiv.org/html/2608.12841#bib.bib13);[37](https://arxiv.org/html/2608.12841#bib.bib30);[44](https://arxiv.org/html/2608.12841#bib.bib31);[22](https://arxiv.org/html/2608.12841#bib.bib52);[46](https://arxiv.org/html/2608.12841#bib.bib53);[27](https://arxiv.org/html/2608.12841#bib.bib54);[49](https://arxiv.org/html/2608.12841#bib.bib55), program\-level synthesis[26](https://arxiv.org/html/2608.12841#bib.bib15), graph\-structured evolution[17](https://arxiv.org/html/2608.12841#bib.bib16), self\-evolving agents with experience memory[42](https://arxiv.org/html/2608.12841#bib.bib17);[47](https://arxiv.org/html/2608.12841#bib.bib20), safety\- and reproducibility\-constrained generation[33](https://arxiv.org/html/2608.12841#bib.bib18), market\-logic modeling[43](https://arxiv.org/html/2608.12841#bib.bib19), and standardized benchmarks[30](https://arxiv.org/html/2608.12841#bib.bib14)\. Newer systems add tree search and chain\-of\-thought prompting[34](https://arxiv.org/html/2608.12841#bib.bib28);[10](https://arxiv.org/html/2608.12841#bib.bib32), multi\-agent generation\-and\-selection pipelines[11](https://arxiv.org/html/2608.12841#bib.bib34);[40](https://arxiv.org/html/2608.12841#bib.bib47), and the fusion of formulaic factors with textual newsflow[18](https://arxiv.org/html/2608.12841#bib.bib35)\. The classic formulaic\-alpha vocabulary[24](https://arxiv.org/html/2608.12841#bib.bib4)underlies all of these\. Our Part I sits within this line and shares its proposal\-first, memory\-driven design\. AQuA places this factor\-discovery system alongside a separate autonomous model\-development system\. The two do not share agents, memory, or search state; rather, each independently uses prior experimental outcomes to improve subsequent proposals in its own domain\.

#### Autonomous research agents\.

Beyond finance, language\-model agents have been built to run the scientific process end to end[29](https://arxiv.org/html/2608.12841#bib.bib1);[32](https://arxiv.org/html/2608.12841#bib.bib27);[55](https://arxiv.org/html/2608.12841#bib.bib37);[45](https://arxiv.org/html/2608.12841#bib.bib36)\. The same agentic pattern has moved into trading, where multi\-agent teams of language models propose and act on investment decisions[31](https://arxiv.org/html/2608.12841#bib.bib38);[35](https://arxiv.org/html/2608.12841#bib.bib39)\. We adopt the same ambition of an autonomous research loop, but study recursive self\-improvement separately in two quantitative settings: factor discovery and model development\. In each setting, validated evidence from one iteration is retained and used to improve later research decisions\. We additionally target the specific failure mode of quantitative work, data leakage, by making the data path and the evaluator unreachable from the agent rather than relying on its judgment\.

#### Deep models for financial time series\.

The model in Part II draws on standard sequence\-modeling components, including convolutional networks for sequences[4](https://arxiv.org/html/2608.12841#bib.bib10), state\-space models[15](https://arxiv.org/html/2608.12841#bib.bib11), and attention[39](https://arxiv.org/html/2608.12841#bib.bib12), and on deep architectures designed for financial data[52](https://arxiv.org/html/2608.12841#bib.bib3);[16](https://arxiv.org/html/2608.12841#bib.bib22)\. A recent wave applies hybrid recurrent, convolutional, and transformer models, often combined with reinforcement learning, graph, or foundation\-model components, to stock\-return and portfolio prediction[23](https://arxiv.org/html/2608.12841#bib.bib40);[28](https://arxiv.org/html/2608.12841#bib.bib41);[36](https://arxiv.org/html/2608.12841#bib.bib42);[2](https://arxiv.org/html/2608.12841#bib.bib43);[41](https://arxiv.org/html/2608.12841#bib.bib44);[1](https://arxiv.org/html/2608.12841#bib.bib45);[25](https://arxiv.org/html/2608.12841#bib.bib46)\. Our contribution is not a new primitive but a separate autonomous loop that lets the agent compose these primitives into models and accumulate evidence across variants\. Its relation to Part I is conceptual—both systems learn from prior experiments—rather than architectural\.

## 3Method Overview

Part I and Part II of AQuA are separate research systems\. They do not share agents, memories, candidate spaces, or research state\. We describe them together here only because both exhibit the same high\-level pattern: each system uses validated evidence from earlier experiments to guide later research decisions in its own domain\.

Within each part, an iteration proceeds through five stages\. It begins with a*hypothesis*: a candidate factor in Part I, or a model configuration in Part II\. The hypothesis is*constructed or trained*into a concrete artifact,*evaluated*on held\-out data, and*validated*against look\-ahead bias and regime dependence\. Surviving artifacts are then*selected or combined*into that part’s running output, a combined factor signal in Part I and a trading model in Part II\. A final*persistent research\-state update*writes the iteration’s evidence to the store belonging to that part, which its next iteration consults before proposing a new hypothesis\. This feedback separates each system from a one\-shot pipeline: successive iterations accumulate and reuse evidence rather than restart from scratch\.

For partp∈\{I,II\}p\\in\\\{\\mathrm\{I\},\\mathrm\{II\}\\\}, letRt\(p\)R\_\{t\}^\{\(p\)\}denote its persistent research state before iterationtt,Ht\(p\)H\_\{t\}^\{\(p\)\}its hypothesis,Ct\(p\)C\_\{t\}^\{\(p\)\}the constructed factor or trained model, andEt\(p\)E\_\{t\}^\{\(p\)\}the validated evidence returned by the experiment\. The within\-part recursion is

Rt\+1\(p\)=𝒰p\(Rt\(p\),Ht\(p\),Ct\(p\),Et\(p\)\),\(Ht\+1\(p\),Ct\+1\(p\)\)∼𝒫p\(⋅∣Rt\+1\(p\)\)\.R\_\{t\+1\}^\{\(p\)\}=\\mathcal\{U\}\_\{p\}\\\!\\left\(R\_\{t\}^\{\(p\)\},H\_\{t\}^\{\(p\)\},C\_\{t\}^\{\(p\)\},E\_\{t\}^\{\(p\)\}\\right\),\\qquad\\left\(H\_\{t\+1\}^\{\(p\)\},C\_\{t\+1\}^\{\(p\)\}\\right\)\\sim\\mathcal\{P\}\_\{p\}\\\!\\left\(\\cdot\\mid R\_\{t\+1\}^\{\(p\)\}\\right\)\.\(1\)The part\-specific update𝒰p\\mathcal\{U\}\_\{p\}converts experimental outcomes into reusable research knowledge, and𝒫p\\mathcal\{P\}\_\{p\}uses that knowledge to guide the next proposal\. The superscript emphasizes the separation: Part I does not updateR\(II\)R^\{\(\\mathrm\{II\}\)\}, and Part II does not updateR\(I\)R^\{\(\\mathrm\{I\}\)\}\. We use*recursive self\-improvement*in this bounded, research\-process\-level sense within each part; neither system updates the underlying language model or the evaluator\.

The commonality is therefore a high\-level research pattern, not a shared implementation\. Construction writes a symbolic factor expression in Part I and trains a hybrid time\-series model in Part II; evaluation reports an information coefficient in both, but under part\-specific conventions, so the two coefficients are not directly comparable and we report them separately\.

### 3\.1Sealed sandbox and constrained iteration

Because the language model proposes and, in places, writes its own experiments, the central failure mode is data leakage: an autonomously generated feature, label, or normalization that consults information unavailable at prediction time\. Backtest overfitting of this kind is well documented and produces simulated performance that does not survive out of sample[6](https://arxiv.org/html/2608.12841#bib.bib8);[5](https://arxiv.org/html/2608.12841#bib.bib9);[20](https://arxiv.org/html/2608.12841#bib.bib23)\. We contain it structurally by fixing a sandbox before any iteration begins\. The data splits𝒟\\mathcal\{D\}, the feature and label definitionsℱ\\mathcal\{F\}andℒ\\mathcal\{L\}, and the evaluator𝒱\\mathcal\{V\}are sealed and human\-authored, and the model never edits them\. Each action of the model is a specificationθ\\thetadrawn from a constrained spaceΘ\\Theta, which the harness compiles and scores through the sealed evaluator,

sk=𝒱⁡\(𝒞⁡\(θk\),𝒮\),θk∈Θ,𝒮=\(𝒟,ℱ,ℒ,𝒱\)​sealed\.s\_\{k\}=\\mathcal\{V\}\\big\(\\mathcal\{C\}\(\\theta\_\{k\}\);\\,\\mathcal\{S\}\\big\),\\qquad\\theta\_\{k\}\\in\\Theta,\\quad\\mathcal\{S\}=\(\\mathcal\{D\},\\mathcal\{F\},\\mathcal\{L\},\\mathcal\{V\}\)\\ \\text\{sealed\}\.\(2\)The spaceΘ\\Thetais defined so that noθ\\thetacan alter𝒮\\mathcal\{S\}\. Leakage cannot enter through generation because generation cannot reach the sealed components\.

Sealing the data path closes leakage through generation, but a second channel remains through selection: a search that can read the metric it will ultimately report will, given enough iterations, learn to select for it\. We close this channel by separating the metric the loop optimizes from the metric it reports\. During search the harness returns to the agent only a score on a validation slice fixed in advance; the scoresks\_\{k\}that ranks candidates and the inner\-validation signal that drives early stopping and checkpoint choice are computed on that slice alone\. Where a part designates a final test window, that window is scored once, after the configuration is frozen, and is never returned to the agent or used to rank candidates\. In Part II this window is the untouched 2021–2025 period, so the reported test coefficient is out of sample with respect to the entire search\.

In Part I, the registry𝒪\\mathcal\{O\}is the standard vocabulary of formulaic\-alpha operators[24](https://arxiv.org/html/2608.12841#bib.bib4);[51](https://arxiv.org/html/2608.12841#bib.bib5)\. Its leaves are raw fields \(open, high, low, close, volume, vwap, returns\); its inner nodes are operators of three kinds: cross\-sectional operators that act across the universe at a fixed timestamp \(rank, z\-score, sector neutralization\), time\-series operators that summarize a trailing window for each entity \(lag, difference, moving correlation and covariance, rolling rank, rolling standard deviation, linear\-decay weighting\), and element\-wise arithmetic and conditionals\. A factor is a composition of these operators over the raw fields, represented as an expression tree[51](https://arxiv.org/html/2608.12841#bib.bib5);[13](https://arxiv.org/html/2608.12841#bib.bib6); the running example

f=rank⁡\(corr⁡\(r1​d,v1​d,20\)\)−rank⁡\(std⁡\(r1​d,20\)\)f=\\operatorname\{rank\}\\\!\\big\(\\operatorname\{corr\}\(r^\{1d\},v^\{1d\},20\)\\big\)\-\\operatorname\{rank\}\\\!\\big\(\\operatorname\{std\}\(r^\{1d\},20\)\\big\)\(3\)is one such tree\. Every time\-series operator reads only its trailing window and every cross\-sectional operator reads only the current timestamp, so causality is closed under composition: any expression the model assembles from𝒪\\mathcal\{O\}is causal by construction, and it cannot introduce a primitive that consults future data\. Operator\-based search of this space has a long line of work, from genetic programming[51](https://arxiv.org/html/2608.12841#bib.bib5);[13](https://arxiv.org/html/2608.12841#bib.bib6)and reinforcement learning[48](https://arxiv.org/html/2608.12841#bib.bib7);[53](https://arxiv.org/html/2608.12841#bib.bib29)to recent language\-model agents[19](https://arxiv.org/html/2608.12841#bib.bib2);[37](https://arxiv.org/html/2608.12841#bib.bib30);[34](https://arxiv.org/html/2608.12841#bib.bib28)\. Before assembling the expression, the agent states each factor as a falsifiable proposal, with a hypothesis, mechanism, predicted direction, and refutation conditions\. Listing[1](https://arxiv.org/html/2608.12841#LST1)shows one such proposal\.

Listing 1:A proposal in Part I is a falsifiable factor hypothesis, with its mechanism, predicted direction, and refutation conditions, stated before any expression is built\. The deployed expression is withheld; the full iteration record is in Appendix[A](https://arxiv.org/html/2608.12841#A1)\.proposal:

hypothesis:\>

Afteraforcedopen\-interestunwind,apricereboundthatisnotconfirmed

byaggressivetakerflowismorelikelytofail\.

mechanism:\>

Deleveragingremovesforcedpressure,butweakbuy\-flowandpoorbasis

recoveryindicateinsufficientdemandbehindtherebound\.

expected\_direction:higher\_signal\_predicts\_lower\_future\_return

expected\_label:\[ret\_open\_open\_h10,ret\_open\_open\_h30\]

falsification\_criteria:

\-noICconcentrationinsidethedeleveragingeventwindow

\-signflipsorvanishesoutofsampleoracrossregimes

\-subsumedbyaprice,open\-interest,ortaker\-flowbaseline

expected\_failure\_modes:

\-reboundstrengthalreadycapturedbyshort\-horizonmomentum

\-taker\-flowgapisnoiseoncethebasishasnormalized

factor\_blueprint:

\-deleveraging\_intensity:rankednegativechangeinopeninterest

\-rebound\_strength:short\-horizonpricerecoveryaftertheevent

\-flow\_gap:lackoftaker\-flowconfirmationduringtherebound

expression:withheld

In Part II, the specification is a configuration in a domain\-specific language whose registry covers the full experiment: the data split selected from the frozen set, the sampler, the architecture blocks, the loss, and the optimizer\. A hypothesis is a single config diff\. The registry is the model’s only surface; it selects and parameterizes registered operators but cannot write the data loader, the split logic, or the evaluator\. Listing[2](https://arxiv.org/html/2608.12841#LST2)shows one such configuration\. The architecture registry composes standard sequence\-modeling primitives, including temporal convolutions[4](https://arxiv.org/html/2608.12841#bib.bib10), state\-space mixers[15](https://arxiv.org/html/2608.12841#bib.bib11), and attention[39](https://arxiv.org/html/2608.12841#bib.bib12), into deep multi\-resolution stacks\.

Listing 2:A hypothesis in Part II is one config diff over the sealed sandbox, the only artifact the model emits\. Sealed components \(splits, features, labels, evaluator\) are referenced by id and never redefined; operator names and values are illustrative\.sandbox:

splits:frozen/sp\_A\#train/val/test

features:frozen/feat\_g12

label:frozen/lbl\_fwd

evaluator:frozen/eval\#held\-out,two\-legcost,walk\-forward

input:\{normalize:per\_entity\_z,stats\_window:train\_only,history:64\}

sampler:\{kind:stratified\_minute,batch:8192\}

arch:\#composedfromthearchitectureregistry

\-conv\_stem:\{scales:\[3,5,15\]\}\#multi\-scale1\-Dconv

\-multi\_resolution:

fine:\{repeat:4,block:\[temporal\_conv,sequence\_mixer/state\_space\]\}

coarse:\{repeat:2,block:\[sequence\_mixer/attention,feedforward\]\}

fuse:cross\_attention

\-block\_repeat:

times:3\#panelinteraction

block:\[cross\_entity\_mixer,temporal\_conv/depthwise,feedforward\]

\-readout:\{gate:\[fine,coarse,panel\],pool:\[last,mean,attn\]\}

loss:\[spearman\_ic,huber\_csz,turnover\_reg\]

optim:\{adamw,lr:3\.0e\-4,schedule:cosine,precision:bf16,ddp:8\}

eval:\{cost\_bps:2,walk\_forward:expanding\}\#held\-out

Formally,ΘII=\{c=\(split,sampler,arch,loss,optim\)\}\\Theta\_\{\\mathrm\{II\}\}=\\\{\\,c=\(\\mathrm\{split\},\\mathrm\{sampler\},\\mathrm\{arch\},\\mathrm\{loss\},\\mathrm\{optim\}\)\\,\\\}, wheresplit\\mathrm\{split\}is chosen from the frozen𝒟\\mathcal\{D\}and cannot be redefined, and the remaining components are drawn from their registries\. The compiler builds the model and training run, and the sealed evaluator reports held\-out IC,R2R^\{2\}, and Sharpe under a two\-leg cost\. Because the data path is sealed, the split is frozen, and every feature is causal, everyc∈ΘIIc\\in\\Theta\_\{\\mathrm\{II\}\}yields a leakage\-free experiment\. One config diff also corresponds to exactly one variant, which keeps variants directly comparable\. Sections[4](https://arxiv.org/html/2608.12841#S4)and[5](https://arxiv.org/html/2608.12841#S5)instantiate these two registries in full\.

## 4Part I: Autonomous Factor Discovery

### 4\.1System architecture

Part I of AQuA is a multi\-agent system that turns a research goal into validated factors\. An AI Manager mediates each run\. It reads the research policy, the accumulated memory, and the record of previous runs, and from these it writes a plan that assigns every downstream agent a concrete task\. The agents never call one another; each handoff passes through the Manager, which keeps a run auditable and reproducible\. Figure[2](https://arxiv.org/html/2608.12841#S4.F2)shows the pipeline\.

![Refer to caption](https://arxiv.org/html/2608.12841v1/figures/figure2_part1_architecture.png)Figure 2:Part I architecture\. An AI Manager orchestrates a six\-agent pipeline \(Data Steward→\\toVisual Analyst→\\toIdea Miner→\\toFactor Evaluator→\\toBacktest Engineer→\\toResearch Librarian\) that turns a research goal into a combined factor signal\. Three nested feedback loops, for direction calibration, falsification\-driven belief update, and cross\-run memory and policy, drive iterative improvement, backed by persistent Memory, Beliefs, and Policy stores\.Six specialist agents run in sequence\. The Data Steward loads and aligns the market data and returns a quality report together with a set of atomic features\. The Visual Analyst searches the history for representative events of a requested type and summarizes them into event profiles\. The Idea Miner proposes candidate factors from those profiles\. The Factor Evaluator scores each candidate, the Backtest Engineer trades it in simulation, and the Research Librarian records the outcome\. The proposal and validation steps are detailed below; the specific factors the system selects are not disclosed\.

### 4\.2The research pipeline

A factor enters the pipeline as a proposal rather than as an expression\. For each candidate the Idea Miner states a hypothesis, the economic mechanism behind it, the direction it is expected to predict, and the conditions under which it should be considered refuted\. This framing requires every factor to carry its own rationale and its own test before it is built, a discipline that the recent agentic\-mining literature has converged on[19](https://arxiv.org/html/2608.12841#bib.bib2);[38](https://arxiv.org/html/2608.12841#bib.bib13);[33](https://arxiv.org/html/2608.12841#bib.bib18);[37](https://arxiv.org/html/2608.12841#bib.bib30);[40](https://arxiv.org/html/2608.12841#bib.bib47)\. The candidate is then assembled from the formulaic\-alpha operator registry of Section[3\.1](https://arxiv.org/html/2608.12841#S3.SS1), where Listing[1](https://arxiv.org/html/2608.12841#LST1)shows the proposal form it emits\.

Evaluation follows the same contract for every proposal\. The Factor Evaluator computes information coefficients across forward\-return labels, monthly stability, held\-out split behavior, market\-regime behavior, turnover, expression complexity, and correlation with the existing factor pool\. It also performs controlled comparisons against simple baselines built from price, volume, open interest, basis, and taker\-flow changes\. When a proposal is tied to an event, its effect inside the event window is compared with a control region outside the event\. A factor is carried forward when its signal is not simply a restatement of a baseline and when its strongest performance appears in the market context predicted by its mechanism\.

The Backtest Engineer then turns the evaluated signal into a simple quantile portfolio\. It tests both the proposed direction and the reversed direction and keeps the cleaner trading interpretation\. This direction\-calibration step matters because a formula can have predictive content even when the initially stated sign is wrong\. The final object stored by the system is therefore not only a factor score, but a factor together with its direction, target horizon, supporting mechanism, evaluation evidence, and trading behavior\.

### 4\.3Cross\-run learning

The system improves across runs because it maintains memory\. After each run, the Research Librarian writes a structured record containing the goal, event type, visual observations, proposed mechanisms, evaluated factors, selected signals, backtest summaries, and updated beliefs\. A belief attaches confidence to a mechanism in a market context, such as “open\-interest crashes followed by weak flow\-confirmed rebounds tend to reverse” or “quiet volume bursts continue only when price acceptance and taker flow agree\.” The next run is planned against this memory\.

This creates three feedback loops\. The first operates within a single backtest, where the system calibrates the direction of a signal\. The second operates within a run, where failed proposals update the belief state and sharpen the interpretation of the event\. The third operates across runs, where the AI Manager reads the accumulated memory and steers the next search toward mechanisms that have earned evidence\. In practice, later runs do not start from a blank prompt\. They inherit prior observations, avoid mechanisms that repeatedly failed, and refine promising mechanisms with new event definitions, horizons, or conditioning variables\.

### 4\.4A worked iteration

To make the loop concrete, Appendix[A](https://arxiv.org/html/2608.12841#A1)records one full iteration, and we summarize it here\. The run asks whether a weak price rebound after an open\-interest crash predicts reversal\. The Manager turns this question into a plan: find deleveraging episodes, inspect price and flow after the unwind, and test whether rebound quality separates continuation from failure\. The Visual Analyst returns event profiles that separate a clean forced\-deleveraging pattern, where open interest falls, volume expands, and price rebounds only briefly, from a healthier reset where taker flow and basis recover with price\. From this the Idea Miner proposes a mechanism in which a rebound left unconfirmed by aggressive taker flow is more likely to fail\. The evaluator clears it against the price, open\-interest, and flow baselines, with its skill concentrated in the event window its mechanism targets\. The strongest factors in this family reach single\-factor information coefficients on the order of0\.0260\.026to0\.0370\.037\. A later run changes the goal to quiet\-market volume expansion and reuses the same machinery without touching the evaluator, showing that the loop explores a new mechanism by re\-planning rather than re\-coding\. The deployed expressions are withheld throughout\.

### 4\.5Results

The output of Part I is a combined factor signal formed from the factors that survive the evaluation and direction\-calibration stages\. Figure[3](https://arxiv.org/html/2608.12841#S4.F3)reports the quality of this combined signal across autonomous\-research iterations\. The combined validation information coefficient rises as the loop accumulates and reuses evidence, reaching approximately0\.1900\.190\. The improvement coincides with the addition of richer event profiles, open\-interest and flow conditioning, crowding\-divergence mechanisms, and regime\-aware factor selection\. As noted in Section[3](https://arxiv.org/html/2608.12841#S3), the Part I information coefficient follows the combined\-factor convention on the crypto five\-minute universe, and is not comparable to the per\-stock coefficient reported in Part II\.

Figure 3:Part I combined\-factor signal across autonomous\-research iterations\. The combined validation IC improves as the loop accumulates and reuses validated evidence, reaching a combined signal IC of approximately0\.1900\.190\.*Part I IC is the combined\-factor Spearman IC on the crypto five\-minute universe, a different convention from Part II; the two numbers should not be compared directly\.*Each mechanism discovered by the loop is an economically grounded signal, typically with an information coefficient around0\.030\.03in absolute value\. This is the expected regime for intraday formulaic signals: the strength comes not from any single rule but from combining a library of mechanism\-grounded factors into a stronger aggregate signal\. The Part I result is therefore not a claim that one discovered expression is sufficient, but that an autonomous, memory\-bearing research harness can repeatedly turn market hypotheses into tested factor evidence and improve the combined signal over iterations\.

## 5Part II: Autonomous Model Development

### 5\.1The research loop

Part II of AQuA develops trading models through the config\-driven loop of Figure[4](https://arxiv.org/html/2608.12841#S5.F4)\. Each iteration begins with a hypothesis written as a single config diff over the sealed sandbox of Section[3\.1](https://arxiv.org/html/2608.12841#S3.SS1): a change to the architecture, the loss, the sampler, or the optimizer\. The harness compiles the diff into a training run, scores it through the sealed evaluator, and writes the result to a knowledge store that the next hypothesis reads\. Because one diff produces exactly one variant and the data path is fixed, two variants differ only in the knobs that changed\. This keeps them directly comparable and keeps the search leakage\-free\.

![Refer to caption](https://arxiv.org/html/2608.12841v1/figures/figure4_part2_loop.png)Figure 4:Part II architecture\. A config\-driven loop, from a hypothesis \(config diff\) through the training framework and the evaluation engine to a knowledge update, iterates over model variants\. The training framework \(architecture DSL, loss, data, and sampler registries, distributed training\) produces a hybrid time\-series model; the evaluation engine reports held\-out per\-stock IC,R2=mean⁡\(I​C2\)R^\{2\}=\\mathrm\{mean\}\(IC^\{2\}\), walk\-forward, and a two\-leg\-cost threshold long/short Sharpe, within a governance contract of frozen splits, a sealed evaluator, and a fixed selection metric\. The model signal is constructed into a trading strategy \(sector\-neutralize, then volatility targeting, then long/short\)\.
### 5\.2The hybrid model

Our predictor is a hybrid model that combines convolutional feature extraction with sequence modeling[52](https://arxiv.org/html/2608.12841#bib.bib3);[23](https://arxiv.org/html/2608.12841#bib.bib40), shown in Figure[5](https://arxiv.org/html/2608.12841#S5.F5)\. The front\-end is a deep convolutional stack: several blocks of one\-dimensional convolutions run over the input history at multiple kernel sizes and dilations, building a rich local representation of each stock’s recent price\-volume dynamics at every timestep\. We leave the precise convolutional configuration unspecified\. These representations feed a temporal\-modeling stage instantiated from sequence models, spanning recurrent networks \(LSTM[21](https://arxiv.org/html/2608.12841#bib.bib24)\), state\-space models \(Mamba[15](https://arxiv.org/html/2608.12841#bib.bib11)\), and attention \(Transformer[39](https://arxiv.org/html/2608.12841#bib.bib12);[4](https://arxiv.org/html/2608.12841#bib.bib10)\); the configuration used in our experiments uses attention\. A cross\-sectional stage then mixes information across the panel of stocks at each timestep, the branches are fused with a gating mechanism, and a pooled readout emits the per\-stock score\.

The whole model is one point in the configuration space of Section[3\.1](https://arxiv.org/html/2608.12841#S3.SS1): the convolutional front\-end, the sequence family, the cross\-sectional mixing, and the fusion are all operators in the registry, so a new model is a new config rather than new code\. The framework trains it under the sealed sandbox, and the exact feature set, normalization, and label construction are part of that sandbox and are not disclosed\.

![Refer to caption](https://arxiv.org/html/2608.12841v1/figures/hybrid_time_series.png)Figure 5:The hybrid time\-series model\. A multi\-scale convolutional front\-end extracts local representations from multi\-horizon return, volatility, and risk\-adjusted momentum features at each timestep\. A configurable temporal backbone, instantiated as a recurrent, state\-space, or attention\-based sequence model, captures longer\-range temporal dynamics\. A backbone\-dependent sequence readout produces a fixed\-dimensional representation, which is concatenated with configured auxiliary information such as ticker embeddings, time features, and optional theme or sector context\. A configurable MLP or SwiGLU prediction head then emits one scalar score for each stock sample\.
### 5\.3Task, features, and baselines

We instantiate Part II on intraday US equity prediction\. The task is to predict each stock’s forward return over the next thirty minutes\. We split the data chronologically: we train on 2010–2019, leave 2020 as an embargo gap that no part of training or selection touches, and report on the untouched 2021–2025 test window\. Model selection \(early stopping and checkpoint choice\) is driven only by an inner\-validation slice taken from the end of the training window, and the test window is used solely for final evaluation, never informing training or selection\. The model input is a short history of pure price\-volume features, standardized per stock\.

No single feature carries the signal on its own\. Table[1](https://arxiv.org/html/2608.12841#S5.T1)reports the single\-feature information coefficient of a representative set of price\-volume features over the held\-out window\. None of these features, nor a ridge linear combination of them, exceeds about0\.030\.03in magnitude; the ridge reaches\+0\.025\+0\.025\. The predictable signal lives in their joint, nonlinear, and temporal structure, which is what the model is built to capture\.

Table 1:Single\-feature information coefficient of representative price\-volume features, held out over 2021–2025 \(per\-stock raw IC\)\. No single feature, or ridge combination of them, exceeds about0\.030\.03in magnitude\.The autonomous loop searched a range of model families on this task, all evaluated identically on the held\-out window\. Table[2](https://arxiv.org/html/2608.12841#S5.T2)reports their per\-stock raw IC, from a linear model through gradient boosting and recurrent networks[21](https://arxiv.org/html/2608.12841#bib.bib24);[12](https://arxiv.org/html/2608.12841#bib.bib25);[8](https://arxiv.org/html/2608.12841#bib.bib26)to our hybrid model\. The progression is clear: linear and tree models capture part of the signal, sequence models more, and our hybrid model is the strongest\.

Table 2:Model comparison on the held\-out window \(2021–2025\), per\-stock raw IC\. Models are trained on identical data and scored by the same evaluator\. Our hybrid model is the strongest\.
### 5\.4Evaluation engine

Every run is scored by the same sealed evaluator\. It reports three quantities on held\-out data: a per\-stock time\-series information coefficient, a per\-stockR2R^\{2\}defined as the cross\-sectional mean of squared IC, and a threshold long/short Sharpe ratio under a two\-leg turnover cost\.

Figure[6](https://arxiv.org/html/2608.12841#S5.F6)summarizes the trading result\. The cumulative return of the volatility\-targeted, dollar\-neutral long/short book climbs steadily across 2021–2025, making new highs into the end of the window and ending well above a Nasdaq\-100 buy\-and\-hold benchmark over the same period without taking its 2022 drawdown\.

Figure 6:Part II strategy equity curve\. Cumulative return of the volatility\-targeted, dollar\-neutral threshold long/short book \(gross exposure11, two\-leg turnover cost22bps\) over 2021–2025, with the Nasdaq\-100 \(QQQ\) buy\-and\-hold return over the same window shown for reference\. The market\-neutral book compounds more smoothly and sidesteps the 2022 index drawdown; QQQ is long\-only and not risk\-matched to the neutral book\.
### 5\.5From signal to strategy

We trace one model from prediction to deployable strategy\. A trained model produces a per\-stock score at each timestamp\. The score is turned into a dollar\-neutral threshold long/short book: stocks above an upper threshold are held long and stocks below a lower threshold short, the book is rebalanced on a fixed cadence, and each leg pays a two\-leg turnover cost\. Two construction steps lift the held\-out Sharpe\. Sector\-neutralizing the score removes common sector exposure and raises the held\-out Sharpe to\+2\.15\+2\.15, with the training and held\-out values nearly equal, which indicates that the construction is not overfit\. A causal volatility\-targeting overlay, which scales daily exposure by an online estimate of trailing volatility toward an expanding\-median target, raises it further to\+2\.50\+2\.50\.

### 5\.6Results

Table[3](https://arxiv.org/html/2608.12841#S5.T3)reports the held\-out metrics for 2021–2025\. The model reaches a per\-stock raw information coefficient of\+0\.0843\+0\.0843and a per\-stockR2R^\{2\}of1\.20%1\.20\\%\. The threshold long/short book reaches a Sharpe of\+2\.15\+2\.15before the volatility overlay and\+2\.50\+2\.50after it, and a fully causal walk\-forward that chooses every parameter from past data alone still reaches\+2\.0\+2\.0\. Table[4](https://arxiv.org/html/2608.12841#S5.T4)breaks the Sharpe down by year\. It is positive in every year from 2021 to 2025, including the held\-out years and the 2022 drawdown, so the strategy is not carried by a single regime\. As stated in Section[3](https://arxiv.org/html/2608.12841#S3), the Part II information coefficient is a per\-stock time\-series quantity, reported on its own and not against Part I\.

Table 3:Part II main results on the held\-out evaluation window \(2021–2025\)\. All figures are out of sample, with model selection on validation only\. The information coefficient is the per\-stock time\-series Pearson IC \(raw\);R2R^\{2\}is the cross\-sectional mean of squared IC; the Sharpe ratio is for a dollar\-neutral threshold long/short book at a two\-leg turnover cost of22bps, with parameters tuned on a training fraction and scored on the held\-out remainder\.*\(IC convention differs from Part I and the two should not be compared\.\)*Table 4:Part II strategy Sharpe@2bp by calendar year, held out, showing performance is not concentrated in any single regime\.

## 6Discussion

Part I and Part II are separate systems rather than a coupled or shared\-state framework\. They use different agents, memories, candidate spaces, and outputs\. Each nevertheless exhibits the same high\-level property: an autonomous loop proposes a candidate, validates it on held\-out data, accumulates the resulting evidence in its own persistent research state, and uses that state to improve later research decisions\. Part I applies this within\-system recursion to economic hypotheses and symbolic factor expressions; Part II applies it independently to model architectures and training configurations\. The commonality is therefore descriptive rather than architectural: in both cases, one experiment changes the design of the next within the same system\.

What makes both systems reliable is where each places trust\. Its sandbox seals the data path and scores the search on a validation metric it cannot confuse with the reported one, so a surviving result is credible from how the environment is built rather than from an audit of the agent’s reasoning\. This lets each autonomous, and at times opaque, search improve its research process without inheriting its capacity to overfit the number it optimizes\.

We found it useful to separate leakage into two channels, and that split is the part of this work most likely to transfer beyond finance\. Generation leakage enters when the agent can define a feature, label, or transform that consults information unavailable at prediction time; we close it by construction, since no admissible specification can reach the sealed data path\. Selection leakage enters when the agent can read the metric it will be judged on and, over enough iterations, learn to select for it; we close it by reporting a metric the loop never optimizes against\. Any autonomous research agent scored by an evaluator faces both channels, whatever the domain, so the sealed\-sandbox and split\-metric construction is a general recipe rather than a finance\-specific trick\.

Coupling the two systems is the natural next step: the factors discovered in Part I are direct inputs to the models trained in Part II\. That coupling introduces a leakage channel neither part has on its own\. If factor discovery and model training draw on the same data, a factor selected for its in\-sample signal can hand the model a subtly overfit input, so the two searches come to share information the sealed metric was meant to keep apart\. Keeping the coupled system honest means sealing the discovered factor set before the model loop begins, and treating the factor library as another frozen component of the Part II sandbox rather than a live search the model can steer\.

## 7Limitations

Two scope limitations bound these results\. Each system is demonstrated on a single market and horizon, crypto at five minutes for Part I and US equities at thirty minutes for Part II, and we do not claim the numbers transfer to other markets or frequencies without re\-tuning\. The loops also run with a human operator who sets the research goal, owns the sandbox, and supervises promotion, so the systems are autonomous within those bounds rather than unattended\. The reported metrics are simulated under a turnover\-cost model and have not been validated in live trading\.

The guarantees are also uneven across the two channels of leakage\. Sealing the data and feature path is structural: no admissible specification can reach past it, so causal correctness holds by construction\. Keeping the final test window out of selection is weaker\. The harness returns only validation scores during search, but the isolation of the test window rests on the sealed protocol and operator discipline rather than a hard technical barrier, and an operator with direct access to the store could in principle consult it\. We therefore treat test isolation as a governance property to be audited over a run, not as a cryptographic guarantee\.

## 8Conclusion

We presented AQuA, which comprises two separate autonomous research systems: one for factor discovery and one for model development\. The two parts operate over different research objects and do not share agents, memories, candidate spaces, or research state\. Both nevertheless implement recursive self\-improvement within their own research process: validated evidence is incorporated into a part\-specific state that guides subsequent hypotheses and candidate designs\. A separate sealed sandbox in each part keeps the data path and evaluator outside this recursive update\. On a crypto universe the factor system reaches a combined signal IC of about0\.1900\.190, and on US equities the model system reaches a per\-stock IC of\+0\.0843\+0\.0843and a regime\-robust threshold long/short Sharpe of up to\+2\.50\+2\.50out of sample\. Coupling the two systems, so that discovered factors feed the model loop, is the natural next direction\.

## References

- B\. Asher MarconiTime Series Foundation Models for Multivariate Financial Time Series Forecasting\.External Links:[Link](https://www.ssrn.com/abstract=6085266),[Document](https://dx.doi.org/10.2139/ssrn.6085266)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px3.p1.1)\.
- Ashrafzadehet al\.\(2025\)M\. Ashrafzadeh, M\. Sadrani, and S\. H\. ZolfaniDeep learning and machine learning models for portfolio optimization: Enhancing return prediction with stock clustering\.Results in Engineering27,pp\. 106263\.External Links:ISSN 25901230,[Link](https://linkinghub.elsevier.com/retrieve/pii/S2590123025023357),[Document](https://dx.doi.org/10.1016/j.rineng.2025.106263)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px3.p1.1)\.
- Atinafu and Cohen \(2026\)Y\. Atinafu and R\. CohenRewardHackingAgents: benchmarking evaluation integrity for LLM ML\-engineering agents\.arXiv\.Note:arXiv:2603\.11337External Links:[Document](https://dx.doi.org/10.48550/arXiv.2603.11337),[Link](https://arxiv.org/abs/2603.11337)Cited by:[§1](https://arxiv.org/html/2608.12841#S1.p4.1)\.
- Baiet al\.\(2018\)S\. Bai, J\. Z\. Kolter, and V\. KoltunAn Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling\.arXiv\.Note:arXiv:1803\.01271External Links:[Link](http://arxiv.org/abs/1803.01271),[Document](https://dx.doi.org/10.48550/arXiv.1803.01271)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px3.p1.1),[§3\.1](https://arxiv.org/html/2608.12841#S3.SS1.p4.1),[§5\.2](https://arxiv.org/html/2608.12841#S5.SS2.p1.1)\.
- Baileyet al\.\(2017\)D\. Bailey, J\. Borwein, M\. López De Prado, and Q\. J\. ZhuThe probability of backtest overfitting\.The Journal of Computational Finance\.External Links:ISSN 14601559,[Link](http://www.risk.net/journal-of-computational-finance/technical-paper/2471206/the-probability-of-backtest-overfitting),[Document](https://dx.doi.org/10.21314/JCF.2016.322)Cited by:[§1](https://arxiv.org/html/2608.12841#S1.p1.1),[§3\.1](https://arxiv.org/html/2608.12841#S3.SS1.p1.1)\.
- Baileyet al\.\(2014\)D\. H\. Bailey, J\. M\. Borwein, M\. López De Prado, and Q\. J\. ZhuPseudo\-Mathematics and Financial Charlatanism: The Effects of Backtest Overfitting on Out\-of\-Sample Performance\.Notices of the American Mathematical Society61\(5\),pp\. 458\.External Links:ISSN 0002\-9920, 1088\-9477,[Link](https://www.ams.org/jourcgi/jour-getitem?pii=noti1105),[Document](https://dx.doi.org/10.1090/noti1105)Cited by:[§1](https://arxiv.org/html/2608.12841#S1.p1.1),[§3\.1](https://arxiv.org/html/2608.12841#S3.SS1.p1.1)\.
- Bakeret al\.\(2025\)B\. Baker, J\. Huizinga, L\. Gao, Z\. Dou, M\. Y\. Guan, A\. Madry, W\. Zaremba, J\. Pachocki, and D\. FarhiMonitoring reasoning models for misbehavior and the risks of promoting obfuscation\.arXiv\.Note:arXiv:2503\.11926External Links:[Document](https://dx.doi.org/10.48550/arXiv.2503.11926),[Link](https://arxiv.org/abs/2503.11926)Cited by:[§1](https://arxiv.org/html/2608.12841#S1.p4.1)\.
- Becket al\.\(2024\)M\. Beck, K\. Pöppel, M\. Spanring, A\. Auer, O\. Prudnikova, M\. Kopp, G\. Klambauer, J\. Brandstetter, and S\. HochreiterxLSTM: Extended Long Short\-Term Memory\.arXiv\.Note:arXiv:2405\.04517External Links:[Link](http://arxiv.org/abs/2405.04517),[Document](https://dx.doi.org/10.48550/arXiv.2405.04517)Cited by:[§5\.3](https://arxiv.org/html/2608.12841#S5.SS3.p3.1)\.
- Blum and Hardt \(2015\)A\. Blum and M\. HardtThe ladder: a reliable leaderboard for machine learning competitions\.InProceedings of the 32nd International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.37,pp\. 1006–1014\.External Links:[Link](https://proceedings.mlr.press/v37/blum15.html)Cited by:[§1](https://arxiv.org/html/2608.12841#S1.p4.1)\.
- Cao \(2025\)L\. CaoChain\-of\-Alpha: Unleashing the Power of Large Language Models for Alpha Mining in Quantitative Trading\.arXiv\.Note:arXiv:2508\.06312External Links:[Link](http://arxiv.org/abs/2508.06312),[Document](https://dx.doi.org/10.48550/arXiv.2508.06312)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1)\.
- Chen and Kawashima \(2025\)Q\. Chen and H\. KawashimaMulti\-Agent LLM Framework for Formulaic Alpha Generation and Selection in Quantitative Trading\.In2025 IEEE International Conference on Big Data \(BigData\),pp\. 7143–7152\.External Links:[Link](https://ieeexplore.ieee.org/document/11400963/),[Document](https://dx.doi.org/10.1109/BigData66926.2025.11400963)Cited by:[§1](https://arxiv.org/html/2608.12841#S1.p2.1),[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1)\.
- Choet al\.\(2014\)K\. Cho, B\. v\. Merrienboer, C\. Gulcehre, D\. Bahdanau, F\. Bougares, H\. Schwenk, and Y\. BengioLearning Phrase Representations using RNN Encoder\-Decoder for Statistical Machine Translation\.arXiv\.Note:arXiv:1406\.1078External Links:[Link](http://arxiv.org/abs/1406.1078),[Document](https://dx.doi.org/10.48550/arXiv.1406.1078)Cited by:[§5\.3](https://arxiv.org/html/2608.12841#S5.SS3.p3.1)\.
- Cuiet al\.\(2021\)C\. Cui, W\. Wang, M\. Zhang, G\. Chen, Z\. Luo, and B\. C\. OoiAlphaEvolve: A Learning Framework to Discover Novel Alphas in Quantitative Investment\.InProceedings of the 2021 International Conference on Management of Data,Virtual Event China,pp\. 2208–2216\.External Links:ISBN 9781450383431,[Link](https://dl.acm.org/doi/10.1145/3448016.3457324),[Document](https://dx.doi.org/10.1145/3448016.3457324)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2608.12841#S3.SS1.p3.1),[§3\.1](https://arxiv.org/html/2608.12841#S3.SS1.p3.2)\.
- Denisonet al\.\(2024\)C\. Denison, M\. MacDiarmid, F\. Barez, D\. Duvenaud, S\. Kravec, S\. Marks, N\. Schiefer, R\. Soklaski, A\. Tamkin, J\. Kaplan, B\. Shlegeris, S\. R\. Bowman, E\. Perez, and E\. HubingerSycophancy to subterfuge: investigating reward\-tampering in large language models\.arXiv\.Note:arXiv:2406\.10162External Links:[Document](https://dx.doi.org/10.48550/arXiv.2406.10162),[Link](https://arxiv.org/abs/2406.10162)Cited by:[§1](https://arxiv.org/html/2608.12841#S1.p4.1)\.
- Gu and Dao \(2023\)A\. Gu and T\. DaoMamba: Linear\-Time Sequence Modeling with Selective State Spaces\.arXiv\.Note:arXiv:2312\.00752External Links:[Link](http://arxiv.org/abs/2312.00752),[Document](https://dx.doi.org/10.48550/arXiv.2312.00752)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px3.p1.1),[§3\.1](https://arxiv.org/html/2608.12841#S3.SS1.p4.1),[§5\.2](https://arxiv.org/html/2608.12841#S5.SS2.p1.1)\.
- Guet al\.\(2020\)S\. Gu, B\. Kelly, and D\. XiuEmpirical Asset Pricing via Machine Learning\.The Review of Financial Studies33\(5\),pp\. 2223–2273\.External Links:ISSN 0893\-9454, 1465\-7368,[Link](https://academic.oup.com/rfs/article/33/5/2223/5758276),[Document](https://dx.doi.org/10.1093/rfs/hhaa009)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px3.p1.1)\.
- Guoet al\.\(2026\)T\. Guo, H\. Shen, J\. Luo, B\. Chen, H\. Ding, J\. Huang, L\. Liu, Y\. Ma, and M\. ZhangAlphaPROBE: Alpha Mining via Principled Retrieval and On\-graph biased evolution\.arXiv\.Note:arXiv:2602\.11917External Links:[Link](http://arxiv.org/abs/2602.11917),[Document](https://dx.doi.org/10.48550/arXiv.2602.11917)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1)\.
- Guo and Hauptmann \(2025\)T\. Guo and E\. HauptmannExploring the Synergy of Quantitative Factors and Newsflow Representations from Large Language Models for Stock Return Prediction\.arXiv\.Note:arXiv:2510\.15691External Links:[Link](http://arxiv.org/abs/2510.15691),[Document](https://dx.doi.org/10.48550/arXiv.2510.15691)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1)\.
- Hanet al\.\(2026\)J\. Han, S\. Zhang, W\. Li, Y\. Dong, T\. Hu, Y\. Zhu, X\. Yu, X\. Guo, Z\. Liu, K\. Wang, J\. Liu, T\. Jiang, R\. An, S\. Hu, Z\. Yang, R\. Che, and H\. WangQuantaAlpha: An Evolutionary Framework for LLM\-Driven Alpha Mining\.arXiv\.Note:arXiv:2602\.07085External Links:[Link](http://arxiv.org/abs/2602.07085),[Document](https://dx.doi.org/10.48550/arXiv.2602.07085)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2608.12841#S3.SS1.p3.2),[§4\.2](https://arxiv.org/html/2608.12841#S4.SS2.p1.1)\.
- Harveyet al\.\(2016\)C\. R\. Harvey, Y\. Liu, and H\. Zhu… And the Cross\-Section of Expected Returns\.Review of Financial Studies29\(1\),pp\. 5–68\.External Links:ISSN 0893\-9454, 1465\-7368,[Link](https://academic.oup.com/rfs/article-lookup/doi/10.1093/rfs/hhv059),[Document](https://dx.doi.org/10.1093/rfs/hhv059)Cited by:[§1](https://arxiv.org/html/2608.12841#S1.p1.1),[§3\.1](https://arxiv.org/html/2608.12841#S3.SS1.p1.1)\.
- Hochreiter and Schmidhuber \(1997\)S\. Hochreiter and J\. SchmidhuberLong Short\-Term Memory\.Neural Computation9\(8\),pp\. 1735–1780\.External Links:ISSN 0899\-7667, 1530\-888X,[Link](https://direct.mit.edu/neco/article/9/8/1735-1780/6109),[Document](https://dx.doi.org/10.1162/neco.1997.9.8.1735)Cited by:[§5\.2](https://arxiv.org/html/2608.12841#S5.SS2.p1.1),[§5\.3](https://arxiv.org/html/2608.12841#S5.SS3.p3.1)\.
- Huanget al\.\(2026\)Y\. Huang, Z\. Fan, K\. Hu, and Y\. YeFrom hypotheses to factors: constrained llm agents in cryptocurrency markets\.arXiv\.Note:arXiv:2604\.26747External Links:[Link](http://arxiv.org/abs/2604.26747),[Document](https://dx.doi.org/10.48550/arXiv.2604.26747)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1)\.
- Kabiret al\.\(2025\)M\. R\. Kabir, D\. Bhadra, M\. Ridoy, and M\. MilanovaLSTM–Transformer\-Based Robust Hybrid Deep Learning Model for Financial Time Series Forecasting\.Sci7\(1\),pp\. 7\.External Links:ISSN 2413\-4155,[Link](https://www.mdpi.com/2413-4155/7/1/7),[Document](https://dx.doi.org/10.3390/sci7010007)Cited by:[§1](https://arxiv.org/html/2608.12841#S1.p2.1),[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px3.p1.1),[§5\.2](https://arxiv.org/html/2608.12841#S5.SS2.p1.1)\.
- Kakushadze \(2016\)Z\. Kakushadze101 Formulaic Alphas\.arXiv\.Note:arXiv:1601\.00991External Links:[Link](http://arxiv.org/abs/1601.00991),[Document](https://dx.doi.org/10.48550/arXiv.1601.00991)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2608.12841#S3.SS1.p3.1)\.
- Kirtac and Germano \(2025\)K\. Kirtac and G\. GermanoLarge language models in finance: estimating financial sentiment for stock prediction\.External Links:[Link](https://www.ssrn.com/abstract=5166656),[Document](https://dx.doi.org/10.2139/ssrn.5166656)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px3.p1.1)\.
- Linet al\.\(2026\)Q\. Lin, R\. Feng, Y\. Feng, Z\. Huang, Y\. Chen, Z\. Yang, L\. Zhou, B\. Fei, J\. Liu, and Y\. LiFactorEngine: A Program\-level Knowledge\-Infused Factor Mining Framework for Quantitative Investment\.arXiv\.Note:arXiv:2603\.16365External Links:[Link](http://arxiv.org/abs/2603.16365),[Document](https://dx.doi.org/10.48550/arXiv.2603.16365)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1)\.
- Liuet al\.\(2025\)F\. Liu, Y\. Huang, S\. Luo, Y\. Wang, Y\. Yang, X\. Li, Z\. Hu, J\. Feng, and Q\. LiuCognitive alpha mining via llm\-driven code\-based evolution\.arXiv\.Note:arXiv:2511\.18850External Links:[Link](http://arxiv.org/abs/2511.18850),[Document](https://dx.doi.org/10.48550/arXiv.2511.18850)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1)\.
- Liu \(2026\)T\. LiuA Comparative Study of Transformer\-Based and Classical Models for Financial Time\-Series Forecasting\.Journal of Risk and Financial Management19\(3\),pp\. 203\.External Links:ISSN 1911\-8074,[Link](https://www.mdpi.com/1911-8074/19/3/203),[Document](https://dx.doi.org/10.3390/jrfm19030203)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px3.p1.1)\.
- Luet al\.\(2024\)C\. Lu, C\. Lu, R\. T\. Lange, J\. Foerster, J\. Clune, and D\. HaThe AI Scientist: Towards Fully Automated Open\-Ended Scientific Discovery\.arXiv\.Note:arXiv:2408\.06292External Links:[Link](http://arxiv.org/abs/2408.06292),[Document](https://dx.doi.org/10.48550/arXiv.2408.06292)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px2.p1.1)\.
- Luoet al\.\(2026\)H\. Luo, H\. T\. Ko, J\. Chen, D\. Sun, Y\. Zhang, and C\. LiuAlphaBench: benchmarking large language models in formulaic alpha factor mining\.InInternational Conference on Learning Representations \(ICLR\),Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1)\.
- Miyazakiet al\.\(2026\)K\. Miyazaki, T\. Kawahara, S\. Roberts, and S\. ZohrenToward Expert Investment Teams: A Multi\-Agent LLM System with Fine\-Grained Trading Tasks\.The Journal of Financial Data Science,pp\. jfds\.2026\.008\.External Links:ISSN 2640\-3943,[Link](http://pm-research.com/lookup/doi/10.3905/jfds.2026.008),[Document](https://dx.doi.org/10.3905/jfds.2026.008)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px2.p1.1)\.
- Romera\-Paredeset al\.\(2024\)B\. Romera\-Paredes, M\. Barekatain, A\. Novikov, M\. Balog, M\. P\. Kumar, E\. Dupont, F\. J\. R\. Ruiz, J\. S\. Ellenberg, P\. Wang, O\. Fawzi, P\. Kohli, and A\. FawziMathematical discoveries from program search with large language models\.Nature625\(7995\),pp\. 468–475\.External Links:ISSN 0028\-0836, 1476\-4687,[Link](https://www.nature.com/articles/s41586-023-06924-6),[Document](https://dx.doi.org/10.1038/s41586-023-06924-6)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px2.p1.1)\.
- Shiet al\.\(2026a\)R\. Shi, S\. Yan, Y\. Cai, and C\. LvHubble: An LLM\-Driven Agentic Framework for Safe, Diverse, and Reproducible Alpha Factor Discovery\.arXiv\.Note:arXiv:2604\.09601External Links:[Link](http://arxiv.org/abs/2604.09601),[Document](https://dx.doi.org/10.48550/arXiv.2604.09601)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1),[§4\.2](https://arxiv.org/html/2608.12841#S4.SS2.p1.1)\.
- Shiet al\.\(2026b\)Y\. Shi, Y\. Duan, and J\. LiNavigating the Alpha Jungle: An LLM\-Powered MCTS Framework for Formulaic Alpha Factor Mining\.Proceedings of the AAAI Conference on Artificial Intelligence40\(2\),pp\. 997–1005\.External Links:ISSN 2374\-3468, 2159\-5399,[Link](https://ojs.aaai.org/index.php/AAAI/article/view/37069),[Document](https://dx.doi.org/10.1609/aaai.v40i2.37069)Cited by:[§1](https://arxiv.org/html/2608.12841#S1.p2.1),[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2608.12841#S3.SS1.p3.2)\.
- Singhi \(2025\)A\. SinghiAn Adaptive Multi\-Agent Bitcoin Trading System\.External Links:[Link](https://www.ssrn.com/abstract=5580590),[Document](https://dx.doi.org/10.2139/ssrn.5580590)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px2.p1.1)\.
- Songet al\.\(2025\)Z\. Song, H\. S\. Tsang, R\. T\. Hsung, Y\. Zhu, and W\. LoFrom Market Volatility to Predictive Insight: An Adaptive Transformer–RL Framework for Sentiment\-Driven Financial Time\-Series Forecasting\.Forecasting7\(4\),pp\. 55\.External Links:ISSN 2571\-9394,[Link](https://www.mdpi.com/2571-9394/7/4/55),[Document](https://dx.doi.org/10.3390/forecast7040055)Cited by:[§1](https://arxiv.org/html/2608.12841#S1.p2.1),[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px3.p1.1)\.
- Tanget al\.\(2025\)Z\. Tang, Z\. Chen, J\. Yang, J\. Mai, Y\. Zheng, K\. Wang, J\. Chen, and L\. LinAlphaAgent: LLM\-Driven Alpha Mining with Regularized Exploration to Counteract Alpha Decay\.InProceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V\.2,pp\. 2813–2822\.External Links:[Link](https://dl.acm.org/doi/10.1145/3711896.3736838),[Document](https://dx.doi.org/10.1145/3711896.3736838)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2608.12841#S3.SS1.p3.2),[§4\.2](https://arxiv.org/html/2608.12841#S4.SS2.p1.1)\.
- Tanget al\.\(2026\)Z\. Tang, X\. Yin, W\. Chen, Z\. Chen, Y\. Zheng, W\. Ye, K\. Wang, and L\. LinAlphaAgentEvo: evolution\-oriented alpha mining via self\-evolving agentic reinforcement learning\.InInternational Conference on Learning Representations \(ICLR\),Note:OpenReview lNmZrawUMuCited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1),[§4\.2](https://arxiv.org/html/2608.12841#S4.SS2.p1.1)\.
- Vaswaniet al\.\(2017\)A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, L\. Kaiser, and I\. PolosukhinAttention Is All You Need\.arXiv\.Note:arXiv:1706\.03762External Links:[Link](http://arxiv.org/abs/1706.03762),[Document](https://dx.doi.org/10.48550/arXiv.1706.03762)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px3.p1.1),[§3\.1](https://arxiv.org/html/2608.12841#S3.SS1.p4.1),[§5\.2](https://arxiv.org/html/2608.12841#S5.SS2.p1.1)\.
- Vuet al\.\(2026\)S\. M\. Vu, T\. T\. Pham, and V\. H\. TranSelf\-Improving Alpha Mining for Quantitative Trading via Multi\-Agent Large Language Models with Knowledge Base Accumulation\.External Links:[Link](https://www.ssrn.com/abstract=6906675),[Document](https://dx.doi.org/10.2139/ssrn.6906675)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1),[§4\.2](https://arxiv.org/html/2608.12841#S4.SS2.p1.1)\.
- Wang \(2025\)T\. WangEnhancing Stock Market Prediction with Temporal Graph Neural Networks and Large Language Model\-Based Explainability\.Procedia Computer Science274,pp\. 147–160\.External Links:ISSN 18770509,[Link](https://linkinghub.elsevier.com/retrieve/pii/S1877050925037366),[Document](https://dx.doi.org/10.1016/j.procs.2025.12.015)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px3.p1.1)\.
- Wanget al\.\(2026\)Y\. Wang, J\. Xu, H\. Zhang, S\. Huang, D\. D\. Sun, and X\. ZhangFactorMiner: A Self\-Evolving Agent with Skills and Experience Memory for Financial Alpha Discovery\.arXiv\.Note:arXiv:2602\.14670External Links:[Link](http://arxiv.org/abs/2602.14670),[Document](https://dx.doi.org/10.48550/arXiv.2602.14670)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1)\.
- Wenget al\.\(2026\)Z\. Weng, S\. Zhang, T\. Wang, and Y\. XiaAlphaLogics: A Market Logic\-Driven Multi\-Agent System for Scalable and Interpretable Alpha Factor Generation\.arXiv\.Note:arXiv:2603\.20247External Links:[Link](http://arxiv.org/abs/2603.20247),[Document](https://dx.doi.org/10.48550/arXiv.2603.20247)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1)\.
- Wuet al\.\(2026\)Y\. Wu, C\. Lou, J\. Zhang, S\. Chen, and Y\. YangEvoAlpha: An LLM\-Enhanced Evolutionary Framework for Formulaic Alpha Mining\.InICASSP 2026 \- 2026 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 18572–18576\.External Links:[Link](https://ieeexplore.ieee.org/document/11463591/),[Document](https://dx.doi.org/10.1109/ICASSP55912.2026.11463591)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1)\.
- Xinet al\.\(2026\)A\. Xin, J\. Siow, J\. Wang, Z\. Yao, F\. Zhang, J\. Song, L\. Hou, and J\. LiEurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery\.arXiv\.Note:arXiv:2606\.13662External Links:[Link](http://arxiv.org/abs/2606.13662),[Document](https://dx.doi.org/10.48550/arXiv.2606.13662)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px2.p1.1)\.
- Yiet al\.\(2026\)J\. Yi, J\. Yang, Y\. Jin, Y\. Li, and J\. LiAlphaSchema: exploring the space of trading semantics for llm\-based alpha mining\.arXiv\.Note:arXiv:2607\.26642External Links:[Link](http://arxiv.org/abs/2607.26642),[Document](https://dx.doi.org/10.48550/arXiv.2607.26642)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1)\.
- Yuet al\.\(2026a\)H\. Yu, Z\. Zheng, J\. Z\. Pan, T\. Liu, Z\. Wang, and F\. HeAlphaMemo: Structured Search\-Process Memory for Self\-Evolving Alpha Mining Agents\.arXiv\.Note:arXiv:2606\.20625External Links:[Link](http://arxiv.org/abs/2606.20625),[Document](https://dx.doi.org/10.48550/arXiv.2606.20625)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1)\.
- Yuet al\.\(2023\)S\. Yu, H\. Xue, X\. Ao, F\. Pan, J\. He, D\. Tu, and Q\. HeGenerating Synergistic Formulaic Alpha Collections via Reinforcement Learning\.arXiv\.Note:arXiv:2306\.12964External Links:[Link](http://arxiv.org/abs/2306.12964),[Document](https://dx.doi.org/10.48550/arXiv.2306.12964)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2608.12841#S3.SS1.p3.2)\.
- Yuet al\.\(2026b\)X\. Yu, Y\. Fu, M\. Fan, E\. Li, Y\. Gao, and S\. XuTowards autonomous formulaic alpha discovery: an evolutionary computation perspective\.arXiv\.Note:arXiv:2608\.01789External Links:[Link](http://arxiv.org/abs/2608.01789),[Document](https://dx.doi.org/10.48550/arXiv.2608.01789)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1)\.
- Zhanget al\.\(2026\)L\. Zhang, T\. Jia, Y\. Zhai, Z\. Xie, C\. Duan, M\. He, P\. S\. Yu, and Y\. LiFrom Feedback Loops to Policy Updates: Reinforcement Fine\-Tuning for LLM\-Based Alpha Factor Discovery\.arXiv\.Note:arXiv:2605\.15412External Links:[Link](http://arxiv.org/abs/2605.15412),[Document](https://dx.doi.org/10.48550/arXiv.2605.15412)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1)\.
- Zhanget al\.\(2020\)T\. Zhang, Y\. Li, Y\. Jin, and J\. LiAutoAlpha: an Efficient Hierarchical Evolutionary Algorithm for Mining Alpha Factors in Quantitative Investment\.arXiv\.Note:arXiv:2002\.08245External Links:[Link](http://arxiv.org/abs/2002.08245),[Document](https://dx.doi.org/10.48550/arXiv.2002.08245)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2608.12841#S3.SS1.p3.1),[§3\.1](https://arxiv.org/html/2608.12841#S3.SS1.p3.2)\.
- Zhanget al\.\(2019\)Z\. Zhang, S\. Zohren, and S\. RobertsDeepLOB: Deep Convolutional Neural Networks for Limit Order Books\.arXiv\.Note:arXiv:1808\.03668External Links:[Link](http://arxiv.org/abs/1808.03668),[Document](https://dx.doi.org/10.48550/arXiv.1808.03668)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px3.p1.1),[§5\.2](https://arxiv.org/html/2608.12841#S5.SS2.p1.1)\.
- Zhaoet al\.\(2025a\)J\. Zhao, C\. Zhang, M\. Qin, and P\. YangQuantFactor REINFORCE: Mining Steady Formulaic Alpha Factors With Variance\-Bounded REINFORCE\.IEEE Transactions on Signal Processing73,pp\. 2448–2463\.External Links:ISSN 1053\-587X, 1941\-0476,[Link](https://ieeexplore.ieee.org/document/11024173/),[Document](https://dx.doi.org/10.1109/TSP.2025.3576781)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2608.12841#S3.SS1.p3.2)\.
- Zhaoet al\.\(2025b\)J\. Zhao, C\. Zhang, C\. Wang, and P\. YangLearning from Expert Factors: Trajectory\-level Reward Shaping for Formulaic Alpha Mining\.arXiv\.Note:arXiv:2507\.20263External Links:[Link](http://arxiv.org/abs/2507.20263),[Document](https://dx.doi.org/10.48550/arXiv.2507.20263)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px1.p1.1)\.
- Zhouet al\.\(2025\)L\. Zhou, H\. Ling, C\. Fu, Y\. Huang, M\. Sun, W\. Yu, X\. Wang, X\. Li, X\. Su, J\. Zhang, X\. Chen, C\. Liang, X\. Qian, H\. Ji, W\. Wang, M\. Zitnik, and S\. JiAutonomous Agents for Scientific Discovery: Orchestrating Scientists, Language, Code, and Physics\.arXiv\.Note:arXiv:2510\.09901External Links:[Link](http://arxiv.org/abs/2510.09901),[Document](https://dx.doi.org/10.48550/arXiv.2510.09901)Cited by:[§2](https://arxiv.org/html/2608.12841#S2.SS0.SSS0.Px2.p1.1)\.

## Appendix APart I iteration record and walkthrough

This appendix gives the full version of the worked iteration summarized in Section[4](https://arxiv.org/html/2608.12841#S4)\. Listing[3](https://arxiv.org/html/2608.12841#LST3)shows the record the system stores for a single iteration, and the text below traces two iterations in detail\.

Listing 3:One Part I iteration record\. The Idea Miner emits a falsifiable proposal rather than a bare formula\. The evaluator then attaches controlled tests, direction calibration, and a memory update\. Variable names are representative; deployed expressions are withheld\.run:

goal:"Afteranopen\-interestcrash,doesaweakreboundpredictreversal?"

event\_type:open\_interest\_crash

universe:BTCUSDT\_5m

manager\_plan:

visual\_search:

trigger:

\-sharp\_drop\(open\_interest\)

\-sharp\_drop\(open\_interest\_value\)

\-elevated\(volume\)

context\_fields:

\-price

\-volume

\-basis\_rate

\-taker\_buy\_sell\_ratio

\-long\_short\_account\_ratio

\-top\_trader\_position\_ratio

target\_labels:\[ret\_open\_open\_h10,ret\_open\_open\_h30\]

instruction:"Testwhetherreboundqualityafterdeleveragingseparatescontinuationfromfailure\."

visual\_observation:

event\_profile:

\-openinterestfallsabruptlyduringahigh\-volumeunwind

\-priceoftenreboundsafterforcedpressurefades

\-reboundswithweaktaker\-flowconfirmationfrequentlystall

\-basisrecoveryandpositioningresetseparatecleanreboundsfromfailedones

proposal:

proposal\_id:oi\_crash\_weak\_rebound\_flow\_gap

hypothesis:\>

Afteraforcedopen\-interestunwind,apricereboundthatisnotconfirmed

byaggressivetakerflowismorelikelytofail\.

mechanism:\>

Deleveragingremovesforcedpressure,butweakbuy\-flowandpoorbasis

recoveryindicateinsufficientdemandaftertherebound\.

expected\_direction:higher\_signal\_predicts\_lower\_future\_return

factor\_blueprint:

\-deleveraging\_intensity:rankednegativechangeinopeninterest

\-rebound\_strength:short\-horizonpricerecoveryaftertheevent

\-flow\_gap:lackoftaker\-flowconfirmationduringtherebound

\-basis\_filter:weakorcompressedbasis\-raterecovery

expression:withheld

evaluation\_contract:

primary\_labels:

\-ret\_open\_open\_h10

\-ret\_open\_open\_h30

baselines:

\-short\_horizon\_price\_momentum

\-open\_interest\_change

\-taker\_flow\_change

\-basis\_rate\_change

controlled\_tests:

\-event\_window\_ic\_vs\_control\_window\_ic

\-full\_factor\_vs\_best\_component

\-monthly\_ic\_stability

\-correlation\_with\_existing\_factor\_pool

\-proposed\_direction\_vs\_reversed\_direction

structured\_observation:

verdict:selected\_for\_factor\_pool

single\_factor\_ic\_range:approximately\_0\.026\_to\_0\.037

finding:\>

Themechanismisstrongestwhentheopen\-interestshockisfollowedby

weakreboundacceptanceandpoortaker\-flowconfirmation\.

memory\_update:

belief:"OIcrashreboundswithoutflowconfirmationaremorelikelytofail\."

action:"Increasepriorityofdeleveraging\-plus\-flow\-gapmechanismsinlaterruns\."

Listing[3](https://arxiv.org/html/2608.12841#LST3)traces one iteration from a research question to a stored belief\. The run asks whether a weak rebound after an open\-interest crash predicts reversal\. The AI Manager converts this question into a concrete plan: search for deleveraging episodes, inspect price and flow behavior after the unwind, and test whether rebound quality separates continuation from failure\.

The Visual Analyst returns event profiles rather than a single chart\. Some episodes show a clean forced\-deleveraging pattern: open interest falls quickly, volume expands, basis compresses, and price rebounds only briefly\. Other episodes show a healthier reset, where taker flow and basis recover with price\. The Idea Miner uses this distinction to propose mechanisms such as weak rebound after deleveraging, flow\-confirmed continuation, and crowded\-position unwind\. The exact expressions are withheld, but the generated factors combine open\-interest shocks, short\-horizon price response, taker\-flow imbalance, basis behavior, and positioning divergence\.

The Factor Evaluator then scores each proposal against multiple horizons\. In this family of runs, the strongest open\-interest\-crash example selected by the system was an “OI crash rebound flow gap” mechanism, which tests whether a rebound after an open\-interest shock is unsupported by aggressive flow\. Related proposals tested whether open\-interest value rises without price acceptance, whether top\-trader positioning diverges from broader account ratios, and whether taker\-flow confirmation changes the sign of short\-horizon price continuation\. The best single\-factor examples in this family reached information coefficients on the order of0\.0260\.026–0\.0370\.037depending on the label and event context, and the selected signals were then passed to the combination layer\.

A second iteration illustrates how the same loop changes research direction\. The goal is changed to quiet\-market volume expansion: the system searches for low\-volatility, low\-activity periods followed by a sudden burst in volume, quote volume, and trade count\. The Visual Analyst rejects generic already\-volatile high\-volume cascades and focuses on true quiet\-to\-active transitions\. The Idea Miner then proposes continuation factors when the burst is accepted by price, taker flow, and open interest, and reversal factors when the burst has a weak candle body, noisy flow, or no basis confirmation\. This shows how the same architecture can explore a new market mechanism without changing the underlying evaluator\.

## Appendix BA Part II failure case: leakage that survived agent review

The sealed sandbox of Section[3\.1](https://arxiv.org/html/2608.12841#S3.SS1)is not the design AQuA started from\. It is the response to concrete failures of an earlier, more permissive loop in which the agent could write feature and factor code directly and a second agent reviewed each candidate for look\-ahead bias\. We record the most instructive failure here, because it is the reason AQuA constrains the agent to a fixed operator registry rather than trusting review\.

In the earlier loop, the agent authored each feature as code and a separate reviewer agent checked it for causality before training\. One proposed feature was an intraday volume\-participation ratio: the volume traded from the open up to the current minute, divided by a daily volume normalizer\. The intent is causal and the description reads as backward\-looking, so the reviewer agent approved it\. The implementation, however, normalized by the current day’s total volume, a sum that runs from the open through the close\. The denominator therefore depended on bars after the current minute, and the feature quietly encoded end\-of\-day information into every intraday timestamp\. The same failure appeared in a multi\-resolution variant whose daily branch aggregated all of the current day’s bars and was then read at mid\-day timestamps\.

The symptom was a held\-out information coefficient far above what comparable price\-volume features produced, and it did not survive a clean re\-split of the evaluation window\. A manual audit traced it to the full\-day denominator\. The reviewer agent had reasoned about the feature’s economic intent, a ratio of past volume, rather than the exact set of bars its implementation touched, a blind spot it shared with the author agent\.

The lesson is that an LLM reviewing LLM\-written code is advisory, not structural: the author and the reviewer share the same failure modes, so a subtle temporal\-footprint bug can pass both\. AQuA’s response is operatorization\. The agent no longer writes feature or factor code\. It composes a fixed registry of causal operators in which every time\-series operator reads only a trailing window ending at the current timestamp and every cross\-sectional operator reads only the current timestamp\. Causality is then closed under composition, as described in Section[3\.1](https://arxiv.org/html/2608.12841#S3.SS1), and a full\-day normalizer is not expressible in the specification space at all\. The guarantee moves from “the reviewer should catch leakage” to “leakage cannot be written\.”

Similar Articles

QuantAgent: Price-Driven Multi-Agent LLMs for High-Frequency Trading

Papers with Code Trending

QuantAgent is a multi-agent LLM framework designed specifically for high-frequency trading, using four specialized agents (Indicator, Pattern, Trend, Risk) to make rapid, risk-aware decisions based on short-horizon signals. In zero-shot evaluations across ten financial instruments including Bitcoin and Nasdaq futures, it outperforms existing neural and rule-based baselines in predictive accuracy and cumulative return.