Algometrics: Forecasting Under Algorithmic Feedback
Summary
This paper introduces algometrics, a framework for time series forecasting under algorithmic feedback, proving that deployment risk differs from historical risk and is not identifiable from passive data alone. It provides methods for estimating deployment risk using interventions or randomized actions.
View Cached Full Text
Cached at: 05/26/26, 08:58 AM
# Algometrics: Forecasting Under Algorithmic Feedback
Source: [https://arxiv.org/html/2605.23978](https://arxiv.org/html/2605.23978)
###### Abstract
In algorithmic markets, predictive models become part of the data\-generating process they aim to forecast\. Once their outputs are converted into trades, allocations, execution schedules, or risk controls, they change the future data on which they are evaluated\. I introduce algometrics, a framework for time series whose evolution depends on the predictive algorithms forecasting them\. The framework distinguishes historical risk, measured under passive forecasting, from deployment risk, measured when forecasts drive actions\. I prove three results\. First, deployment risk is not identifiable from passive historical data alone: even in a one\-step linear feedback model, infinitely many algorithm\-mediated environments induce the same historical law while implying different deployment risks for the same forecaster\. Second, historical model rankings can invert under crowding, so a predictor with lower passive error can have higher deployment error once similar algorithms are adopted\. Third, randomized or instrumented actions identify short\-horizon linear feedback, and I derive a finite\-sample bound for deployment\-risk estimation\. These results suggest that time\-series benchmarks in algorithmic markets should report feedback sensitivity alongside predictive accuracy\.
## 1Introduction
Financial markets are a natural destination for machine learning because they are high\-dimensional, weakly predictive, and noisy\. A growing empirical and theoretical literature argues that model complexity can be valuable in return prediction: many weak signals, interactions, and nonlinearities may be invisible to parsimonious models but useful to modern ML systems\(Guet al\.,[2020](https://arxiv.org/html/2605.23978#bib.bib5); Kellyet al\.,[2024](https://arxiv.org/html/2605.23978#bib.bib4)\)\. Finance is not just a domain where overfitting is easy; it is one where controlled complexity can be productive\.
This paper studies a different source of complexity\. In financial markets, forecasts are rarely inert\. A return prediction becomes an order or a portfolio weight; a credit score changes funding access; a liquidity forecast changes execution; a risk model changes leverage and margin; a widely copied signal changes the price process on which future models train\. The statistical object is a time series partly generated by the algorithms that forecast it\.
I call this problem*algometrics*: the measurement and learning of time series under algorithmic feedback\. It refers to settings in which \(i\) predictive algorithms observe a history, \(ii\) their outputs are converted into actions, \(iii\) actions influence future observations, and \(iv\) a population of related algorithms may create crowding or strategic adaptation\. Financial markets are the central example; the same pattern arises in online advertising, pricing, recommender systems, platform labor, and cybersecurity\.
The paper makes three claims\. First, historical risk and deployment risk are different estimands\. A model can be accurate when it passively forecasts a market and inaccurate when it becomes part of the market’s order flow\. Second, the difference is not a small\-sample nuisance\. Without variation in algorithmic actions, deployment risk is generally not identifiable from passive histories\. Third, evaluation is still possible if the data include interventions, randomized exposure, simulator access with calibrated impact, or other instruments that reveal how outcomes respond to algorithmic actions\.
The contribution is an identification argument\. I formalize an*algorithm\-mediated time series*as a sequence whose transition kernel depends on algorithm\-induced actions\. I define the*feedback gap*between passive and deployment risk\. I then prove two negative results and one positive result\. The first negative result shows non\-identifiability: in a one\-step linear feedback model, all values of the feedback coefficient induce the same passive historical distribution, but they imply different deployment risks\. The second shows ranking inversion: the historically best predictor can become worse than a conservative predictor after crowding\. The positive result shows that randomized actions identify a short\-horizon linear feedback coefficient and yield a finite\-sample deployment\-risk bound\.
The contribution type is theory\. The closed\-form illustration in Section[7](https://arxiv.org/html/2605.23978#S7)visualizes the mechanism; the main claims are mathematical\. The message is that time\-series benchmarks in algorithmic environments should report historical prediction scores together with the assumptions under which those scores extrapolate to deployment\.
## 2From statistical complexity to algorithmic feedback
The case for complex ML in finance is often framed as a response to parsimony\. Classical equity\-premium prediction struggled to beat simple baselines\(Welch and Goyal,[2008](https://arxiv.org/html/2605.23978#bib.bib6)\)\. Recent ML work has revisited that conclusion by exploiting large cross\-sections, nonlinearities, interactions, and regularization\(Guet al\.,[2020](https://arxiv.org/html/2605.23978#bib.bib5)\)\.Kellyet al\.\([2024](https://arxiv.org/html/2605.23978#bib.bib4)\)argue more directly for a “virtue of complexity” in return prediction: complex models can recover predictive structure that small models miss\. Skeptical perspectives note that ML performance in finance is sensitive to economic restrictions\(Avramovet al\.,[2023](https://arxiv.org/html/2605.23978#bib.bib47); Martin and Nagel,[2022](https://arxiv.org/html/2605.23978#bib.bib46)\)and to multiple\-testing concerns\(Harveyet al\.,[2016](https://arxiv.org/html/2605.23978#bib.bib17)\), and tree\-based approaches deliver gains comparable to deep models with greater interpretability\(Bryzgalovaet al\.,[2025](https://arxiv.org/html/2605.23978#bib.bib49)\)\. Algometrics is complementary to both sides of this debate: even if the predictive case for complexity is granted, the deployment\-relevant case additionally requires modeling how the predictor changes the prediction target\.
Algometrics adds an endogenous layer\. A model’s complexity affects not only what it learns from history but also how it acts on the market\. A high\-capacity model may trade more aggressively, coordinate unintentionally with models trained on similar data, or produce signals whose value decays once others imitate them\. These effects are not captured by the usual historical loss
R0\(f\)=𝔼\[ℓ\(f\(Ht\),Yt\+1\)\],R\_\{0\}\(f\)=\\mathbb\{E\}\\big\[\\ell\(f\(H\_\{t\}\),Y\_\{t\+1\}\)\\big\],\(1\)whereHtH\_\{t\}is the observed history andYt\+1Y\_\{t\+1\}is the next return, price change, spread, default indicator, or other target\. The loss treats the target as external to the forecast\.
The concern is related to performative prediction, where predictions influence the distribution of outcomes\(Perdomoet al\.,[2020](https://arxiv.org/html/2605.23978#bib.bib7); Mendler\-Dünneret al\.,[2020](https://arxiv.org/html/2605.23978#bib.bib8); Milleret al\.,[2021](https://arxiv.org/html/2605.23978#bib.bib9)\)\. It also connects to strategic classification\(Hardtet al\.,[2016](https://arxiv.org/html/2605.23978#bib.bib10)\), market impact\(Kyle,[1985](https://arxiv.org/html/2605.23978#bib.bib11); Almgren and Chriss,[2001](https://arxiv.org/html/2605.23978#bib.bib12)\), adaptive markets\(Lo,[2004](https://arxiv.org/html/2605.23978#bib.bib14)\), herding\(Cont and Bouchaud,[2000](https://arxiv.org/html/2605.23978#bib.bib13)\), and agent\-based computational finance\(LeBaron,[2006](https://arxiv.org/html/2605.23978#bib.bib27)\)\. The distinct feature here is the time\-series estimand: historical data are often collected under one algorithmic regime, while deployment occurs under another\.
Table 1:Core objects in algometrics\. The terms are operational: each points to an estimand or diagnostic for time\-series learning under feedback\.This framing does not deny the value of historical benchmarks\. They are essential for detecting leakage, survivorship bias, overfitting, and unrealistic cost assumptions\. The claim is narrower: a benchmark that estimatesR0\(f\)R\_\{0\}\(f\)should not be interpreted as estimatingRm\(f\)R\_\{m\}\(f\)unless the feedback gap is argued to be negligible, bounded, or measured\.
## 3Algorithm\-mediated time series
LetHt=\(Y1,A1,…,Yt−1,At−1,Yt\)H\_\{t\}=\(Y\_\{1\},A\_\{1\},\\ldots,Y\_\{t\-1\},A\_\{t\-1\},Y\_\{t\}\)denote the observed history up to and including the time\-ttoutcome but before the time\-ttaction\. A forecasterf∈ℱf\\in\\mathcal\{F\}maps histories to predictionsY^t\+1=f\(Ht\)\\hat\{Y\}\_\{t\+1\}=f\(H\_\{t\}\)\. A deployment mapπ\\piconverts a forecast into an actionAt=π\(f,Ht,mt\)A\_\{t\}=\\pi\(f,H\_\{t\},m\_\{t\}\), wheremt∈\[0,1\]m\_\{t\}\\in\[0,1\]is the adoption, capital, or traffic share affected by the forecaster\. The convention orders observation, prediction, and action within each period:YtY\_\{t\}is observed first,ffproducesY^t\+1\\hat\{Y\}\_\{t\+1\}fromHtH\_\{t\}, andAtA\_\{t\}is then induced;AtA\_\{t\}entersHt\+1H\_\{t\+1\}, notHtH\_\{t\}\. In a market,AtA\_\{t\}may represent demand, order size, portfolio weight, leverage, cancellation intensity, or an execution schedule\.
###### Definition 1\(Algorithm\-mediated time series\)\.
An algorithm\-mediated time series is a family of transition kernels
Pη\(Yt\+1∈dy,St\+1∈ds∣Ht,St,At,mt\),P\_\{\\eta\}\\big\(Y\_\{t\+1\}\\in dy,S\_\{t\+1\}\\in ds\\mid H\_\{t\},S\_\{t\},A\_\{t\},m\_\{t\}\\big\),\(2\)whereStS\_\{t\}is a latent or observed state,AtA\_\{t\}is an action induced by one or more algorithms,mtm\_\{t\}is an adoption mass, andη\\etaindexes environment parameters such as market impact, liquidity, or strategic response\. The passive regime setsAt=0A\_\{t\}=0andmt=0m\_\{t\}=0\. A deployment regime usesAt=π\(f,Ht,mt\)A\_\{t\}=\\pi\(f,H\_\{t\},m\_\{t\}\)\.
For a lossℓ\\ell, historical risk is
R0\(f\)=𝔼Pη0\[ℓ\(f\(Ht\),Yt\+1\)\],R\_\{0\}\(f\)=\\mathbb\{E\}\_\{P\_\{\\eta\}^\{0\}\}\\left\[\\ell\(f\(H\_\{t\}\),Y\_\{t\+1\}\)\\right\],\(3\)wherePη0P\_\{\\eta\}^\{0\}is the law induced by passive observation\. Deployment risk at adoptionmmis
Rm\(f\)=𝔼Pηf,m\[ℓ\(f\(Ht\),Yt\+1\)\],R\_\{m\}\(f\)=\\mathbb\{E\}\_\{P\_\{\\eta\}^\{f,m\}\}\\left\[\\ell\(f\(H\_\{t\}\),Y\_\{t\+1\}\)\\right\],\(4\)wherePηf,mP\_\{\\eta\}^\{f,m\}is the law induced whenffis acted upon\. The feedback gap is
Γm\(f\)=Rm\(f\)−R0\(f\)\.\\Gamma\_\{m\}\(f\)=R\_\{m\}\(f\)\-R\_\{0\}\(f\)\.\(5\)For reward objectives, such as expected return or Sharpe\-like scores, one can define the gap with the sign reversed\. The mathematical issue is the same: a passive score and a deployment score are different estimands\. The trajectory definition above admits a useful one\-step specialization: the*one\-step deployment risk*is
Rm\(1\)\(f\)=𝔼Ht∼Pη0\[𝔼\[ℓ\(f\(Ht\),Yt\+1\)\|Ht,At=π\(f,Ht,m\)\]\],R\_\{m\}^\{\(1\)\}\(f\)=\\mathbb\{E\}\_\{H\_\{t\}\\sim P\_\{\\eta\}^\{0\}\}\\\!\\left\[\\mathbb\{E\}\\\!\\left\[\\ell\(f\(H\_\{t\}\),Y\_\{t\+1\}\)\\,\\big\|\\,H\_\{t\},\\,A\_\{t\}=\\pi\(f,H\_\{t\},m\)\\right\]\\right\],\(6\)where the outer expectation draws histories from the passive law and only the time\-ttaction is induced by the deployment policy\. Theorems[1](https://arxiv.org/html/2605.23978#Thmtheorem1)and[2](https://arxiv.org/html/2605.23978#Thmtheorem2), and Proposition[1](https://arxiv.org/html/2605.23978#Thmproposition1), are statements aboutRm\(1\)\(f\)R\_\{m\}^\{\(1\)\}\(f\)\. Comparing deployment risks across forecasters is a policy comparison: each forecasterffis evaluated under the data\-generating processPηf,mP\_\{\\eta\}^\{f,m\}induced by its own deployment, not on a common realized label sequence\. This is not a confound; it is the object of study\. In algorithmic environments the future depends on the policy, so a deployment\-risk comparison is necessarily a comparison of policies and of the worlds they induce\.
A useful local diagnostic is algorithmic elasticity\. For a metricddon distributions,
ℰ\(Ht;a,a′\)=d\(Pη\(Yt\+1∣Ht,At=a\),Pη\(Yt\+1∣Ht,At=a′\)\)‖a−a′‖,\\mathcal\{E\}\(H\_\{t\};a,a^\{\\prime\}\)=\\frac\{d\\left\(P\_\{\\eta\}\(Y\_\{t\+1\}\\mid H\_\{t\},A\_\{t\}=a\),\\,P\_\{\\eta\}\(Y\_\{t\+1\}\\mid H\_\{t\},A\_\{t\}=a^\{\\prime\}\)\\right\)\}\{\\left\\lVert a\-a^\{\\prime\}\\right\\rVert\},\(7\)whenevera≠a′a\\neq a^\{\\prime\}, wherePη\(Yt\+1∣Ht,At\)=∫Pη\(Yt\+1∣Ht,St,At\)Pη\(dSt∣Ht\)P\_\{\\eta\}\(Y\_\{t\+1\}\\mid H\_\{t\},A\_\{t\}\)=\\int P\_\{\\eta\}\(Y\_\{t\+1\}\\mid H\_\{t\},S\_\{t\},A\_\{t\}\)\\,P\_\{\\eta\}\(dS\_\{t\}\\mid H\_\{t\}\)marginalizes the latent state under the prevailing regime\. High elasticity means that the target is sensitive to actions induced by forecasts\. In finance, this can be caused by price impact, liquidity withdrawal, signal crowding, or strategic response\. In low\-elasticity regimes, historical risk may be a reasonable proxy for deployment risk\. In high\-elasticity regimes, the proxy is suspect\.
### 3\.1When passive evaluation is justified
The framework does not imply that every historical benchmark is invalid\. It identifies the condition under which historical evaluation is doing more than passive prediction: the action\-induced distributional change must be small relative to the loss\. A simple bound makes this explicit\.
###### Proposition 1\(Small\-feedback bound\)\.
Fix a one\-step evaluation conditional on histories drawn from the passive law\. Fix a ground metricdYd\_\{Y\}on the outcome space and letW1W\_\{1\}denote the corresponding11\-Wasserstein distance\. Suppose the lossℓ\(y^,y\)\\ell\(\\hat\{y\},y\)isLℓL\_\{\\ell\}\-Lipschitz inyywith respect todYd\_\{Y\}, and that
W1\(P\(Yt\+1∣Ht,At=a\),P\(Yt\+1∣Ht,At=0\)\)≤κ\(Ht\)‖a‖W\_\{1\}\\left\(P\(Y\_\{t\+1\}\\mid H\_\{t\},A\_\{t\}=a\),P\(Y\_\{t\+1\}\\mid H\_\{t\},A\_\{t\}=0\)\\right\)\\leq\\kappa\(H\_\{t\}\)\\left\\lVert a\\right\\rVert\(8\)for all feasible actionsaa\. Writingπf\(Ht\):=π\(f,Ht,m\)\\pi\_\{f\}\(H\_\{t\}\):=\\pi\(f,H\_\{t\},m\)for the deployment policy at fixed adoptionmmandΓ\(f\):=Γm\(1\)\(f\)\\Gamma\(f\):=\\Gamma\_\{m\}^\{\(1\)\}\(f\)for the one\-step feedback gap, we have
\|Γ\(f\)\|≤Lℓ𝔼\[κ\(Ht\)‖πf\(Ht\)‖\]\.\|\\Gamma\(f\)\|\\leq L\_\{\\ell\}\\,\\mathbb\{E\}\\left\[\\kappa\(H\_\{t\}\)\\left\\lVert\\pi\_\{f\}\(H\_\{t\}\)\\right\\rVert\\right\]\.\(9\)
The proposition is useful because it separates two questions that are often mixed together\. A model can be statistically complex while having small market footprint, in which case passive evaluation may be adequate\. Conversely, a simple model can have a large feedback gap if it controls enough capital or traffic\. Complexity of representation and complexity of deployment are distinct axes\.
## 4Passive histories do not identify deployment risk
The first result shows that the feedback gap cannot in general be recovered from passive histories\. The theorem is intentionally elementary\. It is stronger to show failure in a one\-step linear model than to rely on complicated market dynamics\.
###### Theorem 1\(Passive non\-identifiability\)\.
Consider the one\-step linear feedback family
Yt\+1=μ\(Ht\)\+βAt\+εt\+1,𝔼\[εt\+1∣Ht,At\]=0,Y\_\{t\+1\}=\\mu\(H\_\{t\}\)\+\\beta A\_\{t\}\+\\varepsilon\_\{t\+1\},\\qquad\\mathbb\{E\}\[\\varepsilon\_\{t\+1\}\\mid H\_\{t\},A\_\{t\}\]=0,\(10\)with squared loss and finite second moments\. In the passive regimeAt=0A\_\{t\}=0\. Fix any forecasterffand deployment policyAt=πf\(Ht\)A\_\{t\}=\\pi\_\{f\}\(H\_\{t\}\)satisfying𝔼\[πf\(Ht\)2\]\>0\\mathbb\{E\}\[\\pi\_\{f\}\(H\_\{t\}\)^\{2\}\]\>0under the passive history law\. Then all values ofβ\\betainduce the same passive distribution over\(Ht,Yt\+1\)\(H\_\{t\},Y\_\{t\+1\}\), but the one\-step deployment risk offfis a nonconstant quadratic function ofβ\\beta\. Consequently, no estimator based only on passive histories can identify deployment risk uniformly over this family\.
The proof is in Appendix[A](https://arxiv.org/html/2605.23978#A1)\. The intuition is that passive data only revealμ\(Ht\)\\mu\(H\_\{t\}\)becauseAtA\_\{t\}never varies\. The coefficientβ\\betais invisible until an action is taken\. The result is structurally the time\-series and feedback analog ofManski \([1993](https://arxiv.org/html/2605.23978#bib.bib28)\)’s reflection problem: observed outcomes alone cannot separate the model’s effect on the world from the world’s effect on the model’s inputs without exogenous variation\.
Figure 1:Closed\-form illustration of Theorem[1](https://arxiv.org/html/2605.23978#Thmtheorem1)\. In panel \(a\), passive observations are generated withA=0A=0, so the conditional lawY∣H,A=0Y\\mid H,A=0is the same for all feedback coefficientsβ\\beta\. Passive data therefore revealμ\(H\)\\mu\(H\)but not the action effect\. In panel \(b\), the same forecasterf\(h\)=hf\(h\)=his deployed with actionA=f\(H\)A=f\(H\)\. The deployment riskRβ\(f\)R\_\{\\beta\}\(f\)varies withβ\\beta, even though the passive law in panel \(a\) is unchanged\. Thus passive historical data do not identify deployment risk\.Figure[1](https://arxiv.org/html/2605.23978#S4.F1)visualizes this concretely: the same observable passive law is consistent with infinitely many feedback coefficients, each producing a different deployment risk for the same forecaster\.
The theorem does not claim that feedback is large in every financial model\. It says that passive data cannot identify the feedback coefficient: if the evaluated action was absent from the history, or present at a different adoption share, additional passive observation will not recover it\. The estimand on offer is the wrong one\.
## 5Historical rankings can invert after crowding
The second result shows that feedback can change model rankings\. This matters because benchmarks usually select models by comparing scores\. If a leaderboard ranks passive risk but users care about deployment risk, the selected model can be the wrong one\.
###### Theorem 2\(Crowding\-induced ranking inversion\)\.
LetH∼N\(0,1\)H\\sim N\(0,1\)andε∼N\(0,σ2\)\\varepsilon\\sim N\(0,\\sigma^\{2\}\)be independent\. Under passive observation,
Y0=H\+ε\.Y^\{0\}=H\+\\varepsilon\.\(11\)Forc∈\[0,1\]c\\in\[0,1\], define the forecasterfc\(H\)=cHf\_\{c\}\(H\)=cHand suppose deployment induces actionA=αfc\(H\)A=\\alpha f\_\{c\}\(H\)with adoption intensityα≥0\\alpha\\geq 0\. Let the deployed target be
Yc,α=H−γαcH\+ε,Y^\{c,\\alpha\}=H\-\\gamma\\alpha cH\+\\varepsilon,\(12\)whereγ\>0\\gamma\>0is a crowding or negative\-impact coefficient\. Then for everyc∈\[0,1\]c\\in\[0,1\],
R0\(fc\)=\(c−1\)2\+σ2,Rα\(fc\)=\(c\(1\+γα\)−1\)2\+σ2\.R\_\{0\}\(f\_\{c\}\)=\(c\-1\)^\{2\}\+\\sigma^\{2\},\\qquad R\_\{\\alpha\}\(f\_\{c\}\)=\\big\(c\(1\+\\gamma\\alpha\)\-1\\big\)^\{2\}\+\\sigma^\{2\}\.\(13\)Fix any conservative predictorfc′f\_\{c^\{\\prime\}\}withc′∈\[0,1\)c^\{\\prime\}\\in\[0,1\)\. Thenf1f\_\{1\}has strictly lower passive risk thanfc′f\_\{c^\{\\prime\}\},
R0\(f1\)=σ2<\(c′−1\)2\+σ2=R0\(fc′\),R\_\{0\}\(f\_\{1\}\)=\\sigma^\{2\}<\(c^\{\\prime\}\-1\)^\{2\}\+\\sigma^\{2\}=R\_\{0\}\(f\_\{c^\{\\prime\}\}\),\(14\)butf1f\_\{1\}has strictly higher deployment risk thanfc′f\_\{c^\{\\prime\}\}whenever
γα\>1−c′1\+c′\.\\gamma\\alpha\>\\frac\{1\-c^\{\\prime\}\}\{1\+c^\{\\prime\}\}\.\(15\)The casec′=0c^\{\\prime\}=0recovers the thresholdγα\>1\\gamma\\alpha\>1\. The casec′=0\.25c^\{\\prime\}=0\.25used in Figure[2](https://arxiv.org/html/2605.23978#S7.F2)gives the thresholdγα\>0\.6\\gamma\\alpha\>0\.6\.
This stylized theorem captures a documented market phenomenon\. Published anomalies systematically lose predictive power once they become widely known\(McLean and Pontiff,[2016](https://arxiv.org/html/2605.23978#bib.bib15); Falcket al\.,[2022](https://arxiv.org/html/2605.23978#bib.bib50)\), and the August 2007 quant crisis is the canonical real\-world example of correlated strategies failing simultaneously when many funds deploy similar signals\(Khandani and Lo,[2011](https://arxiv.org/html/2605.23978#bib.bib44)\)\. The accurate signal is valuable when it is passive and scarce; if enough capital trades on it, the induced demand can move prices against the signal, consume liquidity, or accelerate the decay of the opportunity\. A conservative predictor that looks worse in passive loss can be safer under adoption\. The theorem is not about the particular linear form\. It shows why a single historical score is an incomplete decision rule for model selection\.
## 6A positive result: instrumented feedback estimation
The negative results do not imply that deployment risk is unknowable\. They imply that some variation in actions is needed\. Such variation can come from randomized order slicing, staggered deployment, A/B exposure, exogenous instrumented demand, or calibrated simulators\. The result below is in spirit an instrumented\-identification statement\(Imbens and Angrist,[1994](https://arxiv.org/html/2605.23978#bib.bib41); Harriset al\.,[2022](https://arxiv.org/html/2605.23978#bib.bib3)\), in which exogenous variation in actions identifies a causal feedback coefficient that passive observation cannot recover\. A complementary line of work studies online regret minimization under performative feedback\(Jagadeesanet al\.,[2022](https://arxiv.org/html/2605.23978#bib.bib37); Izzoet al\.,[2021](https://arxiv.org/html/2605.23978#bib.bib34)\), where the learner controls the deployment sequence directly\. I give a finite\-sample version for a short\-horizon linear feedback model\.
###### Assumption 1\(Instrumented linear feedback\)\.
Fori=1,…,ni=1,\\ldots,n, observeYi=zi⊤w\+εiY\_\{i\}=z\_\{i\}^\{\\top\}w\+\\varepsilon\_\{i\}wherezi=\(ϕi,Ai\)∈ℝpz\_\{i\}=\(\\phi\_\{i\},A\_\{i\}\)\\in\\mathbb\{R\}^\{p\},w=\(θ,β\)∈ℝpw=\(\\theta,\\beta\)\\in\\mathbb\{R\}^\{p\}, and‖zi‖2≤L\\left\\lVert z\_\{i\}\\right\\rVert\_\{2\}\\leq Lalmost surely\. Let\{ℱi\}i≥0\\\{\\mathcal\{F\}\_\{i\}\\\}\_\{i\\geq 0\}be a filtration such thatziz\_\{i\}isℱi−1\\mathcal\{F\}\_\{i\-1\}\-measurable and the actionAiA\_\{i\}is randomized or instrumented conditional onℱi−1\\mathcal\{F\}\_\{i\-1\}\. The errors form a martingale difference sequence with𝔼\[εi∣ℱi−1\]=0\\mathbb\{E\}\[\\varepsilon\_\{i\}\\mid\\mathcal\{F\}\_\{i\-1\}\]=0,Var\(εi∣ℱi−1\)=σ2\\mathrm\{Var\}\(\\varepsilon\_\{i\}\\mid\\mathcal\{F\}\_\{i\-1\}\)=\\sigma^\{2\}, andεi∣ℱi−1\\varepsilon\_\{i\}\\mid\\mathcal\{F\}\_\{i\-1\}isσ2\\sigma^\{2\}\-sub\-Gaussian\. The analysis works conditionally on the design event𝒢λ=\{Gn⪰λIp\}\\mathcal\{G\}\_\{\\lambda\}=\\\{G\_\{n\}\\succeq\\lambda I\_\{p\}\\\}, whereGn=n−1∑izizi⊤G\_\{n\}=n^\{\-1\}\\sum\_\{i\}z\_\{i\}z\_\{i\}^\{\\top\}andλ\>0\\lambda\>0\. This MDS formulation is appropriate for autoregressive time\-series settings in whichziz\_\{i\}contains lags ofYYand conditioning on the full design\{zj\}j=1n\\\{z\_\{j\}\\\}\_\{j=1\}^\{n\}would entangle future covariates with past noise\.
###### Theorem 3\(Finite\-sample deployment\-risk estimation\)\.
Under Assumption[1](https://arxiv.org/html/2605.23978#Thmassumption1), letw^\\hat\{w\}be ordinary least squares\. Conditional on the design event𝒢λ\\mathcal\{G\}\_\{\\lambda\}, with probability at least1−δ1\-\\deltaover the regression noise,
‖w^−w‖2≤σLλ2plog\(2p/δ\)n\.\\left\\lVert\\hat\{w\}\-w\\right\\rVert\_\{2\}\\leq\\frac\{\\sigma L\}\{\\lambda\}\\sqrt\{\\frac\{2p\\log\(2p/\\delta\)\}\{n\}\}\.\(16\)For any deployment policy with feature\-action vectorzπ\(H\)z^\{\\pi\}\(H\)satisfying‖zπ\(H\)‖2≤L\\left\\lVert z^\{\\pi\}\(H\)\\right\\rVert\_\{2\}\\leq Lalmost surely, the induced conditional meanzπ\(H\)⊤wz^\{\\pi\}\(H\)^\{\\top\}wis estimated uniformly up to
\|zπ\(H\)⊤\(w^−w\)\|≤ϵn,ϵn:=σL2λ2plog\(2p/δ\)n\.\\left\|z^\{\\pi\}\(H\)^\{\\top\}\(\\hat\{w\}\-w\)\\right\|\\leq\\epsilon\_\{n\},\\qquad\\epsilon\_\{n\}\\;:=\\;\\frac\{\\sigma L^\{2\}\}\{\\lambda\}\\sqrt\{\\frac\{2p\\log\(2p/\\delta\)\}\{n\}\}\.\(17\)Suppose further that the oracle prediction error\|f\(H\)−zπ\(H\)⊤w\|\|f\(H\)\-z^\{\\pi\}\(H\)^\{\\top\}w\|is bounded byBBalmost surely\. Then the plug\-in squared\-loss deployment\-risk estimate
R^m\(f\)=𝔼H\[\(f\(H\)−zπ\(H\)⊤w^\)2\]\+σ2\\widehat\{R\}\_\{m\}\(f\)\\;=\\;\\mathbb\{E\}\_\{H\}\\\!\\left\[\(f\(H\)\-z^\{\\pi\}\(H\)^\{\\top\}\\hat\{w\}\)^\{2\}\\right\]\+\\sigma^\{2\}differs from the oracle deployment risk by at most
\|R^m\(f\)−Rm\(f\)\|≤2Bϵn\+ϵn2\.\\left\|\\widehat\{R\}\_\{m\}\(f\)\-R\_\{m\}\(f\)\\right\|\\;\\leq\\;2B\\epsilon\_\{n\}\+\\epsilon\_\{n\}^\{2\}\.\(18\)
The theorem is modest by design\. It does not solve market simulation or long\-horizon equilibrium\. It says that once actions vary exogenously enough to identify their effect, deployment\-sensitive risk estimation becomes an ordinary statistical problem, of the kind that double/debiased ML methods address routinely\(Chernozhukovet al\.,[2018](https://arxiv.org/html/2605.23978#bib.bib42)\)\. Recent work on causal estimation of performativity in cross\-sectional and non\-autoregressive settings\(Mendler\-Dünneret al\.,[2022](https://arxiv.org/html/2605.23978#bib.bib2); Chenget al\.,[2024](https://arxiv.org/html/2605.23978#bib.bib32)\)pursues a related identifiability question\. The contrast with Theorem[1](https://arxiv.org/html/2605.23978#Thmtheorem1)is the main point: passive histories hide feedback; instrumented histories can reveal it\.
#### Misspecified feedback\.
The linear specification in Assumption[1](https://arxiv.org/html/2605.23978#Thmassumption1)is a baseline\. Empirical microstructure exhibits concave impact, with the square\-root lawΔp∝sign\(Q\)\|Q\|\\Delta p\\propto\\mathrm\{sign\}\(Q\)\\sqrt\{\|Q\|\}as a robust regularity\(Almgrenet al\.,[2005](https://arxiv.org/html/2605.23978#bib.bib1)\)\. If the true conditional mean isg\(z\)g\(z\)with bounded misspecificationρ:=supz\|g\(z\)−z⊤w⋆\|\\rho:=\\sup\_\{z\}\|g\(z\)\-z^\{\\top\}w\_\{\\star\}\|relative to the best linear projectionw⋆w\_\{\\star\}, the parameter bound in Theorem[3](https://arxiv.org/html/2605.23978#Thmtheorem3)carries an additive bias of orderρ/λ\\rho/\\lambda, and the plug\-in deployment\-risk estimate inflates byO\(Bρ\+ρ2\)O\(B\\rho\+\\rho^\{2\}\)\. Sieve regression, kernel methods, or parametric concave\-impact models are natural extensions that recover identification under richer feedback while preserving the basic mechanism\.
## 7Illustrative closed\-form example
Figure[2](https://arxiv.org/html/2605.23978#S7.F2)visualizes the ranking inversion in Theorem[2](https://arxiv.org/html/2605.23978#Thmtheorem2)with conservative parameterc′=0\.25c^\{\\prime\}=0\.25and feedback coefficientγ=1\.35\\gamma=1\.35, so the inversion thresholdγα\>\(1−c′\)/\(1\+c′\)=0\.6\\gamma\\alpha\>\(1\-c^\{\\prime\}\)/\(1\+c^\{\\prime\}\)=0\.6is crossed atα≈0\.44\\alpha\\approx 0\.44\. LetH∼N\(0,1\)H\\sim N\(0,1\)andε∼N\(0,0\.52\)\\varepsilon\\sim N\(0,0\.5^\{2\}\)\. The passive target isY0=H\+εY^\{0\}=H\+\\varepsilon\. The passive\-best model predictsf\(H\)=Hf\(H\)=H\. A conservative model predicts0\.25H0\.25H\. At deployment, adoption intensityα\\alphacreates negative feedbackYc,α=H−1\.35αcH\+εY^\{c,\\alpha\}=H\-1\.35\\alpha cH\+\\varepsilon\. The passive\-best predictor remains superior for smallα\\alpha, but its loss rises rapidly as adoption increases because its own action changes the target\. The conservative predictor has worse passive fit but lower feedback exposure\.
Figure 2:Closed\-form illustration of Theorem[2](https://arxiv.org/html/2605.23978#Thmtheorem2)withσ=0\.5\\sigma=0\.5andγ=1\.35\\gamma=1\.35, comparing the passive\-best predictorf\(h\)=hf\(h\)=hwith the conservative predictorf\(h\)=0\.25hf\(h\)=0\.25h\. Each curve reports deployment risk under the data\-generating process induced by that predictor’s own action; the comparison is therefore a comparison of policies, not forecasts on a common label sequence\. The passive\-best predictor has the lowest historical risk atα=0\\alpha=0, but its deployment risk rises with adoption, producing a ranking inversion atα⋆≈0\.44\\alpha^\{\\star\}\\approx 0\.44\.The example uses no proprietary data and is not intended as a market model\. It is a diagnostic example for benchmark design\. A financial time\-series benchmark that reports only the left endpoint of the figure is answering a passive question\. A deployment\-aware benchmark would report a crowding curve, a feedback assumption, or an instrumented estimate of action sensitivity\.
## 8Implications for financial ML evaluation
The theory suggests several evaluation practices\. Financial ML papers should specify whether their claims concern passive prediction, paper trading, small\-scale deployment, or population\-level deployment, since these are different regimes\. Benchmarks should report adoption\-sensitivity curves when a model’s actions plausibly affect the target, and historical leaderboard rankings should be stress\-tested under impact, crowding, and adaptive\-opponent scenarios\. Existing tooling can support this: agent\-based market simulators such as ABIDES and ABIDES\-Gym make interaction explicit\(Byrdet al\.,[2020](https://arxiv.org/html/2605.23978#bib.bib19); Amrouniet al\.,[2021](https://arxiv.org/html/2605.23978#bib.bib20)\), and platforms such as Qlib and FinRL can be extended with feedback modules rather than treated as historical replay systems alone\(Yanget al\.,[2020](https://arxiv.org/html/2605.23978#bib.bib21); Liuet al\.,[2020](https://arxiv.org/html/2605.23978#bib.bib22),[2022](https://arxiv.org/html/2605.23978#bib.bib23)\)\. Finally, when real deployments are possible, randomized or staggered exposure should be valued as measurement infrastructure, not as product experimentation alone\.
Table[2](https://arxiv.org/html/2605.23978#S8.T2)translates the theory into benchmark diagnostics\. None of these diagnostics requires a perfect market simulator\. The goal is to make the assumed deployment regime explicit and to test whether the main model ranking survives plausible feedback perturbations\.
Table 2:Deployment\-aware diagnostics suggested by the theory\.These suggestions are not limited to trading\. In credit, a model can change default labels by changing access to credit\. In insurance, pricing models can change the risk pool\. In recommender systems, ranking models change attention and future engagement labels\. In cybersecurity, detection models change attacker behavior\. Finance is simply the cleanest laboratory because actions, prices, and feedback are often recorded at fine temporal resolution\.
## 9Related work
#### Financial ML and return prediction\.
Modern empirical asset pricing has shown that flexible models can improve the prediction of returns and characteristics\-based portfolios\(Guet al\.,[2020](https://arxiv.org/html/2605.23978#bib.bib5)\)\. The “virtue of complexity” argument gives a theoretical and empirical rationale for using many parameters in return prediction\(Kellyet al\.,[2024](https://arxiv.org/html/2605.23978#bib.bib4)\)\. Algometrics is complementary: it asks when a learned signal remains valid after the model changes the market state through action\.
#### Performative and strategic prediction\.
Performative prediction formalizes the idea that predictions can change the distribution on which they are evaluated\(Perdomoet al\.,[2020](https://arxiv.org/html/2605.23978#bib.bib7); Mendler\-Dünneret al\.,[2020](https://arxiv.org/html/2605.23978#bib.bib8); Milleret al\.,[2021](https://arxiv.org/html/2605.23978#bib.bib9)\), with extensions to stateful environments\(Brownet al\.,[2022](https://arxiv.org/html/2605.23978#bib.bib31)\), multi\-agent decision\-dependent games\(Naranget al\.,[2023](https://arxiv.org/html/2605.23978#bib.bib48)\), the influence exerted by deployed predictors on the data they observe\(Hardtet al\.,[2022](https://arxiv.org/html/2605.23978#bib.bib36); Mendler\-Dünneret al\.,[2024](https://arxiv.org/html/2605.23978#bib.bib38)\), and the broader agenda surveyed inHardt and Mendler\-Dünner \([2023](https://arxiv.org/html/2605.23978#bib.bib33)\)\. Strategic classification studies agents who adapt their features in response to a classifier\(Hardtet al\.,[2016](https://arxiv.org/html/2605.23978#bib.bib10); Levanon and Rosenfeld,[2021](https://arxiv.org/html/2605.23978#bib.bib40)\)\. A recent application to financial market making is developed inKleitsikaset al\.\([2025](https://arxiv.org/html/2605.23978#bib.bib39)\)\. Relative to this literature, algometrics emphasizes a time\-series identification problem: the adapted object is the sequence of future labels, not an individual feature vector, and historical data are typically collected under a different algorithmic regime than deployment\.
#### Market microstructure and agent\-based finance\.
Market impact, liquidity, and execution costs are central to microstructure\(Kyle,[1985](https://arxiv.org/html/2605.23978#bib.bib11); Almgren and Chriss,[2001](https://arxiv.org/html/2605.23978#bib.bib12)\)\. Agent\-based computational finance studies aggregate outcomes generated by interacting traders\(LeBaron,[2006](https://arxiv.org/html/2605.23978#bib.bib27); Cont and Bouchaud,[2000](https://arxiv.org/html/2605.23978#bib.bib13)\)\. High\-fidelity simulators such as ABIDES make these interactions programmable for ML research\(Byrdet al\.,[2020](https://arxiv.org/html/2605.23978#bib.bib19); Amrouniet al\.,[2021](https://arxiv.org/html/2605.23978#bib.bib20)\)\. The contribution is not a new simulator; it is an identification argument for why passive time\-series benchmarks cannot, by themselves, estimate deployment risk\.
## 10Limitations and scope
The framework is intentionally stylized\. The theorems use one\-step or short\-horizon feedback rather than full equilibrium markets\. Real markets include latent information, heterogeneous objectives, inventory constraints, asymmetric information, regulatory limits, and adversarial behavior\. The positive result assumes randomized or instrumented actions and a linear feedback model; many deployments violate both assumptions\. The closed\-form illustration is a visualization of the mechanism, not evidence about any real asset class\.
The term algometrics should also not be read as a replacement for econometrics, market microstructure, or time\-series analysis\. It names a subset of problems in which algorithms are part of the data\-generating process\. Classical tools remain necessary\. The new requirement is to state when passive historical risk is a deployment\-relevant estimand and when it is only a first\-stage diagnostic\. Finally, deployment\-sensitive evaluation can be misused\. Better estimates of feedback may improve risk management and benchmark validity, but they may also help institutions trade more effectively against less sophisticated participants\. The response is to make assumptions and externalities visible, especially when models are released as reusable financial ML infrastructure\.
## 11Conclusion
The virtue of complexity in financial ML concerns what rich models can learn from high\-dimensional markets\. Algometrics asks what happens after those models act\. The paper formalized this distinction through algorithm\-mediated time series, historical risk, deployment risk, and the feedback gap\. It showed that passive histories do not identify deployment risk in general, that model rankings can invert under crowding, and that instrumented action variation can recover short\-horizon feedback in a linear setting\. In algorithmic markets, time\-series learning should measure how the future responds to the models that predict it\.
## References
- R\. Almgren and N\. Chriss \(2001\)Optimal execution of portfolio transactions\.Journal of Risk3,pp\. 5–39\.Cited by:[§2](https://arxiv.org/html/2605.23978#S2.p3.1),[§9](https://arxiv.org/html/2605.23978#S9.SS0.SSS0.Px3.p1.1)\.
- R\. Almgren, C\. Thum, E\. Hauptmann, and H\. Li \(2005\)Direct estimation of equity market impact\.Risk18\(7\),pp\. 58–62\.Cited by:[§6](https://arxiv.org/html/2605.23978#S6.SS0.SSS0.Px1.p1.6)\.
- S\. Amrouni, A\. Moulin, J\. Vann, S\. Vyetrenko, T\. Balch, and M\. Veloso \(2021\)ABIDES\-Gym: Gym environments for multi\-agent discrete event simulation and application to financial markets\.External Links:2110\.14771Cited by:[§8](https://arxiv.org/html/2605.23978#S8.p1.1),[§9](https://arxiv.org/html/2605.23978#S9.SS0.SSS0.Px3.p1.1)\.
- D\. Avramov, S\. Cheng, and L\. Metzker \(2023\)Machine learning vs\. economic restrictions: evidence from stock return predictability\.Management Science\.Cited by:[§2](https://arxiv.org/html/2605.23978#S2.p1.1)\.
- G\. Brown, S\. Hod, and I\. Kalemaj \(2022\)Performative prediction in a stateful world\.InProceedings of the 25th International Conference on Artificial Intelligence and Statistics \(AISTATS\),Proceedings of Machine Learning Research, Vol\.151,pp\. 6045–6061\.Cited by:[§9](https://arxiv.org/html/2605.23978#S9.SS0.SSS0.Px2.p1.1)\.
- S\. Bryzgalova, M\. Pelger, and J\. Zhu \(2025\)Forest through the trees: building cross\-sections of stock returns\.The Journal of Finance80\(5\),pp\. 2447–2506\.External Links:[Document](https://dx.doi.org/10.1111/jofi.13477)Cited by:[§2](https://arxiv.org/html/2605.23978#S2.p1.1)\.
- D\. Byrd, M\. Hybinette, and T\. H\. Balch \(2020\)ABIDES: towards high\-fidelity multi\-agent market simulation\.InProceedings of the 2020 ACM SIGSIM Conference on Principles of Advanced Discrete Simulation,pp\. 11–22\.Cited by:[§8](https://arxiv.org/html/2605.23978#S8.p1.1),[§9](https://arxiv.org/html/2605.23978#S9.SS0.SSS0.Px3.p1.1)\.
- G\. Cheng, M\. Hardt, and C\. Mendler\-Dünner \(2024\)Causal inference out of control: estimating performativity without treatment randomization\.InProceedings of the 41st International Conference on Machine Learning \(ICML\),Proceedings of Machine Learning Research, Vol\.235,pp\. 8077–8103\.Cited by:[§6](https://arxiv.org/html/2605.23978#S6.p2.1)\.
- V\. Chernozhukov, D\. Chetverikov, M\. Demirer, E\. Duflo, C\. Hansen, W\. Newey, and J\. Robins \(2018\)Double/debiased machine learning for treatment and structural parameters\.The Econometrics Journal21\(1\),pp\. C1–C68\.External Links:[Document](https://dx.doi.org/10.1111/ectj.12097)Cited by:[§6](https://arxiv.org/html/2605.23978#S6.p2.1)\.
- R\. Cont and J\. Bouchaud \(2000\)Herd behavior and aggregate fluctuations in financial markets\.Macroeconomic Dynamics4\(2\),pp\. 170–196\.Cited by:[§2](https://arxiv.org/html/2605.23978#S2.p3.1),[§9](https://arxiv.org/html/2605.23978#S9.SS0.SSS0.Px3.p1.1)\.
- A\. Falck, A\. Rej, and D\. Thesmar \(2022\)When do systematic strategies decay?\.Quantitative Finance22\(11\),pp\. 1955–1969\.External Links:[Document](https://dx.doi.org/10.1080/14697688.2022.2098810)Cited by:[§5](https://arxiv.org/html/2605.23978#S5.p2.1)\.
- S\. Gu, B\. Kelly, and D\. Xiu \(2020\)Empirical asset pricing via machine learning\.The Review of Financial Studies33\(5\),pp\. 2223–2273\.External Links:[Document](https://dx.doi.org/10.1093/rfs/hhaa009)Cited by:[§1](https://arxiv.org/html/2605.23978#S1.p1.1),[§2](https://arxiv.org/html/2605.23978#S2.p1.1),[§9](https://arxiv.org/html/2605.23978#S9.SS0.SSS0.Px1.p1.1)\.
- M\. Hardt, M\. Jagadeesan, and C\. Mendler\-Dünner \(2022\)Performative power\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Vol\.35\.Cited by:[§9](https://arxiv.org/html/2605.23978#S9.SS0.SSS0.Px2.p1.1)\.
- M\. Hardt, N\. Megiddo, C\. Papadimitriou, and M\. Wootters \(2016\)Strategic classification\.InProceedings of the 2016 ACM Conference on Innovations in Theoretical Computer Science,pp\. 111–122\.Cited by:[§2](https://arxiv.org/html/2605.23978#S2.p3.1),[§9](https://arxiv.org/html/2605.23978#S9.SS0.SSS0.Px2.p1.1)\.
- M\. Hardt and C\. Mendler\-Dünner \(2023\)Performative prediction: past and future\.arXiv preprint arXiv:2310\.16608\.Cited by:[§9](https://arxiv.org/html/2605.23978#S9.SS0.SSS0.Px2.p1.1)\.
- K\. Harris, D\. D\. T\. Ngo, L\. Stapleton, H\. Heidari, and Z\. S\. Wu \(2022\)Strategic instrumental variable regression: recovering causal relationships from strategic responses\.InProceedings of the 39th International Conference on Machine Learning \(ICML\),Proceedings of Machine Learning Research, Vol\.162,pp\. 8502–8522\.Cited by:[§6](https://arxiv.org/html/2605.23978#S6.p1.1)\.
- C\. R\. Harvey, Y\. Liu, and H\. Zhu \(2016\)… And the cross\-section of expected returns\.The Review of Financial Studies29\(1\),pp\. 5–68\.Cited by:[§2](https://arxiv.org/html/2605.23978#S2.p1.1)\.
- G\. W\. Imbens and J\. D\. Angrist \(1994\)Identification and estimation of local average treatment effects\.Econometrica62\(2\),pp\. 467–475\.Cited by:[§6](https://arxiv.org/html/2605.23978#S6.p1.1)\.
- Z\. Izzo, L\. Ying, and J\. Zou \(2021\)How to learn when data reacts to your model: performative gradient descent\.InProceedings of the 38th International Conference on Machine Learning \(ICML\),Proceedings of Machine Learning Research, Vol\.139,pp\. 4641–4650\.Cited by:[§6](https://arxiv.org/html/2605.23978#S6.p1.1)\.
- M\. Jagadeesan, T\. Zrnic, and C\. Mendler\-Dünner \(2022\)Regret minimization with performative feedback\.InProceedings of the 39th International Conference on Machine Learning \(ICML\),Proceedings of Machine Learning Research, Vol\.162\.Cited by:[§6](https://arxiv.org/html/2605.23978#S6.p1.1)\.
- B\. T\. Kelly, S\. Malamud, and K\. Zhou \(2024\)The virtue of complexity in return prediction\.The Journal of Finance79\(1\),pp\. 459–503\.External Links:[Document](https://dx.doi.org/10.1111/jofi.13298)Cited by:[§1](https://arxiv.org/html/2605.23978#S1.p1.1),[§2](https://arxiv.org/html/2605.23978#S2.p1.1),[§9](https://arxiv.org/html/2605.23978#S9.SS0.SSS0.Px1.p1.1)\.
- A\. E\. Khandani and A\. W\. Lo \(2011\)What happened to the quants in august 2007? evidence from factors and transactions data\.Journal of Financial Markets14\(1\),pp\. 1–46\.Cited by:[§5](https://arxiv.org/html/2605.23978#S5.p2.1)\.
- C\. Kleitsikas, S\. Leonardos, and C\. Ventre \(2025\)Performative market making\.arXiv preprint arXiv:2508\.04344\.Cited by:[§9](https://arxiv.org/html/2605.23978#S9.SS0.SSS0.Px2.p1.1)\.
- A\. S\. Kyle \(1985\)Continuous auctions and insider trading\.Econometrica53\(6\),pp\. 1315–1335\.Cited by:[§2](https://arxiv.org/html/2605.23978#S2.p3.1),[§9](https://arxiv.org/html/2605.23978#S9.SS0.SSS0.Px3.p1.1)\.
- B\. LeBaron \(2006\)Agent\-based computational finance\.Handbook of Computational Economics2,pp\. 1187–1233\.Cited by:[§2](https://arxiv.org/html/2605.23978#S2.p3.1),[§9](https://arxiv.org/html/2605.23978#S9.SS0.SSS0.Px3.p1.1)\.
- S\. Levanon and N\. Rosenfeld \(2021\)Strategic classification made practical\.InProceedings of the 38th International Conference on Machine Learning \(ICML\),Proceedings of Machine Learning Research, Vol\.139\.Cited by:[§9](https://arxiv.org/html/2605.23978#S9.SS0.SSS0.Px2.p1.1)\.
- X\. Liu, Z\. Xia, J\. Rui, J\. Gao, H\. Yang, M\. Zhu, C\. D\. Wang, Z\. Wang, and J\. Guo \(2022\)FinRL\-Meta: market environments and benchmarks for data\-driven financial reinforcement learning\.Advances in Neural Information Processing Systems Datasets and Benchmarks Track\.Cited by:[§8](https://arxiv.org/html/2605.23978#S8.p1.1)\.
- X\. Liu, H\. Yang, Q\. Chen, R\. Zhang, L\. Yang, B\. Xiao, and C\. D\. Wang \(2020\)FinRL: a deep reinforcement learning library for automated stock trading in quantitative finance\.External Links:2011\.09607Cited by:[§8](https://arxiv.org/html/2605.23978#S8.p1.1)\.
- A\. W\. Lo \(2004\)The adaptive markets hypothesis: market efficiency from an evolutionary perspective\.Journal of Portfolio Management30\(5\),pp\. 15–29\.Cited by:[§2](https://arxiv.org/html/2605.23978#S2.p3.1)\.
- C\. F\. Manski \(1993\)Identification of endogenous social effects: the reflection problem\.The Review of Economic Studies60\(3\),pp\. 531–542\.Cited by:[§4](https://arxiv.org/html/2605.23978#S4.p2.3)\.
- I\. W\. R\. Martin and S\. Nagel \(2022\)Market efficiency in the age of big data\.Journal of Financial Economics145\(1\),pp\. 154–177\.External Links:[Document](https://dx.doi.org/10.1016/j.jfineco.2021.10.006)Cited by:[§2](https://arxiv.org/html/2605.23978#S2.p1.1)\.
- R\. D\. McLean and J\. Pontiff \(2016\)Does academic research destroy stock return predictability?\.The Journal of Finance71\(1\),pp\. 5–32\.Cited by:[§5](https://arxiv.org/html/2605.23978#S5.p2.1)\.
- C\. Mendler\-Dünner, G\. Carovano, and M\. Hardt \(2024\)Measuring the performative power of online search\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Cited by:[§9](https://arxiv.org/html/2605.23978#S9.SS0.SSS0.Px2.p1.1)\.
- C\. Mendler\-Dünner, F\. Ding, and Y\. Wang \(2022\)Anticipating performativity by predicting from predictions\.InAdvances in Neural Information Processing Systems \(NeurIPS\),Vol\.35\.Cited by:[§6](https://arxiv.org/html/2605.23978#S6.p2.1)\.
- C\. Mendler\-Dünner, J\. C\. Perdomo, T\. Zrnic, and M\. Hardt \(2020\)Stochastic optimization for performative prediction\.InAdvances in Neural Information Processing Systems,Vol\.33,pp\. 4929–4939\.Cited by:[§2](https://arxiv.org/html/2605.23978#S2.p3.1),[§9](https://arxiv.org/html/2605.23978#S9.SS0.SSS0.Px2.p1.1)\.
- J\. Miller, J\. C\. Perdomo, and T\. Zrnic \(2021\)Outside the echo chamber: optimizing the performative risk\.InProceedings of the 38th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.139,pp\. 7710–7720\.Cited by:[§2](https://arxiv.org/html/2605.23978#S2.p3.1),[§9](https://arxiv.org/html/2605.23978#S9.SS0.SSS0.Px2.p1.1)\.
- A\. Narang, E\. Faulkner, D\. Drusvyatskiy, M\. Fazel, and L\. J\. Ratliff \(2023\)Multiplayer performative prediction: learning in decision\-dependent games\.Journal of Machine Learning Research24\(202\),pp\. 1–56\.Cited by:[§9](https://arxiv.org/html/2605.23978#S9.SS0.SSS0.Px2.p1.1)\.
- J\. C\. Perdomo, T\. Zrnic, C\. Mendler\-Dünner, and M\. Hardt \(2020\)Performative prediction\.InProceedings of the 37th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.119,pp\. 7599–7609\.Cited by:[§2](https://arxiv.org/html/2605.23978#S2.p3.1),[§9](https://arxiv.org/html/2605.23978#S9.SS0.SSS0.Px2.p1.1)\.
- I\. Welch and A\. Goyal \(2008\)A comprehensive look at the empirical performance of equity premium prediction\.The Review of Financial Studies21\(4\),pp\. 1455–1508\.External Links:[Document](https://dx.doi.org/10.1093/rfs/hhm014)Cited by:[§2](https://arxiv.org/html/2605.23978#S2.p1.1)\.
- X\. Yang, W\. Liu, D\. Zhou, J\. Bian, and T\. Liu \(2020\)Qlib: an AI\-oriented quantitative investment platform\.External Links:2009\.11189Cited by:[§8](https://arxiv.org/html/2605.23978#S8.p1.1)\.
## Appendix AProofs
### A\.1Proof of Proposition[1](https://arxiv.org/html/2605.23978#Thmproposition1)
For a fixed historyHtH\_\{t\}and predictiony^=f\(Ht\)\\hat\{y\}=f\(H\_\{t\}\), the mapy↦ℓ\(y^,y\)y\\mapsto\\ell\(\\hat\{y\},y\)isLℓL\_\{\\ell\}\-Lipschitz\. By the Kantorovich\-Rubinstein dual characterization ofW1W\_\{1\},
\|𝔼\[ℓ\(y^,Y\)∣Ht,At=a\]−𝔼\[ℓ\(y^,Y\)∣Ht,At=0\]\|\\displaystyle\\left\|\\mathbb\{E\}\[\\ell\(\\hat\{y\},Y\)\\mid H\_\{t\},A\_\{t\}=a\]\-\\mathbb\{E\}\[\\ell\(\\hat\{y\},Y\)\\mid H\_\{t\},A\_\{t\}=0\]\\right\|\(19\)≤LℓW1\(P\(Y∣Ht,At=a\),P\(Y∣Ht,At=0\)\)\.\\displaystyle\\hskip 28\.90755pt\\leq L\_\{\\ell\}W\_\{1\}\\left\(P\(Y\\mid H\_\{t\},A\_\{t\}=a\),P\(Y\\mid H\_\{t\},A\_\{t\}=0\)\\right\)\.\(20\)Using the assumed elasticity bound witha=πf\(Ht\)a=\\pi\_\{f\}\(H\_\{t\}\)and taking expectations over passive histories gives
\|Γ\(f\)\|≤Lℓ𝔼\[κ\(Ht\)‖πf\(Ht\)‖\]\.\|\\Gamma\(f\)\|\\leq L\_\{\\ell\}\\mathbb\{E\}\[\\kappa\(H\_\{t\}\)\\left\\lVert\\pi\_\{f\}\(H\_\{t\}\)\\right\\rVert\]\.\(21\)This proves the claim\. ∎
### A\.2Proof of Theorem[1](https://arxiv.org/html/2605.23978#Thmtheorem1)
In the passive regimeAt=0A\_\{t\}=0, the model is
Yt\+1=μ\(Ht\)\+εt\+1\.Y\_\{t\+1\}=\\mu\(H\_\{t\}\)\+\\varepsilon\_\{t\+1\}\.\(22\)The passive conditional law ofYt\+1Y\_\{t\+1\}givenHtH\_\{t\}is therefore independent ofβ\\beta\. Hence any two valuesβ\\betaandβ′\\beta^\{\\prime\}induce the same passive distribution over\(Ht,Yt\+1\)\(H\_\{t\},Y\_\{t\+1\}\)\.
Under deployment,At=πf\(Ht\)A\_\{t\}=\\pi\_\{f\}\(H\_\{t\}\)and
Yt\+1β=μ\(Ht\)\+βπf\(Ht\)\+εt\+1\.Y\_\{t\+1\}^\{\\beta\}=\\mu\(H\_\{t\}\)\+\\beta\\pi\_\{f\}\(H\_\{t\}\)\+\\varepsilon\_\{t\+1\}\.\(23\)The one\-step deployment risk offfunder squared loss is
Rβ\(f\)\\displaystyle R\_\{\\beta\}\(f\)=𝔼\[\(f\(Ht\)−μ\(Ht\)−βπf\(Ht\)−εt\+1\)2\]\\displaystyle=\\mathbb\{E\}\\left\[\\left\(f\(H\_\{t\}\)\-\\mu\(H\_\{t\}\)\-\\beta\\pi\_\{f\}\(H\_\{t\}\)\-\\varepsilon\_\{t\+1\}\\right\)^\{2\}\\right\]\(24\)=𝔼\[\(f\(Ht\)−μ\(Ht\)−εt\+1\)2\]\\displaystyle=\\mathbb\{E\}\\left\[\\left\(f\(H\_\{t\}\)\-\\mu\(H\_\{t\}\)\-\\varepsilon\_\{t\+1\}\\right\)^\{2\}\\right\]\(25\)−2β𝔼\[πf\(Ht\)\(f\(Ht\)−μ\(Ht\)\)\]\+β2𝔼\[πf\(Ht\)2\],\\displaystyle\\quad\-2\\beta\\mathbb\{E\}\\left\[\\pi\_\{f\}\(H\_\{t\}\)\\left\(f\(H\_\{t\}\)\-\\mu\(H\_\{t\}\)\\right\)\\right\]\+\\beta^\{2\}\\mathbb\{E\}\\left\[\\pi\_\{f\}\(H\_\{t\}\)^\{2\}\\right\],\(26\)where the cross\-term involvingεt\+1\\varepsilon\_\{t\+1\}vanishes by the tower property since𝔼\[εt\+1∣Ht,At\]=0\\mathbb\{E\}\[\\varepsilon\_\{t\+1\}\\mid H\_\{t\},A\_\{t\}\]=0\. The coefficient onβ2\\beta^\{2\}is positive by assumption, soRβ\(f\)R\_\{\\beta\}\(f\)is a nonconstant quadratic function ofβ\\beta\. Since the passive data distribution is identical for allβ\\betabut the deployment risk differs for at least two values ofβ\\beta, no estimator that is a measurable function only of passive histories can identify deployment risk uniformly over this family\. ∎
### A\.3Proof of Theorem[2](https://arxiv.org/html/2605.23978#Thmtheorem2)
Forfc\(H\)=cHf\_\{c\}\(H\)=cHunder passive observation,
R0\(fc\)=𝔼\[\(cH−H−ε\)2\]=\(c−1\)2\+σ2\.R\_\{0\}\(f\_\{c\}\)=\\mathbb\{E\}\\left\[\(cH\-H\-\\varepsilon\)^\{2\}\\right\]=\(c\-1\)^\{2\}\+\\sigma^\{2\}\.\(27\)Under deployment offcf\_\{c\}, the target isYc,α=H−γαcH\+εY^\{c,\\alpha\}=H\-\\gamma\\alpha cH\+\\varepsilon, and the forecast iscHcH, so
Rα\(fc\)\\displaystyle R\_\{\\alpha\}\(f\_\{c\}\)=𝔼\[\(cH−\(H−γαcH\+ε\)\)2\]\\displaystyle=\\mathbb\{E\}\\left\[\\left\(cH\-\(H\-\\gamma\\alpha cH\+\\varepsilon\)\\right\)^\{2\}\\right\]\(28\)=𝔼\[\(\(c\(1\+γα\)−1\)H−ε\)2\]\\displaystyle=\\mathbb\{E\}\\left\[\\left\(\(c\(1\+\\gamma\\alpha\)\-1\)H\-\\varepsilon\\right\)^\{2\}\\right\]\(29\)=\(c\(1\+γα\)−1\)2\+σ2,\\displaystyle=\(c\(1\+\\gamma\\alpha\)\-1\)^\{2\}\+\\sigma^\{2\},\(30\)which establishes equation \([13](https://arxiv.org/html/2605.23978#S5.E13)\)\.
Fixc′∈\[0,1\)c^\{\\prime\}\\in\[0,1\)\. The passive comparison is immediate:R0\(f1\)=σ2<\(c′−1\)2\+σ2=R0\(fc′\)R\_\{0\}\(f\_\{1\}\)=\\sigma^\{2\}<\(c^\{\\prime\}\-1\)^\{2\}\+\\sigma^\{2\}=R\_\{0\}\(f\_\{c^\{\\prime\}\}\)\. For deployment risk,
Rα\(f1\)\>Rα\(fc′\)⇔γ2α2\>\(c′\(1\+γα\)−1\)2\.R\_\{\\alpha\}\(f\_\{1\}\)\>R\_\{\\alpha\}\(f\_\{c^\{\\prime\}\}\)\\;\\iff\\;\\gamma^\{2\}\\alpha^\{2\}\>\\big\(c^\{\\prime\}\(1\+\\gamma\\alpha\)\-1\\big\)^\{2\}\.\(31\)It remains to resolve the absolute value on the right\-hand side\.
*Regime I:c′\(1\+γα\)≤1c^\{\\prime\}\(1\+\\gamma\\alpha\)\\leq 1\.*The right\-hand side equals\(1−c′\(1\+γα\)\)2\(1\-c^\{\\prime\}\(1\+\\gamma\\alpha\)\)^\{2\}, and taking nonnegative square roots gives
γα\>1−c′−c′γα,i\.e\.,γα\(1\+c′\)\>1−c′,\\gamma\\alpha\>1\-c^\{\\prime\}\-c^\{\\prime\}\\gamma\\alpha,\\quad\\text\{i\.e\.,\}\\quad\\gamma\\alpha\(1\+c^\{\\prime\}\)\>1\-c^\{\\prime\},\(32\)which is equivalent to \([15](https://arxiv.org/html/2605.23978#S5.E15)\)\.
*Regime II:c′\(1\+γα\)\>1c^\{\\prime\}\(1\+\\gamma\\alpha\)\>1\.*Thenc′\(1\+γα\)−1\>0c^\{\\prime\}\(1\+\\gamma\\alpha\)\-1\>0, and the inequalityγ2α2\>\(c′\(1\+γα\)−1\)2\\gamma^\{2\}\\alpha^\{2\}\>\(c^\{\\prime\}\(1\+\\gamma\\alpha\)\-1\)^\{2\}is equivalent toγα\>c′\(1\+γα\)−1\\gamma\\alpha\>c^\{\\prime\}\(1\+\\gamma\\alpha\)\-1, i\.e\.,γα\(1−c′\)\>−\(1−c′\)\\gamma\\alpha\(1\-c^\{\\prime\}\)\>\-\(1\-c^\{\\prime\}\), which is automatic forc′∈\[0,1\)c^\{\\prime\}\\in\[0,1\)andγα≥0\\gamma\\alpha\\geq 0\. Moreover, Regime II requiresγα\>\(1−c′\)/c′\\gamma\\alpha\>\(1\-c^\{\\prime\}\)/c^\{\\prime\}, which strictly exceeds the threshold\(1−c′\)/\(1\+c′\)\(1\-c^\{\\prime\}\)/\(1\+c^\{\\prime\}\)in \([15](https://arxiv.org/html/2605.23978#S5.E15)\), so the inversion is already established before entering Regime II\.
Settingc′=0c^\{\\prime\}=0collapses Regime II to the empty set and recoversγα\>1\\gamma\\alpha\>1from Regime I\. ∎
### A\.4Proof of Theorem[3](https://arxiv.org/html/2605.23978#Thmtheorem3)
LetZ∈ℝn×pZ\\in\\mathbb\{R\}^\{n\\times p\}have rowszi⊤z\_\{i\}^\{\\top\}, letY=Zw\+εY=Zw\+\\varepsilon, and letw^=\(Z⊤Z\)−1Z⊤Y\\hat\{w\}=\(Z^\{\\top\}Z\)^\{\-1\}Z^\{\\top\}Y\. Then
w^−w=\(Z⊤Z\)−1Z⊤ε\.\\hat\{w\}\-w=\(Z^\{\\top\}Z\)^\{\-1\}Z^\{\\top\}\\varepsilon\.\(33\)SinceGn=n−1Z⊤Z⪰λIpG\_\{n\}=n^\{\-1\}Z^\{\\top\}Z\\succeq\\lambda I\_\{p\}, we have‖\(Z⊤Z/n\)−1‖op≤1/λ\\left\\lVert\(Z^\{\\top\}Z/n\)^\{\-1\}\\right\\rVert\_\{\\mathrm\{op\}\}\\leq 1/\\lambda\. Therefore
‖w^−w‖2≤1λ‖1nZ⊤ε‖2\.\\left\\lVert\\hat\{w\}\-w\\right\\rVert\_\{2\}\\leq\\frac\{1\}\{\\lambda\}\\left\\\|\\frac\{1\}\{n\}Z^\{\\top\}\\varepsilon\\right\\\|\_\{2\}\.\(34\)For coordinatejj,\(n−1Z⊤ε\)j=n−1∑izijεi\(n^\{\-1\}Z^\{\\top\}\\varepsilon\)\_\{j\}=n^\{\-1\}\\sum\_\{i\}z\_\{ij\}\\varepsilon\_\{i\}is a sum of martingale differences:zijz\_\{ij\}isℱi−1\\mathcal\{F\}\_\{i\-1\}\-measurable with\|zij\|≤‖zi‖2≤L\|z\_\{ij\}\|\\leq\\left\\lVert z\_\{i\}\\right\\rVert\_\{2\}\\leq L, andεi∣ℱi−1\\varepsilon\_\{i\}\\mid\\mathcal\{F\}\_\{i\-1\}isσ2\\sigma^\{2\}\-sub\-Gaussian, sozijεi∣ℱi−1z\_\{ij\}\\varepsilon\_\{i\}\\mid\\mathcal\{F\}\_\{i\-1\}iszij2σ2z\_\{ij\}^\{2\}\\sigma^\{2\}\-sub\-Gaussian\. By the Azuma–Hoeffding inequality for sub\-Gaussian martingale differences,∑izijεi\\sum\_\{i\}z\_\{ij\}\\varepsilon\_\{i\}isnσ2L2n\\sigma^\{2\}L^\{2\}\-sub\-Gaussian, so\(n−1Z⊤ε\)j\(n^\{\-1\}Z^\{\\top\}\\varepsilon\)\_\{j\}has proxy variance at mostσ2L2/n\\sigma^\{2\}L^\{2\}/n\. A union bound overppcoordinates gives, with probability at least1−δ1\-\\delta,
‖1nZ⊤ε‖∞≤σL2log\(2p/δ\)n\.\\left\\\|\\frac\{1\}\{n\}Z^\{\\top\}\\varepsilon\\right\\\|\_\{\\infty\}\\leq\\sigma L\\sqrt\{\\frac\{2\\log\(2p/\\delta\)\}\{n\}\}\.\(35\)Since‖v‖2≤p‖v‖∞\\left\\lVert v\\right\\rVert\_\{2\}\\leq\\sqrt\{p\}\\left\\lVert v\\right\\rVert\_\{\\infty\}, the stated parameter bound follows\. For any policy vectorzπ\(H\)z^\{\\pi\}\(H\)with norm at mostLL,
\|zπ\(H\)⊤\(w^−w\)\|≤L‖w^−w‖2=ϵn\.\|z^\{\\pi\}\(H\)^\{\\top\}\(\\hat\{w\}\-w\)\|\\leq L\\left\\lVert\\hat\{w\}\-w\\right\\rVert\_\{2\}=\\epsilon\_\{n\}\.\(36\)For the plug\-in risk bound, fixHHand writeμπ=zπ\(H\)⊤w\\mu\_\{\\pi\}=z^\{\\pi\}\(H\)^\{\\top\}w,μ^π=zπ\(H\)⊤w^\\hat\{\\mu\}\_\{\\pi\}=z^\{\\pi\}\(H\)^\{\\top\}\\hat\{w\}\. Expanding,
\(f\(H\)−μ^π\)2−\(f\(H\)−μπ\)2=2\(f\(H\)−μπ\)\(μπ−μ^π\)\+\(μπ−μ^π\)2\.\(f\(H\)\-\\hat\{\\mu\}\_\{\\pi\}\)^\{2\}\-\(f\(H\)\-\\mu\_\{\\pi\}\)^\{2\}\\;=\\;2\(f\(H\)\-\\mu\_\{\\pi\}\)\(\\mu\_\{\\pi\}\-\\hat\{\\mu\}\_\{\\pi\}\)\+\(\\mu\_\{\\pi\}\-\\hat\{\\mu\}\_\{\\pi\}\)^\{2\}\.\(37\)Using\|f\(H\)−μπ\|≤B\|f\(H\)\-\\mu\_\{\\pi\}\|\\leq Balmost surely and\|μπ−μ^π\|≤ϵn\|\\mu\_\{\\pi\}\-\\hat\{\\mu\}\_\{\\pi\}\|\\leq\\epsilon\_\{n\}from the previous display,
\|\(f\(H\)−μ^π\)2−\(f\(H\)−μπ\)2\|≤2Bϵn\+ϵn2\.\\left\|\(f\(H\)\-\\hat\{\\mu\}\_\{\\pi\}\)^\{2\}\-\(f\(H\)\-\\mu\_\{\\pi\}\)^\{2\}\\right\|\\;\\leq\\;2B\\epsilon\_\{n\}\+\\epsilon\_\{n\}^\{2\}\.\(38\)Taking expectations overHHand noting that the noise varianceσ2\\sigma^\{2\}cancels betweenR^m\(f\)\\widehat\{R\}\_\{m\}\(f\)andRm\(f\)R\_\{m\}\(f\)yields the stated bound\. ∎Similar Articles
Trading Confidence: Comprehensive Uncertainty Estimation in Algorithmic Trading
Proposes an uncertainty-aware reinforcement learning framework for algorithmic trading that integrates distributional, epistemic, and aleatoric uncertainty using SHAP-weighted reconstruction, MC Dropout, and LSTM consensus. Outperforms traditional models on five major US stock indices.
Deployment-Side Adaptiveness in Multi-Horizon Volatility Forecasting
This paper investigates how different inference-time deployment rules (rollout strategies) impact multi-horizon volatility forecasting. It shows that non-default rollout rules often improve performance and that validation-based deployment policies provide a low-cost improvement over standard MIMO deployment, emphasizing the importance of deployment policy alongside model architecture.
Metric-Aware Hybrid Forecasting for the CTF4Science Lorenz Challenge
The paper describes a metric-aware hybrid forecasting system for the CTF4Science Lorenz challenge, combining neural denoisers, ODE fitting, and histogram-tail substitution to optimize different metrics across nine task pairs, achieving a public leaderboard score of 83.85529.
From Heuristics to Analytics: Forecasting Effort and Progress in Online Learning
This paper introduces engagement forecasting for intelligent tutoring systems, predicting weekly minutes practiced and new skills mastered using interaction logs from 425 middle-school students. Feature-based models reduce error by 22-33% over heuristic baselines, offering explainable patterns for tutor-learner goal setting.
AgentForesight: Online Auditing for Early Failure Prediction in Multi-Agent Systems
This paper introduces AgentForesight, a framework for online auditing and early failure prediction in LLM-based multi-agent systems. It presents a new dataset, AFTraj-22K, and a specialized model, AgentForesight-7B, which outperforms leading proprietary models in detecting decisive errors during trajectory execution.