Beyond Forecasting: The Belief-to-Trade Layer in Prediction-Market Agents
Summary
Raven-Agent is the first autonomous trading agent for prediction markets, featuring an explicit belief-to-trade layer. It achieves positive returns on a controlled replay, bridging the gap between calibrated forecasts and profitable trading.
View Cached Full Text
Cached at: 07/07/26, 04:34 AM
# Beyond Forecasting: The Belief-to-Trade Layer in Prediction-Market Agents
Source: [https://arxiv.org/html/2607.03015](https://arxiv.org/html/2607.03015)
###### Abstract
Forecasting future events has attracted growing attention as a testbed for general\-purpose AI\. A natural way to ground this evaluation is let the models trade in the prediction markets\. Trading, however, requires more than forecasting\. Moreover, recent benchmarks report a substantial gap between calibrated probability scores and the trading results\. We propose Raven\-Agent, to the best of our knowledge, the first autonomous trading agent for prediction markets\. On a controlled replay over an archived decision set, our architecture achieves the only positive return and the only positive risk\-adjusted return among all tested policies\. We have released our code in[https://github\.com/Alchemist\-X/predict\-raven](https://github.com/Alchemist-X/predict-raven)\.
AI forecasting, prediction markets, agentic systems, decision theory, risk management
\\icml@noticeprintedtrue††footnotetext:\\forloop@affilnum1\\c@@affilnum¡\\c@@affiliationcounter0AUTHORERR: Missing \\icmlaffiliation\.\. Correspondence to:Yishu Wang <issue\.00\.gui@gmail\.com\>,Yuxuan Wang <yuxwang@pku\.edu\.cn\>,Jiaqi Deng <jiaqideng@connect\.hku\.hk\>,Hanyang Tang <hytangs@mit\.edu\>\. \\Notice@String
## 1Introduction
Forecasting future events has attracted growing attention as a testbed for general\-purpose AI\(Kargeret al\.,[2024](https://arxiv.org/html/2607.03015#bib.bib1); Yanget al\.,[2025](https://arxiv.org/html/2607.03015#bib.bib2)\)\. The task demands retrieval, temporal reasoning, and decision under uncertainty\. It is naturally studied as a joint problem of probability estimation and action under that estimate\.
To make progress on this problem, one line of exisiting works focuses on probability quality and improves LLM forecasters on resolved events\. The Belief\-Logit\-Forecaster \(BLF\) ofMurphy \([2026](https://arxiv.org/html/2607.03015#bib.bib7)\)uses structured belief states and multi\-trial aggregation, currently leading ForecastBench\(Kargeret al\.,[2024](https://arxiv.org/html/2607.03015#bib.bib1)\)\. Yet such systems improve calibration without acting on their forecasts\. On the other hand, prediction market benchmarks score systems by realized trading return as well as by probability\(Yanget al\.,[2025](https://arxiv.org/html/2607.03015#bib.bib2); Chenget al\.,[2026](https://arxiv.org/html/2607.03015#bib.bib6); Zhanget al\.,[2026](https://arxiv.org/html/2607.03015#bib.bib5)\)\. They consistently report that strong probability scores do not translate into profitable trades\. However, they treat the trading rule as a fixed protocol such as a per market bet or a confidence threshold, not as a designable object\. Moreover, position sizing, portfolio exposure, and risk control are either absent or expressed as a prompt instruction in these papers\. We refer to these downstream components as the*trading layer*, and we make it the object of study in this paper\.
Across both lines, the trading layer is either omitted or hardcoded, never treated as a designable object\. A useful trading layer should be explicit and modular, with selection, sizing, and risk control as separate components\. It should be deterministic, so that risk constraints remain in effect even when the model is confidently wrong\. It should also be composable with any forecaster, so that improvements in calibration and in trade execution can be evaluated independently\.
In this paper, we propose Raven\-Agent, an autonomous prediction market agent built around an explicit, deterministic, composable trading layer\. Within the layer, a ranking module addresses selection by scoring each candidate by a time normalized return rate\. The same module addresses sizing with a fractional Kelly bet on the retained positions\. An execution and risk module addresses risk control by enforcing deterministic constraints on stake, exposure, stop loss, and drawdown outside the language model prompt\. We also include an information module for evidence retrieval and a forecasting module that produces per contract probabilities\. We validate Raven\-Agent on a controlled replay protocol\. The results show our method improves return on stake and risk adjusted return over other heuristic policies\.
Our contributions can be summarized in three\-folds:
- •We propose a replay protocol that reuses archived forecasts to isolate the trading layer from the forecaster\.
- •We propose Raven\-Agent, an autonomous agent that makes the trading layer an explicit module composable with a swappable forecaster\.
- •On a controlled replay, the trading layer improves return on stake and risk adjusted return over other policies\.
## 2Problem Setup
### 2\.1Prediction Markets
We consider binary prediction markets in which the selected contract pays one dollar if it resolves YES and zero otherwise\. For each marketii, we observe the realized outcome of the chosen side, writtenzi∈\{0,1\}z\_\{i\}\\in\\\{0,1\\\}\. The entry price isqi∈\(0,1\)q\_\{i\}\\in\(0,1\), and we letpi∈\(0,1\)p\_\{i\}\\in\(0,1\)denote the model probability that the contract pays\. A stake ofsis\_\{i\}dollars buyssi/qis\_\{i\}/q\_\{i\}shares\. The realized profit at resolution is
Πi=si\(ziqi−1\)\.\\Pi\_\{i\}=s\_\{i\}\\left\(\\frac\{z\_\{i\}\}\{q\_\{i\}\}\-1\\right\)\.\(1\)For markets unresolved at the evaluation date, we substitute the mark to market pricemim\_\{i\}forziz\_\{i\}in[Equation1](https://arxiv.org/html/2607.03015#S2.E1)\. Resolved and unresolved positions are reported separately throughout\.
To compare candidates whose payoff dates differ widely, we need a return rate, not just an edge\. The per dollar expected return of a long position ispi/qi−1p\_\{i\}/q\_\{i\}\-1before fees and slippage\. The raw probability edgeei=pi−qie\_\{i\}=p\_\{i\}\-q\_\{i\}does not by itself reflect that capital is locked until resolution\. A small edge that resolves quickly can dominate a larger edge that ties capital for months\. WithTiT\_\{i\}days to resolution, we therefore rank candidates by a time normalized score
ρi=\(piqi−1\)⋅30max\(Ti,1\)\.\\rho\_\{i\}=\\left\(\\frac\{p\_\{i\}\}\{q\_\{i\}\}\-1\\right\)\\cdot\\frac\{30\}\{\\max\(T\_\{i\},1\)\}\.\(2\)*This is the per dollar expected return of a candidate trade in units of one month\.*It is a ranking device, not a claim of exact monthly compounding\.
### 2\.2Trading policies
A trading policy formalizes which action follows from a forecast\. We write it asπ:\(pi,qi,Ti,ℓi,W,𝒮\)↦\(ai,si\)\\pi:\(p\_\{i\},q\_\{i\},T\_\{i\},\\ell\_\{i\},W,\\mathcal\{S\}\)\\mapsto\(a\_\{i\},s\_\{i\}\), whereℓi\\ell\_\{i\}is liquidity,WWis bankroll, and𝒮\\mathcal\{S\}is the current portfolio state\. The actionaia\_\{i\}is either open or skip, andsi≥0s\_\{i\}\\geq 0is the stake\. Two agents with the same forecasts can produce very different outcomes if they differ in their policy\. We hold all positions to resolution in the main replay and defer dynamic exit to future work\.
## 3Methodology
In this section, we instantiate the trading agent Raven\-Agent\. The full loop is summarized in[Algorithm1](https://arxiv.org/html/2607.03015#alg1), and the four modules are described below\.
Algorithm 1Raven\-Agent forecast to trade loop\.0:bankroll
WWand current portfolio state
1:Scan candidate markets and retrieve evidence
2:Estimate
pip\_\{i\}for each candidate
iifrom the candidate lists\.
3:Compute
ρi\\rho\_\{i\}via[Equation2](https://arxiv.org/html/2607.03015#S2.E2); drop candidates below a threshold and rank the rest by
ρi\\rho\_\{i\}\.
4:Size each surviving candidate by a quarter Kelly fraction, clipped by liquidity and a per round batch cap\.
5:Reject any proposal violating exposure, stop loss, or drawdown constraints\.
6:Submit the remaining orders and review the portfolio state, checking if we need adjust any positions\.
7:Repeat as new candidate markets and evidence arrive\.
### 3\.1Information collection
The information module scans candidate markets and writes a timestamped pulse\. We periodically query the Polymarket Gamma API for each candidate, retrieving its price, liquidity, fee schedule, and resolution date\. A liquidity filter drops markets below a threshold, and an evidence retriever pulls recent news for the survivors\. This follows the spirit of retrieval augmented forecasting\(Murphy,[2026](https://arxiv.org/html/2607.03015#bib.bib7)\)\. The pulse serves as a contract between data sources and downstream modules\. New evidence feeds can therefore be added without changing the trading interface\.
### 3\.2Probability analysis
The forecasting module emits the per contract probabilitypip\_\{i\}for the side the agent prefers to hold\. It is a swappable layer, and we currently support two runtimes\. One parses the pulse markdown directly and combines model side estimates with market implied priors\. The other dispatches the pulse to an external language model through a command line bridge\. Both record, alongsidepip\_\{i\}, the chosen side, a confidence level, and a short thesis for later audit\. Forecaster centric systems like BLF\(Murphy,[2026](https://arxiv.org/html/2607.03015#bib.bib7)\)improve the probability itself via belief states and multi trial aggregation\. Our contribution operates downstream of the forecaster, so any such forecaster can be substituted without changing the trading layer\. In the present replay we hold this module fixed and reuse archived probabilities, so any difference in outcome is attributable to the trading layer\.
### 3\.3Ranking and sizing
The ranking module both selects candidates and sets their stakes\. For selection, the module ranks candidates byρi\\rho\_\{i\}from[Equation2](https://arxiv.org/html/2607.03015#S2.E2), drops those below a threshold, and keeps the topK=4K=4\. Ranking byρi\\rho\_\{i\}rather than by the raw edgeei=pi−qie\_\{i\}=p\_\{i\}\-q\_\{i\}accounts for how long capital is locked\. This favors short resolution markets when edges are similar\.
For sizing, the module applies one quarter of the Kelly bet to each retained candidate:
si=14⋅max\(0,pi−qi1−qi\)⋅W,s\_\{i\}=\\frac\{1\}\{4\}\\cdot\\max\\\!\\left\(0,\\frac\{p\_\{i\}\-q\_\{i\}\}\{1\-q\_\{i\}\}\\right\)\\cdot W,\(3\)clipped by available liquidity and a per round batch cap\. The Kelly criterion gives the bet size that maximizes long run log wealth under a known win probability\(Kelly,[1956](https://arxiv.org/html/2607.03015#bib.bib9)\)\. We use one quarter rather than full Kelly becausepip\_\{i\}is itself an estimate\. Scaling down the Kelly bet is the standard practitioner correction for input noise\.
### 3\.4Execution and risk
The execution and risk module enforces deterministic risk constraints on every order before submission\. We implement it as a set of services that check each proposed order against five hard constraints\. A per trade notional cap and an aggregate exposure cap limit single stake and total open notional respectively\. A per event cap limits capital across correlated contracts of the same event\. A position level stop loss closes any position whose mark to market loss exceeds a fixed fraction of its cost basis\. A portfolio level drawdown halt freezes new openings when the equity drawdown\(Wmax−W\)/Wmax\(W\_\{\\max\}\-W\)/W\_\{\\max\}exceeds a threshold\. The module also rejects orders against stale pulses or contracts outside the allowlist, and routes surviving orders as fill or kill market orders\. Prior work shows that prompt\-level risk guidance is unreliable under end\-to\-end LLM control\(Zhanget al\.,[2026](https://arxiv.org/html/2607.03015#bib.bib5)\), and that optimizing prediction accuracy can reduce trading return when the objective is misaligned\(Janget al\.,[2025](https://arxiv.org/html/2607.03015#bib.bib4)\)\. Keeping these checks outside the prompt avoids that failure mode and produces auditable rejection reasons that a learned policy cannot easily match\.
## 4Experiments
### 4\.1Setup
#### Environment\.
We evaluate on a counterfactual replay of archived live decisions from a Polymarket deployment\. The archive records 59 open decisions across 65 pulse runs; 44 reach the executable stage after filtering by liquidity, resolution clarity, and event overlap\. For each candidate, Raven\-Agent estimated the payout probability, compared it with the market price, computed the edge, and proposed a stake; live trades were submitted through Polymarket’s CLOB API\. The replay reuses the archived forecasts, prices, and execution records from these runs \(concrete examples in[AppendixC](https://arxiv.org/html/2607.03015#A3)\)\.
All policies face the same archived markets, timestamps, prices, and probability estimates; the only difference is the rule used to decide which positive\-edge forecasts become trades and how much capital they receive\. Positions are held to resolution where available, and otherwise marked to market at a fixed evaluation date\.
#### Policies\.
We compare five policies\.i\)*Forecast\-only*trades every positive\-edge forecast at a fixed $10 stake\.ii\)*Edge\-proportional*trades the same set but sizes each stake proportionally to the probability edgeei=pi−qie\_\{i\}=p\_\{i\}\-q\_\{i\}, with total stake normalized to match Forecast\-only \($580\) so that ROI differences reflect allocation, not capital deployed\.iii\)*Edge\-filter*keeps only forecasts whose edge exceeds five percentage points, at a fixed $10 stake\.iv\)*Raven\-Agent \(fixed\)*applies the full execution and risk layer of Raven\-Agent at a fixed $10 stake per surviving trade\. It rejects forecasts when market data is stale, model confidence is low, liquidity is insufficient, or event exposure is excessive\.v\)*Raven\-Agent \(full\)*uses the same execution layer with Raven’s deployed quarter\-Kelly sizing, so capital scales with edge, time to resolution, liquidity, and portfolio risk\.
#### Metrics\.
We report three metrics\. Probability quality is measured by the Brier scoreBrier=1M∑i\(pi−zi\)2\\mathrm\{Brier\}=\\tfrac\{1\}\{M\}\\sum\_\{i\}\(p\_\{i\}\-z\_\{i\}\)^\{2\}on resolved rows\(Brier,[1950](https://arxiv.org/html/2607.03015#bib.bib8)\)\. Trading quality is measured by return on stakeROI=∑iΠi/∑isi\\mathrm\{ROI\}=\\sum\_\{i\}\\Pi\_\{i\}/\\sum\_\{i\}s\_\{i\}\. Risk adjusted return is measured by a stake weighted Sharpe ratioSharpew=r¯w/stdw\(ri\)\\mathrm\{Sharpe\}\_\{w\}=\\bar\{r\}\_\{w\}/\\mathrm\{std\}\_\{w\}\(r\_\{i\}\), whererir\_\{i\}is the realized or mark to market trade return andwi=siw\_\{i\}=s\_\{i\}\. We abbreviateSharpew\\mathrm\{Sharpe\}\_\{w\}asSwS\_\{w\}and report the aggregate∑iΠi\\sum\_\{i\}\\Pi\_\{i\}as profit and loss \(PnL\) in the tables\.
### 4\.2Result Analysis
Table 1:Trading layer ablation on the main archive\. Edge\-proportional total stake is normalized to $580 to match Forecast\-only\.†Brier score computed on each policy’s traded subset of resolved rows\. All policies use the same archived probabilities; the full\-archive Brier is 0\.329 for every policy\. See §[4\.2](https://arxiv.org/html/2607.03015#S4.SS2)for discussion\.
Figure 1:Raven\-Agent’s live Polymarket profile and all\-time profit and loss curve\. Across 20 predictions, the agent currently holds $251\.01 in open positions and has accumulated $53\.92 in cumulative profit\.\(a\)Current open positions: seven NO\-side contracts spanning politics, commodities, sports, entertainment, and public health\.
\(b\)Recent trade history: a mix of NO\-side entries \(crude oil, Eurovision, F1, NBA, Venezuela\)\.
Figure 2:Polymarket dashboard snapshots of Raven\-Agent’s open positions and recent trade history during live deployment\.Raven\-Agent \(full\) is the only policy with positive ROI and the only one with positiveSwS\_\{w\}\([Table1](https://arxiv.org/html/2607.03015#S4.T1)\)\. On Polymarket, Raven\-Agent has also accumulated $53\.92 across 20 predictions in live deployment \([Figures1](https://arxiv.org/html/2607.03015#S4.F1)and[2](https://arxiv.org/html/2607.03015#S4.F2)\)\.
The Edge\-proportional policy shows that sizing alone, without selection, is harmful: concentrating capital on confidently wrong high\-edge predictions amplifies losses to−\-55\.5% ROI, far worse than flat sizing \(−\-10\.7%\)\. Raven\-Agent \(full\) \(\+15\.9%\) avoids this by first removing value\-destroying forecasts, suggesting that selection and risk filtering are important complements to informed sizing\. Among the fixed\-stake policies, each additional trading\-layer component improves ROI, and sizing then turns the same selected set from a small loss into a substantial gain\.
The Selected Brier decreases along the policy sequence because more aggressive filtering retains better\-calibrated rows, not because the forecaster improves \(the full\-archive Brier is 0\.329 for every policy\)\. Bootstrap confidence intervals and leave\-one\-out sensitivity are reported in[AppendixD](https://arxiv.org/html/2607.03015#A4); an expanded 132\-row replay on a fused archive is in[AppendixE](https://arxiv.org/html/2607.03015#A5)\.
## 5Conclusion and discussion
### 5\.1Conclusion
We studied the gap between forecasting and trading on live prediction markets\. We presented Raven\-Agent, a system whose trading layer handles selection, sizing, and risk control on top of a swappable forecaster\. On a controlled replay over a fixed forecast archive, Raven\-Agent achieved the only positive return and risk\-adjusted return among five policies\. The edge\-proportional baseline, which sizes by edge without selection, performs substantially worse than flat sizing \(−\-55\.5% vs\.−\-10\.7%\), suggesting that selection and risk filtering are important complements to informed sizing\.
### 5\.2Discussion
The current results are preliminary\. The sample is modest and the forecaster is held fixed, which isolates the trading layer but leaves the forecaster contribution unmeasured\. A stronger forecaster can be plugged into the same replay infrastructure, and larger archives would strengthen the statistical claims\. Despite these caveats, the replay suggests that explicit selection, sizing, and risk controls materially improve observed outcomes even with the forecaster held constant\.
## References
- G\. W\. Brier \(1950\)Verification of forecasts expressed in terms of probability\.Monthly Weather Review78\(1\),pp\. 1–3\.External Links:[Document](https://dx.doi.org/10.1175/1520-0493%281950%29078%3C0001%3AVOFEIT%3E2.0.CO%3B2)Cited by:[§4\.1](https://arxiv.org/html/2607.03015#S4.SS1.SSS0.Px3.p1.8)\.
- P\. Cheng, J\. Liu, and Y\. Long \(2026\)PolyBench: benchmarking LLM forecasting and trading capabilities on live prediction market data\.arXiv preprint arXiv:2604\.14199\.External Links:[Link](https://arxiv.org/abs/2604.14199)Cited by:[Appendix B](https://arxiv.org/html/2607.03015#A2.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2607.03015#S1.p2.1)\.
- Y\. Jang, J\. Kim, and B\. Zhang \(2025\)The losing winner: an LLM agent that predicts the market but loses money\.NeurIPS 2025 Workshop on Generative AI in Finance\.External Links:[Link](https://nips.cc/virtual/2025/loc/san-diego/132549)Cited by:[Appendix B](https://arxiv.org/html/2607.03015#A2.SS0.SSS0.Px3.p1.3),[§3\.4](https://arxiv.org/html/2607.03015#S3.SS4.p1.1)\.
- E\. Karger, H\. Bastani, Y\. Chen, Z\. Jacobs, D\. Halawi, F\. Zhang, and P\. E\. Tetlock \(2024\)ForecastBench: a dynamic benchmark of ai forecasting capabilities\.arXiv preprint arXiv:2409\.19839\.External Links:[Link](https://arxiv.org/abs/2409.19839)Cited by:[Appendix B](https://arxiv.org/html/2607.03015#A2.SS0.SSS0.Px1.p1.1),[Appendix B](https://arxiv.org/html/2607.03015#A2.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2607.03015#S1.p1.1),[§1](https://arxiv.org/html/2607.03015#S1.p2.1)\.
- J\. L\. Kelly \(1956\)A new interpretation of information rate\.The Bell System Technical Journal35\(4\),pp\. 917–926\.External Links:[Document](https://dx.doi.org/10.1002/j.1538-7305.1956.tb03809.x)Cited by:[§3\.3](https://arxiv.org/html/2607.03015#S3.SS3.p2.1)\.
- K\. Murphy \(2026\)Agentic forecasting using sequential bayesian updating of linguistic beliefs\.arXiv preprint arXiv:2604\.18576\.External Links:[Link](https://arxiv.org/abs/2604.18576)Cited by:[Appendix B](https://arxiv.org/html/2607.03015#A2.SS0.SSS0.Px2.p1.1),[§1](https://arxiv.org/html/2607.03015#S1.p2.1),[§3\.1](https://arxiv.org/html/2607.03015#S3.SS1.p1.1),[§3\.2](https://arxiv.org/html/2607.03015#S3.SS2.p1.2)\.
- Q\. Yang, S\. Mahns, S\. Li, A\. Gu, J\. Wu, and H\. Xu \(2025\)LLM\-as\-a\-prophet: understanding predictive intelligence with prophet arena\.arXiv preprint arXiv:2510\.17638\.External Links:[Link](https://arxiv.org/abs/2510.17638)Cited by:[Appendix B](https://arxiv.org/html/2607.03015#A2.SS0.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2607.03015#S1.p1.1),[§1](https://arxiv.org/html/2607.03015#S1.p2.1)\.
- Z\. Zenget al\.\(2025\)FutureX: an advanced live benchmark for LLM agents in future prediction\.arXiv preprint arXiv:2508\.11987\.External Links:[Link](https://arxiv.org/abs/2508.11987)Cited by:[Appendix B](https://arxiv.org/html/2607.03015#A2.SS0.SSS0.Px1.p1.1)\.
- J\. Zhang, G\. Liu, O\. Johansson, H\. Yitayew, K\. Ohly, and G\. Li \(2026\)Prediction arena: benchmarking AI models on real\-world prediction markets\.arXiv preprint arXiv:2604\.07355\.External Links:[Link](https://arxiv.org/abs/2604.07355)Cited by:[Appendix B](https://arxiv.org/html/2607.03015#A2.SS0.SSS0.Px1.p1.1),[Appendix B](https://arxiv.org/html/2607.03015#A2.SS0.SSS0.Px3.p1.3),[Appendix B](https://arxiv.org/html/2607.03015#A2.SS0.SSS0.Px4.p1.1),[§1](https://arxiv.org/html/2607.03015#S1.p2.1),[§3\.4](https://arxiv.org/html/2607.03015#S3.SS4.p1.1)\.
## Appendix ALive deployment and code
The Raven\-Agent deployment discussed in this paper was run on a live Polymarket account\. For transparency, the public Polymarket profile is available at:[https://polymarket\.com/profile/0x6664e32f79aee42639f73633e40b5a842b07614e](https://polymarket.com/profile/0x6664e32f79aee42639f73633e40b5a842b07614e)\. The accompanying code and redacted experiment artifacts are available at:[https://github\.com/Alchemist\-X/predict\-raven](https://github.com/Alchemist-X/predict-raven)\. The repository contains the agent implementation, replay scripts, and derived tables used in the experiments, with private credentials and non\-public runtime paths removed\.
## Appendix BRelated work
#### Forecasting benchmarks\.
Evaluation of LLM forecasting has moved from static question answering toward live prediction market settings\(Kargeret al\.,[2024](https://arxiv.org/html/2607.03015#bib.bib1); Yanget al\.,[2025](https://arxiv.org/html/2607.03015#bib.bib2); Zeng and others,[2025](https://arxiv.org/html/2607.03015#bib.bib3)\)\. Prophet Arena\(Yanget al\.,[2025](https://arxiv.org/html/2607.03015#bib.bib2)\)reports Brier score, calibration, and market return as three complementary axes\. It shows that frontier language models match a market consensus policy on probability quality but do not break even on return\. PolyBench\(Chenget al\.,[2026](https://arxiv.org/html/2607.03015#bib.bib6)\)couples market snapshots with order book state\. It finds that most evaluated models incur losses under realistic execution, even when their stated probabilities are confident\. Prediction Arena\(Zhanget al\.,[2026](https://arxiv.org/html/2607.03015#bib.bib5)\)reports the same pattern under live trading\. Our work builds on this evaluation infrastructure, but shifts the unit of analysis from the forecaster to the trading policy\. We hold archived forecasts fixed and vary the surrounding decision and risk modules\.
#### Forecasting agents\.
A complementary line of work pushes the forecasting layer itself\. BLF\(Murphy,[2026](https://arxiv.org/html/2607.03015#bib.bib7)\)uses structured linguistic belief states, multi trial aggregation, and hierarchical calibration to substantially improve Brier scores onKargeret al\.\([2024](https://arxiv.org/html/2607.03015#bib.bib1)\)\. These methods are orthogonal to our work\. A stronger forecaster can be inserted into the probability analysis module of Raven\-Agent without changing the rest of the system, and the benefits of the two lines compound\. We do not optimize the forecaster in this paper\. We study what should surround it before its outputs control capital\.
#### Risk aware decision making\.
End to end risk aware policy optimization is widely studied in financial reinforcement learning\. Exploration there is sample intensive and tolerates reversible mistakes\. On live prediction markets, every exploratory action consumes real bankroll on a market that resolves within days, which makes large scale policy training infeasible\. We therefore implement risk constraints as deterministic services outside the language model prompt\. This bounds the worst case loss per decision and produces auditable rejection reasons that learned policies cannot easily match\. Recent concurrent work provides direct empirical evidence for this design choice\.Zhanget al\.\([2026](https://arxiv.org/html/2607.03015#bib.bib5)\)deployed six frontier language models as end\-to\-end prediction market agents on Kalshi, each trading with a $10,000 bankroll over 57 days\. All six lost money, with returns ranging from−\-16% to−\-31%\. The only risk controls that held were hard\-coded constraints \(a 15% per\-market concentration cap and solvency checks\); prompt\-level risk guidance was frequently ignored\.Janget al\.\([2025](https://arxiv.org/html/2607.03015#bib.bib4)\)show that even with a deterministic execution rule, optimizing a proxy forecasting objective \(classification accuracy via RLVR\) can improve accuracy while worsening trading return, with their best\-accuracy agent achieving−\-14\.8% return\. Together, these findings motivate separating the trading layer from the language model and enforcing risk constraints outside the prompt\.
#### Comparison with Prediction Arena\.
Prediction Arena\(Zhanget al\.,[2026](https://arxiv.org/html/2607.03015#bib.bib5)\)and Raven\-Agent differ in a key architectural choice\. Prediction Arena gives each language model end\-to\-end control: the model receives market data and portfolio state, reasons freely with tool access, and directly issues trade orders with no external sizing or risk module\. Raven\-Agent separates the forecasting module from the trading layer, so selection, sizing, and risk enforcement are deterministic components that the language model cannot override\. The two setups are therefore not directly comparable on the same leaderboard: Prediction Arena evaluates the LLM as a complete trading agent, while our replay protocol evaluates the trading layer with the forecaster held fixed\. A controlled comparison would require running Prediction Arena agents on our archived decision set, substituting their end\-to\-end decisions for our trading layer while keeping the same market snapshots and price data\. We leave this to future work\.
## Appendix CArchive coverage
The replay archive records the intermediate outputs produced by Raven\-Agent during deployment\. It starts from raw market scans, then applies filtering, candidate persistence, recommendation generation, and final trading decisions\. The table below reports how many artifacts are available at each stage of this pipeline\. These counts are intended to make the replay scope clear: many markets are scanned, a smaller set is kept as candidates, and only a subset becomes executable trading decisions\.
#### Replay procedure\.
In each archived run, the agent first queried Polymarket markets and filtered them by liquidity, resolution clarity, usable prices, and event overlap\. The candidate set included markets such as crude\-oil threshold contracts, political resignation markets, sports futures, and other high\-liquidity event markets\. For example, one archived run scanned 1,544 Polymarket markets, retained 92 after filtering, and selected 12 candidates for deeper analysis\. The final recommendations in that run focused on the No side of crude oil reaching $100 by the end of March, Bitcoin dipping to $65,000 in March, and Netanyahu leaving office by June 30\.
For each retained candidate, Raven\-Agent estimated the probability that the selected contract would pay out\. It compared this probability with the Polymarket price, computed the edge, ranked opportunities, and proposed a stake\. When the live system decided to trade, it submitted real orders through Polymarket’s CLOB API rather than simulating execution inside the model\. The replay uses the archived forecasts, prices, order metadata, and execution records from these live runs\.
Table 2:Archive coverage used by the replay experiments\. Counts are grouped by the replay pipeline stage\.
## Appendix DRobustness analysis
#### Bootstrap confidence intervals\.
We resample trades with replacement 10,000 times and report 95% percentile intervals for ROI andSwS\_\{w\}\([Table3](https://arxiv.org/html/2607.03015#A4.T3)\)\. Raven\-Agent \(full\) is the only policy whose ROI confidence interval lies entirely above zero\. Edge\-proportional’s interval lies entirely below zero, confirming that the negative result is robust and not an artifact of a few extreme trades\.
Table 3:Bootstrap 95% confidence intervals \(10,000 resamples\)\.
#### Leave\-one\-out sensitivity\.
We compute ROI after dropping each trade in turn \([Table4](https://arxiv.org/html/2607.03015#A4.T4)\)\. Raven\-Agent \(full\) ROI stays positive across all deletions, ranging from \+12\.7% to \+19\.0%, indicating that no single trade drives the result\.
Table 4:Leave\-one\-out ROI sensitivity\.
## Appendix EExpanded replay on a fused archive
To probe robustness on a larger sample, we extend the replay to an earlier Raven\-Agent deployment that produced 90 additional selected recommendations\. The combined 132\-row set yields a positive ROI of \+5\.9% under fixed $10 stakes \([Table5](https://arxiv.org/html/2607.03015#A5.T5)\), confirming the main\-archive direction on a larger sample\. Bootstrap 95% CI for the fused set is \[−\-4\.5%, \+16\.1%\] for ROI and \[−\-0\.07, \+0\.27\] forSwS\_\{w\}\. The earlier archive contains only selected trades \(the rejected diagnostic cannot be extended\), and the archived sizing column is omitted because the earlier deployment used a much larger bankroll \(∼\{\\sim\}$100k vs\.∼\{\\sim\}$1k\), making absolute stake and PnL incomparable\.
Table 5:Fused Raven\-Agent selected\-trade replay \(fixed $10 stake\)\.
## Appendix FExample archived agent traces
Raven\-Agent stores intermediate artifacts during deployment, not only final trades\. These artifacts are useful for auditing the belief\-to\-trade pipeline because they record which markets were scanned, which candidates were filtered, what probability estimates were produced, and why a selected market was considered actionable\.[Figure3](https://arxiv.org/html/2607.03015#A6.F3)shows an example daily market pulse after translation and condensation\. The trace begins with a broad market scan, narrows the pool through filtering and candidate selection, and records the rationale for the final Top 3 recommendations\.[Figure4](https://arxiv.org/html/2607.03015#A6.F4)shows the corresponding probability reasoning trace for one selected market\. It preserves the market\-implied probability, the agent probability, evidence adjustments, settlement\-rule checks, confidence notes, and the resulting trading proposal\.
These figures are not additional quantitative results\. They are included to make the archived decision process concrete and to show that the replay is based on structured records left by the deployed agent rather than on reconstructed explanations after the fact\. Private credentials and non\-public runtime paths are omitted\.
Figure 3:Example Raven\-Agent market screening trace\. The artifact records the daily scan, filtering rules, representative rejected candidates, and final selected markets from one archived market pulse\.Figure 4:Example Raven\-Agent probability reasoning trace for a selected crude oil market\. The artifact records the market\-implied probability, Raven\-Agent probability estimate, evidence adjustments, settlement\-rule checks, and proposed action\.Similar Articles
RAVEN: A Regime-Aware Variable-context Expert Network for Financial Time Series Forecasting
The paper proposes RAVEN, a Mixture-of-Experts framework that adaptively determines temporal context windows for each input sample to handle non-stationary financial time series. It achieves state-of-the-art performance on financial and traffic benchmarks.
Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents
Forecast-Dojo is a replayable environment for benchmarking and training LLM forecasting agents, combining resolved prediction-market questions with dated news to enable repeated evaluation and learning from outcomes.
AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets
This paper introduces AI-Trader, the first fully automated live benchmark for evaluating LLMs in financial decision-making across US stocks, A-shares, and cryptocurrencies. It highlights that general intelligence does not guarantee trading success and emphasizes the importance of risk control in autonomous agents.
@JenovaAIAgent: Prediction Market Analyst is an AI agent that helps you turn noisy event-contract prices on Kalshi and Polymarket into …
Prediction Market Analyst is an AI agent that analyzes prices on Kalshi and Polymarket to provide clear probabilities, edges, and risk assessments for better decision-making in prediction markets.
Autonomous AI trading is harder than it looks — deterministic behavior in live markets nearly broke me
The author details the challenges of building a deterministic autonomous trading agent using a Rust execution layer and Python AI layer with Claude/OpenAI, emphasizing the critical role of hard-coded risk management to prevent emotional or inconsistent trading.