The author built an AI trading agent but found that using LLM for direct trading decisions is unstable, leading to a deterministic system where LLM translates strategies into explicit rules for execution.
I've been building a trading-agent project for a while, and the architecture ended up going in a direction I didn't expect. The first version was pretty straightforward: strategy in plain English → LLM → market context → trade decision For example, you give it something like: "Buy strong momentum stocks, don't let any one position get too large, and sell when momentum weakens." On each run, the LLM gets the portfolio, holdings, quotes, open orders, news, watchlist, etc. and decides whether to buy, sell, or hold. It works, but I ran into a problem that became more obvious the more I used it: the strategy isn't really stable. The model is interpreting the strategy again every time it runs. What exactly does "strong momentum" mean today? Is it the same thing tomorrow? If I ask the model what it would have done six months ago, there's no guarantee I get the same interpretation I would have gotten at the time. That also makes backtesting pretty awkward. So I built a second type of agent. Instead: plain English → LLM → explicit strategy → backtest → deterministic execution The LLM is only used when you create the strategy. It turns the description into an actual structured strategy with things like: universe / screening entry conditions exits position sizing ranking rebalancing After that, there is no LLM involved in the trading loop. The strategy can be inspected and backtested, and the execution engine runs the same rules against the same inputs. I also log the decision at each check, including when the result is simply "no trade." One thing I like about this approach is that the universe doesn't have to be a fixed list of tickers. You can describe something like "hold the 3 sector ETFs with the strongest momentum" and have the screen re-evaluated as part of the strategy. I also have a third version connected to a separate Robinhood account where the LLM is making the trading decisions directly. That one has been useful as an experiment because it makes the difference between the two architectures pretty obvious. So I currently have three flavors: LLM agent: LLM makes the trading decision every run. Rules agent: LLM translates the idea into rules once, then a deterministic engine trades. Brokerage agent: LLM makes decisions and executes through a brokerage connection. The funny part is that after building all of this, I'm less convinced that having an LLM make the actual trading decision is the most interesting use of an agent. Maybe the useful part is having the LLM translate something a human can describe into something explicit, testable, and executable. Or maybe that's just a conventional rule-based trading system with a natural-language interface. Curious what other people building agents think: where do you think an LLM actually adds value in a system like this? I built an AI trading agent, then ended up building one that doesn't use AI to trade I've been building WallStreetClaws as a side project for a while. The original idea was pretty simple: give an LLM a $100k paper account and let it trade based on a strategy written in plain English. The first version, Cortex, works roughly like: strategy → LLM → market context → trade decision Every run, the model gets the portfolio, holdings, quotes, open orders, recent trades, news and watchlist, and decides what to do. It can buy, sell, or hold. I then built a separate version that connects to Robinhood and lets the LLM trade real money. That was where some of the problems became much more obvious. The model is interpreting the strategy from scratch every time it runs. If the strategy says: "Buy strong momentum stocks and sell when momentum weakens." What exactly is "strong" this time? Is it going to interpret that the same way six hours later? What about tomorrow? And that makes backtesting strange too. You're not really backtesting a fixed strategy if the thing making the decision can reinterpret the strategy every time. So I ended up building a second type of agent, which I called Flux. Instead of: strategy → LLM → trade it's: plain English → LLM → structured strategy → backtest → deterministic execution The LLM is only involved when the strategy is created. It turns the description into explicit rules for things like: universe / screening entry conditions exits position sizing ranking rebalancing Once the strategy exists, the LLM is completely out of the trading loop. Flux runs those rules on a schedule and makes the same decision given the same inputs. It also records a decision trace, including why it didn't trade. One thing I wanted to avoid was forcing people to specify a fixed list of stocks. The strategy can define a screen instead. For example, you can say something like "hold the 3 sector ETFs with the strongest momentum," and the universe gets re-evaluated as part of the strategy. It's currently a $100k paper account, so you can actually let the strategy run without connecting a brokerage account. I also built backtesting into the workflow so you can see what the generated rules would have done before putting them into the paper environment. So at this point I have three fairly different agents: Cortex LLM makes the trading decisions. $100k paper account. Flux LLM translates the strategy into rules once. Deterministic engine handles the trading and rebalancing afterward. $100k paper account. Robinhood Agent LLM makes the trading decisions and executes them through a separate Robinhood connection. Real money, so this one is obviously a much more experimental setup. Building all three has made me question the original idea a bit. I started out thinking the interesting thing was "can an LLM trade?" Now I'm wondering if the more useful role for an LLM is actually translating something a human can describe into a strategy that is explicit, inspectable, backtestable and executable. At that point, though, maybe it's just a conventional systematic trading engine with a natural-language interface. I'm curious what people building AI agents think. Does having the LLM translate a human strategy into structured rules feel like a useful agent architecture, or is the interesting part actually keeping the LLM in the trading loop? I've been using this project to explore that question rather than claiming I have the answer.
The author details the challenges of building a deterministic autonomous trading agent using a Rust execution layer and Python AI layer with Claude/OpenAI, emphasizing the critical role of hard-coded risk management to prevent emotional or inconsistent trading.
The author describes an AI agent designed to reproduce production Python crashes using LangGraph, featuring a unique architecture where the LLM plans actions but deterministic Python functions generate the final test code to ensure reliability.
The author argues that the reliability of AI agents comes from deterministic code, not the LLM, and shares five key practices for building trustworthy agents on messy real-world data.
This paper introduces AI-Trader, the first fully automated live benchmark for evaluating LLMs in financial decision-making across US stocks, A-shares, and cryptocurrencies. It highlights that general intelligence does not guarantee trading success and emphasizes the importance of risk control in autonomous agents.
The author built PortfolioLab, an AI agent pipeline for trading that stages models through backtesting, paper trading, and read-only API execution to prevent premature exposure to real money. They are seeking feedback on trust patterns in AI agent architectures for high-stakes applications.