Reinforcement Learning as (Discrete) Potential Theory
Summary
This paper explores the deep connection between reinforcement learning and potential theory, viewing core RL representations and algorithms through Laplace's, Poisson's, and diffusion equations to potentially enhance sample efficiency and formal constraints.
View Cached Full Text
Cached at: 08/19/26, 10:24 AM
# Reinforcement Learning as (Discrete) Potential Theory
Source: [https://arxiv.org/html/2608.17181](https://arxiv.org/html/2608.17181)
Chris ConnollyThanks:Elements of this paper were prepared with assistance from SRI’s subscription to Chat\-GPT 5\.2\.Affiliation:Computer Science LaboratoryAffiliation:SRI International
April 28, 2026
###### Abstract
Reinforcement learning \(RL\) theory fundamentally depends on probability theory through the Markov chain\. There is a deep connection between probability theory and potential theory\. This paper reviews that connection and explores the potential\-theoretic viewpoint for core reinforcement learning representations and algorithms under a fixed\-policy assumption\. This viewpoint may offer a path for improved sample efficiency and formal constraints that can be applied to RL\. When the fixed\-policy assumption is relaxed, the linear potential theory framework can be naturally extended to the nonlinear case\.111Although the material here is presented in terms ofdiscretepotential theory, the extension to continuous state spaces follows from the connection to potential theory\.
## 1Introduction
Reinforcement learning \(RL\) theory fundamentally depends on probability theory through the Markov chain\. There is a deep connection between probability theory and potential theory\. This paper is primarily a review that examines RL through the lens of potential theory\. While many of the concepts described here represent decades of research, the equivalences between probability theory and potential theory offer viewpoints that may help improve the state of the art in reinforcement learning, for example sample efficiency and credit assignment\. Potential theory and probability theory \(and therefore RL\) are fundamentally linked in the sense that Markov chain outcome probabilities are equivalent to potentials that are encountered in natural physical phenomena\[[5](https://arxiv.org/html/2608.17181#bib.bib5),[7](https://arxiv.org/html/2608.17181#bib.bib6)\]\. This equivalence was explored by Doob and is perhaps most clearly and simply described by Doyle and Snell in terms of resistive networks\[[6](https://arxiv.org/html/2608.17181#bib.bib11)\]: a Markov chain state corresponds to a node in an equivalent resistive network, and each resistor in that network represents an entry in the Markov chain transition matrix\. With this equivalence in mind, there are three cases of fixed\-policy222We fix policy so that we can maintain linearity and more clearly demonstrate the direct equivalences\. This assumption can be relaxed by expressing the nonlinear \(e\.g\., policy optimization\) problem in terms of a decomposition into multiple linear subproblems\.Reinforcement Learning \(RL\) that can be expressed in terms of potential theory:
1. 1\.Episodic RL333“Episodic” is synonymous with “absorbing”\.with reward only at terminal states⟺\\LongleftrightarrowLaplace’s Equation
2. 2\.Ongoing control with distributed reward⟺\\LongleftrightarrowPoisson’s Equation
3. 3\.RL problems with nonstationary environments⟺\\LongleftrightarrowThe Diffusion \(Heat\) Equation
All three cases admit both discrete and continuous domains\. Finite difference methods, for example, can be seen as discrete potential problems\. The potential\-theoretic view of RL allows us to consider optimal state\-action sequences asstreamlinesin a flow induced by the value function: the flow of value through state space\. For clarity’s sake most of this discussion will focus on discrete problems on graphs, intuitions for which can be found in Doyle and Snell\[[6](https://arxiv.org/html/2608.17181#bib.bib11)\]\.
Case[1](https://arxiv.org/html/2608.17181#S1.I1.i1)is associated with boundary value problems for Laplace’s Equation \(and Markov chains\)\. These problems are defined solely with respect to boundary conditions \(absorbing states\)\. Case[2](https://arxiv.org/html/2608.17181#S1.I1.i2)is associated with non\-absorbing \(ongoing\) control problems, usually unbounded in time withdistributedreward \(e\.g\. the cart\-pole balancing problem\)\. Case[3](https://arxiv.org/html/2608.17181#S1.I1.i3)is associated with Markov decision processes innonstationaryenvironments, i\.e\. with time\-varying transient and boundary conditions\. Cases[1](https://arxiv.org/html/2608.17181#S1.I1.i1)and[2](https://arxiv.org/html/2608.17181#S1.I1.i2)are the equilibrium asymptotes for case[3](https://arxiv.org/html/2608.17181#S1.I1.i3)\.
The connection between Markov chains and potential theory is well known within certain communities: harmonic functions, Poisson equations, and Green’s functions444In potential theory, a Green’s Function maps boundary conditions to interior solutions\.correspond to hitting probabilities, expected costs, and occupancy measures of Markov chains respectively\[[6](https://arxiv.org/html/2608.17181#bib.bib11),[8](https://arxiv.org/html/2608.17181#bib.bib12)\]\. In RL, these structures appear directly in value functions and temporal\-difference learning\[[14](https://arxiv.org/html/2608.17181#bib.bib7),[11](https://arxiv.org/html/2608.17181#bib.bib8)\]\. The potential\-theoretic view has value in identifying formal properties of RL structures like value or Q functions, especially in the absorbing case\. Furthermore, the equivalence suggests that physical \(analog\) substrates offer alternatives to digital solutions\. Features of RL appear to have homologues in the neurobiology of action selection; the most pronounced being the dopamine Reward Prediction Error \(RPE\) signal discovered and researched by Wolfram Schultz\[[12](https://arxiv.org/html/2608.17181#bib.bib13)\]among others\. With this homologue in mind, it is tempting to speculate a relationship between Dayan’s Successor Representation555The Successor Representation is essentially a Green’s Function for a Markov chain\. It allows one to compute the value function by evaluating the product of the SR matrix and the reward function\.\[[4](https://arxiv.org/html/2608.17181#bib.bib9)\]and the presence of electrotonic coupling666A neuroscientific term that includes “resistive coupling”, but more generally a coupling that also admits ions and small molecules\.in the striatum, possibly offering a fast diffusion mechanism for mapping a reward landscape into a value function\. Astrocyte777Astrocytes are glial cells that can mediate ionic stabilization of the extracellular milieu\. Recent research suggests that they may serve a critical computational role in cognition\.networks could provide the required diffusion substrate\.
### 1\.1Preliminaries
#### 1\.1\.1Laplace, Poisson, and Heat Equations
In potential theory, apotential functionv\(x\)v\(x\)\(e\.g\., temperature, electrical potential, fluid potential\) is a function over a variable in some domain𝒮\\mathcal\{S\}that is typically a subset of the reals:v\(x\),x∈S⊂Rnv\(x\),x\\in S\\subset R^\{n\}\. The functionv\(x\)v\(x\)satisfies a partial differential equation \(PDE\) subject to boundary conditions, source conditions, or initial conditions \(or combinations thereof\)\. The Laplacian operator is written:
∇2:=∑i∂2∂xi2\\nabla^\{2\}:=\\sum\_\{i\}\\frac\{\\partial^\{2\}\}\{\\partial x\_\{i\}^\{2\}\}\(1\)and is the spatial diffusion term\.
The three equations we discuss here are Laplace’s Equation:
∇2v\(x\)=0\\nabla^\{2\}v\(x\)=0\(2\)
Poisson’s Equation:
∇2v\(x\)=f\(x\)\\nabla^\{2\}v\(x\)=f\(x\)\(3\)
and the Heat Equation, sometimes known as the Diffusion Equation:
∇2v\(x,t\)=α∂v\(x,t\)∂t\\nabla^\{2\}v\(x,t\)=\\alpha\\frac\{\\partial v\(x,t\)\}\{\\partial t\}\(4\)
Laplace’s equation models physical phenomena when there are no internal sources in the interior of the domain𝒮\\mathcal\{S\}\. Constraints are imposed only at the boundary∂𝒮\\partial\\mathcal\{S\}of the domain\. Laplace’s equation is used to model steady\-state potentials \(electical, fluid, thermal, etc\.\), and yields minimum\-energy, maximum\-entropy functions\. Solutions to Laplace’s equation exist and are unique when proper boundary conditions are established; solutions are also known asharmonic functions\. Harmonic functions exhibit themean\-valueproperty\. This means that the value of the function at any point is the \(possibly weighted\)meanof the values at neighboring points\. One might, for example, establish temperature boundary conditions on the edges of a hot plate\. The interior of the plate will then reach an equilibrium such that the temperature on the plate as a function of position is a harmonic function\. The same is true of resistive networks and irrotational, frictionless fluid potentials\.
Poisson’s equation admits a source functionf\(x\)f\(x\)that results in solutions that are not harmonic, but correspond well to RL problems that involve distributed reward\. Poisson’s equation extends Laplace’s equation by considering physical phenomena that involve ongoing \(but stationary\) perturbation of the diffusion equilibrium, such as current, fluid or heat sources\. Solutions to Poisson’s equation still typically represent steady\-state phenomena\. Solutions to Poisson’s equation exist and are unique up to an additive constant\. One can enforce uniqueness by constraining solutions to be zero\-mean\.
Finally, the Heat Equation allows us to consider time\-varying conditions subject to a diffusivity constantα\\alpha\. When the time\-derivative vanishes, we are left with one of Laplace’s Equation, Poisson’s Equation, or a combination \(superposition\) of the two\. representing the equilibrium solution\.
#### 1\.1\.2The MDP
A discrete\-time Markov Decision Process \(MDP\) is a core structure in RL, and is a mathematical model for sequential decision\-making where the next state depends only on the current state and chosen action \(the Markov property\)\. The MDP can be defined as a tuple:\(𝒮,𝒜,P,r,γ\)\(\\mathcal\{S\},\\mathcal\{A\},P,r,\\gamma\)where:
- •𝒮\\mathcal\{S\}is a set of states\.
- •𝒜\\mathcal\{A\}is a set of actions\.
- •PPis a transition operatorP:𝒮×𝒜→𝒮P:\\mathcal\{S\}\\times\\mathcal\{A\}\\rightarrow\\mathcal\{S\}describes the transition probabilityP\(s′\|s,a\)P\(s^\{\\prime\}\|s,a\)that a process in statessexecuting actionaawill transition to states′s^\{\\prime\}\.
- •rris a per\-state reward\.
- •γ\\gammais a discounting factor for reward expectation\.
We partition𝒮\\mathcal\{S\}into an absorbing set∂𝒮\\partial\\mathcal\{S\}and the set of transient states𝒯=𝒮∖∂𝒮\\mathcal\{T\}=\\mathcal\{S\}\\setminus\\partial\\mathcal\{S\}\. Generally, absorbing sets are useful for episodic RL, while∂𝒮=∅\\partial\\mathcal\{S\}=\\emptysetis possible for ongoing RL control problems\. Hybrid MDPs are also possible\. For clarity we assume a fixed policyπ\\pito establish the basic equivalences, initially without explicit reference to action spaces\. A policyπ\(a\|s\)\\pi\(a\|s\)is the probability of selecting an actionaagiven a statess\. A policy can be deterministic\. SincePPis a stochastic matrix over an augmented state\-action space, we can always marginalize over actions at each state to construct a stochastic transition operator888Recall that each row of astochastic matrixsums to 1\.for the pure state space\.
Within∂𝒮\\partial\\mathcal\{S\}, we can consider two subsets999We choose two subsets for simplicity’s sake, but a much greater variety of boundary conditions can be chosen, based on problem constraints and desired outcomes\.𝒜⊂∂𝒮\\mathcal\{A\}\\subset\\partial\\mathcal\{S\}andℬ⊂∂𝒮\\mathcal\{B\}\\subset\\partial\\mathcal\{S\},𝒜∪ℬ=∂𝒮\\mathcal\{A\}\\cup\\mathcal\{B\}=\\partial\\mathcal\{S\}and the value functionV\(s\)V\(s\)over statess∈𝒮s\\in\\mathcal\{S\}in the MDP\. We assignV\(s\)=1,s∈𝒜V\(s\)=1,s\\in\\mathcal\{A\}andV\(s\)=0,s∈ℬV\(s\)=0,s\\in\\mathcal\{B\}\. These provide theboundary conditionsfor the value functionV\(s\)V\(s\)over all states\. The value function resulting from this set of conditions is exactlythe probability that a Markov process starting at any state s reaches set𝒜\\mathcal\{A\}before reachingℬ\\mathcal\{B\}\.
Transition matrices for both episodic and ongoing RL admit a form of Dayan’s Successor Representation, a linear operator𝐌\\bf\{M\}that maps reward vectors directly to value functionsVV\[[4](https://arxiv.org/html/2608.17181#bib.bib9)\]\. Gradient ascent onVVmaximizes success probability\. When thinking of this problem in terms of success probability, we assume that the reward functionrris normalized to the interval\[0,1\]\[0,1\], and the resulting value function provides the probabiity of success at each state\. If this normalization constraint is relaxed, the value function is an expectedvaluethat might represent a payoff, for example\.
A fixed\-policy Markov Decision Process \(MDP\) is a Markov Reward Process \(MRP\) with transition matrixP=PπP=P\_\{\\pi\}and expected one\-step reward vectorr=rπr=r\_\{\\pi\}\. The value function is the expectation
Vπ\(s\)=𝔼π\[∑t≥0γtRt\+1∣S0=s\]V^\{\\pi\}\(s\)=\\mathbb\{E\}\_\{\\pi\}\\left\[\\sum\_\{t\\geq 0\}\\gamma^\{t\}R\_\{t\+1\}\\mid S\_\{0\}=s\\right\]\(5\)and satisfies the Bellman equation101010Our treatment usesVVfor RL value andvvwhen the context is potential theory:V≡vV\\equiv vis the potential and value function over state\.
V=r\+γPV\.V=r\+\\gamma PV\.\(6\)where the discount factorγ∈\[0,1\]\\gamma\\in\[0,1\]\. Equivalently,
\(I−γP\)V=r\.\(I\-\\gamma P\)V=r\.\(7\)
The linear system in Equation[7](https://arxiv.org/html/2608.17181#S1.E7)is the discrete analogue of Poisson’s equation \(Equation[3](https://arxiv.org/html/2608.17181#S1.E3)\) where the reward functionr\(x\)=f\(x\)r\(x\)=f\(x\)plays the role of the source functionf\(x\)f\(x\)in Poisson’s Equation\. Equation[6](https://arxiv.org/html/2608.17181#S1.E6), while it is assumed to be in matrix form here, is also the formula for a single step of Barto and Sutton’sTemporal Differencing\(TD\) credit assignment algorithm\[[14](https://arxiv.org/html/2608.17181#bib.bib7)\]\.
Figure 1:Discrete harmonic value function on an 80x80 grid domain where absorbing states are shown with P=1 and P=0\. In this example, values at grid edges are allowed to float\.
## 2Case by Case Equivalences
### 2\.1Discrete Laplace Equation: Harmonic Value Functions
Harmonic functions are solutions to Laplace’s equationandare right eigenvectors for absorbing Markov chains\. Laplace’s equation is a model for an equilibrium potential over a domain whose interior has no sources\. For example, Laplace’s equation describes ideal fluid flow in a domain: low\-probability states are sources, and high\-probability states are sinks\. A flow defined by this potential yieldsstreamlines\. Following a streamline maximizes the probability of reaching a sink\. Laplace’s equation constrains to zero the sum of second derivatives of the value functionv\(x\)v\(x\):
subject to boundary conditions\. Whenxxis a discrete lattice, Equation[8](https://arxiv.org/html/2608.17181#S2.E8)can be turned into a finite difference equation, wherev\(xk\)v\(x\_\{k\}\)for somekkis simply the weighted average of neighboring points on the lattice\. If we use a transition matrixPPto represent the averaging kernel, we can define thediscreteLaplacian as:
∇2:=I−P\.\\nabla^\{2\}:=I\-P\.\(9\)
A functionvvis harmonic if
v=Pv⟺∇2v=0\.v=Pv\\quad\\Longleftrightarrow\\quad\\nabla^\{2\}v=0\.\(10\)
Fixed\-point iteration ofPPonvvis equivalent to a Jacobi or Gauss\-Seidel iteration for numerically solving the corresponding potential problem\. The action ofPPreplaces each state’s value with the weighted average of neighboring states’ values\. This process converges by virtue of the fact thatPPis stochastic\.
Note thatvvis the firstrighteigenvector of the transition operatorPP\. In contrast to Equation[10](https://arxiv.org/html/2608.17181#S2.E10), thelefteigenvector:
corresponds to thestationary distributionof the transition operator\. The stationary distribution provides thevisit ratesof a non\-absorbing Markov chain\. Generally, the Poisson case \(distributed reward\) admits a nontrivial stationary distribution but a trivial hitting probability \(unless some states are absorbing\)\. The Laplace case \(absorbing\) admits trivial stationary distributions: visit rates are concentrated at absorbing states \(with transient states’ distributions vanishingly small\), but the hitting probabilities are nontrivial or constant\.
### RL Interpretation \(Terminal\-Only Reward\)
Let𝒮\\mathcal\{S\}be the set of all states of an absorbing Markov chain and let∂𝒮\\partial\\mathcal\{S\}be the set of absorbing states\. Let𝒯=𝒮∖∂𝒮\\mathcal\{T\}=\\mathcal\{S\}\\setminus\\partial\\mathcal\{S\}be the set oftransient\(interior\) states\.111111That is, the set difference of S and its absorbing states\.In an episodic MDP where reward occurs only at terminal states∂𝒮\\partial\\mathcal\{S\},transientstatess∈𝒯s\\in\\mathcal\{T\}satisfy
V\(s\)=∑s′∈𝒮Pπ\(s,s′\)V\(s′\)\.V\(s\)=\\sum\_\{s^\{\\prime\}\\in\\mathcal\{S\}\}P\_\{\\pi\}\(s,s^\{\\prime\}\)V\(s^\{\\prime\}\)\.\(12\)or in vector form:
ThusVV\(orvv\) is harmonic on transient states, with Dirichlet boundary condition
V\(s\)=g\(s\),s∈∂𝒮\.V\(s\)=g\(s\),\\quad s\\in\\partial\\mathcal\{S\}\.\(14\)where g\(s\) is the reward value when absorption occurs\.
This is precisely the discrete Dirichlet problem, where boundary conditions can be interpreted probabilistically\. For example, suppose∂𝒮\\partial\\mathcal\{S\}is partitioned into two sets,WWandLL, such that:
g\(x\)=1,x\\displaystyle g\(x\)=1,\\quad x∈W\\displaystyle\\in W\(15\)g\(y\)=0,y\\displaystyle g\(y\)=0,\\quad y∈L\\displaystyle\\in L\(16\)Then the value function at any transient state provides the probability that a random process will hit a state in setWW\(win\) before hitting any state in setLL\(lose\)\. Markov chain harmonic functions are therefore also called “hitting probabilities”\[[7](https://arxiv.org/html/2608.17181#bib.bib6)\]\. A deterministic policy that seeks to maximize the probability of success will therefore ascend the value gradient to reach setWW\. Effectively, this recipe followsthe flow of valueinstreamlinesthat maximize reward\.
### Formal Properties of Harmonic Functions
Any value function of an absorbing MDP is harmonic\. This establishes some strong necessary conditions for value functions:
1. 1\.Mean Value Property:The value at any transient state is the \(weighted\) mean of values at neighboring states\.
2. 2\.Min\-Max Property:In any neighborhood𝒩\\mathcal\{N\}of a transient state, the value function achieves its minimum and maximum valueonlyon the boundary∂𝒩\\partial\\mathcal\{N\}of that neighborhood\.
It is obvious that all linear functions are also harmonic\. Ifg\(x\)=cg\(x\)=cfor all boundary statesx∈∂𝒮\{x\\in\\partial\\mathcal\{S\}\}, then the resulting value function must also be the constant c\. These properties suggest strong convergence tests for value functions\. In the absorbing case, if a TD\-derived value function exhibits local extrema in the set of transient states, then it has not converged properly\. Likewise, the residualPvn−vn\+1Pv\_\{n\}\-v\_\{n\+1\}provides a kind of convergence score that indicates whether more training is needed\.
### 2\.2Discrete Poisson Equation: Rewards as Sources
The Poisson variant of RL admits distributed rewards that are not necessarily limited to absorbing states\. This corresponds to a physical problem with sources \(charge, fluid, etc\.\) in the interior of the domain\. In the RL case, sources are rewards\. The Poisson case corresponds to ongoing optimal control problems of the sort that are familiar to the RL community in environments like the cart\-pole balancing problem\.
Figure 2:A distributed reward function for a Poisson\-style RL problem\.With per\-step rewardrr, the undiscounted Bellman equation becomes
V=r\+PV\\displaystyle V=r\+PV\(17\)\(I−P\)V=r\\displaystyle\(I\-P\)V=r\(18\)or
∇2V=r\.\\nabla^\{2\}V=r\.\(19\)
This is the discrete Poisson equation: the reward acts as a distributed source term\.
### Discounted Case
Forγ<1\\gamma<1,
\(I−γP\)V=r\.\(I\-\\gamma P\)V=r\.\(20\)The introduction of distributed rewards that do not involve absorbing sets results in a process that is unbounded in time and can be used to solve persistent control problems \(e\.g\., cart\-pole balancing\)\.
Equation[20](https://arxiv.org/html/2608.17181#S2.E20)corresponds to a “killed” or discounted Markov chain\[[8](https://arxiv.org/html/2608.17181#bib.bib12),[11](https://arxiv.org/html/2608.17181#bib.bib8)\]\. This in turn corresponds to a subharmonic function, but begs the interpretation of discounting in the context of potential theory\. For the Laplace case in three dimensions, a point source away from boundary conditions will exhibit a potential that roughly obeys the inverse square law, so a hyperbolic discounting is immediately available in the potential\-theoretic formulation whenγ=1\\gamma=1\.121212This has implications for modeling the exact nature of discounting of future rewards in cognitive neuroscience experiments\.
### 2\.3The Heat Equation and Temporal Differencing
The Heat Equation represents a more general expression of the RL problem, and furthermore admits nonstationary environments\. The classical Heat Equation relates the first time derivative of value to the spatial Laplacian – it says that the difference ofv\(x,t\)v\(x,t\)in time \(v\(x,t\+1\)−v\(x,t\)v\(x,t\+1\)\-v\(x,t\)\) is equal to the spatial difference between the spatial mean ofvvand its current value at all states:
∂v∂t=α∇2v\\frac\{\\partial v\}\{\\partial t\}=\\alpha\\nabla^\{2\}v\(21\)describing smoothing of value over time131313The Black\-Scholes options pricing model is a Heat Equation with an advection term\., whereα\\alphais equivalent tothermal diffusivityof a material\. The value function in this case can be defined subject to boundary conditions, initial conditions, and time\-varying sources\.
For RL, iterative policy evaluation
Vk\+1=r\+γPVkV\_\{k\+1\}=r\+\\gamma PV\_\{k\}\(22\)
is a relaxation process that in the time limitt→∞\{t\\rightarrow\\infty\}converges to the Poisson solution\. Rewriting terms of equation[22](https://arxiv.org/html/2608.17181#S2.E22), this is analogous to a discrete Heat Equation with source andα=1\\alpha=1:
Vk\+1−Vk=r−\(I−γP\)Vk\.V\_\{k\+1\}\-V\_\{k\}=r\-\(I\-\\gamma P\)V\_\{k\}\.\(23\)where the left hand side corresponds to a discrete time derivative, and the right\-hand side corresponds to the Poisson resolvent\[[16](https://arxiv.org/html/2608.17181#bib.bib10)\]\.
In continuous time, lettingL:=P−IL:=P\-I, one obtains
ddtu\(t\)=Lu\(t\)\+r−κu\(t\),\\frac\{d\}\{dt\}u\(t\)=Lu\(t\)\+r\-\\kappa u\(t\),\(24\)whereκ\\kappacorresponds to discounting\. Thus, value functions arise as steady states of forced diffusion with decay\.
## 3Green’s Functions: The Fundamental Matrix and Dayan’s Successor Representation
In a transition matrix, absorbing states∂𝒮\\partial\\mathcal\{S\}are those with no outward transitions, and hence have a single 1 in their rows\. By reordering states and the corresponding rows and columns in the transition matrix, a block structure can be constructed for MDPs\. This helps us find a kind of Green’s function or equivalently, Dayan’s Successor Representation for RL\[[4](https://arxiv.org/html/2608.17181#bib.bib9)\]\. These representations allow us to compute value \(or Q\) given only knowledge of the reward function\. Although computed slightly differently, computation of the “Green’s Function” for episodic and non\-abosorbing MDPs is straightforward\.
### 3\.1Absorbing \(Episodic\) MDPs
Partition states into transient𝒯\\mathcal\{T\}and terminal∂𝒮\\partial\\mathcal\{S\}and sortPPto group these states so that transient states are in the upper left block\. The sorted transition matrix then has the following block form:
P=\[QB0I\]\.P=\\begin\{bmatrix\}Q&B\\\\ 0&I\\end\{bmatrix\}\.\(25\)
In the absorbing case, we can compute value from reward using this relationship:
\(I−Q\)V𝒯=Bg\.\(I\-Q\)V\_\{\\mathcal\{T\}\}=Bg\.\(26\)whereV𝒯V\_\{\\mathcal\{T\}\}is the value function over transient states andg\(s\)g\(s\)is a function over absorbing states that \(when normalized to\[0,1\]\[0,1\]\) represents outcome \(absorption\) probabilities that can be configured depending on the problem to be solved\.
### 3\.2MDPs with Distributed Reward: The Poisson Case \(Step Rewards\)
For ongoing control problems, we consider a non\-absorbing chain\. Under this case,P=QP=Qsince there are no absorbing states \(i\.e\.,QQfillsPP\)\. The discrete Poisson equation then looks like this:
The fundamental matrix is the result of repeated action of Q for progressively more iterations\. It can be expressed as a Neumann series that converges to the fundamental matrix𝒩\\mathcal\{N\}:
𝒩=\(I−Q\)−1=∑t≥0Qt\\mathcal\{N\}=\(I\-Q\)^\{\-1\}=\\sum\_\{t\\geq 0\}Q^\{t\}\(28\)and acts as a Green’s function for the Markov chain potential:
V=𝒩rV=\\mathcal\{N\}r\(29\)
Without other constraints, solutions to Poisson’s equation are only unique up to an additive constant, so fixed\-point iteration for this case can only converge if something like a zero\-mean constraint is imposed\. Define average rewardr¯\\bar\{r\}and differential valuehh:
\(I−P\)h=r−r¯𝟏\.\(I\-P\)h=r\-\\bar\{r\}\\mathbf\{1\}\.\(30\)A normalization condition \(e\.g\.h\(s0\)=0h\(s\_\{0\}\)=0\) fixes the additive constant\[[11](https://arxiv.org/html/2608.17181#bib.bib8)\]and enforces uniqueness of the solution without risking divergence of the solution\.
## 4Successor Representation
Dayan’s\[[4](https://arxiv.org/html/2608.17181#bib.bib9)\]successor representation \(SR\) was developed in the context of reinforcement learning and resembles the fundamental matrix with the twist that exponential discounting can be applied to the transition operatorPPas follows:
M=∑t≥0\(γP\)t=\(I−γP\)−1\.M=\\sum\_\{t\\geq 0\}\(\\gamma P\)^\{t\}=\(I\-\\gamma P\)^\{\-1\}\.\(31\)Then
When discounting is applied through the factorγ∈\[0,1\]\\gamma\\in\[0,1\]:
\(I−γP\)V=r\\displaystyle\(I\-\\gamma P\)V=r\(33\)\(I−γP\)−1r=V\\displaystyle\(I\-\\gamma P\)^\{\-1\}r=V\(34\)This equation has a unique solution forγ<1\\gamma<1by virtue of exponential discounting\.
## 5Eligibility Trace to Eligibility Field
Traditional TD distributes credit \(the value update\) along recent state traces \(trajectories explored during RL learning\)\. This is similar to a Monte Carlo solution to the potential\-theory counterpart\. Eligibility traces reveal state\-space topology along chains of experience–theeligibility traces\. However, the structure of the underlying transition operator can be learned \(estimated\) as an RL agent explores state space\. The potential\-theoretic view suggests that with a learned transition structure, an eligibilityfieldcan be defined that exploits this learned structure to update a significantly larger set of states than TD alone can update\.
Eligibility distributed over a trace is defined by TD\(λ\\lambda\) as:
et\(s\)=γλet−1\(s\)\+𝟏\{St=s\}\.e\_\{t\}\(s\)=\\gamma\\lambda e\_\{t\-1\}\(s\)\+\\mathbf\{1\}\\\{S\_\{t\}=s\\\}\.\(35\)
where𝟏\{St=s\}\\mathbf\{1\}\\\{S\_\{t\}=s\\\}represents the one\-hot encoding of state\.
Eligibility may be generalized to a fielde′e^\{\\prime\}using the partially learned transition operatorP⊤P^\{\\top\}:
et\+1′=γλP⊤et′\+ϕ\(St\+1\),e^\{\\prime\}\_\{t\+1\}=\\gamma\\lambda P^\{\\top\}e^\{\\prime\}\_\{t\}\+\\phi\(S\_\{t\+1\}\),\(36\)whereϕ\(St\+1\)\\phi\(S\_\{t\+1\}\)is an injection vector \(e\.g\., one\-hot atSt\+1S\_\{t\+1\}\)\. Compared to the classic accumulating traceet\+1=γλet\+ϕ\(St\+1\)e\_\{t\+1\}=\\gamma\\lambda e\_\{t\}\+\\phi\(S\_\{t\+1\}\), this inserts a spatial propagation step viaP⊤P^\{\\top\}\. Eligibility not only decays along chains, it can be spread backward along thegraphdefined byP⊤P^\{\\top\}so credit diffuses backward across the learned Markov transition structure\.
## 6Summary and Observations
This discussion has been limited to RL with a fixed policy\. More generally, the max operation in policy iteration \(and the max\-min operation in game theory\) results in nonlinearities that lead to nonlinear potential theory, which is out of the scope of this paper\. Nonetheless, reinforcement learning with fixed policies can be viewed with respect to discrete potential theory:
- •Laplace equation: terminal\-only rewards⇒\\Rightarrowharmonic functions / hitting probabilities\.
- •Poisson equation: rewards as sources\.
- •Heat equation: dynamic relaxation toward equilibrium \(to Laplace or Poisson\)\.
- •Successor representation: Green’s function of the Markov operator\.
When the state space is large enough, harmonic value functions can be prone to barren plateaus, especially near saddle points\. Laplace’s equation in particular allows us to imagine optimal state sequences asdiscrete streamlinesthat maximize goal\-reaching success\. A similar view motivated the use of harmonic functions as potentials for robot motion planning\[[15](https://arxiv.org/html/2608.17181#bib.bib2),[3](https://arxiv.org/html/2608.17181#bib.bib3),[13](https://arxiv.org/html/2608.17181#bib.bib4)\]\.
Formal guarantees for RL:Episodic RL problems yield value functions that are harmonic over transient states\. The mean\-value and min\-max properties of harmonic functions guarantee that when fully converged, harmonic value functions should exhibit no local minima\. This also suggests tests for RL convergence and correctness\. The mean\-value property additionally suggests interpolation strategies\. Harmonic interpolants \(e\.g\., harmonic polynomials\[[2](https://arxiv.org/html/2608.17181#bib.bib1)\]\) can be employed to interpolate value functions and collapse barren plateaus in state space\.
Sample Efficiency:The potential\-theoretic viewpoint suggests possible acceleration of RL: The TD eligibilitytraceover a recent one\-dimensionalchainof states becomes a multidimensional eligibilityfieldover learned transition structures, suggesting faster credit distribution over a larger set of states and thus better sample efficiency\. The viewpoint also admits analog computing substrates for solving RL problems\[[13](https://arxiv.org/html/2608.17181#bib.bib4)\]\. For example, a resistive network effectively computes value functions\. For the episodic \(absorbing\) case, the TD update rule can be used to learn \(modify\) voltages at absorbing states that correspond to terminal reward probabilities\. The network itself distributes credit through the learned transition operator\.
Neurobiology:Although beyond the scope for this paper, there are intriguing parallels between RL theory as described here and biological reinforcement learning as observed in midbrain dopamine signaling\. A major recipient of dopamine signals is a subcortical structure: the striatum\. This structure also receives immense convergent projections from the cortex and has long been associated with action selection\. Dopamine signals in the ventral striatum are associated with the RPE, which resembles temporal difference in RL \(received minus expected reward\[[12](https://arxiv.org/html/2608.17181#bib.bib13)\]\. The RPE as observed by Schultz and colleagues is a transient pulse of dopamine delivered to the striatum during trial\-and\-error or implicit learning and corresponds closely to the TD training signal at terminal states of an MDP\. This transient diminishes and ultimately disappears141414More precisely, the transient migrates backward in time to the appearance of predictive cues\.during habit formation as the expected reward approaches actual reward\.
The relationship between RL and potential theory raises the possibility that an electrical network within the striatum might contribute to fast computation of value functions\. One possible neural substrate for this computation is the electrically\-coupled astrocytic syncytium surrounding striatal projection neurons\[[10](https://arxiv.org/html/2608.17181#bib.bib16)\]\. Once thought of solely as glial support for extracellular ionic stability, astrocytes are now felt to be more directly involved in neural computation\[[9](https://arxiv.org/html/2608.17181#bib.bib14),[1](https://arxiv.org/html/2608.17181#bib.bib15)\]\. In the striatum, astrocytes are in a position to modulate membrane potentials in a way that resembles state values or Q functions in RL\.
## References
- \[1\]A\.R\. R\. Abiero, J\.M\. Gladding, J\.A\. Iredale, H\.R\. Drury, E\.E\. Manning, C\.V\. Dayas, A\. Dhungana, K\. Ganesan, K\. Turner, S\. Becchi, M\.D\. Kendig, C\. Nolan, B\. Balleine, A\. Castorina, L\. Cole, K\.J\. Clemens, and L\.A\. Bradfield\(2026\)Dorsomedial striatal neuroinflammation causes excessive goal\-directed action control by disrupting astrocyte function\.Neuropsychopharmacology51,pp\. 486–496\.Cited by:[§6](https://arxiv.org/html/2608.17181#S6.p6.1)\.
- \[2\]\(1991\)Harmonic function theory\.Graduate Texts in Mathematics, Vol\.137,Springer\-Verlag,New York\.Cited by:[§6](https://arxiv.org/html/2608.17181#S6.p3.1)\.
- \[3\]C\. I\. Connolly and R\. A\. Grupen\(1993\)The applications of harmonic functions to robotics\.Journal of Robotic Systems10\(7\),pp\. 931–946\.Cited by:[§6](https://arxiv.org/html/2608.17181#S6.p2.1)\.
- \[4\]P\. Dayan\(1993\)Improving generalization for temporal difference learning: the successor representation\.Neural Computation\.Cited by:[§1\.1\.2](https://arxiv.org/html/2608.17181#S1.SS1.SSS2.p4.1),[§1](https://arxiv.org/html/2608.17181#S1.p4.1),[§3](https://arxiv.org/html/2608.17181#S3.p1.1),[§4](https://arxiv.org/html/2608.17181#S4.p1.1)\.
- \[5\]J\. Doob\(1984\)Classical potential theory and its probabilistic counterpart\.Springer\-Verlag,New York\.Cited by:[§1](https://arxiv.org/html/2608.17181#S1.p1.1)\.
- \[6\]P\. Doyle and J\. L\. Snell\(1984\)Random walks and electric networks\.Carus Monographs in Mathematics,American Mathematical Society\.Cited by:[§1](https://arxiv.org/html/2608.17181#S1.p1.1),[§1](https://arxiv.org/html/2608.17181#S1.p2.1),[§1](https://arxiv.org/html/2608.17181#S1.p4.1)\.
- \[7\]J\. G\. Kemeny, J\. L\. Snell, and A\. W\. Knapp\(1976\)Denumerable markov chains\.Graduate Texts in Mathematics, Vol\.40,Springer\-Verlag,New York\.Cited by:[§1](https://arxiv.org/html/2608.17181#S1.p1.1),[§2](https://arxiv.org/html/2608.17181#S2.SSx1.p2.2)\.
- \[8\]D\. Levin and Y\. Peres\(2nd ed\., 2017\)Markov chains and mixing times\.Cited by:[§1](https://arxiv.org/html/2608.17181#S1.p4.1),[§2](https://arxiv.org/html/2608.17181#S2.SSx3.p2.1)\.
- \[9\]N\. T\. M\. Santello and A\. Volterra\(2019\)Astrocyte function from information processing to cognition and cognitive impairment\.Nature Neuroscience22,pp\. 154–166\.Cited by:[§6](https://arxiv.org/html/2608.17181#S6.p6.1)\.
- \[10\]J\. Pai, F\. Sogukpinar, K\. Ogasawara, G\. J\. Smith, F\. R\. Fiocchi, Y\. Dai, Y\. Wu, M\. J\. Frank, S\. Ching, F\. Lucantonio, T\. Papouin, M\. Pignatelli, N\. Hiratani, and I\. E\. Monosov\(2026\)Ventral striatal astrocytes contribute to reinforcement learning\.bioRxiv doi: https://doi\.org/10\.1101/2025\.10\.20\.683205\.Cited by:[§6](https://arxiv.org/html/2608.17181#S6.p6.1)\.
- \[11\]M\. Puterman\(1994\)Markov decision processes\.Cited by:[§1](https://arxiv.org/html/2608.17181#S1.p4.1),[§2](https://arxiv.org/html/2608.17181#S2.SSx3.p2.1),[§3\.2](https://arxiv.org/html/2608.17181#S3.SS2.p4.2)\.
- \[12\]W\. Schultz, P\. Dayan, and P\. R\. Montague\(1997\)A neural substrate of prediction and reward\.Science275,pp\. 1593–1599\.Cited by:[§1](https://arxiv.org/html/2608.17181#S1.p4.1),[§6](https://arxiv.org/html/2608.17181#S6.p5.1)\.
- \[13\]M\. R\. Stan, W\. P\. Burleson, C\. I\. Connolly, and R\. A\. Grupen\(1994\)Analog VLSI for robot path planning\.Journal of VLSI Signal Processing8\(1\),pp\. 61–73\.Cited by:[§6](https://arxiv.org/html/2608.17181#S6.p2.1),[§6](https://arxiv.org/html/2608.17181#S6.p4.1)\.
- \[14\]R\. Sutton and A\. Barto\(1998\)Reinforcement learning: an introduction\.MIT Press,Cambridge, MA\.Cited by:[§1\.1\.2](https://arxiv.org/html/2608.17181#S1.SS1.SSS2.p6.1),[§1](https://arxiv.org/html/2608.17181#S1.p4.1)\.
- \[15\]L\. Tarassenko and A\. Blake\(1991\)Analogue computation of collision\-free paths\.InProceedings of the 1991 IEEE International Conference on Robotics and Automation,pp\. 540–545\.Cited by:[§6](https://arxiv.org/html/2608.17181#S6.p2.1)\.
- \[16\]J\. Tsitsiklis and B\. V\. Roy\(1997\)An analysis of temporal\-difference learning with function approximation\.IEEE TAC\.Cited by:[§2\.3](https://arxiv.org/html/2608.17181#S2.SS3.p3.2)\.Similar Articles
@svlevine: Diffusion (or flow) makes for excellent policies, but training them with RL is notoriously hard: BPTT is unstable, RL o…
New paper shows how to optimize flow matching actors for reinforcement learning by approximating the Jacobian of the flow denoising process with the identity matrix, making training feasible.
@svlevine: A new way to do off-policy RL with diffusion: if we have off-policy data, we need to figure out what the diffusion late…
A new method for off-policy reinforcement learning with diffusion models, using flow reversal to handle off-policy data by reversing the diffusion process on it.
From Ticks to Flows: Dynamics of Neural Reinforcement Learning in Continuous Environments
This paper presents a theoretical framework for deep reinforcement learning in continuous environments, modeling it as a continuous-time stochastic process using stochastic control theory. The authors characterize an actor-critic algorithm's dynamics in the infinite width limit of two-layer networks, deriving an equation for infinitesimal changes in state distribution under a vanishingly small learning rate.
@tanayj: https://x.com/tanayj/status/2072766211256119475
This article explores the challenge of applying reinforcement learning to tasks that lack clear verifiability, citing Dario Amodei's prediction about achieving a 'country of geniuses in a data center' and discussing techniques such as RLVR, RLHF, Constitutional AI, and rubric-based rewards from Scale AI.
MechRL: Reinforcement Learning Agents Perform Circuit Discovery for Mechanistic Interpretability
Proposes MechRL, a reinforcement learning approach to automate circuit discovery in transformer language models. A PPO agent trained on multiple tasks discovers attention head circuits that match known canonical circuits and generalizes to a held-out task.