Attribution Markets: A Fisher-Market Formulation for Fractional Credit Assignment Between Planned Tasks and Performed Actions
Summary
This paper proposes formulating the bridge between planned tasks and performed actions as a quasi-linear Fisher market, allowing fractional credit assignment. It introduces instruments for conservation and junk filtering, and extends the model with entropy regularization to handle noise, unifying it with optimal transport.
View Cached Full Text
Cached at: 07/24/26, 05:13 AM
# Attribution Markets: A Fisher-Market Formulation for Fractional Credit Assignment Between Planned Tasks and Performed Actions
Source: [https://arxiv.org/html/2607.20694](https://arxiv.org/html/2607.20694)
###### Abstract
Personal and organizational planning systems maintain two records that drift apart: what was*planned*\(a task’s effort budget\) and what was*done*\(a logged action’s duration and description\)\. Existing systems bridge them with an exclusive, all\-or\-nothing link that strands genuinely related but unlinked effort and reports false stalls on active goals\. We formulate the bridge as a*quasi\-linear Fisher market*: planned tasks are budget\-constrained buyers, performed actions are divisible goods, and a fused text/structural/temporal signal sets each buyer’s valuation\. Two market instruments – a seller reserve price and a buyer cash option – yield conservation, a hard budget cap, and a provable junk filter as theorems\. We extend the market with a concave*completion utility*discounting progress as a task nears its plan; standard convergence theory for the market’s algorithm does not transfer here, resolved by a satiation\-threshold fixed point with existence \(Brouwer\) and local uniqueness under an explicit diagonal\-dominance condition, validated empirically on random and adversarial instances\. A de\-circularized, multi\-seed benchmark – observed affinity corrupted independently of the scored ground truth – surfaces a genuine weak spot: the market’s sharp, zero\-entropy equilibrium is more sensitive to affinity noise than entropy\-regularized optimal transport’s permanently smoothed one\. We resolve this with a one\-parameter entropy\-regularized generalization unifying the two, plus a noise\-adaptive rule for its regularization strength\. We report full reproducibility parameters, discuss limitations candidly, and relate the result to multi\-touch attribution, optimal transport, and online Fisher\-market algorithms\.
###### keywords:
Fisher markets , resource allocation , credit assignment , multi\-touch attribution , portfolio optimization , proportional response dynamics , entropic regularization , decision support systems
## 1Introduction
### 1\.1Motivation
A planning system – personal or organizational – maintains two records that are supposed to describe the same activity from two directions\. The*plan side*holds tasks: a description, an effort budget, and a time window\. The*result side*holds a log of what actually happened: a description, a duration, a timestamp\. Progress reporting requires a bridge between the two: every logged unit of effort should be credited to some task \(or honestly to none\), and a task’s reported progress should be the sum of the credit it received\.
The bridge used in practice is almost always an*exclusive link*: an action is created under a task, in which case all of its duration counts toward that task, or it is not, in which case none of it does\. This all\-or\-nothing rule fails in a specific and common way: work that is genuinely related to a task – preparatory reading, adjacent practice, work in the same domain performed without explicitly opening the task – is invisible to the bridge\. The result is a system that reports zero progress on a goal the person has, in fact, been advancing, and a review process that opens with a false alarm\.
### 1\.2Contribution
We propose to treat the bridge as a market: every performed action is a divisible*good*, every planned task is a budget\-constrained*buyer*, and a fused similarity signal sets each buyer’s valuation for each good\. The underlying object is not new – it is the Fisher market of Eisenberg and Gale\[[1](https://arxiv.org/html/2607.20694#bib.bib1)\]– but, to our knowledge, it has not been applied to this attribution problem, and two instruments placed on it \(a seller reserve price and a buyer cash option\) turn out to deliver exactly the guarantees the problem needs\. Concretely, we \(i\) define a formal*attribution market*\(Section[4](https://arxiv.org/html/2607.20694#S4)\) with three properties proved as theorems – conservation, a hard budget cap, and a provable junk filter – none of which the softmax or optimal\-transport baselines satisfy; \(ii\) add a*completion\-seeking*extension \(Section[5](https://arxiv.org/html/2607.20694#S5)\) whose saturating utility breaks the standard convergence proof for the market’s algorithm, show that the obvious naive fix is vacuous, and resolve it with a*satiation\-threshold*fixed point \(Algorithm[2](https://arxiv.org/html/2607.20694#alg2)\) whose existence we prove via Brouwer’s theorem and whose local convergence we prove under an explicit sufficient condition, both validated numerically; \(iii\) surface and explain a genuine*weak spot*\(Section[6\.4](https://arxiv.org/html/2607.20694#S6.SS4)\) – a de\-circularized, multi\-seed benchmark shows the market is more sensitive to affinity noise than entropy\-regularized optimal transport, which we trace to the distinction between a permanently regularized fixed point and a vanishing\-regularization solution path and resolve with a one\-parameter entropic generalization and a noise\-adaptive rule for its strength; and \(iv\) report a fully reproducible, de\-circularized synthetic evaluation \(Section[6](https://arxiv.org/html/2607.20694#S6)\) with an explicit reproducibility table, stated threats to validity, and a candid account of where the guarantees are weaker than they first appear \(Section[7\.3](https://arxiv.org/html/2607.20694#S7.SS3)\)\.
This is, deliberately, a formulation\-and\-application paper rather than a claim of new mathematics: the Fisher equilibrium, its computation by proportional response dynamics, and the mean–variance machinery we draw on are all classical\. What is new is treating planned effort as the currency of a market that clears logged effort – a move that converts two informal requirements \(“do not over\-credit a task,” “do not force unrelated work onto the nearest task”\) into theorems about a well\-understood equilibrium, together with the analysis of exactly where the formulation’s natural extensions break existing convergence theory and how those breaks are repaired\. In the tradition of applied operations\-research and decision\-support work, we import an established equilibrium concept into a new domain, analyze rigorously what transfers and what does not, and validate the result\. Two further extensions – temporal dynamics under a settlement horizon and a forward\-looking market that steers future planning – are summarized in Section[5\.4](https://arxiv.org/html/2607.20694#S5.SS4)and developed in full in a companion technical report\[[2](https://arxiv.org/html/2607.20694#bib.bib2)\]; the scope of the present paper is the formulation, its guarantees, their computation, and their honest empirical evaluation\.
### 1\.3Paper organization
Section[2](https://arxiv.org/html/2607.20694#S2)reviews related work\. Section[3](https://arxiv.org/html/2607.20694#S3)gives the data model and the affinity fusion signal, including a concrete fitting procedure\. Section[4](https://arxiv.org/html/2607.20694#S4)defines the base attribution market and proves its guarantees\. Section[5](https://arxiv.org/html/2607.20694#S5)develops the completion\-seeking extension and its convergence resolution, the risk\-aware valuation, and, in summary, the temporal and forward\-looking extensions\. Section[6](https://arxiv.org/html/2607.20694#S6)reports the de\-circularized empirical evaluation, including the noise\-sensitivity weak spot it surfaces, the theoretical explanation for it, and its resolution\. Section[7](https://arxiv.org/html/2607.20694#S7)discusses when simpler alternatives suffice \(Section[7\.1](https://arxiv.org/html/2607.20694#S7.SS1)\), the broader thesis that planning is describable as a market – of which the Fisher model here is one instance, opening a dynamic, uncertainty\-aware, and learning\-based program \(Section[7\.2](https://arxiv.org/html/2607.20694#S7.SS2)\) – and the model’s limitations \(Section[7\.3](https://arxiv.org/html/2607.20694#S7.SS3)\)\. Section[8](https://arxiv.org/html/2607.20694#S8)concludes\.
## 2Related work
The attribution problem sits at the intersection of several literatures, none of which solves it alone; Table[1](https://arxiv.org/html/2607.20694#S2.T1)summarizes the mapping we develop in this section\.
#### Credit assignment and multi\-touch attribution\.
The general credit\-assignment problem – given a composite outcome, which contributing decisions deserve credit – goes back to Minsky\[[3](https://arxiv.org/html/2607.20694#bib.bib3)\]; reinforcement learning develops its temporal form\[[4](https://arxiv.org/html/2607.20694#bib.bib4),[5](https://arxiv.org/html/2607.20694#bib.bib5)\]\. The closest existing*problem*to ours is multi\-touch attribution in marketing: a conversion is credited fractionally across advertising touchpoints\. Early data\-driven models used bagged logistic regression\[[6](https://arxiv.org/html/2607.20694#bib.bib6)\]and Markov\-chain removal effects\[[7](https://arxiv.org/html/2607.20694#bib.bib7)\]; the field’s recent turn is causal, reweighting journeys to remove confounding before attributing\[[8](https://arxiv.org/html/2607.20694#bib.bib8)\]\. This literature supplies the fractional\-credit framing but, without a budget or a price system, offers no mechanism analogous to our budget\-cap guarantee\.
#### Soft assignment and optimal transport\.
Mixture models fitted by expectation–maximization\[[9](https://arxiv.org/html/2607.20694#bib.bib9)\]and attention\[[10](https://arxiv.org/html/2607.20694#bib.bib10)\]are per\-item fractional weighting schemes, equivalent in our hierarchy \(Section[4\.4](https://arxiv.org/html/2607.20694#S4.SS4)\) to a single, uncoupled Sinkhorn step\. Optimal transport\[[11](https://arxiv.org/html/2607.20694#bib.bib11),[12](https://arxiv.org/html/2607.20694#bib.bib12),[13](https://arxiv.org/html/2607.20694#bib.bib13)\], and its unbalanced generalization\[[14](https://arxiv.org/html/2607.20694#bib.bib14)\], couples the columns \(actions\) through conservation and the rows \(tasks\) through capacity; low\-rank and unbalanced solvers\[[15](https://arxiv.org/html/2607.20694#bib.bib15)\]have made it efficient at scale\. This is the strongest*price\-free*baseline in our comparison, and, as our experiments show \(Section[6](https://arxiv.org/html/2607.20694#S6)\), it is the more*noise\-robust*one – a fact we explain theoretically in Section[6\.4](https://arxiv.org/html/2607.20694#S6.SS4)\.
#### Market\-based allocation\.
The Fisher market equilibrium and its convex\-program characterization for linear utilities\[[1](https://arxiv.org/html/2607.20694#bib.bib1),[16](https://arxiv.org/html/2607.20694#bib.bib16)\], its computation by proportional response dynamics\[[17](https://arxiv.org/html/2607.20694#bib.bib17),[18](https://arxiv.org/html/2607.20694#bib.bib18)\], and its fairness interpretation as proportionally fair\[[19](https://arxiv.org/html/2607.20694#bib.bib19)\]and as the Nash bargaining solution\[[20](https://arxiv.org/html/2607.20694#bib.bib20)\]are the mathematical core of our base model\. Approximate competitive equilibrium from equal incomes has been deployed for course\-seat assignment\[[21](https://arxiv.org/html/2607.20694#bib.bib21)\], and pacing equilibria are used for budget\-constrained ad auctions\[[22](https://arxiv.org/html/2607.20694#bib.bib22)\]; both demonstrate that taking a real allocation problem to a market equilibrium is an established, production\-tested move, which is the move we make here\. Recent work extends Fisher\-market computation online, with regret and statistical\-inference guarantees\[[23](https://arxiv.org/html/2607.20694#bib.bib23),[24](https://arxiv.org/html/2607.20694#bib.bib24)\]; we discuss this as the natural path to a streaming version of our model in Section[5\.4](https://arxiv.org/html/2607.20694#S5.SS4)\.
#### Portfolio optimization\.
The Eisenberg–Gale objective is, formally, a budget\-weighted sum of log\-returns – the objective of log\-optimal \(Kelly\) portfolio growth\[[25](https://arxiv.org/html/2607.20694#bib.bib25),[26](https://arxiv.org/html/2607.20694#bib.bib26)\]– and Eisenberg and Gale’s original derivation was as the pari\-mutuel method for betting markets, making "tasks as bettors staking budget on evidence" the model’s literal ancestry rather than a decorative analogy\. Mean–variance portfolio selection\[[27](https://arxiv.org/html/2607.20694#bib.bib27)\]and its robustification\[[28](https://arxiv.org/html/2607.20694#bib.bib28),[29](https://arxiv.org/html/2607.20694#bib.bib29)\]inform the risk\-aware extension we use in Section[5](https://arxiv.org/html/2607.20694#S5)\.
#### Desktop task\-context tracking and record linkage\.
TaskTracer and TaskPredictor\[[30](https://arxiv.org/html/2607.20694#bib.bib30),[31](https://arxiv.org/html/2607.20694#bib.bib31)\]associated low\-level desktop activity with declared tasks two decades ago, using hard classification – exactly the all\-or\-nothing bridge whose failure motivates this paper\. Record linkage theory\[[32](https://arxiv.org/html/2607.20694#bib.bib32)\]supplies the log\-linear evidence\-fusion template we use for the affinity signal \(Section[3\.2](https://arxiv.org/html/2607.20694#S3.SS2)\); LLM\-based entity matching\[[33](https://arxiv.org/html/2607.20694#bib.bib33)\]and standardized embedding benchmarks\[[34](https://arxiv.org/html/2607.20694#bib.bib34)\]are the modern instantiation of that sensor\.
Table 1:Requirements of the attribution problem and the literature each is drawn from\. No single strand supplies all of them; the contribution of this paper is the assembly and its analysis\.
## 3Problem formulation
### 3\.1Data model and notation
At any point in time the system holds tasksi=1,…,mi=1,\\dots,m, each with a text description, an effort budgetbi\>0b\_\{i\}\>0\(planned hours not yet settled\), and a planning window\[σi,τi\]\[\\sigma\_\{i\},\\tau\_\{i\}\]; and actionsj=1,…,nj=1,\\dots,n, each with a text description, a durationdj\>0d\_\{j\}\>0\(logged hours\), a timestamptjt\_\{j\}, and an optional explicit linkℓj∈\{1,…,m\}∪\{∅\}\\ell\_\{j\}\\in\\\{1,\\dots,m\\\}\\cup\\\{\\varnothing\\\}\. A distinguished indexi=0i=0, the*float*, holds unattributed effort\. The object to be computed is a*share matrix*W∈\[0,1\]\(m\+1\)×nW\\in\[0,1\]^\{\(m\+1\)\\times n\}with∑i=0mwij=1\\sum\_\{i=0\}^\{m\}w\_\{ij\}=1for everyjj; task progress isPi=∑jwijdjP\_\{i\}=\\sum\_\{j\}w\_\{ij\}d\_\{j\}and quality\-adjusted progress isVi=∑jqijwijdjV\_\{i\}=\\sum\_\{j\}q\_\{ij\}w\_\{ij\}d\_\{j\}, whereqijq\_\{ij\}is the affinity defined next\.
### 3\.2The affinity signal
Following the record\-linkage template\[[32](https://arxiv.org/html/2607.20694#bib.bib32)\], we fuse three independent evidence sources in log\-odds space and pass the result through a logistic link:
qij=σ\(λsemcos\(E\(xiT\),E\(xjA\)\)\+λlink1\[ℓj=i\]\+λtimeκ\(tj;σi,τi\)\)γ,q\_\{ij\}\\;=\\;\\sigma\\Big\(\\lambda\_\{\\mathrm\{sem\}\}\\cos\\\!\\big\(E\(x^\{\\mathrm\{T\}\}\_\{i\}\),E\(x^\{\\mathrm\{A\}\}\_\{j\}\)\\big\)\+\\lambda\_\{\\mathrm\{link\}\}\\,\\mathbb\{1\}\[\\ell\_\{j\}=i\]\+\\lambda\_\{\\mathrm\{time\}\}\\,\\kappa\(t\_\{j\};\\sigma\_\{i\},\\tau\_\{i\}\)\\Big\)^\{\\gamma\},\(1\)whereEEis a text embedding model,κ\\kappais a window kernel \(full weight inside the task’s window, discounted outside\),σ\(⋅\)\\sigma\(\\cdot\)is the logistic function, andγ\>1\\gamma\>1sharpens moderate similarities toward the extremes\. Unlike the exploratory treatment in our earlier technical report\[[2](https://arxiv.org/html/2607.20694#bib.bib2)\], we here give a concrete, reproducible fitting procedure forλ=\(λsem,λlink,λtime\)\\lambda=\(\\lambda\_\{\\mathrm\{sem\}\},\\lambda\_\{\\mathrm\{link\}\},\\lambda\_\{\\mathrm\{time\}\}\)\.
#### Fitting procedure\.
Given a labelled correction set𝒟=\{\(i,j,yij\)\}\\mathcal\{D\}=\\\{\(i,j,y\_\{ij\}\)\\\}, whereyij∈\{0,1\}y\_\{ij\}\\in\\\{0,1\\\}records a user\-confirmed match \(11\) or non\-match \(0\) between taskiiand actionjj,λ\\lambdais fit byℓ2\\ell\_\{2\}\-regularized logistic regression: writingsij\(λ\)s\_\{ij\}\(\\lambda\)for the pre\-logistic score in \([1](https://arxiv.org/html/2607.20694#S3.E1)\),
λ^=argminλ−∑\(i,j,y\)∈𝒟\[ylogσ\(sij\)\+\(1−y\)log\(1−σ\(sij\)\)\]\+η2∥λ∥22,\\hat\{\\lambda\}\\;=\\;\\operatorname\*\{arg\\,min\}\_\{\\lambda\}\\;\-\\\!\\\!\\sum\_\{\(i,j,y\)\\in\\mathcal\{D\}\}\\Big\[y\\log\\sigma\(s\_\{ij\}\)\+\(1\-y\)\\log\\big\(1\-\\sigma\(s\_\{ij\}\)\\big\)\\Big\]\\;\+\\;\\frac\{\\eta\}\{2\}\\lVert\\lambda\\rVert\_\{2\}^\{2\},\(2\)a convex problem with closed\-form gradient∇λ=−∑\(y−σ\(sij\)\)∇λsij\+ηλ\\nabla\_\{\\lambda\}=\-\\sum\(y\-\\sigma\(s\_\{ij\}\)\)\\,\\nabla\_\{\\lambda\}s\_\{ij\}\+\\eta\\lambda, solvable by any off\-the\-shelf convex solver \(Newton or L\-BFGS converge in a handful of iterations at this dimensionality\)\. We expect a fittedλ^\\hat\{\\lambda\}to reflect the intuition that an explicit link should dominate a moderate semantic similarity; the exact values are deployment\-specific and are not load\-bearing for any result in this paper\.
### 3\.3Design desiderata
The model is required to satisfy four properties, which Section[4](https://arxiv.org/html/2607.20694#S4)establishes as theorems rather than validated heuristics:
D1 \(conservation\)\.No logged hour is credited more than once\.
D2 \(budget respect\)\.No task can be credited more progress than it planned\.
D3 \(honest abstention\)\.Evidence below a quality floor is left unattributed rather than forced onto the nearest task\.
D4 \(explainability\)\.Every share is attributable to an auditable quantity, not a black\-box score\.
Figure[1](https://arxiv.org/html/2607.20694#S3.F1)shows how the model delivers these four properties end to end, and locates each desideratum at the stage that enforces it\.
Figure 1:The attribution pipeline, read left to right in four stages\.\(1\) Data model:planned tasks \(text, effort budgetbib\_\{i\}, time window\) and logged actions \(text, durationdjd\_\{j\}, timestamp, optional link\)\.\(2\) Affinity fusion\([1](https://arxiv.org/html/2607.20694#S3.E1)\) combines three evidence sources – semantic embedding similarity, the explicit structural link, and a temporal window kernel – into a single fused affinityqij∈\[0,1\]q\_\{ij\}\\in\[0,1\]\.\(3\) Market clearing:qijq\_\{ij\}parameterizes the attribution market of Section[4](https://arxiv.org/html/2607.20694#S4), whose proportional\-response clearing \(with reserve rateρ\\rhoand cash rateu0u\_\{0\}\) is what enforces conservation, the budget cap, and honest abstention – desiderataD1–D3\.\(4\) Outputs:fractional shareswijw\_\{ij\}, per\-task progressPiP\_\{i\}andViV\_\{i\}, and pricespjp\_\{j\}, which give an auditable explanation of every share \(D4\)\.
## 4The base attribution market
### 4\.1Market definition
Figure[2](https://arxiv.org/html/2607.20694#S4.F2)gives the market’s dictionary: tasks are investors with a budget of planned hours, actions are companies whose equity is their logged hours, and a share is the fraction of an action’s hours credited to a task\.
Figure 2:The attribution market\. Two tasks \(investors\) hold shares of three actions \(companies\); the float holds the reserve on every action, which becomes the action’s full owner if no task’s valuation clears it\.###### Definition 1\(Attribution market\)\.
Given affinitiesqq, durationsdd, budgetsbb, a reserve rateρ\>0\\rho\>0, and a cash utilityu0\>0u\_\{0\}\>0, the*attribution market*is the quasi\-linear Fisher market in which taskiiis a buyer with budgetbib\_\{i\}and linear valuationuij=qijdju\_\{ij\}=q\_\{ij\}d\_\{j\}for the whole of actionjj\(so a sharewwof actionjjis worthqijdjwq\_\{ij\}d\_\{j\}wto taskii\); taskiimay also hold cash, valued atu0u\_\{0\}per unit; and actionjjcarries a standing reserve bid ofρdj\\rho d\_\{j\}from the float\.
An equilibrium is a price vectorp∈ℝ\>0np\\in\\mathbb\{R\}\_\{\>0\}^\{n\}and spendingfij≥0f\_\{ij\}\\geq 0, cashci≥0c\_\{i\}\\geq 0with∑jfij\+ci=bi\\sum\_\{j\}f\_\{ij\}\+c\_\{i\}=b\_\{i\}, such that shareswij=fij/pjw\_\{ij\}=f\_\{ij\}/p\_\{j\},w0j=ρdj/pjw\_\{0j\}=\\rho d\_\{j\}/p\_\{j\}clear every action \(pj=ρdj\+∑ifijp\_\{j\}=\\rho d\_\{j\}\+\\sum\_\{i\}f\_\{ij\}\) and every task’s money earns the maximal available rate of return:
fij\>0⇒qijdjpj=βi,ci\>0⇒βi=u0,βi=max\{u0,maxjqijdjpj\}\.f\_\{ij\}\>0\\;\\Rightarrow\\;\\frac\{q\_\{ij\}d\_\{j\}\}\{p\_\{j\}\}=\\beta\_\{i\},\\qquad c\_\{i\}\>0\\;\\Rightarrow\\;\\beta\_\{i\}=u\_\{0\},\\qquad\\beta\_\{i\}=\\max\\Big\\\{u\_\{0\},\\;\\max\_\{j\}\\frac\{q\_\{ij\}d\_\{j\}\}\{p\_\{j\}\}\\Big\\\}\.\(3\)Without the reserve and cash terms, Definition[1](https://arxiv.org/html/2607.20694#Thmdefinition1)is exactly the Eisenberg–Gale equilibrium\[[1](https://arxiv.org/html/2607.20694#bib.bib1)\]; existence follows from the standard convex\-programming formulation\[[1](https://arxiv.org/html/2607.20694#bib.bib1),[16](https://arxiv.org/html/2607.20694#bib.bib16)\], and the reserve \(a fixed exogenous bid\) and cash \(a standard quasi\-linear extension\) preserve that structure\.
### 4\.2Guarantees
###### Proposition 1\(Conservation\)\.
For every actionjj,∑i=0mwij=1\\sum\_\{i=0\}^\{m\}w\_\{ij\}=1; hence∑i≥1Pi≤∑jdj\\sum\_\{i\\geq 1\}P\_\{i\}\\leq\\sum\_\{j\}d\_\{j\}\.
###### Proof\.
Immediate fromwij=fij/pjw\_\{ij\}=f\_\{ij\}/p\_\{j\},w0j=ρdj/pjw\_\{0j\}=\\rho d\_\{j\}/p\_\{j\}, and market clearingpj=ρdj\+∑ifijp\_\{j\}=\\rho d\_\{j\}\+\\sum\_\{i\}f\_\{ij\}\. ∎
###### Proposition 2\(Budget cap\)\.
Every equilibrium satisfiesPi≤bi/ρP\_\{i\}\\leq b\_\{i\}/\\rho\.
###### Proof\.
Sincepj≥ρdjp\_\{j\}\\geq\\rho d\_\{j\}for everyjj,Pi=∑j\(fij/pj\)dj≤1ρ∑jfij≤bi/ρP\_\{i\}=\\sum\_\{j\}\(f\_\{ij\}/p\_\{j\}\)d\_\{j\}\\leq\\frac\{1\}\{\\rho\}\\sum\_\{j\}f\_\{ij\}\\leq b\_\{i\}/\\rho\. ∎
###### Proposition 3\(Junk filter\)\.
Ifqij<u0ρq\_\{ij\}<u\_\{0\}\\rhothenwij=0w\_\{ij\}=0\.
###### Proof\.
Iffij\>0f\_\{ij\}\>0then by \([3](https://arxiv.org/html/2607.20694#S4.E3)\)qijdj/pj=βi≥u0q\_\{ij\}d\_\{j\}/p\_\{j\}=\\beta\_\{i\}\\geq u\_\{0\}, sopj≤qijdj/u0p\_\{j\}\\leq q\_\{ij\}d\_\{j\}/u\_\{0\}; butpj≥ρdjp\_\{j\}\\geq\\rho d\_\{j\}, givingqij≥u0ρq\_\{ij\}\\geq u\_\{0\}\\rho\. ∎
###### Proposition 4\(Proportional fairness\)\.
In theρ,u0→0\\rho,u\_\{0\}\\to 0limit, the equilibrium allocation maximizes∑ibilogVi\\sum\_\{i\}b\_\{i\}\\log V\_\{i\}– the Nash bargaining solution among tasks with bargaining power proportional tobib\_\{i\}\[[20](https://arxiv.org/html/2607.20694#bib.bib20),[19](https://arxiv.org/html/2607.20694#bib.bib19)\]\.
Proposition[2](https://arxiv.org/html/2607.20694#Thmproposition2)and[3](https://arxiv.org/html/2607.20694#Thmproposition3)give D2 and D3 as theorems; Proposition[1](https://arxiv.org/html/2607.20694#Thmproposition1)gives D1; pricespjp\_\{j\}give D4 \(a high price marks contested, scarce evidence\)\. None of the four hold for softmax attention attribution, and only conservation and a soft version of the budget cap hold for entropic optimal transport \(Section[4\.4](https://arxiv.org/html/2607.20694#S4.SS4)\)\.
### 4\.3Computation
Equilibrium is computed by*proportional response dynamics*\(PRD\)\[[17](https://arxiv.org/html/2607.20694#bib.bib17),[18](https://arxiv.org/html/2607.20694#bib.bib18)\]: each task repeatedly re\-splits its budget over goods and cash in proportion to the utility each earned in the previous round \(Algorithm[1](https://arxiv.org/html/2607.20694#alg1)\)\. PRD is stateless per iteration, distributed across tasks, and provably converges to the equilibrium of Definition[1](https://arxiv.org/html/2607.20694#Thmdefinition1)for the linear utilities used here\[[17](https://arxiv.org/html/2607.20694#bib.bib17),[18](https://arxiv.org/html/2607.20694#bib.bib18)\]; each iteration is oneO\(mn\)O\(mn\)elementwise pass\.
Algorithm 1Proportional response dynamics for the base attribution market1:affinities
q∈\[0,1\]m×nq\\in\[0,1\]^\{m\\times n\}, durations
d∈ℝ\>0nd\\in\\mathbb\{R\}\_\{\>0\}^\{n\}, budgets
b∈ℝ\>0mb\\in\\mathbb\{R\}\_\{\>0\}^\{m\}, reserve
ρ≥0\\rho\\geq 0, cash rate
u0\>0u\_\{0\}\>0, tolerance
tol\\mathrm\{tol\}, iteration cap
TmaxT\_\{\\max\}, safety constant
ε=10−9\\varepsilon=10^\{\-9\}
2:
vij←qijdjv\_\{ij\}\\leftarrow q\_\{ij\}d\_\{j\}⊳\\trianglerightvalue matrix, computed once
3:
fij←0\.7bi\(vij\+ε\)/∑j′\(vij′\+ε\)f\_\{ij\}\\leftarrow 0\.7\\,b\_\{i\}\\,\(v\_\{ij\}\+\\varepsilon\)\\big/\\textstyle\\sum\_\{j^\{\\prime\}\}\(v\_\{ij^\{\\prime\}\}\+\\varepsilon\);
ci←0\.3bic\_\{i\}\\leftarrow 0\.3\\,b\_\{i\}⊳\\trianglerightε\\varepsilonguards a task withqi⋅≡0q\_\{i\\cdot\}\\equiv 0
4:for
t=1,…,Tmaxt=1,\\dots,T\_\{\\max\}do
5:
pj←ρdj\+∑ifijp\_\{j\}\\leftarrow\\rho d\_\{j\}\+\\sum\_\{i\}f\_\{ij\}⊳\\trianglerightone column sum,O\(mn\)O\(mn\)
6:
wij←fij/pjw\_\{ij\}\\leftarrow f\_\{ij\}/p\_\{j\}
7:
gij←vijwijg\_\{ij\}\\leftarrow v\_\{ij\}\\,w\_\{ij\};
gi0←u0cig\_\{i0\}\\leftarrow u\_\{0\}c\_\{i\}⊳\\trianglerightrealized utility of last bid
8:
Σi←∑j′gij′\+gi0\+ε′\\Sigma\_\{i\}\\leftarrow\\textstyle\\sum\_\{j^\{\\prime\}\}g\_\{ij^\{\\prime\}\}\+g\_\{i0\}\+\\varepsilon^\{\\prime\}⊳\\trianglerightε′=10−12\\varepsilon^\{\\prime\}=10^\{\-12\}guards an all\-zero row
9:
fijnew←bigij/Σif\_\{ij\}^\{\\mathrm\{new\}\}\\leftarrow b\_\{i\}\\,g\_\{ij\}/\\Sigma\_\{i\};
cinew←bigi0/Σic\_\{i\}^\{\\mathrm\{new\}\}\\leftarrow b\_\{i\}\\,g\_\{i0\}/\\Sigma\_\{i\}
10:if
maxij\|fijnew−fij\|<tol\\max\_\{ij\}\|f\_\{ij\}^\{\\mathrm\{new\}\}\-f\_\{ij\}\|<\\mathrm\{tol\}then
11:
f←fnewf\\leftarrow f^\{\\mathrm\{new\}\},
c←cnewc\\leftarrow c^\{\\mathrm\{new\}\};break
12:endif
13:
f←fnewf\\leftarrow f^\{\\mathrm\{new\}\},
c←cnewc\\leftarrow c^\{\\mathrm\{new\}\}
14:endfor
15:return
ww,
pp
#### Vectorized form\.
Every step in Algorithm[1](https://arxiv.org/html/2607.20694#alg1)is a whole\-array operation, not a loop over\(i,j\)\(i,j\)pairs\. Withv=q⊙dv=q\\odot d\(elementwise product,ddbroadcast across themmtask\-rows\) computed once outside the loop, one round costs a column sum \(pp\), a broadcasted division \(ww\), an elementwise product \(gg\), a row sum \(Σ\\Sigma\), and a broadcasted division \(fnewf^\{\\mathrm\{new\}\}\) – fiveO\(mn\)O\(mn\)array operations, with no branching, no sorting, and no linear system to solve\. This is four to six lines in any array language \(NumPy, PyTorch, plain typed arrays\); no library beyond elementwise arithmetic and axis\-wise sums is required\.
#### Initialization and numerical safety\.
The safety constantε=10−9\\varepsilon=10^\{\-9\}added tovijv\_\{ij\}before the first normalization exists for one reason: a task withqi⋅≡0q\_\{i\\cdot\}\\equiv 0\(zero affinity to every logged action, e\.g\. a brand\-new task\) would otherwise divide0/00/0on line 2\. Withε\>0\\varepsilon\>0such a task starts with its0\.7bi0\.7\\,b\_\{i\}spread*uniformly*across all actions rather than crashing, and PRD’s own dynamics push that spend toward cash within the first few rounds once it earns near\-zero utility everywhere \(Proposition[3](https://arxiv.org/html/2607.20694#Thmproposition3)\)\. The second constantε′=10−12\\varepsilon^\{\\prime\}=10^\{\-12\}insideΣi\\Sigma\_\{i\}guards the symmetric case at every later round; withu0\>0u\_\{0\}\>0as required,Σi\\Sigma\_\{i\}never actually reaches zero, so the guard costs nothing but makes the implementation total \(no exception path\) rather than partial\. Both constants are four to seven orders of magnitude below the coarsest quantity that matters \(a logged minute is≈1\.7×10−2\\approx 1\.7\\times 10^\{\-2\}h\), so neither perturbs the reported shares\.
#### Stopping rule in practice\.
Table[3](https://arxiv.org/html/2607.20694#S6.T3)liststol=10−9\\mathrm\{tol\}=10^\{\-9\}\(absolute, in hours, onmaxij\|Δfij\|\\max\_\{ij\}\|\\Delta f\_\{ij\}\|\) andTmax=400T\_\{\\max\}=400\. On the main benchmark \(Section[6](https://arxiv.org/html/2607.20694#S6):m=7m=7tasks,≈170\\approx 170actions,4545seed/noise combinations\)*none*of the4545runs reach10−910^\{\-9\}within400400rounds – at this scale the loop always exits through the iteration cap, not the tolerance\. This does not mean the allocation is unstable: across the same4545runs,maxij\|Δfij\|\\max\_\{ij\}\|\\Delta f\_\{ij\}\|falls below10−210^\{\-2\}h after a mean of6262rounds \(median5959, worst case123123\) and below10−310^\{\-3\}h after a mean of246246rounds \(median240240\) – both far below the duration of a single logged action, so the returned shares are already stable well beyond the resolution that matters long before the cap is hit\. An implementation that wants a tolerance the loop can actually satisfy in practice should use something in the10−310^\{\-3\}–10−210^\{\-2\}h range and treatTmaxT\_\{\\max\}, nottol\\mathrm\{tol\}, as the real governor of worst\-case runtime\.
#### Complexity and cost\.
Each round isO\(mn\)O\(mn\)time and the live state \(ff,ww,gg\) isO\(mn\)O\(mn\)memory\. At DoPlan’s scale \(tens of tasks, a few hundred logged actions per attribution epoch\) a full400400\-round run is on the order of10510^\{5\}–10610^\{6\}elementary array operations – comfortably under a second in an interpreted array language, so the whole computation can run synchronously inside a request rather than as a background job\.
#### Worked example\.
A fully worked two\-task, three\-action trace of Algorithm[1](https://arxiv.org/html/2607.20694#alg1)– exact to the precision shown, and usable to unit\-test an independent implementation without the full benchmark – is given in[B](https://arxiv.org/html/2607.20694#A2)\.
### 4\.4Relation to alternative attribution rules
Four rules of increasing coupling can be applied to the same affinity matrix \(Table[2](https://arxiv.org/html/2607.20694#S4.T2)\):*hard assignment*\(winner\-take\-all above a threshold\),*softmax*\(per\-action fractional weighting, uncoupled across actions\),*entropic optimal transport*\(softmax columns coupled by conservation and soft row capacities, computed by Sinkhorn scaling\[[12](https://arxiv.org/html/2607.20694#bib.bib12)\]\), and the*market*\(the same coupling with an endogenous price system\)\. Each step adds a coupling at the cost of algorithmic complexity; Section[6](https://arxiv.org/html/2607.20694#S6)quantifies what each is worth\.
Table 2:The hierarchy of attribution rules\. Only the market has theorem\-level guarantees; entropic OT is included as the strongest price\-free, noise\-robust baseline \(Section[6\.4](https://arxiv.org/html/2607.20694#S6.SS4)\)\.
## 5The completion\-seeking extension
### 5\.1Motivation and the completion utility
The base market gives every task a*linear*valuation: an hour of well\-matched progress is worth the same whether the task has just started or is nearly done\. A task approaching its plan should value further progress*less*– the diminishing returns of completion\. We replace the linear valuation with a concave, increasing*completion utility*of quality\-adjusted progress,
Ui\(V\)=Ti\(1−e−V/Ti\),Ui′\(V\)=e−V/Ti,U\_\{i\}\(V\)\\;=\\;T\_\{i\}\\big\(1\-e^\{\-V/T\_\{i\}\}\\big\),\\qquad U\_\{i\}^\{\\prime\}\(V\)=e^\{\-V/T\_\{i\}\},\(4\)indexed by the task’s planned totalTiT\_\{i\};Ui′\(0\)=1U\_\{i\}^\{\\prime\}\(0\)=1and decays toward completion\. Taskiinow solves, given pricespp,
maxfi≥0,ci≥0Ui\(∑jqijfijpjdj\)\+u0cis\.t\.∑jfij\+ci=bi\.\\max\_\{f\_\{i\}\\geq 0,\\,c\_\{i\}\\geq 0\}\\;U\_\{i\}\\\!\\Big\(\\textstyle\\sum\_\{j\}q\_\{ij\}\\tfrac\{f\_\{ij\}\}\{p\_\{j\}\}d\_\{j\}\\Big\)\+u\_\{0\}c\_\{i\}\\quad\\text\{s\.t\.\}\\quad\\textstyle\\sum\_\{j\}f\_\{ij\}\+c\_\{i\}=b\_\{i\}\.\(5\)
###### Proposition 5\(Junk filter sharpens\)\.
In any completion\-market equilibrium, taskiiholds a positive share of actionjjonly ifqij≥u0ρ/Ui′\(Vi\)q\_\{ij\}\\geq u\_\{0\}\\rho/U\_\{i\}^\{\\prime\}\(V\_\{i\}\); becauseUi′U\_\{i\}^\{\\prime\}is non\-increasing, this threshold rises asViV\_\{i\}grows\.
###### Proof\.
Iffij\>0f\_\{ij\}\>0, optimality of \([5](https://arxiv.org/html/2607.20694#S5.E5)\) equalizes bang\-per\-buck atβi≥u0\\beta\_\{i\}\\geq u\_\{0\}, soUi′\(Vi\)qijdj/pj=βi≥u0U\_\{i\}^\{\\prime\}\(V\_\{i\}\)q\_\{ij\}d\_\{j\}/p\_\{j\}=\\beta\_\{i\}\\geq u\_\{0\}; withpj≥ρdjp\_\{j\}\\geq\\rho d\_\{j\}this givesqij≥u0ρ/Ui′\(Vi\)q\_\{ij\}\\geq u\_\{0\}\\rho/U\_\{i\}^\{\\prime\}\(V\_\{i\}\)\. ∎
The budget cap \(Proposition[2](https://arxiv.org/html/2607.20694#Thmproposition2)\) is unaffected – its proof uses onlypj≥ρdjp\_\{j\}\\geq\\rho d\_\{j\}and∑jfij≤bi\\sum\_\{j\}f\_\{ij\}\\leq b\_\{i\}, independent of utility shape\. What is lost is the closed\-form proportional\-fairness characterization of Proposition[4](https://arxiv.org/html/2607.20694#Thmproposition4), which relied on the linear/log\-aggregate structure\.
### 5\.2The satiation\-threshold fixed point
The correct fix follows from \([3](https://arxiv.org/html/2607.20694#S4.E3)\): for a task withUi′\(Vi\)U\_\{i\}^\{\\prime\}\(V\_\{i\}\)held at a fixed scalarμi\\mu\_\{i\}, the ranking of goods by bang\-per\-buckμiqijdj/pj\\mu\_\{i\}q\_\{ij\}d\_\{j\}/p\_\{j\}is*identical*to the ranking byqijdj/pjq\_\{ij\}d\_\{j\}/p\_\{j\}, sinceμi\\mu\_\{i\}multiplies every candidate equally – but the comparison against the cash floor,μiqijdj/pj≥u0\\mu\_\{i\}q\_\{ij\}d\_\{j\}/p\_\{j\}\\geq u\_\{0\}, is*not*invariant: it is equivalent toqijdj/pj≥u0/μiq\_\{ij\}d\_\{j\}/p\_\{j\}\\geq u\_\{0\}/\\mu\_\{i\}\. That is, taskii’s best response under a fixed satiation levelμi\\mu\_\{i\}is*exactly*the base linear market’s best response with a*task\-specific cash threshold*u0,i=u0/μiu\_\{0,i\}=u\_\{0\}/\\mu\_\{i\}– a trivial, already\-linear generalization of Definition[1](https://arxiv.org/html/2607.20694#Thmdefinition1)\(which already has a scalaru0u\_\{0\}; making it heterogeneous across tasks does not affect linearity, existence, or PRD convergence for the inner solve\)\.
Algorithm 2Satiation\-threshold fixed point for the completion market1:affinities
qq, durations
dd, budgets
bb, targets
TT,
ρ\\rho,
u0u\_\{0\}
2:
μi←1\\mu\_\{i\}\\leftarrow 1for all
ii
3:repeat
4:
u0,i←min\(u0/μi,20u0\)u\_\{0,i\}\\leftarrow\\min\(u\_\{0\}/\\mu\_\{i\},\\,20u\_\{0\}\)⊳\\trianglerightcapped satiation\-adjusted cash rate
5:
\(w,p\)←\(w,p\)\\leftarrowAlgorithm[1](https://arxiv.org/html/2607.20694#alg1)with per\-task cash rate
u0,iu\_\{0,i\}⊳\\trianglerightunmodified inner solve
6:
Vi←∑jqijwijdjV\_\{i\}\\leftarrow\\sum\_\{j\}q\_\{ij\}w\_\{ij\}d\_\{j\}⊳\\trianglerighttrue quality\-adjusted progress
7:
μinew←exp\(−Vi/Ti\)\\mu\_\{i\}^\{\\mathrm\{new\}\}\\leftarrow\\exp\(\-V\_\{i\}/T\_\{i\}\)
8:until
maxi\|μinew−μi\|<tol\\max\_\{i\}\|\\mu\_\{i\}^\{\\mathrm\{new\}\}\-\\mu\_\{i\}\|<\\mathrm\{tol\}
9:return
μ\\mu,
ww,
pp
Figure[3](https://arxiv.org/html/2607.20694#S5.F3)lays the loop out schematically: the inner box is exactly Algorithm[1](https://arxiv.org/html/2607.20694#alg1), called unchanged on every outer iteration with only its per\-task cash rate perturbed, so nothing about the base market’s guarantees or convergence behavior is touched by wrapping it\.
Figure 3:Algorithm[2](https://arxiv.org/html/2607.20694#alg2): an outer fixed\-point loop over the satiation multipliersμi\\mu\_\{i\}wraps the unmodified inner linear market solve, which retains every guarantee and convergence property of Section[4](https://arxiv.org/html/2607.20694#S4)\.###### Proposition 6\(Existence\)\.
A completion\-market equilibrium exists\.
###### Proof\.
For fixedμ∈\(0,1\]m\\mu\\in\(0,1\]^\{m\}, the inner problem is a linear Fisher market with heterogeneous cash ratesu0,i=u0/μiu\_\{0,i\}=u\_\{0\}/\\mu\_\{i\}; its equilibrium utility valuesV\(μ\)V\(\\mu\)are unique \(a classical fact about linear Fisher markets\) and, by Berge’s Maximum Theorem applied to this parametric convex program, continuous inμ\\mu\. HenceΦ:μ↦U′\(V\(μ\)\)=exp\(−V\(μ\)/T\)\\Phi:\\mu\\mapsto U^\{\\prime\}\(V\(\\mu\)\)=\\exp\(\-V\(\\mu\)/T\)is a continuous map from the compact convex cube\(0,1\]m\(0,1\]^\{m\}to itself, and Brouwer’s fixed\-point theorem guarantees a fixed pointμ∗=Φ\(μ∗\)\\mu^\{\\ast\}=\\Phi\(\\mu^\{\\ast\}\), which is a completion\-market equilibrium by construction\. ∎
###### Proposition 7\(Sufficient condition for local convergence\)\.
If the cross\-task sensitivity of the outer map dominates the own\-task sensitivity by no more than a factor summing to less than one in each row of the Jacobian ofΦ\\Phi\(diagonal dominance: for everyii,∑k≠i\|∂Φi/∂μk\|<1−\|∂Φi/∂μi\|\\sum\_\{k\\neq i\}\|\\partial\\Phi\_\{i\}/\\partial\\mu\_\{k\}\|<1\-\|\\partial\\Phi\_\{i\}/\\partial\\mu\_\{i\}\|in a neighborhood of a fixed point\), thenΦ\\Phiis a contraction in max\-norm near that fixed point, the fixed point is locally unique, and Algorithm[2](https://arxiv.org/html/2607.20694#alg2)converges to it geometrically from any sufficiently close start, by the Banach fixed\-point theorem\.
### 5\.3Risk\-aware valuation
Affinity is estimated, not observed\. Treatingqijq\_\{ij\}as a random variable with per\-task meanμiaff\\mu^\{\\mathrm\{aff\}\}\_\{i\}and covarianceΣi\\Sigma\_\{i\}over the live actions, and writingxix\_\{i\}for taskii’s booked hours, a mean–variance valuation\[[27](https://arxiv.org/html/2607.20694#bib.bib27)\]replaces linear progress with
maxxi≥0\(μiaff\)⊤xi−γi2xi⊤Σixis\.t\.∑j\(pj/dj\)xij≤bi\.\\max\_\{x\_\{i\}\\geq 0\}\\;\\;\(\\mu^\{\\mathrm\{aff\}\}\_\{i\}\)^\{\\top\}x\_\{i\}\-\\frac\{\\gamma\_\{i\}\}\{2\}x\_\{i\}^\{\\top\}\\Sigma\_\{i\}x\_\{i\}\\qquad\\text\{s\.t\.\}\\qquad\\textstyle\\sum\_\{j\}\(p\_\{j\}/d\_\{j\}\)x\_\{ij\}\\leq b\_\{i\}\.\(6\)This per\-task subproblem is a concave quadratic program over a simplex\-like polytope\. WhenΣi≻0\\Sigma\_\{i\}\\succ 0, the objective isγiλmin\(Σi\)\\gamma\_\{i\}\\lambda\_\{\\min\}\(\\Sigma\_\{i\}\)\-strongly concave and its gradient isγiλmax\(Σi\)\\gamma\_\{i\}\\lambda\_\{\\max\}\(\\Sigma\_\{i\}\)\-Lipschitz, so*projected gradient ascent*with the exactO\(klogk\)O\(k\\log k\)Euclidean simplex projection algorithm\[[35](https://arxiv.org/html/2607.20694#bib.bib35)\]converges at the linear rate\(1−λmin\(Σi\)/λmax\(Σi\)\)k\\big\(1\-\\lambda\_\{\\min\}\(\\Sigma\_\{i\}\)/\\lambda\_\{\\max\}\(\\Sigma\_\{i\}\)\\big\)^\{k\}\[[36](https://arxiv.org/html/2607.20694#bib.bib36),[37](https://arxiv.org/html/2607.20694#bib.bib37)\], a concretely implementable, off\-the\-shelf solver for the per\-task subproblem\. The*joint*multi\-task market under this valuation, however, no longer inherits the Eisenberg–Gale convex\-program equivalence \(the aggregate no longer has a single concave potential\): existence of a joint equilibrium is instead a monotone variational\-inequality question\[[38](https://arxiv.org/html/2607.20694#bib.bib38)\], holding under a diagonal\-dominance condition on the cross\-task price coupling analogous to Proposition[7](https://arxiv.org/html/2607.20694#Thmproposition7)\. We flag efficient computation of the general joint case as an open algorithmic question and use the per\-task subproblem, with prices held fixed within a PRD round, as the practical approximation in this paper\.
### 5\.4Temporal and forward\-looking extensions \(summary\)
The market re\-clears at discrete epochs; shares older than a settlement horizonHHfreeze into a ledger, bounding both churn and computation, and within the live window warm\-started, damped re\-clearing keeps consecutive equilibria from flapping – a form of transaction\-cost portfolio rebalancing\[[39](https://arxiv.org/html/2607.20694#bib.bib39)\]in which the settlement horizon is a no\-trade band\. A second, coupled market – distinguished by a separate variableπid\\pi\_\{id\}, future budget allocated across*directions*rather than past\-hour shareswijw\_\{ij\}of individual actions – lets a task’s investment steer the planner’s future scheduling \(an endogenous\-supply feedback\) and lets the task pivot away from a direction whose realized advancement decays, using the same multiplicative\-weights family as Algorithm[1](https://arxiv.org/html/2607.20694#alg1)\(proportional response is, in this family, the same update as exponential weights\[[40](https://arxiv.org/html/2607.20694#bib.bib40)\], the EXP3 bandit\[[41](https://arxiv.org/html/2607.20694#bib.bib41)\], and universal portfolios\[[26](https://arxiv.org/html/2607.20694#bib.bib26)\]\)\. The endogenous\-supply loop creates a lock\-in hazard familiar from the non\-stationary bandit literature\[[42](https://arxiv.org/html/2607.20694#bib.bib42),[41](https://arxiv.org/html/2607.20694#bib.bib41)\], resolved by maintaining an exploration reserve\. Both extensions – their full formalization, a reinforcement\-learning construction for the forward market, and simulation results – are developed at length in a companion technical report\[[2](https://arxiv.org/html/2607.20694#bib.bib2)\]; we summarize them here only to situate the present paper’s scope and do not evaluate them empirically in this work\.
## 6Experimental validation
### 6\.1De\-circularized benchmark design
A benchmark in which the affinity every method*consumes*is computed from the same latent embedding geometry that*generates*the ground truth is close to circular: it tests whether a method recovers structure planted using the method’s own similarity assumption\. We avoid this by decoupling the two explicitly\. We generate synthetic instances withm=7m=7tasks and roughly170170actions over a6363\-day horizon, with generative ground\-truth shares \(mixtures of on\-plan, partially\-adjacent, late\-emerging, and distractor actions\) set directly by the generator and*independent*of the affinity any method observes\. Every method instead observes a*corrupted*Qobs=clip\(Qclean\+σ𝒩\(0,1\)\+confuser noise,0,1\)Q\_\{\\mathrm\{obs\}\}=\\mathrm\{clip\}\(Q\_\{\\mathrm\{clean\}\}\+\\sigma\\,\\mathcal\{N\}\(0,1\)\+\\text\{confuser noise\},0,1\), whereQcleanQ\_\{\\mathrm\{clean\}\}is the latent geometric affinity andσ\\sigmais the observation\-noise level, independent per random seed of the ground truth it is scored against; we repeat over1515independent random seeds at each of three noise levels and report mean±\\pmstandard deviation\. This design, absent from a naively self\-generating benchmark, is what surfaces the noise\-sensitivity weak spot of Section[6\.4](https://arxiv.org/html/2607.20694#S6.SS4)below\. Table[3](https://arxiv.org/html/2607.20694#S6.T3)lists every parameter needed to reproduce the result; code is available as supplementary material\.
Table 3:Full reproducibility parameters\.
### 6\.2Metrics
*Total\-variation \(TV\) error*: half theℓ1\\ell\_\{1\}distance between predicted and true share vectors of an action, averaged over actions – the fraction of an action misallocated\.*Ghost credit*: hours credited to tasks with zero true involvement\.*Missed credit*: true hours not credited to their task\.*Starving\-task recovery*: credited over true hours, in percent, for a task whose direct on\-plan work is withdrawn after day1010and whose only true progress thereafter flows through partial stakes in adjacent actions – the motivating scenario for this work\.*Budget violations*: fraction of instances in which any task’s credited hours exceed its budget \(counted above a10−610^\{\-6\}\-hour threshold\)\.*Worst overshoot*: the largest amount, in hours, by which any task’s credited hours exceed its budget, across all tasks and instances at a noise level\. We report the magnitude alongside the count because the count alone cannot distinguish a solver\-tolerance artifact from a material failure – a distinction that, as Table[4](https://arxiv.org/html/2607.20694#S6.T4)shows, reverses the apparent ranking of the rules on this metric\.
Because the1515random seeds are shared across rules within a noise level, all between\-rule comparisons are*paired*; wherever the text below calls a difference significant or a tie, the claim is backed by a two\-sided Wilcoxon signed\-rank test on the per\-seed TV errors at that noise level \(n=15n=15pairs, so the smallest attainableppis6\.1×10−56\.1\\times 10^\{\-5\}, reached exactly when all1515seeds agree in sign\)\.
### 6\.3Results
Table 4:De\-circularized, multi\-seed benchmark \(n=15n=15instances per cell, mean±\\pmstandard deviation\)\. Bold marks the lowest mean per column within each noise block; a bolded mean is not necessarily a significant win – e\.g\. the TV\-error margin between hard assignment and Sinkhorn atσ=0\.30\\sigma=0\.30is a statistical tie under the paired Wilcoxon test \(p=0\.85p=0\.85\), whereas the market’s TV\-error deficit against Sinkhorn is significant at every noise level \(p=6\.1×10−5p=6\.1\\times 10^\{\-5\}, all1515paired seeds agreeing\)\. The market’s zero\-violation column is a consequence of Proposition[2](https://arxiv.org/html/2607.20694#Thmproposition2), not an empirical observation – it is exact by construction at every noise level, unlike the other three rules’ soft or absent caps\.*Viol\.*is the mean number of budget\-violating tasks per instance;*Over*is the worst overshoot in hours\. The overshoot column shows why the violation*count*alone would mislead: Sinkhorn’s counted violations never exceed5×10−45\\times 10^\{\-4\}h – a solver\-tolerance artifact, flat across noise levels – while softmax’s and hard assignment’s overshoots are real hours that grow with noise; ranked by magnitude, Sinkhorn’s soft cap is tolerance\-tight in practice and only softmax and hard assignment are materially unsafe\.Table[4](https://arxiv.org/html/2607.20694#S6.T4)shows a genuine trade\-off rather than a clean win for any rule\. Atσ=0\\sigma=0the market is competitive on TV error \(0\.3190\.319against hard assignment’s0\.3310\.331– a statistical tie,p=0\.095p=0\.095– and second only to Sinkhorn’s0\.2370\.237, a significant gap,p=6\.1×10−5p=6\.1\\times 10^\{\-5\}\) while recovering90\.7%90\.7\\%of the starving task’s true progress against hard assignment’s44\.0%44\.0\\%– the motivating pathology of this work, reproduced quantitatively\. As noise rises, Sinkhorn’s entropic smoothing dominates raw share accuracy at moderate noise \(best TV atσ=0\.15\\sigma=0\.15:0\.3800\.380vs\. the market’s0\.4620\.462,p=6\.1×10−5p=6\.1\\times 10^\{\-5\}\); at the highest noise level tested, hard assignment edges ahead of Sinkhorn by an insignificant margin \(0\.5520\.552vs\.0\.5550\.555,p=0\.85p=0\.85\) – a statistical tie, not a reversal of the trend – while the market is significantly worse than both at both nonzero noise levels \(all four pairwise tests atp=6\.1×10−5p=6\.1\\times 10^\{\-5\}, with all1515paired seeds agreeing in sign in each case\)\. Meanwhile the market is the*only*rule whose budget\-cap guarantee holds exactly at every noise level\. The worst\-overshoot column qualifies what that is worth in practice: Sinkhorn’s soft cap is tolerance\-tight \(worst overshoot5×10−45\\times 10^\{\-4\}h, noise\-independent\), so the market’s practical safety margin over Sinkhorn is small – exact\-by\-theorem versus approximate\-to\-solver\-tolerance – whereas softmax and hard assignment incur genuine multi\-hour overshoots under noise \(11\.211\.2h and5\.95\.9h worst\-case atσ=0\.30\\sigma=0\.30\)\. This is a real trade\-off, not an artifact: the market becomes less accurate exactly as it stays maximally safe\. Section[6\.4](https://arxiv.org/html/2607.20694#S6.SS4)explains why from first principles and Section[6\.5](https://arxiv.org/html/2607.20694#S6.SS5)gives a resolution\.
00\.10\.10\.20\.20\.30\.30\.20\.20\.30\.30\.40\.40\.50\.50\.60\.6observation noiseσ\\sigmaTV error\(a\) share accuracy degrades with noiseHardSoftmaxSinkhornMarket
00\.10\.10\.20\.20\.30\.300\.20\.20\.40\.40\.60\.6observation noiseσ\\sigmamean budget violations / instance\(b\) only the market’s cap is exact
Figure 4:\(a\) Total\-variation share error rises with observation noise for every rule; Sinkhorn is most robust\. \(b\) Mean budget violations per instance, counted above a10−610^\{\-6\}\-hour threshold: the market alone stays exactly at zero \(Proposition[2](https://arxiv.org/html/2607.20694#Thmproposition2)\)\. The count overstates Sinkhorn’s line: its violations never exceed5×10−45\\times 10^\{\-4\}h \(solver tolerance\), whereas softmax’s and hard assignment’s reach real hours \(Table[4](https://arxiv.org/html/2607.20694#S6.T4), worst\-overshoot column\)\.
### 6\.4Explanation
The distinction is between two different kinds of fixed point\. PRD is proved to converge to the*exact*Eisenberg–Gale competitive equilibrium of the unregularized linear market\[[17](https://arxiv.org/html/2607.20694#bib.bib17),[18](https://arxiv.org/html/2607.20694#bib.bib18)\]– a*sharp*, zero\-entropy equilibrium in which small differences in affinity are, at convergence, resolved by genuine price competition, which can amplify small input perturbations into large share differences\. Sinkhorn’s entropic optimal transport, in contrast, converges to the minimizer of an*entropy\-regularized*transport objective\[[12](https://arxiv.org/html/2607.20694#bib.bib12),[13](https://arxiv.org/html/2607.20694#bib.bib13)\]: its smoothing parameterε\\varepsilonis baked permanently into what the algorithm converges*to*, not merely into the path it takes to get there\. The market’s implicit multiplicative\-update dynamics can also be read as a form of mirror descent with a KL\-regularization term\[[18](https://arxiv.org/html/2607.20694#bib.bib18)\], but that regularization is relative to the*previous iterate*and vanishes along the solution path – it disciplines convergence, not the destination\. This is a real, previously undocumented asymmetry between the two rules, grounded entirely in results already cited in Section[4\.4](https://arxiv.org/html/2607.20694#S4.SS4)\.
### 6\.5Resolution: a one\-parameter unification
We generalize Definition[1](https://arxiv.org/html/2607.20694#Thmdefinition1)by adding an entropy regularizer, with strengthτ≥0\\tau\\geq 0, to each task’s per\-round allocation problem:
maxfi≥0∑jqijdjfijpj−τ∑jfijbilogfijbis\.t\.∑jfij≤bi\.\\max\_\{f\_\{i\}\\geq 0\}\\;\\;\\sum\_\{j\}q\_\{ij\}d\_\{j\}\\frac\{f\_\{ij\}\}\{p\_\{j\}\}\\;\-\\;\\tau\\sum\_\{j\}\\frac\{f\_\{ij\}\}\{b\_\{i\}\}\\log\\frac\{f\_\{ij\}\}\{b\_\{i\}\}\\qquad\\text\{s\.t\.\}\\qquad\\sum\_\{j\}f\_\{ij\}\\leq b\_\{i\}\.\(7\)For fixed pricespp, \([7](https://arxiv.org/html/2607.20694#S6.E7)\) is strictly concave \(linear objective plus strictly concave negative entropy\) over a compact polytope, so it has a unique maximizer for everyτ\>0\\tau\>0, continuous inτ\\tau; atτ=0\\tau=0it recovers the base market’s linear best response exactly, and asτ→∞\\tau\\to\\inftyit recovers the uniform allocation\. This places the market and a Sinkhorn\-like smoothed regime on one dial rather than treating them as a discrete choice\. We do not claim a fully engineered, tuned algorithm for the general\-τ\\taucase in this paper – an efficient alternating scaling scheme analogous to Sinkhorn’s is the natural next algorithmic step, which we leave open – but the existence, uniqueness, and interpolation properties above are established analytically and are sufficient to support a concrete practical prescription:*selectτ\\tau\(or, in the un\-engineered two\-point case, choose between the market and Sinkhorn OT\) in proportion to the estimated affinity\-observation noiseσ^\\hat\{\\sigma\}*, exactly analogous to selecting a ridge penalty from an estimated noise level\. Atσ^→0\\hat\{\\sigma\}\\to 0this recovers the sharp, maximally accurate and maximally explanatory market; asσ^\\hat\{\\sigma\}grows, it degrades gracefully toward the more robust, entropy\-smoothed regime – the correct, principled reading of the bias–variance trade\-off our benchmark exposes\.
### 6\.6Completion\-market convergence
Algorithm[2](https://arxiv.org/html/2607.20694#alg2)converged \(residual<10−8<10^\{\-8\}\) in30/3030/30random instances, in a mean of4\.04\.0outer iterations \(min22, max2222\)\. On an adversarial design – two near\-duplicate tasks \(affinity correlation0\.970\.97\) contending for a small, scarce pool, chosen to stress Proposition[7](https://arxiv.org/html/2607.20694#Thmproposition7)’s sufficient condition –9/109/10instances still converged, at a substantially higher mean of24\.124\.1outer iterations, with one instance failing to reach tolerance within the200200\-iteration cap\. This is the honest empirical boundary of the sufficient condition: typical instances converge geometrically and fast; highly contended, near\-duplicate instances converge more slowly and are not guaranteed to\. Figure[5](https://arxiv.org/html/2607.20694#S6.F5)shows representative traces, and we verified across all4040instances that the budget cap \(Proposition[2](https://arxiv.org/html/2607.20694#Thmproposition2)\) held exactly throughout, including during non\-converged iterations\.
022446688−10\-10−5\-50outer iterationkklog10maxi\|μi\(k\+1\)−μi\(k\)\|\\log\_\{10\}\\max\_\{i\}\|\\mu\_\{i\}^\{\(k\+1\)\}\-\\mu\_\{i\}^\{\(k\)\}\|random 1random 2random 3random 4adversarialFigure 5:Outer\-loop convergence: log\-residual versus iteration\. Random instances \(blue\) converge in22–44iterations with a steep geometric slope; the adversarial near\-duplicate instance \(red\) still converges geometrically but visibly more slowly, consistent with Proposition[7](https://arxiv.org/html/2607.20694#Thmproposition7)’s prediction that weaker diagonal dominance slows, but does not always break, convergence\.
### 6\.7Where the market is preferable
#### Share sparsity\.
Define the*sparsity*of a rule as the fraction of task–entry shareswijw\_\{ij\}\(i≥1i\\geq 1, excluding the float row\) that are numerically zero \(below10−910^\{\-9\}\); Table[5](https://arxiv.org/html/2607.20694#S6.T5)reports it\. Softmax and Sinkhorn are*dense by construction*– soft\-max and entropic\-transport allocations are strictly positive, so every task holds a nonzero sliver of every entry \(sparsity≈0\\approx 0\)\. Hard assignment is sparser still \(≈0\.9\\approx 0\.9\) but is not fractional: it cannot split a genuinely shared entry at all\. The market is the*only rule that is both fractional and sparse*\(≈0\.85\\approx 0\.85\): its junk filter \(Proposition[3](https://arxiv.org/html/2607.20694#Thmproposition3)\) drives every sub\-threshold share to*exactly*zero, so it alone can report that an entry contributed*nothing*to a task, rather than a diffuse0\.3%0\.3\\%that a reader must mentally threshold\. For a user\-facing attribution – the intended application – an exact zero is the difference between a legible result and a dense matrix of noise\.
Table 5:Share sparsity: mean fraction of task–entry shares that are numerically zero \(n=15n=15instances per cell\)\. The market is the only*fractional*rule that is also sparse; softmax and Sinkhorn are dense by construction, and hard assignment is sparser only because it is all\-or\-nothing\.
#### Exact, anytime budget cap\.
The market’s capPi≤bi/ρP\_\{i\}\\leq b\_\{i\}/\\rho\(Proposition[2](https://arxiv.org/html/2607.20694#Thmproposition2)\) holds*exactly*and at every proportional\-response iterate, not only in the converged limit: each round’s bids sum to at most the budget and prices never fall below the reserve, so the bound is structural rather than asymptotic\. Sinkhorn’s soft cap is a property of the converged optimum only; Table[4](https://arxiv.org/html/2607.20694#S6.T4)found it tolerance\-tight in practice, but “exact at every iterate” and “approximately correct once converged” are different guarantees precisely when a solve is stopped early, warm\-started, or fed adversarial input\.
#### Budget\-weighted contested splits\.
When two tasks genuinely tie on affinity for an entry, the market splits it in proportion to their planned budgets – the budget\-weighted Nash bargaining solution of Proposition[4](https://arxiv.org/html/2607.20694#Thmproposition4)– whereas entropic transport splits ties toward the uniform allocation irrespective of how much each task planned\. Assigning the larger share to the larger commitment is defensible in a way an even split is not, and only the market makes it a property of the solution rather than an artifact of the smoother\.
#### Extensible buyer\-side demand\.
Every extension developed in this paper – the completion\-seeking utility \(Section[5](https://arxiv.org/html/2607.20694#S5)\), risk\-aware valuation \(Section[5\.3](https://arxiv.org/html/2607.20694#S5.SS3)\), and the forward\-looking market \(Section[5\.4](https://arxiv.org/html/2607.20694#S5.SS4)\) – modifies a task’s*demand*, an object entropic transport does not possess: optimal transport sees only a cost matrix and fixed marginals\. The market is a model onto which task\-side structure can be grown; Sinkhorn is a fixed smoother\. This, rather than any single benchmark number, is the strongest long\-run argument for the market formulation\.
Raw accuracy and noise\-robustness favour Sinkhorn \(Table[4](https://arxiv.org/html/2607.20694#S6.T4)\); the market’s higher starving\-task recovery at zero noise \(90\.7%90\.7\\%vs\.70\.6%70\.6\\%\) sits within overlapping standard deviations and is confounded at higher noise \(Section[6\.8](https://arxiv.org/html/2607.20694#S6.SS8)\), so we do not claim it as a robust advantage\. The market’s territory is guarantees, interpretable sparsity, principled contested splits, and extensibility – not point accuracy\.
### 6\.8Threats to validity
The evaluation is synthetic; no claim is made about performance on real user data, only about the internal correctness and comparative behavior of the four rules under a controlled, de\-circularized generative process\. The generative process, noise model, and parameter choices \(Table[3](https://arxiv.org/html/2607.20694#S6.T3)\) were fixed by reasonableness rather than tuned by search, and results are reported at personal\-planner scale \(tens of tasks, hundreds of actions\); scaling behavior at organizational scale is not tested\. The starving\-task recovery metric becomes less interpretable at high noise, where it is partly inflated by the same ghost credit that Figure[4](https://arxiv.org/html/2607.20694#S6.F4)\(a\) shows growing – it should be read jointly with TV error and ghost credit, not in isolation\.
## 7Discussion
### 7\.1When simpler alternatives suffice
A method that adds machinery should survive the question “is there a cheaper way?” At personal\-planner scale the market’s PRD solve costs microseconds \(Section[4\.3](https://arxiv.org/html/2607.20694#S4.SS3)\) – computation is not the expense; implementation complexity, external dependencies, and user trust are\. Table[4](https://arxiv.org/html/2607.20694#S6.T4)itself makes the point starkest: Sinkhorn beats the market on raw share accuracy at two of the three noise levels tested\. Table[6](https://arxiv.org/html/2607.20694#S7.T6)situates the market among eight routes to the same stalled\-progress problem: human link suggestions\[[43](https://arxiv.org/html/2607.20694#bib.bib43)\], schedule\-context heuristics, LLM\-judged percentages\[[44](https://arxiv.org/html/2607.20694#bib.bib44)\], constrained softmax\[[35](https://arxiv.org/html/2607.20694#bib.bib35)\], Sinkhorn OT\[[45](https://arxiv.org/html/2607.20694#bib.bib45),[12](https://arxiv.org/html/2607.20694#bib.bib12)\], the attribution market, Shapley values\[[46](https://arxiv.org/html/2607.20694#bib.bib46),[47](https://arxiv.org/html/2607.20694#bib.bib47)\], and outcome check\-ins\[[48](https://arxiv.org/html/2607.20694#bib.bib48)\]\.
Table 6:Eight routes to the stalled\-progress problem\. Every column reads “more==better” \(\+\+\+\+best,−−\-\-worst,∘\\circneutral or not applicable\); the three cost columns are rated as*cheapness*, so\+\+\+\+means inexpensive\.*Guarantees*= conservation\+\+budget capping\. Accuracy entries for rows 4–6 are grounded in Table[4](https://arxiv.org/html/2607.20694#S6.T4); the market’s guarantees column is a theorem \(Propositions[1](https://arxiv.org/html/2607.20694#Thmproposition1)–[3](https://arxiv.org/html/2607.20694#Thmproposition3)\), not an observation\. Each row’s citation \(Amershi et al\.\[[43](https://arxiv.org/html/2607.20694#bib.bib43)\], Zheng et al\.\[[44](https://arxiv.org/html/2607.20694#bib.bib44)\], Duchi et al\.\[[35](https://arxiv.org/html/2607.20694#bib.bib35)\], Sinkhorn–Knopp and Cuturi\[[45](https://arxiv.org/html/2607.20694#bib.bib45),[12](https://arxiv.org/html/2607.20694#bib.bib12)\], Shapley and Lundberg–Lee\[[46](https://arxiv.org/html/2607.20694#bib.bib46),[47](https://arxiv.org/html/2607.20694#bib.bib47)\], Harkin et al\.\[[48](https://arxiv.org/html/2607.20694#bib.bib48)\]\) supports the general characterization of that technique, not the specific mark in this table – the marks are this paper’s own comparative judgment\. Row 2 \(schedule\-context rule\) is an internal planner heuristic with no published analogue, so it carries no citation by construction, not by omission\.∗Row 8 addresses the stalled\-goal*symptom*directly but does not attribute hours, hence the neutral marks\.No row is uniformly best\. Rows 1–2 are cheap because a human or the existing plan does the work; rows 4–5 are lean automatic engines with weak or soft guarantees, and Sinkhorn wins on raw accuracy\. The market earns its place when the guarantees column matters more than the accuracy column – when a downstream consumer \(a completion percentage, a scheduler\) needs a number it can trust never to exceed a task’s plan, not merely one that is usually close\.
### 7\.2Markets as a modeling substrate
The quasi\-linear Fisher market of this paper is one*instance*of a broader stance: that the relationship between a plan and the effort that realizes it is naturally a*market*, with the equilibrium concept – here, Eisenberg–Gale competitive equilibrium – a replaceable choice rather than the essence\. What is invariant is a scarce resource \(logged or future effort\), budgeted claimants \(tasks\), valuations \(the affinity signal\), and an*endogenous price*that clears supply against demand; our guarantees follow from that clearing structure, not the utility’s linear form\. That the equilibrium concept is a dial, not a commitment, is already visible here: Section[6\.5](https://arxiv.org/html/2607.20694#S6.SS5)places the market and entropy\-regularized optimal transport on one parameter, and the same substrate admits budget\-pacing\[[22](https://arxiv.org/html/2607.20694#bib.bib22)\], Nash\-bargaining \(Proposition[4](https://arxiv.org/html/2607.20694#Thmproposition4)\), and matching\- or exchange\-economy readings of the identical problem\. The market view thus sits*above*any single solution concept\.
Accepting the description imports a mature toolbox\. A real planner clears*repeatedly*, so the honest object is a price*process*, developed as transaction\-cost rebalancing\[[39](https://arxiv.org/html/2607.20694#bib.bib39)\]and by the online\-Fisher\-market literature\[[23](https://arxiv.org/html/2607.20694#bib.bib23),[24](https://arxiv.org/html/2607.20694#bib.bib24)\]\. Two further layers attach:*uncertainty quantification*– propagating the estimated affinity’s covariance \(already in the valuation, Section[5\.3](https://arxiv.org/html/2607.20694#S5.SS3)\) through the equilibrium map to distributions over allocations, of which our noise analysis \(Section[6\.4](https://arxiv.org/html/2607.20694#S6.SS4)\) is a first\-order case and for which generalized polynomial chaos\[[49](https://arxiv.org/html/2607.20694#bib.bib49)\]is standard machinery – and*forecasting*, where we conjecture the forward market’s continuum limit is a mean\-field\-game\[[50](https://arxiv.org/html/2607.20694#bib.bib50)\]/ stochastic\-PDE system, solvable at scale by neural operators\[[51](https://arxiv.org/html/2607.20694#bib.bib51)\], deep backward\-SDE methods\[[52](https://arxiv.org/html/2607.20694#bib.bib52)\], or reinforcement learning\. None of this dynamic, uncertainty\-quantification, or deep\-learning program is developed or evaluated here; it is the direction the formulation opens, and the mean\-field\-game and stochastic\-PDE bridges are conjectures, not results\.
### 7\.3Limitations
Several limitations bound these results\. The market’s noise sensitivity is real: it is second\-best at zero noise and the*least*accurate of the four rules at both nonzero noise levels \(Table[4](https://arxiv.org/html/2607.20694#S6.T4)\), even though its budget cap stays exact, and the general entropy\-regularized market \(Eq\. \([7](https://arxiv.org/html/2607.20694#S6.E7)\)\) that would resolve this is analyzed but not tuned into an efficient algorithm here\. Convergence is only conditional – Proposition[7](https://arxiv.org/html/2607.20694#Thmproposition7)’s sufficient condition can fail, and did in one of ten adversarial instances \(Section[6\.6](https://arxiv.org/html/2607.20694#S6.SS6)\), so a deployment should monitor the outer\-loop residual and fall back to the linear market \(μ≡1\\mu\\equiv 1\) if it does not shrink\. Like every rule in Table[2](https://arxiv.org/html/2607.20694#S4.T2), the market can over\-credit genuinely close tasks with progress a stricter causal notion would not – the same regime that stresses Proposition[7](https://arxiv.org/html/2607.20694#Thmproposition7)\. The mean–variance extension is only partial: Section[5\.3](https://arxiv.org/html/2607.20694#S5.SS3)solves the per\-task subproblem but leaves the joint equilibrium an open variational\-inequality question\. Finally, evaluation is synthetic \(Section[6](https://arxiv.org/html/2607.20694#S6)\) and confined to personal\-planner scale; validation at deployment and organizational scale, and of the affinity fit against real user corrections, is future work\.
## 8Conclusion
We formulated fractional attribution of performed actions to planned tasks as a quasi\-linear Fisher market, in which a seller reserve price and a buyer cash option turn three informal design requirements into theorems: conservation, a hard budget cap, and a provable junk filter\. Extending the market with a completion\-seeking utility exposed a genuine gap in the standard convergence theory for its algorithm; we resolved it with a satiation\-threshold fixed\-point procedure, proved its existence via Brouwer’s theorem, proved a sufficient condition for its local geometric convergence, and validated both properties – including their honest boundary – across random and adversarial instances\. A de\-circularized, multi\-seed empirical evaluation, designed specifically to avoid testing the method against structure planted by its own similarity assumption, surfaced a further genuine weak spot: the market’s price\-competitive, zero\-entropy equilibrium is more sensitive to affinity\-observation noise than entropy\-regularized optimal transport\. We traced this to a precise, citable distinction between a permanently regularized fixed point and a vanishing\-regularization solution path, and resolved it with a one\-parameter entropy\-regularized generalization together with a noise\-adaptive prescription for its use\. The practical implication is a choice, not a verdict: where raw share accuracy is what matters, Sinkhorn – or, at extreme noise, even hard assignment – is the better tool \(Section[7\.1](https://arxiv.org/html/2607.20694#S7.SS1)summarizes when simpler alternatives suffice\); the market earns its place specifically where a downstream consumer needs a number that is guaranteed, not merely usually accurate, never to exceed what a task actually planned\. Every reported result is accompanied by its reproducibility parameters and its threats to validity, and Section[7\.3](https://arxiv.org/html/2607.20694#S7.SS3)states plainly where the model’s guarantees are weaker than they first appear\. Temporal dynamics and a forward\-looking, reinforcement\-learning\-amenable extension are summarized here and developed fully in a companion technical report\[[2](https://arxiv.org/html/2607.20694#bib.bib2)\]; the natural next steps are a working implementation of the entropy\-regularized market’s efficient algorithm, an empirical study of the joint mean–variance equilibrium’s variational\-inequality conditions, and validation of the complete pipeline against real user correction data\. Beyond these concrete steps, Section[7\.2](https://arxiv.org/html/2607.20694#S7.SS2)argues that the deeper contribution is the market*description*itself: the Fisher model here is one instance of a stance under which planning becomes a market, opening a dynamic, uncertainty\-aware, and learning\-based program we sketch but do not develop\.
## CRediT authorship contribution statement
Salavat Ishbulatov: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Writing – original draft, Writing – review & editing, Visualization\.
## Declaration of competing interest
The author declares no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper\.
## Declaration of generative AI in scientific writing
During the preparation of this work the author used a large\-language\-model assistant \(Claude, Anthropic\) to help draft text, develop and verify derivations, and produce the analysis scripts and figures\. The author reviewed and edited all content and takes full responsibility for the content of the published article\.
## Funding
This research did not receive any specific grant from funding agencies in the public, commercial, or not\-for\-profit sectors\.
## Appendix ANotation
Table 7:Notation used throughout the paper\.
## Appendix BWorked example
Takem=2m=2tasks,n=3n=3actions, with
q=\[0\.90\.20\.10\.10\.80\.6\],d=\[231\]h,b=\[45\]h,ρ=0\.25,u0=0\.30\.q=\\begin\{bmatrix\}0\.9&0\.2&0\.1\\\\ 0\.1&0\.8&0\.6\\end\{bmatrix\},\\quad d=\\begin\{bmatrix\}2&3&1\\end\{bmatrix\}\\text\{h\},\\quad b=\\begin\{bmatrix\}4\\\\ 5\\end\{bmatrix\}\\text\{h\},\\quad\\rho=0\.25,\\ u\_\{0\}=0\.30\.Thenv=q⊙d=\[1\.80\.60\.10\.22\.40\.6\]v=q\\odot d=\\begin\{bmatrix\}1\.8&0\.6&0\.1\\\\ 0\.2&2\.4&0\.6\\end\{bmatrix\}, and initialization gives
f\(0\)=\[2\.01600\.67200\.11200\.21882\.62500\.6563\],c\(0\)=\[1\.20001\.5000\]\.f^\{\(0\)\}=\\begin\{bmatrix\}2\.0160&0\.6720&0\.1120\\\\ 0\.2188&2\.6250&0\.6563\\end\{bmatrix\},\\quad c^\{\(0\)\}=\\begin\{bmatrix\}1\.2000\\\\ 1\.5000\\end\{bmatrix\}\.One round of Algorithm[1](https://arxiv.org/html/2607.20694#alg1)gives pricesp\(1\)=\[2\.7348,4\.0470,1\.0183\]p^\{\(1\)\}=\[2\.7348,\\,4\.0470,\\,1\.0183\]and
w\(1\)=\[0\.73720\.16600\.11000\.08000\.64860\.6445\],f\(1\)=\[2\.95270\.22170\.02450\.03323\.23050\.8025\],c\(1\)=\[0\.80110\.9338\]\.w^\{\(1\)\}=\\begin\{bmatrix\}0\.7372&0\.1660&0\.1100\\\\ 0\.0800&0\.6486&0\.6445\\end\{bmatrix\},\\quad f^\{\(1\)\}=\\begin\{bmatrix\}2\.9527&0\.2217&0\.0245\\\\ 0\.0332&3\.2305&0\.8025\\end\{bmatrix\},\\quad c^\{\(1\)\}=\\begin\{bmatrix\}0\.8011\\\\ 0\.9338\\end\{bmatrix\}\.A second round givesp\(2\)=\[3\.4859,4\.2022,1\.0769\]p^\{\(2\)\}=\[3\.4859,\\,4\.2022,\\,1\.0769\],
f\(2\)=\[3\.39020\.07040\.00510\.00373\.58370\.8684\],c\(2\)=\[0\.5344,0\.5442\]\.f^\{\(2\)\}=\\begin\{bmatrix\}3\.3902&0\.0704&0\.0051\\\\ 0\.0037&3\.5837&0\.8684\\end\{bmatrix\},\\quad c^\{\(2\)\}=\[0\.5344,\\,0\.5442\]\.Task 1 is pulling share away from actions 2 and 3 toward its clear best match \(action 1,q11=0\.9q\_\{11\}=0\.9\) exactly as the junk filter \(Proposition[3](https://arxiv.org/html/2607.20694#Thmproposition3)\) predicts, while its cash holding shrinks each round as it finds genuine matches worth more thanu0u\_\{0\}\. These numbers are exact to the precision shown and can be used to unit\-test an independent implementation of Algorithm[1](https://arxiv.org/html/2607.20694#alg1)without needing the full benchmark\.
## References
- \[1\]E\. Eisenberg, D\. Gale,[Consensus of subjective probabilities: The pari\-mutuel method](http://www.jstor.org/stable/2237130), The Annals of Mathematical Statistics 30 \(1\) \(1959\) 165–168\.
- \[2\]S\. Ishbulatov, The attribution market: A theory of dynamic shared credit between planned tasks and performed actions, Technical report \(2026\)\.
- \[3\]M\. Minsky, Steps toward artificial intelligence, Proceedings of the IRE 49 \(1\) \(1961\) 8–30\.[doi:10\.1109/JRPROC\.1961\.287775](https://doi.org/10.1109/JRPROC.1961.287775)\.
- \[4\]R\. S\. Sutton, Temporal credit assignment in reinforcement learning, Ph\.D\. thesis, University of Massachusetts Amherst \(1984\)\.
- \[5\]E\. Pignatelli, J\. Ferret, M\. Geist, T\. Mesnard, H\. van Hasselt, O\. Pietquin, L\. Toni,[A survey of temporal credit assignment in deep reinforcement learning](https://arxiv.org/abs/2312.01072)\(2024\)\.[arXiv:2312\.01072](http://arxiv.org/abs/2312.01072)\.
- \[6\]X\. Shao, L\. Li,[Data\-driven multi\-touch attribution models](https://doi.org/10.1145/2020408.2020453), in: Proceedings of the 17th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’11, Association for Computing Machinery, New York, NY, USA, 2011, p\. 258–264\.[doi:10\.1145/2020408\.2020453](https://doi.org/10.1145/2020408.2020453)\.
- \[7\]E\. Anderl, I\. Becker, F\. von Wangenheim, J\. H\. Schumann,[Mapping the customer journey: Lessons learned from graph\-based online attribution modeling](https://www.sciencedirect.com/science/article/pii/S0167811616300349), International Journal of Research in Marketing 33 \(3\) \(2016\) 457–474\.[doi:10\.1016/j\.ijresmar\.2016\.03\.001](https://doi.org/10.1016/j.ijresmar.2016.03.001)\.
- \[8\]D\. Yao, C\. Gong, L\. Zhang, S\. Chen, J\. Bi,[Causalmta: Eliminating the user confounding bias for causal multi\-touch attribution](https://doi.org/10.1145/3534678.3539108), in: Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD ’22, Association for Computing Machinery, New York, NY, USA, 2022, p\. 4342–4352\.[doi:10\.1145/3534678\.3539108](https://doi.org/10.1145/3534678.3539108)\.
- \[9\]A\. P\. Dempster, N\. M\. Laird, D\. B\. Rubin,[Maximum likelihood from incomplete data via the em algorithm](https://rss.onlinelibrary.wiley.com/doi/abs/10.1111/j.2517-6161.1977.tb01600.x), Journal of the Royal Statistical Society: Series B \(Methodological\) 39 \(1\) \(1977\) 1–22\.[doi:10\.1111/j\.2517\-6161\.1977\.tb01600\.x](https://doi.org/10.1111/j.2517-6161.1977.tb01600.x)\.
- \[10\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, I\. Polosukhin, Attention is all you need, in: I\. Guyon, U\. V\. Luxburg, S\. Bengio, H\. Wallach, R\. Fergus, S\. Vishwanathan, R\. Garnett \(Eds\.\), Advances in Neural Information Processing Systems, Vol\. 30, Curran Associates, Inc\., 2017\.
- \[11\]L\. V\. Kantorovich, On the translocation of masses, Doklady Akademii Nauk SSSR 37 \(1942\) 199–201\.
- \[12\]M\. Cuturi,[Sinkhorn distances: Lightspeed computation of optimal transport](https://proceedings.neurips.cc/paper_files/paper/2013/file/af21d0c97db2e27e13572cbf59eb343d-Paper.pdf), in: C\. Burges, L\. Bottou, M\. Welling, Z\. Ghahramani, K\. Weinberger \(Eds\.\), Advances in Neural Information Processing Systems, Vol\. 26, Curran Associates, Inc\., 2013\.
- \[13\]G\. Peyré, M\. Cuturi,[Computational optimal transport with applications to data sciences](https://doi.org/10.1561/2200000073), Foundations and Trends in Machine Learning 11 \(5\-6\) \(2019\) 355–607\.[doi:10\.1561/2200000073](https://doi.org/10.1561/2200000073)\.
- \[14\]L\. Chizat, G\. Peyré, B\. Schmitzer, F\.\-X\. Vialard, Scaling algorithms for unbalanced optimal transport problems, Mathematics of Computation 87 \(2018\) 2563–2609\.
- \[15\]M\. Scetbon, M\. Klein, G\. Palla, M\. Cuturi,[Unbalanced low\-rank optimal transport solvers](https://proceedings.neurips.cc/paper_files/paper/2023/file/a439259e78294c38d157a51a2c40486b-Paper-Conference.pdf), in: A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, S\. Levine \(Eds\.\), Advances in Neural Information Processing Systems, Vol\. 36, Curran Associates, Inc\., 2023, pp\. 52312–52325\.
- \[16\]N\. Nisan, T\. Roughgarden, É\. Tardos, V\. V\. Vazirani \(Eds\.\), Algorithmic Game Theory, Cambridge University Press, 2007\.
- \[17\]L\. Zhang,[Proportional response dynamics in the fisher market](https://www.sciencedirect.com/science/article/pii/S0304397510003555), Theoretical Computer Science 412 \(24\) \(2011\) 2691–2698, selected Papers from 36th International Colloquium on Automata, Languages and Programming \(ICALP 2009\)\.[doi:10\.1016/j\.tcs\.2010\.06\.021](https://doi.org/10.1016/j.tcs.2010.06.021)\.
- \[18\]B\. Birnbaum, N\. R\. Devanur, L\. Xiao,[Distributed algorithms via gradient descent for fisher markets](https://doi.org/10.1145/1993574.1993594), in: Proceedings of the 12th ACM Conference on Electronic Commerce, EC ’11, Association for Computing Machinery, New York, NY, USA, 2011, p\. 127–136\.[doi:10\.1145/1993574\.1993594](https://doi.org/10.1145/1993574.1993594)\.
- \[19\]F\. P\. Kelly, A\. K\. Maulloo, D\. K\. H\. Tan,[Rate control for communication networks: shadow prices, proportional fairness and stability](https://doi.org/10.1057/palgrave.jors.2600523), Journal of the Operational Research Society 49 \(3\) \(1998\) 237–252\.[doi:10\.1057/palgrave\.jors\.2600523](https://doi.org/10.1057/palgrave.jors.2600523)\.
- \[20\]J\. F\. Nash, The bargaining problem, Econometrica 18 \(2\) \(1950\) 155–162\.
- \[21\]E\. Budish,[The combinatorial assignment problem: Approximate competitive equilibrium from equal incomes](https://www.journals.uchicago.edu/doi/abs/10.1086/664613), Journal of Political Economy 119 \(6\) \(2011\) 1061–1103\.[doi:10\.1086/664613](https://doi.org/10.1086/664613)\.
- \[22\]V\. Conitzer, C\. Kroer, E\. Sodomka, N\. E\. Stier\-Moses, Multiplicative pacing equilibria in auction markets, Operations Research 70 \(2\) \(2022\) 963–989\.
- \[23\]Y\. Gao, A\. Peysakhovich, C\. Kroer,[Online market equilibrium with application to fair division](https://proceedings.neurips.cc/paper_files/paper/2021/file/e562cd9c0768d5464b64cf61da7fc6bb-Paper.pdf), in: M\. Ranzato, A\. Beygelzimer, Y\. Dauphin, P\. Liang, J\. W\. Vaughan \(Eds\.\), Advances in Neural Information Processing Systems, Vol\. 34, Curran Associates, Inc\., 2021, pp\. 27305–27318\.
- \[24\]L\. Liao, Y\. Gao, C\. Kroer,[Statistical inference for fisher market equilibrium](https://arxiv.org/abs/2209.15422)\(2025\)\.[arXiv:2209\.15422](http://arxiv.org/abs/2209.15422)\.
- \[25\]J\. L\. Kelly, Jr\., A new interpretation of information rate, Bell System Technical Journal 35 \(4\) \(1956\) 917–926\.
- \[26\]T\. M\. Cover,[Universal portfolios](https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1467-9965.1991.tb00002.x), Mathematical Finance 1 \(1\) \(1991\) 1–29\.[doi:10\.1111/j\.1467\-9965\.1991\.tb00002\.x](https://doi.org/10.1111/j.1467-9965.1991.tb00002.x)\.
- \[27\]H\. Markowitz, Portfolio selection, The Journal of Finance 7 \(1\) \(1952\) 77–91\.
- \[28\]F\. Black, R\. Litterman,[Global portfolio optimization](https://doi.org/10.2469/faj.v48.n5.28), Financial Analysts Journal 48 \(5\) \(1992\) 28–43\.[doi:10\.2469/faj\.v48\.n5\.28](https://doi.org/10.2469/faj.v48.n5.28)\.
- \[29\]O\. Ledoit, M\. Wolf, Honey, I shrunk the sample covariance matrix, Journal of Portfolio Management 30 \(4\) \(2004\) 110–119\.
- \[30\]A\. N\. Dragunov, T\. G\. Dietterich, K\. Johnsrude, M\. McLaughlin, L\. Li, J\. L\. Herlocker,[Tasktracer: a desktop environment to support multi\-tasking knowledge workers](https://doi.org/10.1145/1040830.1040855), in: Proceedings of the 10th International Conference on Intelligent User Interfaces, IUI ’05, Association for Computing Machinery, New York, NY, USA, 2005, p\. 75–82\.[doi:10\.1145/1040830\.1040855](https://doi.org/10.1145/1040830.1040855)\.
- \[31\]J\. Shen, L\. Li, T\. G\. Dietterich, J\. L\. Herlocker,[A hybrid learning system for recognizing user tasks from desktop activities and email messages](https://doi.org/10.1145/1111449.1111473), in: Proceedings of the 11th International Conference on Intelligent User Interfaces, IUI ’06, Association for Computing Machinery, New York, NY, USA, 2006, p\. 86–92\.[doi:10\.1145/1111449\.1111473](https://doi.org/10.1145/1111449.1111473)\.
- \[32\]I\. P\. Fellegi, A\. B\. Sunter, A theory for record linkage, Journal of the American Statistical Association 64 \(328\) \(1969\) 1183–1210\.
- \[33\]R\. Peeters, A\. Steiner, C\. Bizer,[Entity matching using large language models](https://arxiv.org/abs/2310.11244)\(2024\)\.[arXiv:2310\.11244](http://arxiv.org/abs/2310.11244)\.
- \[34\]N\. Muennighoff, N\. Tazi, L\. Magne, N\. Reimers,[MTEB: Massive text embedding benchmark](https://aclanthology.org/2023.eacl-main.148/), in: A\. Vlachos, I\. Augenstein \(Eds\.\), Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, Association for Computational Linguistics, Dubrovnik, Croatia, 2023, pp\. 2014–2037\.[doi:10\.18653/v1/2023\.eacl\-main\.148](https://doi.org/10.18653/v1/2023.eacl-main.148)\.
- \[35\]J\. Duchi, S\. Shalev\-Shwartz, Y\. Singer, T\. Chandra,[Efficient projections onto the l1\-ball for learning in high dimensions](https://doi.org/10.1145/1390156.1390191), in: Proceedings of the 25th International Conference on Machine Learning, ICML ’08, Association for Computing Machinery, New York, NY, USA, 2008, p\. 272–279\.[doi:10\.1145/1390156\.1390191](https://doi.org/10.1145/1390156.1390191)\.
- \[36\]Y\. Nesterov, Introductory Lectures on Convex Optimization: A Basic Course, Springer, 2004\.
- \[37\]S\. Bubeck,[Convex optimization: Algorithms and complexity](https://doi.org/10.1561/2200000050), Foundations and Trends in Machine Learning 8 \(3\-4\) \(2015\) 231–357\.[doi:10\.1561/2200000050](https://doi.org/10.1561/2200000050)\.
- \[38\]F\. Facchinei, J\.\-S\. Pang, Finite\-Dimensional Variational Inequalities and Complementarity Problems, Springer, 2003\.
- \[39\]M\. H\. A\. Davis, A\. R\. Norman, Portfolio selection with transaction costs, Mathematics of Operations Research 15 \(4\) \(1990\) 676–713\.[doi:10\.1287/moor\.15\.4\.676](https://doi.org/10.1287/moor.15.4.676)\.
- \[40\]Y\. Freund, R\. E\. Schapire,[A decision\-theoretic generalization of on\-line learning and an application to boosting](https://www.sciencedirect.com/science/article/pii/S002200009791504X), Journal of Computer and System Sciences 55 \(1\) \(1997\) 119–139\.[doi:10\.1006/jcss\.1997\.1504](https://doi.org/10.1006/jcss.1997.1504)\.
- \[41\]P\. Auer, N\. Cesa\-Bianchi, Y\. Freund, R\. E\. Schapire,[The nonstochastic multiarmed bandit problem](https://doi.org/10.1137/S0097539701398375), SIAM Journal on Computing 32 \(1\) \(2002\) 48–77\.[doi:10\.1137/S0097539701398375](https://doi.org/10.1137/S0097539701398375)\.
- \[42\]T\. Lai, H\. Robbins,[Asymptotically efficient adaptive allocation rules](https://doi.org/10.1016/0196-8858(85)90002-8), Adv\. Appl\. Math\. 6 \(1\) \(1985\) 4–22\.[doi:10\.1016/0196\-8858\(85\)90002\-8](https://doi.org/10.1016/0196-8858(85)90002-8)\.
- \[43\]S\. Amershi, M\. Cakmak, W\. B\. Knox, T\. Kulesza,[Power to the people: The role of humans in interactive machine learning](https://onlinelibrary.wiley.com/doi/abs/10.1609/aimag.v35i4.2513), AI Magazine 35 \(4\) \(2014\) 105–120\.[doi:10\.1609/aimag\.v35i4\.2513](https://doi.org/10.1609/aimag.v35i4.2513)\.
- \[44\]L\. Zheng, W\.\-L\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. Xing, H\. Zhang, J\. Gonzalez, I\. Stoica,[Judging llm\-as\-a\-judge with mt\-bench and chatbot arena](https://proceedings.neurips.cc/paper_files/paper/2023/file/91f18a1287b398d378ef22505bf41832-Paper-Datasets_and_Benchmarks.pdf), in: A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, S\. Levine \(Eds\.\), Advances in Neural Information Processing Systems, Vol\. 36, Curran Associates, Inc\., 2023, pp\. 46595–46623\.
- \[45\]R\. Sinkhorn, P\. Knopp, Concerning nonnegative matrices and doubly stochastic matrices, Pacific Journal of Mathematics 21 \(2\) \(1967\) 343–348\.
- \[46\]L\. S\. Shapley, A value fornn\-person games, in: Contributions to the Theory of Games II, Princeton University Press, 1953, pp\. 307–317\.
- \[47\]S\. M\. Lundberg, S\.\-I\. Lee, A unified approach to interpreting model predictions, in: Advances in Neural Information Processing Systems \(NeurIPS\), 2017\.
- \[48\]B\. Harkin, T\. L\. Webb, B\. P\. I\. Chang, A\. Prestwich, M\. Conner, I\. Kellar, Y\. Benn, P\. Sheeran,[Does monitoring goal progress promote goal attainment? a meta\-analysis of the experimental evidence\.](https://doi.org/10.1037/bul0000025), Psychological Bulletin 142 \(2\) \(2016\) 198–229\.[doi:10\.1037/bul0000025](https://doi.org/10.1037/bul0000025)\.
- \[49\]D\. Xiu, G\. E\. Karniadakis,[The wiener–askey polynomial chaos for stochastic differential equations](https://doi.org/10.1137/S1064827501387826), SIAM Journal on Scientific Computing 24 \(2\) \(2002\) 619–644\.[doi:10\.1137/S1064827501387826](https://doi.org/10.1137/S1064827501387826)\.
- \[50\]J\.\-M\. Lasry, P\.\-L\. Lions,[Mean field games](https://doi.org/10.1007/s11537-007-0657-8), Japanese Journal of Mathematics 2 \(1\) \(2007\) 229–260\.[doi:10\.1007/s11537\-007\-0657\-8](https://doi.org/10.1007/s11537-007-0657-8)\.
- \[51\]Z\. Li, N\. Kovachki, K\. Azizzadenesheli, B\. Liu, K\. Bhattacharya, A\. Stuart, A\. Anandkumar,[Fourier neural operator for parametric partial differential equations](https://arxiv.org/abs/2010.08895), 2021\.[arXiv:2010\.08895](http://arxiv.org/abs/2010.08895)\.
- \[52\]W\. E, J\. Han, A\. Jentzen,[Deep learning\-based numerical methods for high\-dimensional parabolic partial differential equations and backward stochastic differential equations](https://doi.org/10.1007/s40304-017-0117-6), Communications in Mathematics and Statistics 5 \(4\) \(2017\) 349–380\.[doi:10\.1007/s40304\-017\-0117\-6](https://doi.org/10.1007/s40304-017-0117-6)\.Similar Articles
Optimal Adaptive Market Making: A Theoretical Framework for High-Yield Liquidity Provision in Perpetual Futures Markets
This paper develops a rigorous theoretical framework for optimal market making in perpetual futures markets with zero maker fees, deriving conditions for high annualized returns and unifying classical models.
Reducing Credit Assignment Variance via Counterfactual Reasoning Paths
Introduces Implicit Behavior Policy Optimization (IBPO), a counterfactual comparison-based credit assignment framework that improves training stability and performance in multi-step reasoning tasks for large language models by converting sparse terminal rewards into step-sensitive learning signals.
Solving the Credit Assignment Problem in Multi-Agent Systems (CANTANTE Framework)
CANTANTE is an open-source framework that solves the credit assignment problem in multi-agent systems by converting system-level rewards into per-agent update signals, outperforming DSPy-based baselines on coding and math reasoning benchmarks.
Learning from Environmental Feedback: Credit Assignment across Multiple Timescales for Agentic Reinforcement Learning
This paper introduces EFCA, a multi-timescale credit assignment method for agentic reinforcement learning that uses short-term feedback and medium-term state-history signals from environment interaction to improve task success and quality on ALFWorld and WebShop.
Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents
Gated-BEPO is a new credit assignment method for LLM agents that derives step-level credit from empirical rollout graphs using Bellman fixed-point estimation and adaptively fuses it with episode-level credit via a confidence gate. Experiments on WebShop, ALFWorld, and visual Sokoban show consistent improvements over existing critic-free methods.