Designing Agentic AI Workflow Portfolios under Imperfect Selection and Compute Cost

arXiv cs.AI Papers

Summary

This paper proposes a framework for designing workflow portfolios in agentic AI systems to optimize selection and compute allocation, demonstrating improved accuracy through mathematical optimization and empirical evaluations on datasets.

arXiv:2609.18126v1 Announce Type: new Abstract: Agentic AI systems often approach the same task through multiple workflows that differ in reasoning strategy, verification structure, and compute cost. A natural deployment policy is to use the workflow with the highest average performance, but this can be suboptimal because different workflows may succeed on different instances. We study a portfolio-and-selector paradigm in which a firm runs multiple workflow executions and selects the final answer after observing their outputs. Additional executions may uncover correct answers that the best standalone workflow misses, but they consume compute and introduce plausible distractors that complicate final selection. We formulate this as a workflow portfolio problem in which the firm jointly chooses run size and allocation across workflow types. We summarize selector quality through an odds-lift index and derive sharp bounds on the value of workflow variety. For finite workflow pools, we develop exact formulations, linear programming relaxations, randomized rounding procedures, and computable performance certificates. For large implicit workflow classes, we derive a finite-dimensional dual and an ellipsoid method using a pricing oracle to identify workflows with high weighted accuracy net of recurring compute cost. Under a weak condition, the method obtains a near-optimal solution to the relaxation with polynomially many oracle calls. We evaluate the framework on three datasets: ABCD, Schema-Guided Dialogue, and HotpotQA. Relative to the best standalone workflow, portfolio optimization improves held-out selector accuracy by 3.1, 7.5, and 0.9 percentage points, respectively. Dual-guided workflow generation adds 3.5 points on ABCD and 24.1 on HotpotQA, with no additional gain on Schema-Guided Dialogue.
Original Article
View Cached Full Text

Cached at: 09/17/26, 09:33 AM

# Designing Agentic AI Workflow Portfolios under Imperfect Selection and Compute Cost
Source: [https://arxiv.org/html/2609.18126](https://arxiv.org/html/2609.18126)
###### Abstract

Agentic AI systems often approach the same task through multiple workflows that differ in reasoning strategy, verification structure, and compute cost\. A natural deployment policy is to use the workflow with the highest average performance, but this can be suboptimal because different workflows may succeed on different instances\. We study a portfolio\-and\-selector paradigm in which a firm runs multiple workflow executions and selects the final answer after observing their outputs\. Additional executions may uncover correct answers that the best standalone workflow misses, but they consume compute and introduce plausible distractors that complicate final selection\. We formulate this as a workflow portfolio problem in which the firm jointly chooses run size and allocation across workflow types\. We summarize selector quality through an odds\-lift index and derive sharp bounds on the value of workflow variety\. For finite workflow pools, we develop exact formulations, linear programming relaxations, randomized rounding procedures, and computable performance certificates\. For large implicit workflow classes, we derive a finite\-dimensional dual and an ellipsoid method using a pricing oracle to identify workflows with high weighted accuracy net of recurring compute cost\. Under a weak condition, the method obtains a near\-optimal solution to the relaxation with polynomially many oracle calls\. We evaluate the framework on three datasets: ABCD, Schema\-Guided Dialogue, and HotpotQA\. Relative to the best standalone workflow, portfolio optimization improves held\-out selector accuracy by 3\.1, 7\.5, and 0\.9 percentage points, respectively\. Dual\-guided workflow generation adds 3\.5 points on ABCD and 24\.1 on HotpotQA, with no additional gain on Schema\-Guided Dialogue\.

###### keywords

agentic AI; workflow portfolios; large language models; endogenous portfolio size; compute allocation; selector design; selector strength; odds lift; random utility; ellipsoid method; separation oracle; repeated execution

††manuscriptno:MS\-0000\-0000\.00††runningauthor:Abdolmaleki, Jasin and Wang††runningtitle:Workflow Portfolios under Imperfect Selection and Compute Cost††authors:Ross School of Business, University of Michigan, United States,[mojtabaa@umich\.edu](mailto:[email protected])Ross School of Business, University of Michigan, United States,[sjasin@umich\.edu](mailto:[email protected])TrueFoundry,[boyu\.wang@truefoundry\.com](mailto:[email protected])††affiliation:††affiliation:††affiliation:††affiliation:††affiliation:††affiliation:††affiliation:School of Management, University of San Francisco, United States,[bwang62@usfca\.edu](mailto:[email protected]>)## 1Introduction

Artificial intelligence is moving beyond individual productivity tools and into the operating core of organizations\. Enterprise software vendors are embedding task\-specific AI capabilities into business applications, while firms are beginning to delegate not only information retrieval and drafting, but also parts of operational decision processes to AI systems\([Gartner 2025a](https://arxiv.org/html/2609.18126#bib.bib1),[Gartner 2025b](https://arxiv.org/html/2609.18126#bib.bib2)\)\. Recent business surveys report that many organizations are experimenting with or deploying agentic AI, while also emphasizing that realizing value requires redesigning workflows, establishing governance, and building reliable orchestration layers rather than simply giving employees access to models\([McKinsey & Company 2025](https://arxiv.org/html/2609.18126#bib.bib3),[McKinsey & Company 2026](https://arxiv.org/html/2609.18126#bib.bib4),[Ransbotham et al\. 2025](https://arxiv.org/html/2609.18126#bib.bib5),[Harvard Business Review Analytic Services 2026](https://arxiv.org/html/2609.18126#bib.bib6)\)\. This shift raises a natural operations\-management question: once AI workflows become part of recurring business processes, how should firms decide which workflows to deploy and how much compute to devote to them?

The setting we study arises in service operations, compliance, claims processing, internal knowledge work, and decision support\. A customer\-service system may classify a ticket, retrieve relevant policies, identify the customer’s need, draft a response, check the response for compliance, and decide whether to escalate the case to a human\. A claims\-processing system may summarize a claim, compare it with contract language, identify relevant exceptions, and recommend whether to approve or deny it\. A compliance assistant may retrieve applicable rules, analyze a transaction, flag inconsistencies, and recommend an appropriate action\. In each of these settings, the relevant operational unit is not a single model call, but a workflow: a sequence or graph of prompts, model calls, retrieval steps, tools, verifiers, and aggregation rules that produces a candidate answer\.

At first glance, a natural approach is to identify the workflow with the highest average accuracy and deploy it alone\. Average performance, however, can conceal substantial heterogeneity across tasks: different workflows may succeed on very different cases\. Relying on a single workflow can therefore leave important gaps that another workflow could fill\. A firm may have many workflows that it could develop and maintain\. Some are cheap and fast, while others are more expensive but more careful\. Some rely on retrieval, some on decomposition or repeated reasoning, and others on verification or critique\. A workflow that performs poorly on average may still be valuable if it solves cases that stronger workflows miss, whereas a high\-performing workflow may add little if it succeeds mainly on cases that are already easy\. The firm must therefore decide which workflow types to maintain, how many times to execute each type, and which candidate answer to select as the final output, while accounting for the recurring compute cost of these decisions\.

Much of the existing literature follows a plan, then execute paradigm\([Shen et al\. 2023](https://arxiv.org/html/2609.18126#bib.bib56),[Liang et al\. 2024](https://arxiv.org/html/2609.18126#bib.bib57),[Lu et al\. 2023](https://arxiv.org/html/2609.18126#bib.bib58),[Zhang et al\. 2025b](https://arxiv.org/html/2609.18126#bib.bib32),[Hu et al\. 2025](https://arxiv.org/html/2609.18126#bib.bib15),[Zhang et al\. 2025a](https://arxiv.org/html/2609.18126#bib.bib33)\)\. A large model, router, or planner first selects or designs a single workflow, which is then executed to produce the final answer\. This plan–then–execute paradigm is illustrated in[Figure1](https://arxiv.org/html/2609.18126#S1.F1)\. This approach is natural when the system can reliably predict, before execution, which workflow is best suited to the task\. In many operational settings, however, the firm may not know in advance which workflow will succeed\. It may be easier to recognize a correct answer after observing the candidate outputs and their supporting evidence, such as execution traces, retrieved sources, agreement patterns, confidence signals, or verifier results, than to identify ex ante the single workflow that will generate it\.

Large modelplans oneworkflowDecomposerSolverVerifierSolverVerifierFinaloutputStep one: expensive planningStep two: execute the single planned workflow with cheaper agentsFigure 1:Traditional orchestration paradigm: a large model synthesizes one workflow, which is then executed by cheaper specialized agents\.In this paper, we propose a portfolio and selector paradigm\. The firm first builds and maintains a set of workflow types\. For each task class or deployment setting, it decides how many execution slots to use and how to allocate those slots across the available workflows\. The same workflow may receive more than one slot\. The selected executions are then run in parallel, and a selector chooses the final answer from the resulting candidates\. This operating model creates a fundamental tradeoff\. An additional execution may produce a correct answer that would otherwise be unavailable, or it may increase the representation of a workflow that performs well on many tasks\. At the same time, every execution consumes tokens, latency, tool calls, and money\. If its output is wrong, it also adds another plausible distractor for the selector\. More executions are therefore not necessarily better\.

Workflow AWorkflow BWorkflow C…Select workflows:accuracy–compute optimization1\. Portfolio design and allocationDecomposeSolveVerifySolveVerifyAnswer BDecomposeSolveVerifyAnswer ADecomposeRetrieveSolveVerifyAnswer CABC2\. Execute selected workflows in parallelSelectormodelFinaloutput3\. Post\-output selectionFigure 2:Portfolio\-based orchestration: the decision maker selects workflow types and allocates execution slots, runs the selected workflows in parallel, and chooses the final answer after observing their realized outputs\.The central question of this paper is:*How should a firm decide how many workflow executions to run and how to allocate them across workflow types, when additional executions can improve the quality of the candidate set but also increase selector confusion and compute cost?*

Related literature\.Our work brings together several related streams, which we discuss in detail in[Section2](https://arxiv.org/html/2609.18126#S2)\. One stream studies how to design effective agentic workflows and inference time architectures\([Zhang et al\. 2025b](https://arxiv.org/html/2609.18126#bib.bib32),[Hu et al\. 2025](https://arxiv.org/html/2609.18126#bib.bib15),[Zhuge et al\. 2024](https://arxiv.org/html/2609.18126#bib.bib34),[Saad\-Falcon et al\. 2025](https://arxiv.org/html/2609.18126#bib.bib27),[Zhang et al\. 2025a](https://arxiv.org/html/2609.18126#bib.bib33)\)\. A second stream studies how to route queries across models or workflows to balance quality and cost\([Chen et al\. 2024](https://arxiv.org/html/2609.18126#bib.bib10),[Ong et al\. 2025](https://arxiv.org/html/2609.18126#bib.bib25),[Hu et al\. 2024](https://arxiv.org/html/2609.18126#bib.bib16),[Huang et al\. 2025](https://arxiv.org/html/2609.18126#bib.bib17)\)\. A third stream generates multiple candidate answers and uses ranking, fusion, self consistency, or multi\-agent aggregation to choose among them\([Wang et al\. 2023](https://arxiv.org/html/2609.18126#bib.bib30),[Jiang et al\. 2023](https://arxiv.org/html/2609.18126#bib.bib18),[Wang et al\. 2025](https://arxiv.org/html/2609.18126#bib.bib31)\)\. Together, these studies address important parts of the deployment problem, namely workflow design, pre\-execution allocation, and post\-output aggregation\. Once these decisions are considered jointly, however, the problem becomes one of choosing and allocating a portfolio of alternatives under limited resources and imperfect final selection\. This perspective relates to several established areas in operations research, including choice modeling, assortment optimization, coverage, submodular optimization, column generation, and optimization through separation oracles\([Luce 1959](https://arxiv.org/html/2609.18126#bib.bib21),[McFadden 1974](https://arxiv.org/html/2609.18126#bib.bib22),[van Ryzin and Mahajan 1999](https://arxiv.org/html/2609.18126#bib.bib29),[Talluri and van Ryzin 2004](https://arxiv.org/html/2609.18126#bib.bib28),[Bertsimas and Mišić 2019](https://arxiv.org/html/2609.18126#bib.bib68),[Dong et al\. 2023](https://arxiv.org/html/2609.18126#bib.bib67),[Wang et al\. 2024](https://arxiv.org/html/2609.18126#bib.bib60),[Nemhauser et al\. 1978](https://arxiv.org/html/2609.18126#bib.bib23),[Barman et al\. 2021](https://arxiv.org/html/2609.18126#bib.bib7),[Alaei et al\. 2010](https://arxiv.org/html/2609.18126#bib.bib61),[Dantzig and Wolfe 1960](https://arxiv.org/html/2609.18126#bib.bib12),[Desaulniers et al\. 2005](https://arxiv.org/html/2609.18126#bib.bib13),[Grötschel et al\. 1981](https://arxiv.org/html/2609.18126#bib.bib54),[Grötschel et al\. 1988](https://arxiv.org/html/2609.18126#bib.bib55)\)\. What remains missing is an integrated framework for designing workflow portfolios when both execution decisions and final answer selection affect performance and cost\.

The challenge is not limited to choosing among workflows that have already been evaluated\. In practice, the feasible workflow space is large and cannot usually be enumerated in advance\. We therefore need a way to model how the firm accesses workflows outside its current evaluated pool\. In this paper, we assume that new workflows can be proposed by an oracle, which may take the form of a teacher model, an automated workflow generator, a structured search procedure, or a human engineering team\. The firm guides the oracle toward generating workflows that perform well on tasks that the current portfolio handles poorly, while also accounting for the recurring cost of executing those workflows\. This raises two questions: how should the firm use the oracle to search the implicit workflow space, and how can it certify the quality of the resulting portfolio?

We address these questions through a selector aware workflow portfolio model\. The firm decides how many times to execute each workflow type, which jointly determines the composition and total size of the candidate set\. Each execution produces a candidate answer, and a selector chooses the final output after observing the resulting candidates\. This portfolio\-and\-selector paradigm is illustrated in[Figure2](https://arxiv.org/html/2609.18126#S1.F2)\. We summarize selector performance through recovery curves that relate the number of correct candidates in the set to the probability that the final selected answer is correct\. The firm’s objective is to maximize deployed decision quality net of recurring compute cost\.

Our contributions\.Our first contribution develops a general theory of selector strength\. We introduce an odds lift index that measures how much the selector increases the odds of returning a correct answer relative to random selection\. A finite bound on this index places the selector’s recovery curve below a common concave envelope\. This envelope gives a sharp limit on how much any workflow portfolio can improve on the best single workflow, regardless of how many workflows are available or how complementary they appear\. It also yields simple screening rules that identify when selector quality and execution cost make a multi\-workflow portfolio unattractive, so that the best decision is to use a single workflow or not deploy the system\.

Our second contribution develops an optimization framework for finite workflow pools when the total number of workflow executions, which we call the*run size*, is itself a decision\. The framework determines both the run size and the allocation of execution slots across workflow types, recognizing that an additional execution may improve accuracy on some tasks while increasing selector confusion on others\. For any fixed run size, the problem becomes a concave coverage problem\. This structure leads to exact integer programming formulations, linear programming relaxations, randomized rounding procedures, and computable optimality certificates\. We then develop a sparse grid method for choosing the run size\. Rather than solving a separate optimization problem for every possible size, the method evaluates only a small collection of candidate values while retaining a provable approximation guarantee\. This provides a practical approach for jointly choosing how many workflow executions to run and how to allocate them across workflow types\.

Our third contribution develops an optimization method over the*implicit workflow class*, by which we mean the full set of feasible workflows that the oracle can search over but the firm cannot enumerate in advance\. Because the same workflow type may occupy multiple execution slots, the linear program for a fixed run size does not require an upper bound on the number of times each type can be used\. Although this linear program contains one variable for every workflow type in the implicit class, its dual has only finitely many task price variables, together with one constraint for each workflow\. A global pricing oracle can therefore identify a workflow whose constraint is violated, or certify that no important violation remains\. Using the classical equivalence between optimization and separation, the ellipsoid method computes a solution withinε\\varepsilonof the full linear programming relaxation using a number of oracle calls that is polynomial in the problem encoding andlog⁡\(1/ε\)\\log\(1/\\varepsilon\)\. Combining this optimization error with the losses from the cardinality grid and randomized rounding yields a high probability approximation guarantee relative to the entire implicit workflow class, without requiring that class to be enumerated\.

Our fourth contribution is empirical\. We first calibrate selector recovery across the ABCD and Schema\-Guided Dialogue service\-operations tasks\([Chen et al\. 2021a](https://arxiv.org/html/2609.18126#bib.bib44),[Rastogi et al\. 2020](https://arxiv.org/html/2609.18126#bib.bib45)\)and the HotpotQA question\-answering benchmark\([Yang et al\. 2018](https://arxiv.org/html/2609.18126#bib.bib46)\)\. Selector strength is substantially above random selection in the two service domains and remains positive, although weaker, on HotpotQA\. We then evaluate the complete stochastic workflow\- generation and portfolio\-optimization pipeline in all three domains\. On fresh held\-out tasks, actual selector accuracy rises from 43\.625% for the best initial singleton to 50\.250% for the final ABCD portfolio, from 85\.250% to 92\.750% on Schema\-Guided Dialogue, and from 30\.250% to 55\.250% on HotpotQA\. Dual\-guided generation incorporates four workflows on ABCD, none on Schema\-Guided Dialogue, and one on HotpotQA\. This variation is consistent with the model’s central implication: a generated workflow creates value only when its incremental coverage and selection benefit justify its recurring compute cost\.

The central message is that deploying agentic AI is neither a search for the single best workflow nor a simple rule that more candidates are always better\. Workflow variety creates value only when the selector can identify useful outputs reliably enough to justify the additional compute cost and selection difficulty\. Effective deployment therefore requires firms to manage workflow generation, run size, the allocation of executions across workflow types, selector strength, and compute expenditure as a joint decision\.

Organization of the paper\.The rest of the paper is organized as follows\.[Section2](https://arxiv.org/html/2609.18126#S2)reviews the related literature\.[Section3](https://arxiv.org/html/2609.18126#S3)formulates the portfolio and selector problem and introduces the oracle interface for searching the implicit workflow class\.[Section4](https://arxiv.org/html/2609.18126#S4)studies how selector strength limits the value of workflow variety\.[Section5](https://arxiv.org/html/2609.18126#S5)develops the finite pool optimization framework, including exact integer programs, linear programming relaxations, randomized rounding, and optimality certificates\.[Section6](https://arxiv.org/html/2609.18126#S6)derives the dual formulation for the implicit workflow class and uses a workflow pricing oracle together with the ellipsoid method to obtain performance guarantees without enumerating all feasible workflows\.[Section7](https://arxiv.org/html/2609.18126#S7)extends the model and its algorithmic guarantees to stochastic workflow execution\.[Section8](https://arxiv.org/html/2609.18126#S8)presents the numerical experiments, and[Section9](https://arxiv.org/html/2609.18126#S9)concludes the paper\. Unless otherwise noted, all proofs are provided in the appendix\.

## 2Related Literature

Our work relates to several streams of literature\. The first studies the automated design of agentic workflows, compound AI systems, and LLM generated algorithms or heuristics\. The second examines model routing, cascades, and methods that generate and aggregate multiple candidate outputs\. The third includes work on random utility, assortment choice, and selection among competing alternatives\. The fourth concerns coverage, submodular optimization, and optimization over large implicitly represented decision spaces\. Finally, our paper contributes to the emerging operations management literature on the deployment, allocation, and governance of AI systems\. We discuss these streams in turn and explain how they inform, but do not resolve, the joint problem of workflow generation, execution allocation, and final answer selection studied here\.

#### Automated design of agentic workflows and compound AI systems\.

A growing computer science literature studies how to design and optimize multistep language model systems\. AFlow searches over workflows represented as code, using execution feedback and tree search\([Zhang et al\. 2025b](https://arxiv.org/html/2609.18126#bib.bib32)\)\. Automated Design of Agentic Systems treats agent design as a search problem in which a meta agent proposes new agentic systems\([Hu et al\. 2025](https://arxiv.org/html/2609.18126#bib.bib15)\)\. GPTSwarm represents interacting language agents as a graph whose structure can be optimized\([Zhuge et al\. 2024](https://arxiv.org/html/2609.18126#bib.bib34)\)\. Archon searches over inference architectures that combine repeated sampling, ranking, critique, verification, fusion, and model choice\([Saad\-Falcon et al\. 2025](https://arxiv.org/html/2609.18126#bib.bib27)\)\. MaAS constructs an agentic supernet and selects task dependent multiagent architectures to balance performance and resource use\([Zhang et al\. 2025a](https://arxiv.org/html/2609.18126#bib.bib33)\)\. LLMSelector studies model assignment within a compound AI system, asking which language model should be used in each module of a fixed multicall architecture\([Chen et al\. 2025](https://arxiv.org/html/2609.18126#bib.bib11)\)\. Collectively, these studies show that workflow structure, module choice, and automated search can have substantial effects on system performance\. Our work treats all these methods as potential mechanisms for generating candidate workflows, but studies a different operational decision\. The primary objective in automated workflow design is typically to discover a strong workflow, architecture, or assignment of models to modules\. We instead ask which workflow types should be retained, how many execution slots should be allocated to each, and how much compute should be spent when the final answer is chosen only after the candidate outputs are observed\. Under this perspective, a workflow that is not the strongest on average may still be valuable if it solves cases missed by the existing portfolio and the selector can recognize its contribution\. The same workflow may be harmful if it mainly produces costly distractors that make final selection more difficult\.

#### LLM based algorithm design and heuristic portfolio generation\.

A related literature studies the use of large language models to design algorithms and heuristics\. Recent surveys organize this work around the roles of language models as optimizers, designers, predictors, and components of algorithmic search\([Liu et al\. 2026b](https://arxiv.org/html/2609.18126#bib.bib20)\)\. AlphaEvolve combines language model generated code modifications with evaluator feedback and evolutionary search to improve algorithms\([Novikov et al\. 2025](https://arxiv.org/html/2609.18126#bib.bib24)\)\. EoH\-S moves beyond the search for a single heuristic and instead constructs a small set of complementary heuristics\([Liu et al\. 2026a](https://arxiv.org/html/2609.18126#bib.bib19)\)\. The recurring process of generating candidates, evaluating their performance, and reoptimizing the system also has methodological parallels in self adjusting control and joint learning and optimization\([Jasin 2014](https://arxiv.org/html/2609.18126#bib.bib40),[Chen et al\. 2019](https://arxiv.org/html/2609.18126#bib.bib41),[Chen et al\. 2021b](https://arxiv.org/html/2609.18126#bib.bib42),[Zhang and Jasin 2022](https://arxiv.org/html/2609.18126#bib.bib43),[Agrawal et al\. 2014](https://arxiv.org/html/2609.18126#bib.bib62),[Keskin and Zeevi 2014](https://arxiv.org/html/2609.18126#bib.bib63),[Elmachtoub and Grigas 2022](https://arxiv.org/html/2609.18126#bib.bib64),[Faradonbeh and Faradonbeh 2023](https://arxiv.org/html/2609.18126#bib.bib66)\)\. The above literature is closely related to our workflow generation layer\. In our setting, however, the generated objects are executable AI workflows drawn from an implicit workflow class, and their value depends on how they are deployed together\. The objective therefore accounts not only for the performance of each generated workflow, but also for its contribution to a portfolio under imperfect final selection and its recurring execution cost\.

#### LLM routing, model cascades, and efficient inference\.

Another stream studies how to allocate queries across models with different cost and quality profiles\. FrugalGPT develops adaptive strategies, including model cascades, to reduce inference cost while preserving or improving performance\([Chen et al\. 2024](https://arxiv.org/html/2609.18126#bib.bib10)\)\. RouteLLM learns routers that choose between stronger and weaker models using preference data\([Ong et al\. 2025](https://arxiv.org/html/2609.18126#bib.bib25)\), while RouterBench provides a benchmark for evaluating systems that route queries across multiple language models\([Hu et al\. 2024](https://arxiv.org/html/2609.18126#bib.bib16)\)\. ThriftLLM studies the cost conscious selection of LLM ensembles for classification tasks\([Huang et al\. 2025](https://arxiv.org/html/2609.18126#bib.bib17)\)\. Recent operations research also treats heterogeneous LLM use as a resource allocation problem\.[Dean et al\. \(2026\)](https://arxiv.org/html/2609.18126#bib.bib37)optimize parallel query counts under reliability constraints,[Li et al\. \(2026\)](https://arxiv.org/html/2609.18126#bib.bib38)study sequential model choice and stopping with query and waiting costs, and[Guo et al\. \(2026\)](https://arxiv.org/html/2609.18126#bib.bib39)analyze static cascades under congestion and latency\. Together, these papers determine how model calls should be allocated, sequenced, or stopped\. Our paper studies a complementary problem\. Rather than deciding which model to query next, we choose how many workflow executions to run, how to allocate those executions across workflow types, and which answer to select from the resulting candidates\. Routing and cascade methods govern the process of acquiring model outputs, while our framework governs the design and evaluation of the candidate portfolio presented to the selector\.

#### Output ensembling, ranking, self consistency, and inference time aggregation\.

A related literature improves language model performance by generating multiple candidate outputs and then selecting, combining, or refining them\. Self consistency samples several reasoning paths and returns the answer with the greatest agreement across paths\([Wang et al\. 2023](https://arxiv.org/html/2609.18126#bib.bib30)\)\. LLM Blender ranks outputs from multiple language models and combines the highest quality candidates through generative fusion\([Jiang et al\. 2023](https://arxiv.org/html/2609.18126#bib.bib18)\)\. Mixture of Agents uses a layered architecture in which agents at later stages build on outputs produced by agents in earlier stages\([Wang et al\. 2025](https://arxiv.org/html/2609.18126#bib.bib31)\)\. Archon treats ranking, fusion, critique, verification, and related inference techniques as components of an architecture that can be selected through search\([Saad\-Falcon et al\. 2025](https://arxiv.org/html/2609.18126#bib.bib27)\)\. Our paper abstracts from the details of the selector itself\. The selector may use ranking, fusion, consistency, critique, verification, or some other mechanism\. We instead study the upstream portfolio decision: which workflows should generate candidates, how many executions should be run, and when an additional candidate is worth its compute cost and its effect on final selection\.

#### Choice modeling, coverage, and implicit optimization\.

Our selector model draws on random utility and assortment choice\. Multinomial logit and Plackett–Luce provide tractable models of selection and ranking\([Luce 1959](https://arxiv.org/html/2609.18126#bib.bib21),[McFadden 1974](https://arxiv.org/html/2609.18126#bib.bib22),[Plackett 1975](https://arxiv.org/html/2609.18126#bib.bib26),[van Ryzin and Mahajan 1999](https://arxiv.org/html/2609.18126#bib.bib29),[Talluri and van Ryzin 2004](https://arxiv.org/html/2609.18126#bib.bib28)\)\. Related work estimates preference parameters from observed choices and response times\([Echenique et al\. 2025](https://arxiv.org/html/2609.18126#bib.bib69)\)\. In our setting, the alternatives are workflow outputs rather than products, and their value depends on whether they are correct\. A wrong output enlarges the selector’s choice set without adding correct attraction\. We use Plackett–Luce as an estimable specialization, while our main variety bound relies only on a general odds lift envelope\. This choice structure leads directly to the optimization problem studied in the paper\. Once the run size is fixed, portfolio composition affects selector performance only through the number of correct outputs on each task, yielding a concave coverage problem and allowing us to use established linear programming and rounding tools\([Nemhauser et al\. 1978](https://arxiv.org/html/2609.18126#bib.bib23),[Barman et al\. 2021](https://arxiv.org/html/2609.18126#bib.bib7),[Ageev and Sviridenko 2004](https://arxiv.org/html/2609.18126#bib.bib35),[Chekuri et al\. 2010](https://arxiv.org/html/2609.18126#bib.bib36)\)\. When run size is endogenous, however, adding an execution changes both the correct count and the total number of candidates, so the global objective is neither monotone nor generally submodular\. We handle this additional decision through a sparse cardinality grid\. Finally, when the feasible workflow class cannot be enumerated, a global workflow pricing oracle separates the workflow indexed dual constraints, allowing the ellipsoid method to optimize over the implicit class\([Grötschel et al\. 1981](https://arxiv.org/html/2609.18126#bib.bib54),[Grötschel et al\. 1988](https://arxiv.org/html/2609.18126#bib.bib55)\)\.

#### Operations management and AI deployment\.

An emerging operations management literature studies how AI systems should be adopted, allocated, priced, and governed\. Recent work examines query allocation across heterogeneous language models\([Dean et al\. 2026](https://arxiv.org/html/2609.18126#bib.bib37),[Li et al\. 2026](https://arxiv.org/html/2609.18126#bib.bib38),[Guo et al\. 2026](https://arxiv.org/html/2609.18126#bib.bib39)\), the timing of AI adoption and access\([Abdolmaleki and Duenyas 2026](https://arxiv.org/html/2609.18126#bib.bib47)\), pricing and delay in agentic AI services\([Abdolmaleki et al\. 2026](https://arxiv.org/html/2609.18126#bib.bib48)\), and the provenance of language model generated content\([Radvand et al\. 2026](https://arxiv.org/html/2609.18126#bib.bib49)\)\.[Baek et al\. \(2026\)](https://arxiv.org/html/2609.18126#bib.bib65)study complementarity among humans, large language models, and operations research algorithms in inventory control\. We contribute to this literature by studying workflow portfolio design\. Our focus is how a firm should choose the number and allocation of workflow executions, generate new workflow types, and select among their outputs when performance depends jointly on workflow complementarity, selector strength, and recurring compute cost\.

## 3Model

This section develops the operational model for selector\-aware workflow portfolio design\. We first define the task environment, workflow executions, selector recovery, and the firm’s objective of choosing both the composition and size of the deployed portfolio\. We then distinguish optimization over a finite evaluated workflow pool from search over the larger implicit workflow class and formalize the cost\-aware pricing interface used to generate new workflows\.

### 3\.1Deployment setting

We study an operational setting in which a firm handles a stream of similar tasks\. Although the tasks share a common objective, individual instances may differ substantially in the information, reasoning, and verification required to solve them\. A workflow designed for one source of difficulty may therefore perform poorly on cases that require a different approach\. This heterogeneity creates a role for workflow portfolios: by running several complementary workflows and selecting among their outputs, the firm may obtain a correct answer on cases that no single workflow handles reliably\. As a running example, consider customer\-service routing\. A customer initiates a service conversation, and the system observes the customer’s initial messages together with relevant account information, order status, and policy context\. Determining the appropriate route may require the system to infer the customer’s objective, identify the current service state, retrieve the applicable policy, and verify that the proposed action is feasible\. The final decision may be, for example, to initiate an order cancellation, address a refund\-status inquiry, modify a subscription, begin troubleshooting, or escalate the case to a human agent\. Different workflows may emphasize different parts of this reasoning process\. Running multiple workflows can increase the likelihood that at least one produces the correct route, but it also incurs additional compute cost and creates more candidate outputs for the selector to distinguish\. The model below formalizes this tradeoff\. \(Remark[3\.1](https://arxiv.org/html/2609.18126#S3.Thmtheorem1)discusses the extension to settings with multiple task types\.\)

Tasks\.LetX∈𝒳X\\in\\mathcal\{X\}denote a task instance drawn from a population𝒟\\mathcal\{D\}, and letY⋆∈𝒴Y^\{\\star\}\\in\\mathcal\{Y\}denote the correct answer, target decision, or benchmark solution\. The output space𝒴\\mathcal\{Y\}may be finite, numerical, structured, or textual\. In the baseline model, a returned answeryyreceives payoff𝟏\{y=Y⋆\}\.\\mathbf\{1\}\\\{y=Y^\{\\star\}\\\}\.In the customer\-service routing example,XXcontains the observed conversation, account information, order status, and relevant context, while the targetY⋆Y^\{\\star\}is the correct routing label, such as “cancel order,” “refund status,” “subscription change,” or “escalate to human\.”

Workflows\.A workflowg∈𝒢g\\in\\mathcal\{G\}is an executable AI procedure that maps a task instance to a candidate answer and an observable trace\. Workflows can take many forms: a single prompt, a chain of specialized agents, a retrieval\-augmented routine, a verify\-and\-revise loop, or a graph of agents and tools\. The feasible space𝒢\\mathcal\{G\}contains all workflows consistent with the firm’s operational constraints, such as allowed models, tools, retrieval sources, context length, latency, and termination rules\. The main analysis does not require a particular graph representation of workflows; it only uses each workflow’s evaluated correctness vector and running cost\. Appendix[A](https://arxiv.org/html/2609.18126#A1)gives one possible graph\-based representation of𝒢\\mathcal\{G\}for readers who want a concrete finite workflow class\.

When workflowggis run on taskXX, it produces

Og​\(X\)=\(Ag​\(X\),Tg​\(X\)\),O\_\{g\}\(X\)=\\bigl\(A\_\{g\}\(X\),T\_\{g\}\(X\)\\bigr\),\(1\)whereAg​\(X\)∈𝒴A\_\{g\}\(X\)\\in\\mathcal\{Y\}is the candidate answer andTg​\(X\)T\_\{g\}\(X\)is the observable execution trace, potentially including intermediate messages, retrieved documents, tool outputs, confidence scores, verifier signals, token counts, and latency\. We suppress the dependence onXXwhen it is clear from context\. Because𝒢\\mathcal\{G\}is typically too large to enumerate, the firm must search over an “implicit” workflow space rather than selecting from a fixed list; we return to this search layer in[Sections3\.2](https://arxiv.org/html/2609.18126#S3.SS2)and[3\.3](https://arxiv.org/html/2609.18126#S3.SS3)\. In the routing example, a workflow might first retrieve relevant policy text, infer the customer’s intent, and propose a routing label\. A verifier then checks the proposal against the conversation and retrieved policy\. If the proposal fails verification, a revision agent updates it using the verifier’s feedback, and the verify–revise loop continues until the proposal passes verification or a prespecified iteration limit is reached\.

CustomerticketRetrievepolicyInferintentProposerouteVerifyReviseproposalCandidateanswerpassfailfeedbackFigure 3:A customer\-service routing workflow with a bounded verify–revise loop\.Workflow accuracy\.Each workflow has a recurring per\-task running costcg≥0c\_\{g\}\\geq 0\. This cost may represent dollars, tokens, latency\-adjusted compute, tool\-call charges, or an additive resource index\. The cost is recurring because it is paid every time the workflow is run on an incoming task\. Let\(xi,yi⋆\)i=1n\(x\_\{i\},y\_\{i\}^\{\\star\}\)\_\{i=1\}^\{n\}be a labeled development sample, i\.e\., the training or validation tasks used to evaluate candidate workflows before deployment\. The portfolio optimization problem below is defined on this training sample: workflow correctness, and portfolio values are computed from\(xi,yi⋆\)i=1n\(x\_\{i\},y\_\{i\}^\{\\star\}\)\_\{i=1\}^\{n\}\. For workflowgg, define its task\-level correctness indicator by

ai​g=𝟏\{Ag\(xi\)=yi⋆\},i=1,…,n,g∈𝒢\.a\_\{ig\}=\\mathbf\{1\}\\\{A\_\{g\}\(x\_\{i\}\)=y\_\{i\}^\{\\star\}\\\},\\qquad i=1,\\ldots,n,\\quad g\\in\\mathcal\{G\}\.\(2\)Thus,ai​g=1a\_\{ig\}=1if workflowggreturns the correct routing label for taskii, andai​g=0a\_\{ig\}=0otherwise\. Throughout the main analysis, we treatai​ga\_\{ig\}as deterministic, so each workflow–task pair has a fixed evaluated correctness outcome on the development sample\.[Section7](https://arxiv.org/html/2609.18126#S7)extends the model to stochastic workflow execution, whereai​g∈\[0,1\]a\_\{ig\}\\in\[0,1\]denotes the probability that one execution of workflowggis correct on taskii\.

Portfolio and execution cost\.We useg∈𝒢g\\in\\mathcal\{G\}to index workflow types\. For each incoming task, the firm chooses an execution\-count vector

𝒎=\(mg\)g∈𝒢∈ℤ\+𝒢\\bm\{m\}=\(m\_\{g\}\)\_\{g\\in\\mathcal\{G\}\}\\in\\mathbb\{Z\}\_\{\+\}^\{\\mathcal\{G\}\}with finite support, wheremgm\_\{g\}is the number of execution slots assigned to workflow typegg\. It may appear natural to imposemg≤1m\_\{g\}\\leq 1for every workflow type\. Under an imperfect selector, however, repeated execution of a strong workflow may raise the share of correct candidate outputs on the tasks it already solves and thereby improve selector recovery; see Appendix[C](https://arxiv.org/html/2609.18126#A3)\. Under stochastic execution, repeated runs can also yield different outcomes, making multiplicity even more natural;[Section7](https://arxiv.org/html/2609.18126#S7)extends most of our results to that setting\.

The total number of workflow executions is

k⁡\(𝒎\)=∑g∈𝒢mg,k\(\\bm\{m\}\)=\\sum\_\{g\\in\\mathcal\{G\}\}m\_\{g\},\(3\)and must satisfyk⁡\(𝒎\)≤Kmaxk\(\\bm\{m\}\)\\leq K\_\{\\max\}, whereKmaxK\_\{\\max\}reflects operational limits such as latency, context\-window capacity, compute budget, or risk policy\.

Each execution occupies a separate slot, incurs costcgc\_\{g\}, and produces a separate candidate for the selector\. In the deterministic development\-sample model, all executions of workflow typeggshare the evaluated correctness vector\(a1​g,…,an​g\)\(a\_\{1g\},\\ldots,a\_\{ng\}\)\. The zero vector𝟎\\bm\{0\}represents the outside option of not deploying the AI system and has value zero\. Forg∈𝒢g\\in\\mathcal\{G\}, let𝒆g\\bm\{e\}\_\{g\}denote the execution\-count vector containing one execution of workflowggand zero executions of every other workflow\.

For a feasible execution\-count vector𝒎\\bm\{m\}, let

ri​\(𝒎\)=∑g∈𝒢ai​g​mgr\_\{i\}\(\\bm\{m\}\)=\\sum\_\{g\\in\\mathcal\{G\}\}a\_\{ig\}m\_\{g\}\(4\)denote the number of correct candidate outputs produced on taskii, and let

C⁡\(𝒎\)=∑g∈𝒢cg​mgC\(\\bm\{m\}\)=\\sum\_\{g\\in\\mathcal\{G\}\}c\_\{g\}m\_\{g\}\(5\)denote the total recurring execution cost\.

Selector\.Consider a feasible nonzero execution\-count vector𝒎\\bm\{m\}\. Choose any labelingg1,…,g\_\{1\},\\ldots,gk⁡\(𝒎\)g\_\{k\(\\bm\{m\}\)\}of its execution slots satisfying\|\{j:gj=g\}\|=mg\|\\\{j:g\_\{j\}=g\\\}\|=m\_\{g\}for allg∈𝒢\.g\\in\\mathcal\{G\}\.After all workflow executions have been completed, the firm observes

O𝒎=\(Ogj:j=1,…,k\(𝒎\)\)\.O\_\{\\bm\{m\}\}=\\bigl\(O\_\{g\_\{j\}\}:j=1,\\ldots,k\(\\bm\{m\}\)\\bigr\)\.\(6\)A post\-output selector then chooses one of the candidate outputs:

Sel\(X,O𝒎\)∈\{Agj:j=1,…,k\(𝒎\)\}\.\\operatorname\{Sel\}\(X,O\_\{\\bm\{m\}\}\)\\in\\left\\\{A\_\{g\_\{j\}\}:j=1,\\ldots,k\(\\bm\{m\}\)\\right\\\}\.\(7\)The selector may be an LLM judge, a verifier, a ranking model, a rule\-based policy checker, a human reviewer, or a combination of these mechanisms\. In the customer\-service routing example, the selector observes the proposed routing decisions and their execution traces and chooses the route submitted to the service system\.

We summarize selector performance through a family of*recovery curves*\{ψk​\(r\)\}k=1Kmax\\\{\\psi\_\{k\}\(r\)\\\}\_\{k=1\}^\{K\_\{\\max\}\}, one for each candidate\-set sizekk\. Each curve

ψk:\{0,1,…,k\}→\[0,1\],ψk​\(0\)=0,ψk​\(k\)=1,\\psi\_\{k\}:\\\{0,1,\\ldots,k\\\}\\to\[0,1\],\\qquad\\psi\_\{k\}\(0\)=0,\\quad\\psi\_\{k\}\(k\)=1,\(8\)maps the number of correct candidate outputsrrto the probability that the selector returns a correct final output\. The endpoint conditions are natural: if no candidate is correct, the selector cannot recover one, and if all candidates are correct, any choice succeeds\.

The empirical selector\-aware accuracy of a nonzero execution\-count vector𝒎\\bm\{m\}is

Jψ​\(𝒎\)=1n​∑i=1nψk⁡\(𝒎\)​\(ri​\(𝒎\)\),J\_\{\\psi\}\(\\bm\{m\}\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\psi\_\{k\(\\bm\{m\}\)\}\\bigl\(r\_\{i\}\(\\bm\{m\}\)\\bigr\),\(9\)withJψ​\(𝟎\)=0J\_\{\\psi\}\(\\bm\{0\}\)=0by convention\.

Deployment objective\.The firm’s goal is to choose an execution\-count vector𝒎\\bm\{m\}that maximizes selector\-aware accuracy on the development sample net of recurring execution cost\. To place accuracy and cost on a common scale, letγ≥0\\gamma\\geq 0denote the cost price\. If a correct decision is worthv\>0v\>0monetary units relative to an incorrect decision, dividing monetary value byvvgivesγ=1/v\\gamma=1/v\. The net deployed value of𝒎\\bm\{m\}is

Πψ,γ​\(𝒎\)=Jψ​\(𝒎\)−γ​C​\(𝒎\),Πψ,γ​\(𝟎\)=0\.\\Pi\_\{\\psi,\\gamma\}\(\\bm\{m\}\)=J\_\{\\psi\}\(\\bm\{m\}\)\-\\gamma C\(\\bm\{m\}\),\\qquad\\Pi\_\{\\psi,\\gamma\}\(\\bm\{0\}\)=0\.\(10\)The ideal deployment objective is

OPTψ,γ​\(𝒢\)=max𝒎∈ℤ\+𝒢𝒎​has finite supportk⁡\(𝒎\)≤Kmax⁡Πψ,γ​\(𝒎\)\.\\mathrm\{OPT\}\_\{\\psi,\\gamma\}\(\\mathcal\{G\}\)=\\max\_\{\\begin\{subarray\}\{c\}\\bm\{m\}\\in\\mathbb\{Z\}\_\{\+\}^\{\\mathcal\{G\}\}\\\\ \\bm\{m\}\\text\{ has finite support\}\\\\ k\(\\bm\{m\}\)\\leq K\_\{\\max\}\\end\{subarray\}\}\\Pi\_\{\\psi,\\gamma\}\(\\bm\{m\}\)\.\(11\)This is the best net value achievable over the full workflow space𝒢\\mathcal\{G\}, and serves as the benchmark for the algorithms developed in subsequent sections\.

### 3\.2Restricted pools and the inner–outer view

Solving \([11](https://arxiv.org/html/2609.18126#S3.E11)\) directly is generally infeasible because the workflow class𝒢\\mathcal\{G\}is large and implicit\. We therefore distinguish two objects\. The first is a finite evaluated pool of workflows, over which the portfolio problem can be solved directly\. The second is the larger implicit workflow class, over which the algorithm must search for additional useful workflows\.

For any finite evaluated poolℳ⊆𝒢\\mathcal\{M\}\\subseteq\\mathcal\{G\}, define the restricted benchmark

OPTψ,γ​\(ℳ\)=max𝒎∈ℤ\+ℳk⁡\(𝒎\)≤Kmax⁡Πψ,γ​\(𝒎\)\.\\mathrm\{OPT\}\_\{\\psi,\\gamma\}\(\\mathcal\{M\}\)=\\max\_\{\\begin\{subarray\}\{c\}\\bm\{m\}\\in\\mathbb\{Z\}\_\{\+\}^\{\\mathcal\{M\}\}\\\\ k\(\\bm\{m\}\)\\leq K\_\{\\max\}\\end\{subarray\}\}\\Pi\_\{\\psi,\\gamma\}\(\\bm\{m\}\)\.\(12\)Here,𝒎\\bm\{m\}is the execution\-count vector restricted to workflow types inℳ\\mathcal\{M\}\. Every workflow typeg∈ℳg\\in\\mathcal\{M\}has already been evaluated on the development sample, so its correctness vector\(a1​g,…,an​g\)\(a\_\{1g\},\\ldots,a\_\{ng\}\)and recurring execution costcgc\_\{g\}are known\. The restricted problem \([12](https://arxiv.org/html/2609.18126#S3.E12)\) is the best net value achievable using only these evaluated workflow types\. Whenℳ=𝒢\\mathcal\{M\}=\\mathcal\{G\}, it coincides with the full deployment benchmark \([11](https://arxiv.org/html/2609.18126#S3.E11)\); otherwise it is only a finite\-pool approximation\.

The role of the outer layer is to decide where to search next in the implicit class𝒢\\mathcal\{G\}\. The guiding principle is marginal value\. Once a finite\-pool linear program \(LP\) relaxation is solved, the optimization problem assigns dual prices to development tasks and to execution slots\. A high task price means that an additional correct candidate on that task would be valuable\. A high execution\-slot price means that a new workflow must deliver enough marginal value to justify occupying one of the limited run slots\. These prices are passed to a workflow\-pricing oracle, which searches the implicit class for a workflow with high price\-weighted correctness net of recurring cost\. Thus, the outer layer repeatedly asks a simple operational question:*is there a workflow, not yet available to the optimizer, whose expected marginal contribution is large enough to matter?*

Workflow generator oracle\.We do not answer the above question by searching𝒢\\mathcal\{G\}ourselves\. Instead, we assume access to a workflow generator, which we call the pricing oracle, and which may be a teacher model, an automated workflow\-search procedure, or a human engineering team\. The optimizer supplies the current prices, and the oracle proposes a workflow that scores well against them\. If the oracle returns such a workflow, the workflow is evaluated on the development sample and becomes available to the finite\-pool optimizer\. If the oracle can certify that no workflow has sufficiently large price\-adjusted value, then the current LP relaxation has no important missing workflow at that query tolerance\. We discuss this in more detail in[Section3\.3](https://arxiv.org/html/2609.18126#S3.SS3)\.

The overall information flow is summarized in[Figure4](https://arxiv.org/html/2609.18126#S3.F4)\.

Inner: finite\-pool optimizationOuter: workflow generationSolve finite\-poolrelaxationCompute task andslot pricesSearch𝒢\\mathcal\{G\}for ahigh\-value workflowEvaluate workflowand update the poolFigure 4:Inner–outer workflow\-generation loop\. The dashed regions distinguish the finite\-pool inner optimization from the outer workflow\-generation and pool\-update steps\.The next subsection formalizes this workflow\-pricing interface\.[Section5\.4](https://arxiv.org/html/2609.18126#S5.SS4)derives the corresponding task and slot prices from the LP relaxation, and[Section6](https://arxiv.org/html/2609.18126#S6)uses these prices to optimize over the implicit workflow class through an ellipsoid\-based method\.

### 3\.3Cost\-aware generation over an implicit workflow class

To search beyond the currently evaluated workflow pool, the optimizer needs a way to identify promising workflows in the implicit class𝒢\\mathcal\{G\}\. As noted above, this search is carried out by the pricing oracle, and it is guided by nonnegative*task prices*wiw\_\{i\}supplied by the optimizer\. These prices come from the current dual solution: a largerwiw\_\{i\}means that an additional correct output on development taskiiwould create greater marginal value in the current LP relaxation\. The prices therefore direct the oracle toward the*residual tasks*, meaning those the current workflow pool still handles poorly and on which an additional correct output would create the most value\.

Given task pricesw=\(w1,…,wn\)w=\(w\_\{1\},\\ldots,w\_\{n\}\)and the cost priceγ\\gamma, define the global workflow\-pricing value

P⁡\(w,γ\)=maxg∈𝒢⁡\{∑i=1nwi​ai​g−γ​cg\}\.P\(w,\\gamma\)=\\max\_\{g\\in\\mathcal\{G\}\}\\left\\\{\\sum\_\{i=1\}^\{n\}w\_\{i\}a\_\{ig\}\-\\gamma c\_\{g\}\\right\\\}\.\(13\)The term∑iwi​ai​g\\sum\_\{i\}w\_\{i\}a\_\{ig\}rewards a workflow for being correct on high\-price tasks, whileγ​cg\\gamma c\_\{g\}penalizes its recurring execution cost\. We refer to the difference

sg​\(w\)=∑i=1nwi​ai​g−γ​cgs\_\{g\}\(w\)=\\sum\_\{i=1\}^\{n\}w\_\{i\}a\_\{ig\}\-\\gamma c\_\{g\}as the*score*of workflowggat pricesww, so thatP⁡\(w,γ\)=maxg∈𝒢⁡sg​\(w\)P\(w,\\gamma\)=\\max\_\{g\\in\\mathcal\{G\}\}s\_\{g\}\(w\)\. Accordingly, the pricing problem searches for a workflow with the greatest price\-weighted correctness net of compute cost\.

Letθ\\thetadenote the current dual price of occupying an execution slot \(i\.e\., the*slot price*\)\. The largest violation among the workflow\-indexed dual constraints is

Φ⁡\(w,θ,γ\)=P⁡\(w,γ\)−θ=maxg∈𝒢⁡\{sg​\(w\)−θ\}\.\\Phi\(w,\\theta,\\gamma\)=P\(w,\\gamma\)\-\\theta=\\max\_\{g\\in\\mathcal\{G\}\}\\left\\\{s\_\{g\}\(w\)\-\\theta\\right\\\}\.\(14\)Hence,Φ⁡\(w,θ,γ\)\>0\\Phi\(w,\\theta,\\gamma\)\>0means that some workflow creates more price\-weighted value than the slot price\. The pricing interface asks the oracle to identify workflows that perform well on tasks with high current prices, while screening out candidates whose recurring execution costs outweigh their potential contribution\. In the customer\-service routing example, high task prices may concentrate on cancellation, refund, or escalation cases for which an additional correct workflow output would be especially valuable\.

Because exact maximization over𝒢\\mathcal\{G\}may itself be difficult, we allow an approximate stochastic pricing procedure\. A\(δ,β\)\(\\delta,\\beta\)*weak global pricing oracle*returns a workflowg~∈𝒢\\widetilde\{g\}\\in\\mathcal\{G\}such that, conditional on the query history, with probability at least1−β1\-\\beta,

sg~​\(w\)≥P⁡\(w,γ\)−δ\.s\_\{\\widetilde\{g\}\}\(w\)\\geq P\(w,\\gamma\)\-\\delta\.\(15\)Given a proposed slot priceθ\\theta, exactly one of two things happens\. If the returned workflow has score strictly greater thanθ\\theta, then the proposed price cannot be right: a workflow in𝒢\\mathcal\{G\}is worth more than the slot it would occupy, even after its recurring execution cost is charged, so the optimizer has found a better workflow than it currently has\. If, however, the returned workflow has score at mostθ\\theta, then, on the event that the guarantee in \([15](https://arxiv.org/html/2609.18126#S3.E15)\) holds, no workflow in the entire class scores more thanδ\\deltaabove the proposed price:

P⁡\(w,γ\)≤θ\+δ\.P\(w,\\gamma\)\\leq\\theta\+\\delta\.Thus a single call to the oracle either produces a workflow that beats the current slot price or certifies, up to toleranceδ\\delta, that no workflow in𝒢\\mathcal\{G\}can do so\. This is precisely the weak\-separation information required by the ellipsoid method, which[Section6](https://arxiv.org/html/2609.18126#S6)uses to optimize over the implicit class𝒢\\mathcal\{G\}without enumerating it\.

Recall that the oracle may be implemented by a teacher model, an automated workflow\-search procedure, an explicit search over a prespecified workflow grammar, or a human engineering team\. The theoretical guarantee developed later requires the oracle to satisfy the global approximation property in \([15](https://arxiv.org/html/2609.18126#S3.E15)\)\. In particular, the mere failure of a search procedure to find a workflow with score aboveθ\\thetadoes not certify that no such workflow exists\.

## 4Selector Strength and the Value of Workflow Variety

This section studies how selector quality limits the value of running multiple workflows\. Even if the firm can generate many diverse candidate workflows, variety is only beneficial when the selector can reliably identify correct answers from the realized candidate set\. The results in this section formalize this limitation: they bound how much any multi\-workflow portfolio can outperform the best single workflow, as a function of selector strength and workflow cost\. These bounds provide pre\-optimization screens that can inform the firm whether expanding the portfolio or improving the selector is the more valuable investment\.

The remainder of this section develops these bounds in three steps\. Section[4\.1](https://arxiv.org/html/2609.18126#S4.SS1)defines the odds\-lift indexΛ\\Lambda, a single scalar that summarizes how much the selector improves the odds of returning a correct output relative to picking a candidate at random\. Section[4\.2](https://arxiv.org/html/2609.18126#S4.SS2)usesΛ\\Lambdato bound the net value of any workflow pool relative to the best available singleton workflow, producing a screening rule that can rule out portfolio expansion before solving any optimization problem\. Section[4\.3](https://arxiv.org/html/2609.18126#S4.SS3)specializes this bound to the Plackett–Luce choice model, in which the odds\-lift index reduces to a single discrimination parameter that can be estimated directly from data\.

### 4\.1Recovery envelopes and selector odds lift

The recovery curves\{ψk\}\\\{\\psi\_\{k\}\\\}introduced in Section[3](https://arxiv.org/html/2609.18126#S3)may vary with candidate\-set sizekkand need not follow any particular random\-utility model\. To obtain a single structural bound that applies across all portfolio sizes, we impose the following assumption\.

\{assumption\}

\[Common fraction\-based recovery envelope\] There exists a functionH:\[0,1\]→\[0,1\]H:\[0,1\]\\to\[0,1\]withH⁡\(0\)=0H\(0\)=0andH⁡\(1\)=1H\(1\)=1such that

ψk\(r\)≤H\(rk\),k=1,…,Kmax,r=0,…,k\.\\psi\_\{k\}\(r\)\\leq H\\left\(\\frac\{r\}\{k\}\\right\),\\qquad k=1,\\ldots,K\_\{\\max\},\\quad r=0,\\ldots,k\.\(16\)
[Section4\.1](https://arxiv.org/html/2609.18126#S4.SS1)states that selector performance can be bounded by a function of the*fraction*of correct candidates alone, rather than the countrrand sizekkseparately\. This reflects the premise that a selector’s task difficulty is governed primarily by how concentrated the correct answer is among the alternatives it is shown, a premise that holds across many selector implementations\. It is worth noting that[Section4\.1](https://arxiv.org/html/2609.18126#S4.SS1)is mild: some envelope always exists\. Specifically, given any finite collection of curves\{ψk\}k=1Kmax\\\{\\psi\_\{k\}\\\}\_\{k=1\}^\{K\_\{\\max\}\}, the choice

H\(p\):=max\{ψk\(r\):1≤k≤Kmax,0≤r≤k,r/k≤p\}H\(p\):=\\max\\\{\\psi\_\{k\}\(r\):1\\leq k\\leq K\_\{\\max\},\\ 0\\leq r\\leq k,\\ r/k\\leq p\\\}\(17\)is a valid envelope, sinceψk​\(0\)=0\\psi\_\{k\}\(0\)=0andψk​\(k\)=1\\psi\_\{k\}\(k\)=1for everykk\.

We now introduce an index that measures how much the selector improves the odds of a correct final output relative to random selection\.

###### Definition 4\.1\(Selector odds\-lift index\)

For a recovery envelopeHH, define

Λ⁡\(H\)=supp∈\(0,1\)H⁡\(p\)/\(1−H⁡\(p\)\)p/\(1−p\),\\Lambda\(H\)=\\sup\_\{p\\in\(0,1\)\}\\frac\{H\(p\)/\(1\-H\(p\)\)\}\{p/\(1\-p\)\},\(18\)with the convention thatΛ⁡\(H\)=\+∞\\Lambda\(H\)=\+\\inftyifH⁡\(p\)=1H\(p\)=1for somep∈\(0,1\)p\\in\(0,1\)\.

The indexΛ⁡\(H\)\\Lambda\(H\)admits a direct interpretation as an odds ratio\. At any correct fractionpp,p/\(1−p\)p/\(1\-p\)is the odds of returning a correct output under uniform random selection, andH⁡\(p\)/\(1−H⁡\(p\)\)H\(p\)/\(1\-H\(p\)\)is the corresponding odds under the selector’s recovery envelope\. Their ratio measures the selector’s odds improvement over random chance at that fraction, andΛ⁡\(H\)\\Lambda\(H\)takes the largest such improvement, i\.e\., the supremum, over all fractionsp∈\(0,1\)p\\in\(0,1\)\.

Three cases illustrate the range ofΛ⁡\(H\)\\Lambda\(H\)\. A selector no better than random chance hasH⁡\(p\)=pH\(p\)=pat every fraction, givingΛ⁡\(H\)=1\\Lambda\(H\)=1\. A selector that systematically improves on random chance, without ever achieving certainty in returning the correct output unless all candidates are correct, hasΛ⁡\(H\)∈\(1,∞\)\\Lambda\(H\)\\in\(1,\\infty\)\. A selector that recovers the correct output with certainty at some interior fractionp∈\(0,1\)p\\in\(0,1\), that is,H⁡\(p\)=1H\(p\)=1whilep<1p<1, has infinite odds\-lift,Λ⁡\(H\)=\+∞\\Lambda\(H\)=\+\\infty, by the convention in Definition[4\.1](https://arxiv.org/html/2609.18126#S4.Thmtheorem1)\. This last case is excluded whenever we assumeΛ⁡\(H\)\\Lambda\(H\)is finite, as we do throughout the results that follow\.

The next lemma converts the boundΛ⁡\(H\)≤Λ\\Lambda\(H\)\\leq\\Lambdainto an explicit closed\-form envelope forHH\.

###### Lemma 4\.2\(Odds\-lift envelope\)

IfΛ⁡\(H\)≤Λ<∞\\Lambda\(H\)\\leq\\Lambda<\\inftyfor someΛ≥1\\Lambda\\geq 1, then

H⁡\(p\)≤hΛ​\(p\):=Λ​p1\+\(Λ−1\)​p,p∈\[0,1\]\.H\(p\)\\leq h\_\{\\Lambda\}\(p\):=\\frac\{\\Lambda p\}\{1\+\(\\Lambda\-1\)p\},\\qquad p\\in\[0,1\]\.\(19\)

The functionhΛh\_\{\\Lambda\}is the tight recovery envelope implied by an odds\-lift boundΛ\\Lambda: no selector with odds lift at mostΛ\\Lambdacan exceedhΛh\_\{\\Lambda\}at any correct fraction\. We call this a Luce\-shaped envelope\. \(The formal Plackett–Luce specialization, for which this same functional form arises exactly, is introduced in[Section4\.3](https://arxiv.org/html/2609.18126#S4.SS3)\.\) The next property, the concavity ofhΛh\_\{\\Lambda\}, is what makes the envelope useful for comparing portfolios of different sizes\.

###### Lemma 4\.3\(Shape of the odds\-lift envelope\)

For everyΛ≥1\\Lambda\\geq 1, the functionhΛh\_\{\\Lambda\}is increasing and concave on\[0,1\]\[0,1\]\.

The key here is that the realized recovery curvesψk\\psi\_\{k\}need not themselves be concave\. Lemmas[4\.2](https://arxiv.org/html/2609.18126#S4.Thmtheorem2)and[4\.3](https://arxiv.org/html/2609.18126#S4.Thmtheorem3)show that a finite odds\-lift bound nonetheless forces every such curve to lie below a common concave, Luce\-shaped, envelope\.

### 4\.2The value of workflow variety

We now boundOPTψ,γ​\(ℳ\)\\mathrm\{OPT\}\_\{\\psi,\\gamma\}\(\\mathcal\{M\}\), the best net value attainable using workflow types from the evaluated poolℳ\\mathcal\{M\}\. For each workflow typeg∈ℳg\\in\\mathcal\{M\}, define its*standalone accuracy*as its average correctness on the development sample:

a¯g=1n​∑i=1nai​g\.\\bar\{a\}\_\{g\}=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}a\_\{ig\}\.\(20\)
Because repeated assignments are allowed, the extremal accuracy and cost quantities for a size\-kkportfolio can be defined directly from the workflow types\. For eachk≤Kmaxk\\leq K\_\{\\max\}, define

Ak=maxg∈ℳ⁡a¯g,Ck=k​ming∈ℳ​cg\.A\_\{k\}=\\max\_\{g\\in\\mathcal\{M\}\}\\bar\{a\}\_\{g\},\\qquad C\_\{k\}=k\\min\_\{g\\in\\mathcal\{M\}\}c\_\{g\}\.\(21\)Indeed, allkkexecution slots may be assigned to the same workflow type\. Thus,AkA\_\{k\}is the largest average standalone accuracy that can be assigned tokkexecution slots, whereasCkC\_\{k\}is the smallest total execution cost of anykkslots\. Consequently, every execution\-count vector𝒎∈ℤ\+ℳ\\bm\{m\}\\in\\mathbb\{Z\}\_\{\+\}^\{\\mathcal\{M\}\}satisfyingk⁡\(𝒎\)=kk\(\\bm\{m\}\)=kobeys

1k​∑g∈ℳa¯g​mg≤AkandC⁡\(𝒎\)≥Ck\.\\frac\{1\}\{k\}\\sum\_\{g\\in\\mathcal\{M\}\}\\bar\{a\}\_\{g\}m\_\{g\}\\leq A\_\{k\}\\qquad\\text\{and\}\\qquad C\(\\bm\{m\}\)\\geq C\_\{k\}\.The two bounds are computed separately and need not be attained by the same workflow type\.

Combining these quantities with the concave odds\-lift envelopehΛh\_\{\\Lambda\}from[Section4\.1](https://arxiv.org/html/2609.18126#S4.SS1)yields an upper bound on the net value of every size\-kkportfolio and, consequently, onOPTψ,γ​\(ℳ\)\\mathrm\{OPT\}\_\{\\psi,\\gamma\}\(\\mathcal\{M\}\)\.

###### Theorem 4\.4\(Selector\-limited value of workflow variety\)

Suppose[Section4\.1](https://arxiv.org/html/2609.18126#S4.SS1)holds andΛ⁡\(H\)≤Λ<∞\\Lambda\(H\)\\leq\\Lambda<\\inftyfor someΛ≥1\\Lambda\\geq 1\. For a finite evaluated workflow poolℳ\\mathcal\{M\}, letV1=max\{0,V\_\{1\}=\\max\\big\\\{0,maxg∈ℳ\{a¯g−γcg\}\}\\max\_\{g\\in\\mathcal\{M\}\}\\\{\\bar\{a\}\_\{g\}\-\\gamma c\_\{g\}\\\}\\big\\\}denote the best net value achievable by a singleton portfolio or the outside option\. Then

V1≤OPTψ,γ​\(ℳ\)≤UΛ,γ​\(ℳ\):=max⁡\{0,max1≤k≤Kmax⁡\[hΛ​\(Ak\)−γ​Ck\]\}\.V\_\{1\}\\leq\\mathrm\{OPT\}\_\{\\psi,\\gamma\}\(\\mathcal\{M\}\)\\leq U\_\{\\Lambda,\\gamma\}\(\\mathcal\{M\}\):=\\max\\left\\\{0,\\max\_\{1\\leq k\\leq K\_\{\\max\}\}\\left\[h\_\{\\Lambda\}\(A\_\{k\}\)\-\\gamma C\_\{k\}\\right\]\\right\\\}\.\(22\)In particular, whenγ=0\\gamma=0, lettinga⋆=maxg⁡a¯ga^\{\\star\}=\\max\_\{g\}\\bar\{a\}\_\{g\}yields

a⋆≤OPTψ,0​\(ℳ\)≤hΛ​\(a⋆\)\.a^\{\\star\}\\leq\\mathrm\{OPT\}\_\{\\psi,0\}\(\\mathcal\{M\}\)\\leq h\_\{\\Lambda\}\(a^\{\\star\}\)\.\(23\)We call the upper boundOPTψ,0​\(ℳ\)≤hΛ​\(a⋆\)\\mathrm\{OPT\}\_\{\\psi,0\}\(\\mathcal\{M\}\)\\leq h\_\{\\Lambda\}\(a^\{\\star\}\)in \([23](https://arxiv.org/html/2609.18126#S4.E23)\) the no\-cost upper bound: it applies when the recurring workflow running costs are ignored, i\.e\., whenγ=0\\gamma=0\. This no\-cost upper bound is tight over the class of selectors with odds lift at mostΛ\\Lambda\.

The lower boundV1≤OPTψ,γ​\(ℳ\)V\_\{1\}\\leq\\mathrm\{OPT\}\_\{\\psi,\\gamma\}\(\\mathcal\{M\}\)immediately holds because both the singleton portfolios and the outside option are feasible\. Thus, the main content of the theorem is really the upper bound, which gives a quick test for whether any genuinely multi\-workflow portfolio can be worth considering\. Specifically, if

V1≥max2≤k≤Kmax⁡\[hΛ​\(Ak\)−γ​Ck\],V\_\{1\}\\geq\\max\_\{2\\leq k\\leq K\_\{\\max\}\}\\left\[h\_\{\\Lambda\}\(A\_\{k\}\)\-\\gamma C\_\{k\}\\right\],\(24\)then no multi\-workflow portfolio in the evaluated pool can outperform the best singleton or the outside option\. \(Note that the maximum in \([24](https://arxiv.org/html/2609.18126#S4.E24)\) starts atk=2k=2becausek=0k=0andk=1k=1are already accounted for byV1V\_\{1\}\.\)

###### Corollary 4\.5\(Selector\-strength and uniform\-cost implications\)

Leta⋆=maxg⁡a¯ga^\{\\star\}=\\max\_\{g\}\\bar\{a\}\_\{g\}be as in[Theorem4\.4](https://arxiv.org/html/2609.18126#S4.Thmtheorem4)\. The following hold:

1. 1\.Ifγ=0\\gamma=0, the maximum gain over the best workflow satisfies OPTψ,0​\(ℳ\)−a⋆≤hΛ​\(a⋆\)−a⋆=\(Λ−1\)​a⋆​\(1−a⋆\)1\+\(Λ−1\)​a⋆≤Λ−1Λ\+1\.\\mathrm\{OPT\}\_\{\\psi,0\}\(\\mathcal\{M\}\)\-a^\{\\star\}\\leq h\_\{\\Lambda\}\(a^\{\\star\}\)\-a^\{\\star\}=\\frac\{\(\\Lambda\-1\)a^\{\\star\}\(1\-a^\{\\star\}\)\}\{1\+\(\\Lambda\-1\)a^\{\\star\}\}\\leq\\frac\{\\sqrt\{\\Lambda\}\-1\}\{\\sqrt\{\\Lambda\}\+1\}\.\(25\)
2. 2\.If every workflow has the same execution costcc, and γ​c≥hΛ​\(a⋆\)−a⋆,\\gamma c\\geq h\_\{\\Lambda\}\(a^\{\\star\}\)\-a^\{\\star\},\(26\)then the best nonempty portfolio is a singleton\. With the outside option, the global optimum is either that singleton or no deployment\.

The corollary highlights the limit of adding more workflows\. WhenΛ=1\\Lambda=1, the envelope ish1​\(p\)=ph\_\{1\}\(p\)=p, so no combination of workflows can outperform the best individual workflow\. AsΛ\\Lambdagrows, stronger selection can support more value from complementary workflows, but this value is capped by the selector\-strength term in \([25](https://arxiv.org/html/2609.18126#S4.E25)\)\. The finite poolℳ\\mathcal\{M\}can be arbitrarily large: in the no\-cost bound, its size affects the cap only through the best singleton accuracya⋆a^\{\\star\}, not directly through\|ℳ\|\|\\mathcal\{M\}\|\. Thus a larger pool cannot by itself overcome a weak selector\.

### 4\.3Plackett–Luce specialization

The general bound in[Section4\.2](https://arxiv.org/html/2609.18126#S4.SS2)does not require a random\-utility model\. We now specialize to the Plackett–Luce model, a standard model for choice and ranking in which each alternative has an attraction weight and is selected with probability proportional to that weight\. Beyond its foundations in choice theory, the model has been used extensively in computer science for Bayesian ranking, label ranking, rank aggregation, and online preference learning\([Luce 1959](https://arxiv.org/html/2609.18126#bib.bib21),[Plackett 1975](https://arxiv.org/html/2609.18126#bib.bib26),[Guiver and Snelson 2009](https://arxiv.org/html/2609.18126#bib.bib50),[Cheng et al\. 2010](https://arxiv.org/html/2609.18126#bib.bib51),[Hajek et al\. 2014](https://arxiv.org/html/2609.18126#bib.bib52),[Szörényi et al\. 2015](https://arxiv.org/html/2609.18126#bib.bib53)\)\. The model is useful here because it gives an interpretable one\-parameter measure of selector strength, closed\-form fixed\-size increments, and a tractable marginal\-value formula for the generation loop\.

Recall from Section[3](https://arxiv.org/html/2609.18126#S3)that, for a nonzero execution\-count vector𝒎\\bm\{m\}, the selectorSel⁡\(X,O𝒎\)\\operatorname\{Sel\}\(X,O\_\{\\bm\{m\}\}\)observes the task, the candidate outputs, and their execution traces and returns one candidate as the final output\. To model how it makes this choice, condition on a task instance and letZj=𝟏\{Aj=Y⋆\}Z\_\{j\}=\\mathbf\{1\}\\\{A\_\{j\}=Y^\{\\star\}\\\}indicate whether candidatejjis correct\. We posit that the selector assigns each candidatejja latent score

Sj=η​Zj\+εj,S\_\{j\}=\\eta Z\_\{j\}\+\\varepsilon\_\{j\},\(27\)and returns the candidate with the highest score\. Hereη≥0\\eta\\geq 0is the selector’s*discrimination advantage*: it measures how much more attractive a correct candidate is to the selector than an incorrect one, on average\. Whenη=0\\eta=0, correct and incorrect candidates are equally attractive and the selector behaves like a coin flip; largerη\\etameans the selector systematically favors correctness\. The termsεj\\varepsilon\_\{j\}are i\.i\.d\. standard Gumbel shocks representing idiosyncratic factors, unrelated to correctness, that also influence the selector’s choice\. The selector never observesZjZ\_\{j\}directly; \([27](https://arxiv.org/html/2609.18126#S4.E27)\) is a statistical description of the relationship between true correctness and the selector’s realized preference, not a claim about the selector’s internal reasoning\.

Letλ=exp⁡\(η\)≥1\\lambda=\\exp\(\\eta\)\\geq 1\. By the Gumbel\-max identity, choosing the candidate with the highest scoreSjS\_\{j\}is equivalent to choosing candidatejjwith probability proportional toexp⁡\(η​Zj\)\\exp\(\\eta Z\_\{j\}\): each correct candidate receives attraction weightλ=exp⁡\(η\)\\lambda=\\exp\(\\eta\)and each incorrect candidate receives weight1=exp⁡\(0\)1=\\exp\(0\), and the selector returns a given candidate with probability equal to its weight divided by the sum of all weights\. Summing this probability over all correct candidates, ifrrofkkcandidates are correct, the probability that the selector returns a correct answer is

ψk,λPL\(r\)=λ​rλ​r\+k−r=λ​rk\+\(λ−1\)​r,r=0,1,…,k\.\\psi^\{\\mathrm\{PL\}\}\_\{k,\\lambda\}\(r\)=\\frac\{\\lambda r\}\{\\lambda r\+k\-r\}=\\frac\{\\lambda r\}\{k\+\(\\lambda\-1\)r\},\\qquad r=0,1,\\ldots,k\.\(28\)This is the Plackett–Luce recovery curve\. Define

hλ​\(p\)=λ​p1\+\(λ−1\)​p,p∈\[0,1\]\.h\_\{\\lambda\}\(p\)=\\frac\{\\lambda p\}\{1\+\(\\lambda\-1\)p\},\\qquad p\\in\[0,1\]\.\(29\)Dividing numerator and denominator of \([28](https://arxiv.org/html/2609.18126#S4.E28)\) bykkshowsψk,λPL​\(r\)=hλ​\(r/k\)\\psi^\{\\mathrm\{PL\}\}\_\{k,\\lambda\}\(r\)=h\_\{\\lambda\}\(r/k\): the Plackett–Luce recovery probability depends onrrandkkonly through the correct fractionp=r/kp=r/k, exactly the fraction\-based structure assumed in Assumption[4\.1](https://arxiv.org/html/2609.18126#S4.SS1)\. Moreover,hλh\_\{\\lambda\}has the same functional form as the envelopehΛh\_\{\\Lambda\}from Lemma[4\.2](https://arxiv.org/html/2609.18126#S4.Thmtheorem2), and a direct calculation confirms

hλ​\(p\)/\(1−hλ​\(p\)\)p/\(1−p\)=λ,p∈\(0,1\)\.\\frac\{h\_\{\\lambda\}\(p\)/\(1\-h\_\{\\lambda\}\(p\)\)\}\{p/\(1\-p\)\}=\\lambda,\\qquad p\\in\(0,1\)\.\(30\)Thus the Plackett–Luce discrimination parameterλ\\lambdais exactly its corresponding selector odds\-lift index: a Plackett–Luce selector with strengthλ\\lambdasatisfiesΛ⁡\(H\)=λ\\Lambda\(H\)=\\lambdawithH=hλH=h\_\{\\lambda\}, so the general bounds of Section[4\.2](https://arxiv.org/html/2609.18126#S4.SS2)apply to it with equality atΛ=λ\\Lambda=\\lambda, not merely as an upper bound\.

The continuous concavity ofhλh\_\{\\lambda\}follows immediately from[Lemma4\.3](https://arxiv.org/html/2609.18126#S4.Thmtheorem3)by settingΛ=λ\\Lambda=\\lambda\. The next lemma records the additional result needed for the inner optimization problem in[Section5](https://arxiv.org/html/2609.18126#S5)\.

###### Lemma 4\.6\(Fixed\-size Plackett–Luce increments\)

For everyk≥1k\\geq 1andλ≥1\\lambda\\geq 1,ψk,λPL\\psi^\{\\mathrm\{PL\}\}\_\{k,\\lambda\}is nondecreasing and discrete concave\. Its increments are

dℓ,kPL=ψk,λPL\(ℓ\)−ψk,λPL\(ℓ−1\)=λ​k\(k\+\(λ−1\)​ℓ\)​\(k\+\(λ−1\)​\(ℓ−1\)\),ℓ=1,…,k\.d^\{\\mathrm\{PL\}\}\_\{\\ell,k\}=\\psi^\{\\mathrm\{PL\}\}\_\{k,\\lambda\}\(\\ell\)\-\\psi^\{\\mathrm\{PL\}\}\_\{k,\\lambda\}\(\\ell\-1\)=\\frac\{\\lambda k\}\{\\left\(k\+\(\\lambda\-1\)\\ell\\right\)\\left\(k\+\(\\lambda\-1\)\(\\ell\-1\)\\right\)\},\\quad\\ell=1,\\ldots,k\.\(31\)

The estimation ofλ\\lambdafrom validation data is described in[Section8\.1](https://arxiv.org/html/2609.18126#S8.SS1)\. Onceλ\\lambdais computed, the general portfolio objective from Section[3](https://arxiv.org/html/2609.18126#S3)has a closed\-form expression\. For every nonzero execution\-count vector𝒎\\bm\{m\}, define

Jλ​\(𝒎\)=1n​∑i=1nhλ​\(ri​\(𝒎\)k⁡\(𝒎\)\),Πλ,γ​\(𝒎\)=Jλ​\(𝒎\)−γ​C​\(𝒎\),J\_\{\\lambda\}\(\\bm\{m\}\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}h\_\{\\lambda\}\\left\(\\frac\{r\_\{i\}\(\\bm\{m\}\)\}\{k\(\\bm\{m\}\)\}\\right\),\\qquad\\Pi\_\{\\lambda,\\gamma\}\(\\bm\{m\}\)=J\_\{\\lambda\}\(\\bm\{m\}\)\-\\gamma C\(\\bm\{m\}\),\(32\)with both functions set to zero at𝒎=𝟎\\bm\{m\}=\\bm\{0\}\. HereJλ​\(𝒎\)J\_\{\\lambda\}\(\\bm\{m\}\)plays the role ofJψ​\(𝒎\)J\_\{\\psi\}\(\\bm\{m\}\)from Section[3](https://arxiv.org/html/2609.18126#S3), but with the recovery curveψk⁡\(𝒎\)\\psi\_\{k\(\\bm\{m\}\)\}replaced by its Plackett–Luce formhλh\_\{\\lambda\}evaluated at the correct fractionri​\(𝒎\)/k​\(𝒎\)r\_\{i\}\(\\bm\{m\}\)/k\(\\bm\{m\}\)\. This is the objective used in the optimization algorithms of Sections[5](https://arxiv.org/html/2609.18126#S5)and[6](https://arxiv.org/html/2609.18126#S6)\.

To study the structure ofJλJ\_\{\\lambda\}, we represent each execution slot as a distinct labeled copy of its workflow type\. Thus, an execution\-count vector𝒎\\bm\{m\}can be viewed as an ordinary subset of the ground set𝒢×\{1,…,Kmax\}\\mathcal\{G\}\\times\\\{1,\\ldots,K\_\{\\max\}\\\}containingmgm\_\{g\}labeled copies of each workflow typegg, with the second coordinate serving only to distinguish repeated executions\. A natural question is whetherJλJ\_\{\\lambda\}, viewed as a set function under this representation, has the familiar properties of a coverage objective commonly found in submodular optimization\.

The next proposition shows that neither property holds in general\. The endogenous\-size objective can decrease when an execution is added and can exhibit both increasing and decreasing marginal returns\. Standard monotone\-submodular optimization tools therefore do not apply directly, which motivates our development in[Section5](https://arxiv.org/html/2609.18126#S5)\.

###### Proposition 4\.7\(More workflows need not be better\)

For everyλ≥1\\lambda\\geq 1,JλJ\_\{\\lambda\}is generally nonmonotone with respect to adding executions: there exist a feasible𝐦\\bm\{m\}and workflow typeggsuch that selector\-aware accuracy decreasesJλ​\(𝐦\+𝐞g\)<Jλ​\(𝐦\)\.J\_\{\\lambda\}\(\\bm\{m\}\+\\bm\{e\}\_\{g\}\)<J\_\{\\lambda\}\(\\bm\{m\}\)\.Moreover,JλJ\_\{\\lambda\}is neither submodular nor supermodular\. Subtracting the modular workflow\-cost termγ​C​\(𝐦\)\\gamma C\(\\bm\{m\}\)preserves these failures, soΠλ,γ\\Pi\_\{\\lambda,\\gamma\}has the same general structure\.

The reason is a selection externality: increasing one componentmgm\_\{g\}by one changes not only the number of correct candidatesri​\(𝒎\)r\_\{i\}\(\\bm\{m\}\)but also the total number of candidatesk⁡\(𝒎\)k\(\\bm\{m\}\)faced by the selector\. If the added execution is incorrect on taskii, thenri​\(𝒎\)r\_\{i\}\(\\bm\{m\}\)is unchanged whilek⁡\(𝒎\)k\(\\bm\{m\}\)increases, this lowers the correct fractionri​\(𝒎\)/k​\(𝒎\)r\_\{i\}\(\\bm\{m\}\)/k\(\\bm\{m\}\), and sincehλh\_\{\\lambda\}is increasing, strictly lowers recovery probability on that task; this is whyJλJ\_\{\\lambda\}can decrease when a workflow is added, ruling out monotonicity\. Conversely, becausehλh\_\{\\lambda\}is concave, the marginal gain from adding a correct workflow depends on the composition of the existing portfolio\. A correct workflow can be more valuable after an incorrect workflow has entered the portfolio and diluted the correct fraction, producing increasing marginal returns and violating submodularity\. Supermodularity would require the opposite inequality: the marginal gain from adding a workflow must weakly increase as the base portfolio becomes larger\. This property also fails111For example, letw1w\_\{1\}andw2w\_\{2\}be incorrect workflows and letc1c\_\{1\}be correct\. Addingc1c\_\{1\}to\{w1\}\\\{w\_\{1\}\\\}produces the marginal gainhλ​\(1/2\)h\_\{\\lambda\}\(1/2\), whereas adding it to the larger set\{w1,w2\}\\\{w\_\{1\},w\_\{2\}\\\}produces onlyhλ​\(1/3\)h\_\{\\lambda\}\(1/3\)\. Sincehλh\_\{\\lambda\}is increasing,hλ​\(1/2\)\>hλ​\(1/3\)h\_\{\\lambda\}\(1/2\)\>h\_\{\\lambda\}\(1/3\), which violates increasing marginal returns\. Thus some configurations exhibit increasing marginal returns and others exhibit decreasing marginal returns, soJλJ\_\{\\lambda\}is neither submodular nor supermodular\.\. In Appendix[D](https://arxiv.org/html/2609.18126#A4), we construct explicit examples exhibiting both failures\.

## 5The Inner Optimization Problem

This section studies the portfolio problem for a given nonempty evaluated workflow\-type poolℳ⊆𝒢\\mathcal\{M\}\\subseteq\\mathcal\{G\}, allowing repeated execution of the same workflow type\. The resulting finite\-pool formulations serve two purposes\. They produce exact and approximate deployment plans for the current workflow pool, and they reveal the dual\-price structure used to search over the implicit workflow class in[Section6](https://arxiv.org/html/2609.18126#S6)\.

The restricted problem decomposes exactly by run size:

OPTψ,γ​\(ℳ\)=max⁡\{0,max1≤k≤Kmax⁡OPTk​\(ℳ\)\},OPTk​\(ℳ\)=max𝒎∈ℤ\+ℳk⁡\(𝒎\)=k⁡Πψ,γ​\(𝒎\)\.\\mathrm\{OPT\}\_\{\\psi,\\gamma\}\(\\mathcal\{M\}\)=\\max\\left\\\{0,\\max\_\{1\\leq k\\leq K\_\{\\max\}\}\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\)\\right\\\},\\qquad\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\)=\\max\_\{\\begin\{subarray\}\{c\}\\bm\{m\}\\in\\mathbb\{Z\}\_\{\+\}^\{\\mathcal\{M\}\}\\\\ k\(\\bm\{m\}\)=k\\end\{subarray\}\}\\Pi\_\{\\psi,\\gamma\}\(\\bm\{m\}\)\.\(33\)This decomposition isolates the source of difficulty created by endogenous run size\. As shown in[Proposition4\.7](https://arxiv.org/html/2609.18126#S4.Thmtheorem7), adding an execution changes both the number of correct candidates and the total number of alternatives faced by the selector, so the global objective need not be monotone or submodular\. Within a fixed\-size problem, however, every feasible execution\-count vector satisfiesk⁡\(𝒎\)=kk\(\\bm\{m\}\)=k\. The selector therefore faces the same number of candidates across all portfolios in that slice, and portfolio composition affects recovery only through the correct countsri​\(𝒎\)r\_\{i\}\(\\bm\{m\}\)\. Under[Section5\.2](https://arxiv.org/html/2609.18126#S5.SS2)below, each fixed\-size accuracy objective is a*concave\-coverage*function in the sense of[Barman et al\. \(2021\)](https://arxiv.org/html/2609.18126#bib.bib7)\. We formulate the fixed\-size problem exactly as an integer program, derive an LP relaxation and its dual prices, and develop a randomized\-rounding procedure with an explicit performance certificate\([Ageev and Sviridenko 2004](https://arxiv.org/html/2609.18126#bib.bib35),[Chekuri et al\. 2010](https://arxiv.org/html/2609.18126#bib.bib36),[Calinescu et al\. 2011](https://arxiv.org/html/2609.18126#bib.bib9)\)\.

The remainder of this section proceeds in five steps\.[Section5\.1](https://arxiv.org/html/2609.18126#S5.SS1)shows that a sparse, geometrically spaced cardinality grid certifies a near\-optimal run size without solving the fixed\-size problem for everykk\.[Section5\.2](https://arxiv.org/html/2609.18126#S5.SS2)then fixes a run size and imposes a concavity assumption on eachψk\\psi\_\{k\}, introducing a curvature parameter that will later control the tightness of the rounding guarantee\.[Section5\.3](https://arxiv.org/html/2609.18126#S5.SS3)formulates an exact integer program for each fixed size and shows its optimal value coincides withOPTk​\(ℳ\)\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\)\.[Section5\.4](https://arxiv.org/html/2609.18126#S5.SS4)relaxes this integer program to an LP whose dual prices identify the residual tasks most worth targeting with new workflows\.[Section5\.5](https://arxiv.org/html/2609.18126#S5.SS5)rounds the LP solution into a feasible portfolio and bounds how far its value can fall short of the true fixed\-size optimum, yielding a certificate that holds across the entire cardinality grid\.

### 5\.1Sparse geometric grids for the optimal run size

[Proposition4\.7](https://arxiv.org/html/2609.18126#S4.Thmtheorem7)rules out simply deploying the largest feasible portfolio because more workflows need not increase net value\. Thus, finding the best run size requires comparing acrossk=1,…,Kmaxk=1,\\ldots,K\_\{\\max\}, not assumingk=Kmaxk=K\_\{\\max\}is optimal\. Doing this by solvingOPTk​\(ℳ\)\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\)for every candidatekkcan be expensive, especially since the inner problem must be resolved at every generation round as the pool grows\. This subsection shows that when a single concave function generates every recovery curve, checking only a sparse, geometrically spaced set of sizes is enough to certify a near\-optimal size, without solving the fixed\-size problem at everykk\. Formally, this is the case when the fraction\-based envelope of[Section4\.1](https://arxiv.org/html/2609.18126#S4.SS1)holds with equality rather than merely as an upper bound\. Recall that[Section4\.3](https://arxiv.org/html/2609.18126#S4.SS3)showed that Plackett–Luce recovery has exactly this form\.

###### Theorem 5\.1\(Geometric cardinality approximation\)

Suppose the recovery curves are generated by a common increasing concave functionh:\[0,1\]→\[0,1\]h:\[0,1\]\\to\[0,1\]withh⁡\(0\)=0h\(0\)=0, so thatψk​\(r\)=h⁡\(r/k\)\\psi\_\{k\}\(r\)=h\(r/k\),r=0,…,kr=0,\\ldots,k\. LetOPTk\+​\(ℳ\)=max⁡\{0,OPTk​\(ℳ\)\}\\mathrm\{OPT\}\_\{k\}^\{\+\}\(\\mathcal\{M\}\)=\\max\\\{0,\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\)\\\}\. Then, for every1≤b≤k≤Kmax1\\leq b\\leq k\\leq K\_\{\\max\},

OPTb\+​\(ℳ\)≥bk​OPTk\+​\(ℳ\)\.\\mathrm\{OPT\}\_\{b\}^\{\+\}\(\\mathcal\{M\}\)\\geq\\frac\{b\}\{k\}\\mathrm\{OPT\}\_\{k\}^\{\+\}\(\\mathcal\{M\}\)\.\(34\)Consequently, if a cardinality grid𝒦⊆\{1,…,Kmax\}\\mathcal\{K\}\\subseteq\\\{1,\\ldots,K\_\{\\max\}\\\}has coverage ratioϱ≥1\\varrho\\geq 1, meaning that for everyk≤Kmaxk\\leq K\_\{\\max\}there existsb∈𝒦b\\in\\mathcal\{K\}withb≤k≤ϱ​bb\\leq k\\leq\\varrho b, then

maxb∈𝒦⁡OPTb\+​\(ℳ\)≥1ϱ​OPTψ,γ​\(ℳ\)\.\\max\_\{b\\in\\mathcal\{K\}\}\\mathrm\{OPT\}\_\{b\}^\{\+\}\(\\mathcal\{M\}\)\\geq\\frac\{1\}\{\\varrho\}\\mathrm\{OPT\}\_\{\\psi,\\gamma\}\(\\mathcal\{M\}\)\.\(35\)In particular, the dyadic grid gives a1/21/2\-approximation, and a grid with ratioϱ=1/\(1−ηgrid\)\\varrho=1/\(1\-\\eta\_\{\\mathrm\{grid\}\}\)gives a\(1−ηgrid\)\(1\-\\eta\_\{\\mathrm\{grid\}\}\)\-approximation usingO⁡\(ηgrid−1​log⁡Kmax\)O\(\\eta\_\{\\mathrm\{grid\}\}^\{\-1\}\\log K\_\{\\max\}\)fixed\-size solves\.

Inequality \([34](https://arxiv.org/html/2609.18126#S5.E34)\) says that a portfolio of sizeb≤kb\\leq kcan always secure at least ab/kb/kfraction of the value achievable at sizekk: shrinking the run size costs at most proportionally, never more, because concavity ofhhrules out returns to scale that vanish faster than linearly\. This is what makes a coarse grid safe: whichever true optimal sizek⋆k^\{\\star\}exhaustive search would find, some grid pointbbwithin a factorϱ\\varrhoofk⋆k^\{\\star\}still retains at least a1/ϱ1/\\varrhoshare of the optimal value, which is exactly \([35](https://arxiv.org/html/2609.18126#S5.E35)\)\.

Two concrete grids turn this general guarantee into the specific ratios\. The*dyadic grid*

𝒦dyad=\{2j:j=0,1,…,⌈log2Kmax⌉\}∩\{1,…,Kmax\}\\mathcal\{K\}\_\{\\mathrm\{dyad\}\}=\\\{2^\{j\}:j=0,1,\\ldots,\\lceil\\log\_\{2\}K\_\{\\max\}\\rceil\\\}\\cap\\\{1,\\ldots,K\_\{\\max\}\\\}has coverage ratioϱ=2\\varrho=2andO⁡\(log⁡Kmax\)O\(\\log K\_\{\\max\}\)points, giving the1/21/2\-approximation\. More generally, for anyϱ\>1\\varrho\>1, the*ratio\-ϱ\\varrhogrid*

𝒦ϱ=\{⌈ϱj⌉:j=0,1,…,⌈logϱKmax⌉\}∩\{1,…,Kmax\}\\mathcal\{K\}\_\{\\varrho\}=\\\{\\lceil\\varrho^\{j\}\\rceil:j=0,1,\\ldots,\\lceil\\log\_\{\\varrho\}K\_\{\\max\}\\rceil\\\}\\cap\\\{1,\\ldots,K\_\{\\max\}\\\}has coverage ratioϱ\\varrhoandO⁡\(log⁡Kmax\)O\(\\log K\_\{\\max\}\)points\. Takingϱ=1/\(1−ηgrid\)\\varrho=1/\(1\-\\eta\_\{\\mathrm\{grid\}\}\)gives the\(1−ηgrid\)\(1\-\\eta\_\{\\mathrm\{grid\}\}\)\-approximation usingO⁡\(ηgrid−1​log⁡Kmax\)O\(\\eta\_\{\\mathrm\{grid\}\}^\{\-1\}\\log K\_\{\\max\}\)fixed\-size solves\. The full grid𝒦=\{1,…,Kmax\}\\mathcal\{K\}=\\\{1,\\ldots,K\_\{\\max\}\\\}corresponds to the caseϱ=1\\varrho=1, where every size is its own anchor and the guarantee in \([35](https://arxiv.org/html/2609.18126#S5.E35)\) becomes exact\. Small pools can afford this exhaustive search\. Large generation loops resolve the inner problem at every round, so a sparse grid, dyadic or otherwise, cuts the number of LP solves and oracle calls\.

### 5\.2Concave recovery and selector curvature at fixed size

We now fixkkand build towards solvingOPTk​\(ℳ\)\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\): exactly, through the integer program in[Section5\.3](https://arxiv.org/html/2609.18126#S5.SS3), and at scale, through the rounding certificate in[Section5\.5](https://arxiv.org/html/2609.18126#S5.SS5)\. This subsection lays the groundwork with the following condition on the fixed\-size recovery curveψk\\psi\_\{k\}\.

\{assumption\}

\[Concave recovery at each fixed size\] For everyk=1,…,Kmaxk=1,\\ldots,K\_\{\\max\},ψk\\psi\_\{k\}is nondecreasing and discrete concave:

d1,k≥d2,k≥⋯≥dk,k≥0,dℓ,k=ψk​\(ℓ\)−ψk​\(ℓ−1\)\.d\_\{1,k\}\\geq d\_\{2,k\}\\geq\\cdots\\geq d\_\{k,k\}\\geq 0,\\qquad d\_\{\\ell,k\}=\\psi\_\{k\}\(\\ell\)\-\\psi\_\{k\}\(\\ell\-1\)\.\(36\)Whend1,k\>0d\_\{1,k\}\>0, define

cψ,k=1−dk,kd1,k\.c\_\{\\psi,k\}=1\-\\frac\{d\_\{k,k\}\}\{d\_\{1,k\}\}\.\(37\)
The curvaturecψ,k∈\[0,1\]c\_\{\\psi,k\}\\in\[0,1\]measures how much recovery flattens as correct candidates accumulate: it is zero when marginal recovery gains are constant and approaches one when the last correct candidate contributes much less than the first\. The requirements used below are nested\. The exact integer\-program result in[Proposition5\.2](https://arxiv.org/html/2609.18126#S5.Thmtheorem2)requires only thatψk\\psi\_\{k\}be nondecreasing, equivalently, thatdℓ,k≥0d\_\{\\ell,k\}\\geq 0for everyℓ\\ell\. This is weaker than[Section5\.2](https://arxiv.org/html/2609.18126#S5.SS2), which additionally requires the diminishing\-increment inequalitiesd1,k≥⋯≥dk,kd\_\{1,k\}\\geq\\cdots\\geq d\_\{k,k\}\. The stronger condition is used only to identify the fixed\-size objective as concave coverage and to establish the LP\-rounding certificate\. Diminishing increments are natural when the first correct candidate provides most of the selector’s recoverable signal and additional correct candidates are partly redundant\. For example, once the candidate set already contains a clearly correct answer, a second or third correct answer may provide less additional help to the selector than the first\. Plackett–Luce recovery satisfies the condition exactly by[Lemma4\.6](https://arxiv.org/html/2609.18126#S4.Thmtheorem6)\.

For the rounding analysis in[Section5\.5](https://arxiv.org/html/2609.18126#S5.SS5), it is useful to extend the fixed\-size curve beyond the deployed ranger=0,…,kr=0,\\ldots,k:

ψ~k​\(r\)=dk,k​r\+∑t=1k−1\(dt,k−dt\+1,k\)​min⁡\{r,t\},r≥0\.\\widetilde\{\\psi\}\_\{k\}\(r\)=d\_\{k,k\}r\+\\sum\_\{t=1\}^\{k\-1\}\\left\(d\_\{t,k\}\-d\_\{t\+1,k\}\\right\)\\min\\\{r,t\\\},\\qquad r\\geq 0\.\(38\)This continuation expresses the fixed\-size accuracy function as a modular linear term plus a nonnegative combination of threshold coverage functionsmin⁡\{r,t\}\\min\\\{r,t\\\}, the form used in the rounding proof, and it agrees withψk\\psi\_\{k\}at the deployed integer pointsr=0,…,kr=0,\\ldots,k, so it changes nothing about the value of any actual size\-kkportfolio\.

### 5\.3Exact integer programs

For a fixed run sizekk, the selector\-aware contribution of taskiiunder an execution\-count vector𝒎\\bm\{m\}isψk​\(ri​\(𝒎\)\)\\psi\_\{k\}\(r\_\{i\}\(\\bm\{m\}\)\), which is generally nonlinear in the execution counts\. We linearize this term using binary tier indicators\.

Letzg∈ℤ\+z\_\{g\}\\in\\mathbb\{Z\}\_\{\+\}denote the number of execution slots assigned to workflow typegg\. For each taskiiand tierℓ=1,…,k\\ell=1,\\ldots,k, letyi​ℓ∈\{0,1\}y\_\{i\\ell\}\\in\\\{0,1\\\}indicate whether the execution plan produces at leastℓ\\ellcorrect candidates on taskii\. Because

ψk​\(r\)=∑ℓ=1rdℓ,k,dℓ,k=ψk​\(ℓ\)−ψk​\(ℓ−1\),\\psi\_\{k\}\(r\)=\\sum\_\{\\ell=1\}^\{r\}d\_\{\\ell,k\},\\qquad d\_\{\\ell,k\}=\\psi\_\{k\}\(\\ell\)\-\\psi\_\{k\}\(\\ell\-1\),an integral allocationzzinduces the execution\-count vector withmg=zgm\_\{g\}=z\_\{g\}\. Weightingyi​ℓy\_\{i\\ell\}bydℓ,kd\_\{\\ell,k\}therefore reproducesψk​\(ri​\(𝒎\)\)\\psi\_\{k\}\(r\_\{i\}\(\\bm\{m\}\)\)exactly\. The formulation below enforces that the active tiers form a prefix and that their total number cannot exceed the number of correct candidate executions produced on the task\. For eachkk, solve

Ik​\(ℳ\)=maxy,z\\displaystyle I\_\{k\}\(\\mathcal\{M\}\)=\\max\_\{y,z\}\\quad1n​∑i=1n∑ℓ=1kdℓ,k​yi​ℓ−γ​∑g∈ℳcg​zg\\displaystyle\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\sum\_\{\\ell=1\}^\{k\}d\_\{\\ell,k\}y\_\{i\\ell\}\-\\gamma\\sum\_\{g\\in\\mathcal\{M\}\}c\_\{g\}z\_\{g\}\(39\)s\.t\.∑ℓ=1kyi​ℓ≤∑g∈ℳai​g​zg,\\displaystyle\\sum\_\{\\ell=1\}^\{k\}y\_\{i\\ell\}\\leq\\sum\_\{g\\in\\mathcal\{M\}\}a\_\{ig\}z\_\{g\},i=1,…,n,\\displaystyle i=1,\\ldots,n,yi​ℓ≤yi,ℓ−1,\\displaystyle y\_\{i\\ell\}\\leq y\_\{i,\\ell\-1\},i=1,…,n,ℓ=2,…,k,\\displaystyle i=1,\\ldots,n,\\quad\\ell=2,\\ldots,k,∑g∈ℳzg=k,\\displaystyle\\sum\_\{g\\in\\mathcal\{M\}\}z\_\{g\}=k,zg∈ℤ\+,\\displaystyle z\_\{g\}\\in\\mathbb\{Z\}\_\{\+\},g∈ℳ,\\displaystyle g\\in\\mathcal\{M\},yi​ℓ∈\{0,1\},\\displaystyle y\_\{i\\ell\}\\in\\\{0,1\\\},i=1,…,n,ℓ=1,…,k\.\\displaystyle i=1,\\ldots,n,\\quad\\ell=1,\\ldots,k\.
Solving \([39](https://arxiv.org/html/2609.18126#S5.E39)\) is not an approximation\. The next result shows its optimal value coincides exactly with the fixed\-size optimumOPTk​\(ℳ\)\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\), and it does so under only a nondecreasing recovery curve, without invoking the concavity assumed in[Section5\.2](https://arxiv.org/html/2609.18126#S5.SS2)\.

###### Proposition 5\.2\(Exactness of the fixed\-size inner problem\)

Ifℳ\\mathcal\{M\}is nonempty andψk\\psi\_\{k\}is nondecreasing, thenIk​\(ℳ\)=OPTk​\(ℳ\)I\_\{k\}\(\\mathcal\{M\}\)=\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\)\. Consequently,

OPTψ,γ​\(ℳ\)=max⁡\{0,max1≤k≤Kmax⁡Ik​\(ℳ\)\}\.\\mathrm\{OPT\}\_\{\\psi,\\gamma\}\(\\mathcal\{M\}\)=\\max\\left\\\{0,\\max\_\{1\\leq k\\leq K\_\{\\max\}\}I\_\{k\}\(\\mathcal\{M\}\)\\right\\\}\.\(40\)

Equation \([40](https://arxiv.org/html/2609.18126#S5.E40)\) is the size decomposition \([33](https://arxiv.org/html/2609.18126#S5.E33)\) with eachOPTk​\(ℳ\)\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\)replaced by its exact integer program valueIk​\(ℳ\)I\_\{k\}\(\\mathcal\{M\}\), so endogenous run size does not require a monolithic nonlinear formulation: exact enumeration solves \([39](https://arxiv.org/html/2609.18126#S5.E39)\) for everykk, while[Theorem5\.1](https://arxiv.org/html/2609.18126#S5.Thmtheorem1)permits a sparse grid instead when a controlled approximation is sufficient\.

### 5\.4LP relaxation, dual prices, and reduced costs

The exact integer program in[Section5\.3](https://arxiv.org/html/2609.18126#S5.SS3)solves the inner problem for the current pool, but it says nothing about which workflows outside the pool are worth generating next\. Extracting that information requires relaxing integrality and reading off dual prices\.222[Chen and Chua \(2024\)](https://arxiv.org/html/2609.18126#bib.bib70)likewise use shadow prices to coordinate primal resource allocations, although their focus is joint differential privacy rather than search over an implicit workflow class\.

Relaxing integrality in \([39](https://arxiv.org/html/2609.18126#S5.E39)\) gives

Lk​\(ℳ\)=maxy,z\\displaystyle L\_\{k\}\(\\mathcal\{M\}\)=\\max\_\{y,z\}\\quad1n​∑i=1n∑ℓ=1kdℓ,k​yi​ℓ−γ​∑g∈ℳcg​zg\\displaystyle\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\sum\_\{\\ell=1\}^\{k\}d\_\{\\ell,k\}y\_\{i\\ell\}\-\\gamma\\sum\_\{g\\in\\mathcal\{M\}\}c\_\{g\}z\_\{g\}\(41\)s\.t\.∑ℓ=1kyi​ℓ≤∑g∈ℳai​g​zg,\\displaystyle\\sum\_\{\\ell=1\}^\{k\}y\_\{i\\ell\}\\leq\\sum\_\{g\\in\\mathcal\{M\}\}a\_\{ig\}z\_\{g\},i=1,…,n,\\displaystyle i=1,\\ldots,n,yi​ℓ≤yi,ℓ−1,\\displaystyle y\_\{i\\ell\}\\leq y\_\{i,\\ell\-1\},i=1,…,n,ℓ=2,…,k,\\displaystyle i=1,\\ldots,n,\\quad\\ell=2,\\ldots,k,∑g∈ℳzg=k,\\displaystyle\\sum\_\{g\\in\\mathcal\{M\}\}z\_\{g\}=k,zg≥0,\\displaystyle z\_\{g\}\\geq 0,g∈ℳ,\\displaystyle g\\in\\mathcal\{M\},0≤yi​ℓ≤1,\\displaystyle 0\\leq y\_\{i\\ell\}\\leq 1,i=1,…,n,ℓ=1,…,k\.\\displaystyle i=1,\\ldots,n,\\quad\\ell=1,\\ldots,k\.For fixedzz, letxi​\(z\)=∑gai​g​zgx\_\{i\}\(z\)=\\sum\_\{g\}a\_\{ig\}z\_\{g\}, now possibly fractional\. Under[Section5\.2](https://arxiv.org/html/2609.18126#S5.SS2), the marginal weightsdℓ,kd\_\{\\ell,k\}are already sorted in decreasing order, so filling theyyslots in index order up toxi​\(z\)x\_\{i\}\(z\)is optimal; the accuracy component then equalsψ~k​\(xi​\(z\)\)\\widetilde\{\\psi\}\_\{k\}\(x\_\{i\}\(z\)\), the same analytical continuation defined in \([38](https://arxiv.org/html/2609.18126#S5.E38)\), now evaluated at a possibly fractional argument rather than an integer one\.

Letμi≥0\\mu\_\{i\}\\geq 0denote the dual variable associated with the task\-iicoverage constraint, letνi​ℓ≥0\\nu\_\{i\\ell\}\\geq 0correspond to the prefix constraint for tierℓ\\ell, letσi​ℓ≥0\\sigma\_\{i\\ell\}\\geq 0correspond to the upper boundyi​ℓ≤1y\_\{i\\ell\}\\leq 1, and letθ\\thetabe the unrestricted dual variable associated with the fixed\-cardinality constraint∑gzg=k\\sum\_\{g\}z\_\{g\}=k\. Because repeated execution is allowed, the primal variableszgz\_\{g\}have no upper bounds; hence, the dual contains no workflow\-specific variables corresponding to constraints of the formzg≤1z\_\{g\}\\leq 1\. Using the boundary conventionνi​1=νi,k\+1=0\\nu\_\{i1\}=\\nu\_\{i,k\+1\}=0, a dual formulation of \([41](https://arxiv.org/html/2609.18126#S5.E41)\) is

Dk​\(ℳ\)=minμ,ν,θ,σ\\displaystyle D\_\{k\}\(\\mathcal\{M\}\)=\\min\_\{\\mu,\\nu,\\theta,\\sigma\}\\quadk​θ\+∑i=1n∑ℓ=1kσi​ℓ\\displaystyle k\\theta\+\\sum\_\{i=1\}^\{n\}\\sum\_\{\\ell=1\}^\{k\}\\sigma\_\{i\\ell\}\(42\)s\.t\.μi\+νi​ℓ−νi,ℓ\+1\+σi​ℓ≥dℓ,kn,\\displaystyle\\mu\_\{i\}\+\\nu\_\{i\\ell\}\-\\nu\_\{i,\\ell\+1\}\+\\sigma\_\{i\\ell\}\\geq\\frac\{d\_\{\\ell,k\}\}\{n\},i=1,…,n,ℓ=1,…,k,\\displaystyle i=1,\\ldots,n,\\quad\\ell=1,\\ldots,k,∑i=1nai​g​μi−γ​cg≤θ,\\displaystyle\\sum\_\{i=1\}^\{n\}a\_\{ig\}\\mu\_\{i\}\-\\gamma c\_\{g\}\\leq\\theta,g∈ℳ,\\displaystyle g\\in\\mathcal\{M\},μi,σi​ℓ≥0,\\displaystyle\\mu\_\{i\},\\sigma\_\{i\\ell\}\\geq 0,i=1,…,n,ℓ=1,…,k,\\displaystyle i=1,\\ldots,n,\\quad\\ell=1,\\ldots,k,νi​ℓ≥0,\\displaystyle\\nu\_\{i\\ell\}\\geq 0,i=1,…,n,ℓ=2,…,k,\\displaystyle i=1,\\ldots,n,\\quad\\ell=2,\\ldots,k,θ​free\.\\displaystyle\\theta\\ \\text\{free\}\.
For a workflow typeg∈𝒢∖ℳg\\in\\mathcal\{G\}\\setminus\\mathcal\{M\}, define its reduced\-cost score by

Γk​\(g,μ,θ\)=∑i=1nμi​ai​g−γ​cg−θ\.\\Gamma\_\{k\}\(g;\\mu,\\theta\)=\\sum\_\{i=1\}^\{n\}\\mu\_\{i\}a\_\{ig\}\-\\gamma c\_\{g\}\-\\theta\.\(43\)The first term is the task\-price\-weighted value of the workflow’s correct outputs,γ​cg\\gamma c\_\{g\}is its recurring execution cost, andθ\\thetais the dual price of one execution slot\. Thus,Γk​\(g,μ,θ\)\>0\\Gamma\_\{k\}\(g;\\mu,\\theta\)\>0means that workflowggviolates its full\-dual constraint∑i=1nμi​ai​g−γ​cg≤θ\.\\sum\_\{i=1\}^\{n\}\\mu\_\{i\}a\_\{ig\}\-\\gamma c\_\{g\}\\leq\\theta\.To see why all workflow\-indexed constraints can be checked through one pricing problem, recall the score functionsg​\(μ\)=∑i=1nμi​ai​g−γ​cg\.s\_\{g\}\(\\mu\)=\\sum\_\{i=1\}^\{n\}\\mu\_\{i\}a\_\{ig\}\-\\gamma c\_\{g\}\.Because repeated\-execution multiplicities are uncapped, we have

max⁡∑g∈𝒢zg≥0,g∈𝒢∑g∈𝒢zg=k⁡sg​\(μ\)​zg=k​maxg∈𝒢​sg​\(μ\)\.\\max\_\{\\begin\{subarray\}\{c\}z\_\{g\}\\geq 0,\\ g\\in\\mathcal\{G\}\\\\ \\sum\_\{g\\in\\mathcal\{G\}\}z\_\{g\}=k\\end\{subarray\}\}\\sum\_\{g\\in\\mathcal\{G\}\}s\_\{g\}\(\\mu\)z\_\{g\}=k\\max\_\{g\\in\\mathcal\{G\}\}s\_\{g\}\(\\mu\)\.\(44\)Hence it is enough to solve the single global pricing problem

maxg∈𝒢⁡sg​\(μ\)\.\\max\_\{g\\in\\mathcal\{G\}\}s\_\{g\}\(\\mu\)\.If its value exceedsθ\\theta, the maximizing workflow identifies a missing workflow that could improve the current LP solution and should therefore be added to the evaluated pool\. If its value is at mostθ\\theta, no workflow in the implicit class has enough price\-weighted accuracy, net of execution cost, to improve the current relaxation at these dual prices\.

### 5\.5Cost\-preserving randomized rounding and certificates

The integer program in[Section5\.3](https://arxiv.org/html/2609.18126#S5.SS3), when exactly solved, already returns a deployable portfolio, with no rounding required\. The obstacle is scale, not exactness: this is a concave\-coverage problem, and coverage problems are known to be NP\-hard in general\([Nemhauser et al\. 1978](https://arxiv.org/html/2609.18126#bib.bib23),[Feige 1998](https://arxiv.org/html/2609.18126#bib.bib14)\), so solving the integer program to optimality by branch\-and\-bound can become expensive whenℳ\\mathcal\{M\}ornnis large\. This subsection trades exactness for scalability: it solves the efficient LP relaxation \([41](https://arxiv.org/html/2609.18126#S5.E41)\), rounds the fractional solution into a feasible portfolio, and certifies how close the result comes to the true fixed\-size optimumOPTk​\(ℳ\)\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\)\.

For any feasible fractional execution vectorz≥0z\\geq 0satisfying∑g∈ℳzg=k\\sum\_\{g\\in\\mathcal\{M\}\}z\_\{g\}=k, letxi​\(z\)=∑g∈ℳai​g​zg,x\_\{i\}\(z\)=\\sum\_\{g\\in\\mathcal\{M\}\}a\_\{ig\}z\_\{g\},be the fractional number of correct executions assigned to taskii\. We separate the relaxed objective into three components:

Qk​\(z\)=1n​∑i=1nψ~k​\(xi​\(z\)\),Rk​\(z\)=1n​∑i=1nxi​\(z\),Ck​\(z\)=∑g∈ℳcg​zg\.Q\_\{k\}\(z\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\widetilde\{\\psi\}\_\{k\}\\bigl\(x\_\{i\}\(z\)\\bigr\),\\qquad R\_\{k\}\(z\)=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}x\_\{i\}\(z\),\\qquad C\_\{k\}\(z\)=\\sum\_\{g\\in\\mathcal\{M\}\}c\_\{g\}z\_\{g\}\.\(45\)The corresponding relaxed net value isFk​\(z\)=Qk​\(z\)−γ​Ck​\(z\)\.F\_\{k\}\(z\)=Q\_\{k\}\(z\)\-\\gamma C\_\{k\}\(z\)\.

Letzk⋆z^\{k\\star\}be an optimal solution of the size\-kkLP, and writeQk⋆=Qk\(zk⋆\),Q\_\{k\}^\{\\star\}=Q\_\{k\}\(z^\{k\\star\}\),Rk⋆=Rk\(zk⋆\),R\_\{k\}^\{\\star\}=R\_\{k\}\(z^\{k\\star\}\),andCk⋆=Ck\(zk⋆\)\.C\_\{k\}^\{\\star\}=C\_\{k\}\(z^\{k\\star\}\)\.ThenLk\(ℳ\)=Fk\(zk⋆\)=Qk⋆−γCk⋆\.L\_\{k\}\(\\mathcal\{M\}\)=F\_\{k\}\(z^\{k\\star\}\)=Q\_\{k\}^\{\\star\}\-\\gamma C\_\{k\}^\{\\star\}\.To obtain an integral execution plan, consider any feasible fractional execution vectorzzsatisfying∑g∈ℳzg=k\\sum\_\{g\\in\\mathcal\{M\}\}z\_\{g\}=k, and define

qg​\(z\)=zgk,g∈ℳ\.q\_\{g\}\(z\)=\\frac\{z\_\{g\}\}\{k\},\\qquad g\\in\\mathcal\{M\}\.Because∑gzg=k\\sum\_\{g\}z\_\{g\}=k, the vectorq⁡\(z\)q\(z\)is a probability distribution over workflow types\. DrawG1,…,Gk∼iidq⁡\(z\),G\_\{1\},\\ldots,G\_\{k\}\\stackrel\{\{\\scriptstyle\\mathrm\{iid\}\}\}\{\{\\sim\}\}q\(z\),and define the random execution\-count vector

Mk,gRR\(z\)=∑j=1k𝟏\{Gj=g\},𝑴kRR\(z\)=\(Mk,gRR\(z\):g∈ℳ\)\.M\_\{k,g\}^\{\\mathrm\{RR\}\}\(z\)=\\sum\_\{j=1\}^\{k\}\\mathbf\{1\}\\\{G\_\{j\}=g\\\},\\qquad\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\(z\)=\\bigl\(M\_\{k,g\}^\{\\mathrm\{RR\}\}\(z\):g\\in\\mathcal\{M\}\\bigr\)\.Thus, each of thekkexecution slots independently receives a workflow type according to the fractional allocation, and

k⁡\(𝑴kRR​\(z\)\)=kalmost surely,𝔼⁡\[Mk,gRR​\(z\)\]=zg\.k\\bigl\(\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\(z\)\\bigr\)=k\\quad\\text\{almost surely\},\\qquad\\mathbb\{E\}\\\!\\left\[M\_\{k,g\}^\{\\mathrm\{RR\}\}\(z\)\\right\]=z\_\{g\}\.Hence the rounding procedure preserves every workflow multiplicity, and therefore total execution cost, in expectation\. Moreover, for taskii, the random number of correct candidates satisfies

RiRR​\(z\):=ri​\(𝑴kRR​\(z\)\)∼Binomial⁡\(k,xi​\(z\)k\)\.R\_\{i\}^\{\\mathrm\{RR\}\}\(z\):=r\_\{i\}\\bigl\(\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\(z\)\\bigr\)\\sim\\mathrm\{Binomial\}\\\!\\left\(k,\\frac\{x\_\{i\}\(z\)\}\{k\}\\right\)\.Whenz=zk⋆z=z^\{k\\star\}, abbreviate

𝑴kRR:=𝑴kRR\(zk⋆\),Mk,gRR:=Mk,gRR\(zk⋆\),RiRR:=RiRR\(zk⋆\)\.\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}:=\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\(z^\{k\\star\}\),\\qquad M\_\{k,g\}^\{\\mathrm\{RR\}\}:=M\_\{k,g\}^\{\\mathrm\{RR\}\}\(z^\{k\\star\}\),\\qquad R\_\{i\}^\{\\mathrm\{RR\}\}:=R\_\{i\}^\{\\mathrm\{RR\}\}\(z^\{k\\star\}\)\.
The next theorem bounds this rounding scheme for any feasible fractional vectorzz, not only an exact LP optimizer\.

###### Theorem 5\.3\(Fixed\-size net\-value rounding certificate\)

Suppose[Section5\.2](https://arxiv.org/html/2609.18126#S5.SS2)holds for run sizekkandd1,k\>0d\_\{1,k\}\>0\. Letzzbe any feasible fractional execution vector for the repeated\-execution LP, and let𝐌kRR​\(z\)\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\(z\)be the size\-kkexecution multiset obtained by the with\-replacement rounding procedure described above\. Then

𝔼⁡\[Πψ,γ​\(𝑴kRR​\(z\)\)\]\\displaystyle\\mathbb\{E\}\\left\[\\Pi\_\{\\psi,\\gamma\}\\bigl\(\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\(z\)\\bigr\)\\right\]≥\(1−1e\)​Qk​\(z\)\+dk,ke​Rk​\(z\)−γ​Ck​\(z\),\\displaystyle\\geq\\left\(1\-\\frac\{1\}\{e\}\\right\)Q\_\{k\}\(z\)\+\\frac\{d\_\{k,k\}\}\{e\}R\_\{k\}\(z\)\-\\gamma C\_\{k\}\(z\),\(46\)Fk​\(z\)−𝔼⁡\[Πψ,γ​\(𝑴kRR​\(z\)\)\]\\displaystyle F\_\{k\}\(z\)\-\\mathbb\{E\}\\left\[\\Pi\_\{\\psi,\\gamma\}\\bigl\(\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\(z\)\\bigr\)\\right\]≤1e​\{Qk​\(z\)−dk,k​Rk​\(z\)\}≤cψ,ke​Qk​\(z\)\.\\displaystyle\\leq\\frac\{1\}\{e\}\\left\\\{Q\_\{k\}\(z\)\-d\_\{k,k\}R\_\{k\}\(z\)\\right\\\}\\leq\\frac\{c\_\{\\psi,k\}\}\{e\}Q\_\{k\}\(z\)\.\(47\)In particular, ifz=zk⋆z=z^\{k\\star\}is an optimal solution of the size\-kkLP, then

OPTk​\(ℳ\)−𝔼⁡\[Πψ,γ​\(𝑴kRR\)\]≤1e​\(Qk⋆−dk,k​Rk⋆\)≤cψ,ke​Qk⋆\.\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\)\-\\mathbb\{E\}\\left\[\\Pi\_\{\\psi,\\gamma\}\\bigl\(\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\\bigr\)\\right\]\\leq\\frac\{1\}\{e\}\\left\(Q\_\{k\}^\{\\star\}\-d\_\{k,k\}R\_\{k\}^\{\\star\}\\right\)\\leq\\frac\{c\_\{\\psi,k\}\}\{e\}Q\_\{k\}^\{\\star\}\.\(48\)

The final bound has a uniform interpretation\. Because0≤yi​ℓk⋆≤10\\leq y\_\{i\\ell\}^\{k\\star\}\\leq 1and∑ℓ=1kdℓ,k=ψk​\(k\)−ψk​\(0\)=1,\\sum\_\{\\ell=1\}^\{k\}d\_\{\\ell,k\}=\\psi\_\{k\}\(k\)\-\\psi\_\{k\}\(0\)=1,the LP accuracy component satisfies0≤Qk⋆≤1\.0\\leq Q\_\{k\}^\{\\star\}\\leq 1\.Consequently,

0≤OPTk​\(ℳ\)−𝔼⁡\[Πψ,γ​\(𝑴kRR\)\]≤cψ,ke​Qk⋆≤cψ,ke≤1e\.0\\leq\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\)\-\\mathbb\{E\}\\left\[\\Pi\_\{\\psi,\\gamma\}\(\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\)\\right\]\\leq\\frac\{c\_\{\\psi,k\}\}\{e\}Q\_\{k\}^\{\\star\}\\leq\\frac\{c\_\{\\psi,k\}\}\{e\}\\leq\\frac\{1\}\{e\}\.Thuscψ,k/ec\_\{\\psi,k\}/eis a directly interpretable worst\-case additive loss in accuracy\-equivalent units, while the bound involvingQk⋆Q\_\{k\}^\{\\star\}can be strictly sharper for the realized LP solution \(see also[Remark5\.5](https://arxiv.org/html/2609.18126#S5.Thmtheorem5)at the end of this section\)\.

The recurring execution\-cost term does not incur any additional rounding loss\. Because with\-replacement rounding preserves each workflow multiplicity in expectation,

𝔼\[C\(𝑴kRR\)\]=∑g∈ℳcg𝔼\[Mk,gRR\]=∑g∈ℳcgzgk⋆=Ck⋆\.\\mathbb\{E\}\\left\[C\\bigl\(\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\\bigr\)\\right\]=\\sum\_\{g\\in\\mathcal\{M\}\}c\_\{g\}\\mathbb\{E\}\\left\[M\_\{k,g\}^\{\\mathrm\{RR\}\}\\right\]=\\sum\_\{g\\in\\mathcal\{M\}\}c\_\{g\}z\_\{g\}^\{k\\star\}=C\_\{k\}^\{\\star\}\.Thus, all loss in the certificate arises from rounding the nonlinear selector\-aware accuracy term, not from execution cost\. The with\-replacement procedure is the natural rounding scheme for the repeated\-execution model, because it permits the same workflow type to be assigned to multiple execution slots while preserving expected multiplicities and execution cost\.

Specializing \([37](https://arxiv.org/html/2609.18126#S5.E37)\) to Plackett–Luce recovery gives a closed form for the curvature:

ck,λPL=1−k\+λ−1λ⁡\{k\+\(λ−1\)​\(k−1\)\}\.c\_\{k,\\lambda\}^\{\\mathrm\{PL\}\}=1\-\\frac\{k\+\\lambda\-1\}\{\\lambda\\\{k\+\(\\lambda\-1\)\(k\-1\)\\\}\}\.\(49\)This expression depends only on the run sizekkand selector strengthλ\\lambda\. It satisfiesc1,λPL=0,ck,1PL=0\.c\_\{1,\\lambda\}^\{\\mathrm\{PL\}\}=0,c\_\{k,1\}^\{\\mathrm\{PL\}\}=0\.Fork≥2k\\geq 2andλ\>1\\lambda\>1, the curvature is nondecreasing in bothkkandλ\\lambda\. Moreover, for fixedλ\\lambda,limk→∞ck,λPL=1−1λ2\.\\lim\_\{k\\to\\infty\}c\_\{k,\\lambda\}^\{\\mathrm\{PL\}\}=1\-\\frac\{1\}\{\\lambda^\{2\}\}\.Thus the rounding certificate is exact for a singleton portfolio and for a random selector\. As the portfolio grows or the selector becomes more discriminating, the recovery curve becomes more curved: the first correct candidate accounts for a larger share of the total recovery gain, and the worst\-case additive rounding certificate becomes looser\. This does not mean that a stronger selector reduces portfolio value; it means only that a linear relaxation may approximate the more strongly curved recovery objective less tightly\.

For a cardinality grid𝒦\\mathcal\{K\}, define the grid LP benchmark

ULP​\(𝒦,ℳ\):=max⁡\{0,maxk∈𝒦⁡Lk​\(ℳ\)\}\.U\_\{\\mathrm\{LP\}\}\(\\mathcal\{K\};\\mathcal\{M\}\):=\\max\\left\\\{0,\\max\_\{k\\in\\mathcal\{K\}\}L\_\{k\}\(\\mathcal\{M\}\)\\right\\\}\.\(50\)Under Plackett–Luce recovery,ck,λPLc\_\{k,\\lambda\}^\{\\mathrm\{PL\}\}is nondecreasing inkk\. Letk¯=max⁡𝒦\\bar\{k\}=\\max\\mathcal\{K\}and define the worst\-case rounding loss on the grid by

εrnd​\(𝒦,λ\):=ck¯,λPLe=1e​maxk∈𝒦​ck,λPL\.\\varepsilon\_\{\\mathrm\{rnd\}\}\(\\mathcal\{K\},\\lambda\):=\\frac\{c\_\{\\bar\{k\},\\lambda\}^\{\\mathrm\{PL\}\}\}\{e\}=\\frac\{1\}\{e\}\\max\_\{k\\in\\mathcal\{K\}\}c\_\{k,\\lambda\}^\{\\mathrm\{PL\}\}\.
###### Corollary 5\.4\(Grid\-level Plackett–Luce rounding guarantee\)

Suppose the selector follows the Plackett–Luce recovery curve with strengthλ≥1\\lambda\\geq 1\. For eachk∈𝒦k\\in\\mathcal\{K\}, let𝐌kRR\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}be obtained by applying the with\-replacement rounding procedure to an optimal solution of the size\-kkLP in \([41](https://arxiv.org/html/2609.18126#S5.E41)\)\. Assume the firm retains a feasible baseline execution\-count vector𝐦0∈ℤ\+ℳ\\bm\{m\}^\{0\}\\in\\mathbb\{Z\}\_\{\+\}^\{\\mathcal\{M\}\}, such as a validated incumbent deployment, satisfyingΠλ,γ​\(𝐦0\)≥V¯\>0\.\\Pi\_\{\\lambda,\\gamma\}\(\\bm\{m\}^\{0\}\)\\geq\\underline\{V\}\>0\.After rounding, choose the best realized portfolio among the baseline and the rounded grid candidates:𝐌^∈arg​maxS∈\{𝐦0\}∪\{𝐌kRR:k∈𝒦\}Πλ,γ\(S\)\.\\widehat\{\\bm\{M\}\}\\in\\operatorname\*\{arg\\,max\}\_\{S\\in\\\{\\bm\{m\}^\{0\}\\\}\\cup\\\{\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}:k\\in\\mathcal\{K\}\\\}\}\\Pi\_\{\\lambda,\\gamma\}\(S\)\.Then, we have

𝔼⁡\[Πλ,γ​\(𝑴^\)\]≥max⁡\{V¯,ULP​\(𝒦,ℳ\)−εrnd​\(𝒦,λ\)\}≥V¯V¯\+εrnd​\(𝒦,λ\)​ULP​\(𝒦,ℳ\)\.\\mathbb\{E\}\\\!\\big\[\\Pi\_\{\\lambda,\\gamma\}\(\\widehat\{\\bm\{M\}\}\)\\big\]\\geq\\max\\left\\\{\\underline\{V\},\\,U\_\{\\mathrm\{LP\}\}\(\\mathcal\{K\};\\mathcal\{M\}\)\-\\varepsilon\_\{\\mathrm\{rnd\}\}\(\\mathcal\{K\},\\lambda\)\\right\\\}\\geq\\frac\{\\underline\{V\}\}\{\\underline\{V\}\+\\varepsilon\_\{\\mathrm\{rnd\}\}\(\\mathcal\{K\},\\lambda\)\}U\_\{\\mathrm\{LP\}\}\(\\mathcal\{K\};\\mathcal\{M\}\)\.\(51\)Moreover, if the size\-kkinteger programs are solved exactly instead, the same guarantee holds without the expectation\.

Note that becauseLk​\(ℳ\)L\_\{k\}\(\\mathcal\{M\}\)is an LP relaxation of the size\-kkproblem, we haveULP​\(𝒦,ℳ\)≥max⁡\{0,maxk∈𝒦⁡OPTk​\(ℳ\)\}\.U\_\{\\mathrm\{LP\}\}\(\\mathcal\{K\};\\mathcal\{M\}\)\\geq\\max\\big\\\{0,\\max\_\{k\\in\\mathcal\{K\}\}\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\)\\big\\\}\.Thus,

𝔼⁡\[Πλ,γ​\(𝑴^\)\]≥V¯V¯\+εrnd​\(𝒦,λ\)​maxk∈𝒦​OPTk​\(ℳ\)\.\\displaystyle\\mathbb\{E\}\\\!\\big\[\\Pi\_\{\\lambda,\\gamma\}\(\\widehat\{\\bm\{M\}\}\)\\big\]\\,\\geq\\,\\frac\{\\underline\{V\}\}\{\\underline\{V\}\+\\varepsilon\_\{\\mathrm\{rnd\}\}\(\\mathcal\{K\},\\lambda\)\}\\,\\max\_\{k\\in\\mathcal\{K\}\}\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\)\.
The first inequality in \([51](https://arxiv.org/html/2609.18126#S5.E51)\) combines two safeguards\. The retained baseline guarantees value at leastV¯\\underline\{V\}, while the rounded grid solution achieves the LP benchmark up to the uniform additive rounding lossεrnd​\(𝒦,λ\)\\varepsilon\_\{\\mathrm\{rnd\}\}\(\\mathcal\{K\},\\lambda\)\. The second inequality converts these two additive guarantees into a multiplicative bound relative to the grid LP benchmark\.

The resulting factor depends only on the baseline value and the largest selector curvature among the grid sizes\. Since0≤εrnd​\(𝒦,λ\)≤1/e,0\\leq\\varepsilon\_\{\\mathrm\{rnd\}\}\(\\mathcal\{K\},\\lambda\)\\leq 1/e,the factor is always strictly positive wheneverV¯\>0\\underline\{V\}\>0\. Moreover, the rounding loss vanishes when the fixed\-size recovery curve is linear, in which case the LP solution is preserved in expectation by the rounding procedure\.

## 6Ellipsoid\-Based Optimization Over an Implicit Workflow Class

The finite\-pool formulations in[Section5](https://arxiv.org/html/2609.18126#S5)introduce one primal variable for each evaluated workflow type\. When the feasible workflow class𝒢\\mathcal\{G\}is large and represented only implicitly, the corresponding repeated\-execution LP may contain too many variables to enumerate\. The dual, however, has only finitely many price variables\. Its only implicit component is a family of constraints indexed by workflow types\. This is exactly the setting in which the ellipsoid method can optimize through separation rather than explicit enumeration\([Grötschel et al\. 1981](https://arxiv.org/html/2609.18126#bib.bib54),[Grötschel et al\. 1988](https://arxiv.org/html/2609.18126#bib.bib55)\)\.

Fix a run sizekk\. An equivalent projected dual of the full LP relaxation is

Ek​\(𝒢\)=minμ,ξ,θ\\displaystyle E\_\{k\}\(\\mathcal\{G\}\)=\\min\_\{\\mu,\\xi,\\theta\}\\quad∑i=1nξi\+k​θ\\displaystyle\\sum\_\{i=1\}^\{n\}\\xi\_\{i\}\+k\\theta\(52\)s\.t\.ξi\+r​μi≥ψk​\(r\)n,\\displaystyle\\xi\_\{i\}\+r\\mu\_\{i\}\\geq\\frac\{\\psi\_\{k\}\(r\)\}\{n\},i=1,…,n,r=0,…,k,\\displaystyle i=1,\\ldots,n,\\quad r=0,\\ldots,k,0≤μi≤d1,kn,0≤ξi≤1n,\\displaystyle 0\\leq\\mu\_\{i\}\\leq\\frac\{d\_\{1,k\}\}\{n\},\\qquad 0\\leq\\xi\_\{i\}\\leq\\frac\{1\}\{n\},i=1,…,n,\\displaystyle i=1,\\ldots,n,−γ​cmax≤θ≤d1,k,\\displaystyle\-\\gamma c\_\{\\max\}\\leq\\theta\\leq d\_\{1,k\},∑i=1nμi​ai​g−γ​cg≤θ,\\displaystyle\\sum\_\{i=1\}^\{n\}\\mu\_\{i\}a\_\{ig\}\-\\gamma c\_\{g\}\\leq\\theta,g∈𝒢\.\\displaystyle g\\in\\mathcal\{G\}\.Appendix[D\.11](https://arxiv.org/html/2609.18126#A4.SS11)derives \([52](https://arxiv.org/html/2609.18126#S6.E52)\) and proves thatEk​\(𝒢\)=Lk​\(𝒢\)\.E\_\{k\}\(\\mathcal\{G\}\)=L\_\{k\}\(\\mathcal\{G\}\)\.

### 6\.1Single\-workflow separation

[Section3\.3](https://arxiv.org/html/2609.18126#S3.SS3)introduced the global workflow\-pricing value

P⁡\(w,γ\)=maxg∈𝒢⁡\{∑i=1nwi​ai​g−γ​cg\}P\(w,\\gamma\)=\\max\_\{g\\in\\mathcal\{G\}\}\\left\\\{\\sum\_\{i=1\}^\{n\}w\_\{i\}a\_\{ig\}\-\\gamma c\_\{g\}\\right\\\}and explained how it compares with the priceθ\\thetaof one execution slot\. The dual in \([52](https://arxiv.org/html/2609.18126#S6.E52)\) now provides the formal origin of these quantities: the task priceswiw\_\{i\}are precisely the dual variablesμi\\mu\_\{i\}, and the workflow\-indexed constraints requireP⁡\(μ,γ\)≤θ\.P\(\\mu,\\gamma\)\\leq\\theta\.Thus, the pricing interface from[Section3\.3](https://arxiv.org/html/2609.18126#S3.SS3)is exactly the separation oracle needed to optimize the implicit dual\. We do not solveP⁡\(μ,γ\)P\(\\mu,\\gamma\)by explicitly enumerating or optimizing over𝒢\\mathcal\{G\}\. Instead, at each candidate dual solution, we pass the task prices and execution\-cost penalty to the weak global pricing oracle, which searches for a high\-scoring workflow according to \([15](https://arxiv.org/html/2609.18126#S3.E15)\)\.

In particular, if the oracle returns a workflowg~\\widetilde\{g\}satisfying∑i=1nμi​ai​g~−γ​cg~\>θ,\\sum\_\{i=1\}^\{n\}\\mu\_\{i\}a\_\{i\\widetilde\{g\}\}\-\\gamma c\_\{\\widetilde\{g\}\}\>\\theta,then the constraint indexed byg~\\widetilde\{g\}in \([52](https://arxiv.org/html/2609.18126#S6.E52)\) is violated\. That constraint supplies a separating hyperplane, which the ellipsoid method uses to exclude the current candidate dual point and continue its search\. Conversely, suppose the oracle isδ\\delta\-accurate in the sense of \([15](https://arxiv.org/html/2609.18126#S3.E15)\) and returns a workflow whose score is at mostθ\\theta\. The guarantee established in[Section3\.3](https://arxiv.org/html/2609.18126#S3.SS3)then impliesP⁡\(μ,γ\)≤θ\+δ\.P\(\\mu,\\gamma\)\\leq\\theta\+\\delta\.Hence all workflow\-indexed constraints are satisfied after increasing the slot price fromθ\\thetatoθ\+δ\\theta\+\\delta\. Becauseθ\\thetahas coefficientkkin the dual objective, this adjustment increases the objective by at mostk​δk\\delta\.

The pricing oracle therefore implements weak separation for the implicit dual: it either produces a violated workflow constraint or certifies feasibility of the entire workflow\-indexed constraint family up to an objective error ofk​δk\\delta\.

### 6\.2Batched stochastic pricing

The pricing oracle may itself be stochastic\. We assume that there exists a constantporc\>0p\_\{\\mathrm\{orc\}\}\>0such that every primitive oracle call, conditional on the full query history, satisfies the additive weak\-pricing guarantee in \([15](https://arxiv.org/html/2609.18126#S3.E15)\) with probability at leastporcp\_\{\\mathrm\{orc\}\}\. Thus, when queried at task pricesμ\\muwith toleranceδ\\delta, a primitive call returns a workflowg~\\widetilde\{g\}satisfying

∑i=1nμi​ai​g~−γ​cg~≥P⁡\(μ,γ\)−δ\\sum\_\{i=1\}^\{n\}\\mu\_\{i\}a\_\{i\\widetilde\{g\}\}\-\\gamma c\_\{\\widetilde\{g\}\}\\geq P\(\\mu,\\gamma\)\-\\deltawith conditional probability at leastporcp\_\{\\mathrm\{orc\}\}\.

A single unsuccessful pricing call cannot safely be used to certify that no violated workflow constraint exists\. More seriously, an invalid separator can undermine the correctness of the entire ellipsoid routine\. We therefore amplify the primitive oracle at each separation query\. Specifically, the algorithm makesmmpricing calls at the same dual prices and retains the returned workflow with the highest score\. The batch is unsuccessful only if none of itsmmprimitive calls satisfies the weak\-pricing guarantee\. Therefore, conditional on the history before the batch,

ℙ⁡\{batch fails∣history\}≤\(1−porc\)m≤e−porc​m\.\\mathbb\{P\}\\\{\\text\{batch fails\}\\mid\\text\{history\}\\\}\\leq\(1\-p\_\{\\mathrm\{orc\}\}\)^\{m\}\\leq e^\{\-p\_\{\\mathrm\{orc\}\}m\}\.\(53\)Thus, batching converts a primitive stochastic pricing procedure with constant success probability into a weak\-separation oracle whose failure probability decays exponentially in the batch size\.

### 6\.3Ellipsoid algorithm and finite\-call guarantee

We now combine the projected dual, the batched pricing oracle, and the fixed\-size rounding procedure into an end\-to\-end algorithm\. The method solves the implicit LP separately for each run size on the cardinality grid, recovers a fractional execution allocation supported on workflows discovered during separation, and rounds that allocation into a deployable portfolio\. We writeεell\>0\\varepsilon\_\{\\mathrm\{ell\}\}\>0for the prescribed ellipsoid\-method tolerance, where the subscript “ell\\mathrm\{ell\}” denotes the ellipsoid method\.

Algorithm 1Ellipsoid\-Based Dual\-Guided Workflow Optimization1:Input: cardinality grid

𝒦\\mathcal\{K\}; recovery curves

\{ψk\}\\\{\\psi\_\{k\}\\\}; cost price

γ\\gamma; implicit\-LP tolerance

εell\>0\\varepsilon\_\{\\mathrm\{ell\}\}\>0; batch size

mm; primitive stochastic pricing oracle; feasible baseline execution\-count vector

𝒎0\\bm\{m\}^\{0\}\.

2:for

k∈𝒦k\\in\\mathcal\{K\}do

3:Set the pricing tolerance

τk←εell2​k\.\\tau\_\{k\}\\leftarrow\\frac\{\\varepsilon\_\{\\mathrm\{ell\}\}\}\{2k\}\.
4:Initialize

𝒢krec←∅\\mathcal\{G\}\_\{k\}^\{\\mathrm\{rec\}\}\\leftarrow\\varnothingand initialize the ellipsoid routine for \([52](https://arxiv.org/html/2609.18126#S6.E52)\) with optimization tolerance

εell/2\\varepsilon\_\{\\mathrm\{ell\}\}/2; set

t←0t\\leftarrow 0\.

5:whilethe ellipsoid routine has not met its stopping criteriondo

6:Let

\(μt,ξt,θt\)\(\\mu^\{t\},\\xi^\{t\},\\theta^\{t\}\)be the current ellipsoid iterate\.

7:ifan explicit constraint in \([52](https://arxiv.org/html/2609.18126#S6.E52)\) is violatedthen

8:Supply one such violated explicit constraint to the ellipsoid routine as a separating hyperplane\.

9:else

10:Call the primitive pricing oracle

mmtimes at

\(μt,γ,τk\)\(\\mu^\{t\},\\gamma,\\tau\_\{k\}\), obtaining

g~t,1,…,g~t,m\\widetilde\{g\}^\{t,1\},\\ldots,\\widetilde\{g\}^\{t,m\}, and set

gt∈arg⁡maxj∈\[m\]​\{∑i=1nμit​ai​g~t,j−γ​cg~t,j\}\.g^\{t\}\\in\\arg\\max\_\{j\\in\[m\]\}\\left\\\{\\sum\_\{i=1\}^\{n\}\\mu\_\{i\}^\{t\}a\_\{i\\widetilde\{g\}^\{t,j\}\}\-\\gamma c\_\{\\widetilde\{g\}^\{t,j\}\}\\right\\\}\.
11:Record the returned workflow:

𝒢krec←𝒢krec∪\{gt\}\.\\mathcal\{G\}\_\{k\}^\{\\mathrm\{rec\}\}\\leftarrow\\mathcal\{G\}\_\{k\}^\{\\mathrm\{rec\}\}\\cup\\\{g^\{t\}\\\}\.
12:if

∑i=1nμit​ai​gt−γ​cgt\>θt\\sum\_\{i=1\}^\{n\}\\mu\_\{i\}^\{t\}a\_\{ig^\{t\}\}\-\\gamma c\_\{g^\{t\}\}\>\\theta^\{t\}then

13:Supply the violated workflow constraint

∑i=1nμi​ai​gt−γ​cgt≤θ\\sum\_\{i=1\}^\{n\}\\mu\_\{i\}a\_\{ig^\{t\}\}\-\\gamma c\_\{g^\{t\}\}\\leq\\thetato the ellipsoid routine as a separating hyperplane\.

14:else

15:Supply the weak\-separation certificate

P⁡\(μt,γ\)≤θt\+τkP\(\\mu^\{t\},\\gamma\)\\leq\\theta^\{t\}\+\\tau\_\{k\}to the ellipsoid routine\.

16:endif

17:endif

18:Let the ellipsoid routine perform its update and stopping test; set

t←t\+1t\\leftarrow t\+1\.

19:endwhile

20:Perform primal recovery by solving the restricted fixed\-size LP relaxation \([41](https://arxiv.org/html/2609.18126#S5.E41)\) with

ℳ=𝒢krec\\mathcal\{M\}=\\mathcal\{G\}\_\{k\}^\{\\mathrm\{rec\}\}\. Let

\(yk,zk\)\(y^\{k\},z^\{k\}\)be the resulting recovered primal solution, extended by

zgk=0z\_\{g\}^\{k\}=0for

g∉𝒢krecg\\notin\\mathcal\{G\}\_\{k\}^\{\\mathrm\{rec\}\}, such that

zgk≥0,∑g∈𝒢kreczgk=k,Fk​\(zk\)≥Lk​\(𝒢\)−εell\.z\_\{g\}^\{k\}\\geq 0,\\sum\_\{g\\in\\mathcal\{G\}\_\{k\}^\{\\mathrm\{rec\}\}\}z\_\{g\}^\{k\}=k,F\_\{k\}\(z^\{k\}\)\\geq L\_\{k\}\(\\mathcal\{G\}\)\-\\varepsilon\_\{\\mathrm\{ell\}\}\.
21:Set

qgk=zgk/kq\_\{g\}^\{k\}=z\_\{g\}^\{k\}/k, draw

G1k,…,Gkk∼iidqk,G\_\{1\}^\{k\},\\ldots,G\_\{k\}^\{k\}\\stackrel\{\{\\scriptstyle\\mathrm\{iid\}\}\}\{\{\\sim\}\}q^\{k\},and define

Mk,gRR=∑j=1k𝟏\{Gjk=g\},𝑴kRR=\(Mk,gRR:g∈𝒢krec\)\.M\_\{k,g\}^\{\\mathrm\{RR\}\}=\\sum\_\{j=1\}^\{k\}\\mathbf\{1\}\\\{G\_\{j\}^\{k\}=g\\\},\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}=\\bigl\(M\_\{k,g\}^\{\\mathrm\{RR\}\}:g\\in\\mathcal\{G\}\_\{k\}^\{\\mathrm\{rec\}\}\\bigr\)\.
22:endfor

23:Return the execution\-count vector

𝑴^\\widehat\{\\bm\{M\}\}with the largest realized net value among the outside option

𝟎\\bm\{0\}, the baseline

𝒎0\\bm\{m\}^\{0\}, and the rounded vectors

\{𝑴kRR:k∈𝒦\}\\\{\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}:k\\in\\mathcal\{K\}\\\}\.

For eachk∈𝒦k\\in\\mathcal\{K\}, letNell,k​\(εell\)N\_\{\\mathrm\{ell\},k\}\(\\varepsilon\_\{\\mathrm\{ell\}\}\)be a deterministic upper bound on the number of iterations of the while\-loop in Algorithm[1](https://arxiv.org/html/2609.18126#alg1)before the ellipsoid routine reaches optimization toleranceεell/2\\varepsilon\_\{\\mathrm\{ell\}\}/2\. Thus, in the run corresponding to sizekk, the iteration index takes the valuest=0,…,Tk−1,for some​Tk≤Nell,k​\(εell\)\.t=0,\\ldots,T\_\{k\}\-1,\\text\{for some\}T\_\{k\}\\leq N\_\{\\mathrm\{ell\},k\}\(\\varepsilon\_\{\\mathrm\{ell\}\}\)\.Classical ellipsoid\-method complexity bounds imply333HereBBdenotes an upper bound on the binary encoding length of the rational problem data; see[Grötschel et al\. \(1981\)](https://arxiv.org/html/2609.18126#bib.bib54),[Grötschel et al\. \(1988\)](https://arxiv.org/html/2609.18126#bib.bib55)\.Nell​\(εell\):=∑k∈𝒦Nell,k​\(εell\)=∑k∈𝒦poly⁡\(n,k,B,log⁡1εell\)\.N\_\{\\mathrm\{ell\}\}\(\\varepsilon\_\{\\mathrm\{ell\}\}\):=\\sum\_\{k\\in\\mathcal\{K\}\}N\_\{\\mathrm\{ell\},k\}\(\\varepsilon\_\{\\mathrm\{ell\}\}\)=\\sum\_\{k\\in\\mathcal\{K\}\}\\operatorname\{poly\}\\\!\\left\(n,k,B,\\log\\frac\{1\}\{\\varepsilon\_\{\\mathrm\{ell\}\}\}\\right\)\.Because each iteration invokes at most one batch ofmmprimitive pricing calls, Algorithm[1](https://arxiv.org/html/2609.18126#alg1)makes at mostT:=m​Nell​\(εell\)T:=m\\,N\_\{\\mathrm\{ell\}\}\(\\varepsilon\_\{\\mathrm\{ell\}\}\)primitive pricing calls\.

###### Theorem 6\.1\(High\-probability implicit\-class guarantee\)

Suppose𝒢\\mathcal\{G\}is a finite, implicitly represented workflow\-type class whose rational data have encoding length at mostBB, and supposecg≤cmaxc\_\{g\}\\leq c\_\{\\max\}for everyg∈𝒢g\\in\\mathcal\{G\}\. Suppose the selector follows the Plackett–Luce recovery curve with strengthλ≥1\\lambda\\geq 1, and let the cardinality grid𝒦\\mathcal\{K\}have coverage ratioϱ≥1\\varrho\\geq 1\. Retain a feasible baseline portfolio𝐦0\\bm\{m\}^\{0\}satisfyingΠλ,γ​\(𝐦0\)≥V¯\>0\.\\Pi\_\{\\lambda,\\gamma\}\(\\bm\{m\}^\{0\}\)\\geq\\underline\{V\}\>0\.Suppose further that every primitive pricing call satisfies \([15](https://arxiv.org/html/2609.18126#S3.E15)\), conditional on the query history, with probability at leastporc\>0p\_\{\\mathrm\{orc\}\}\>0\. Run Algorithm[1](https://arxiv.org/html/2609.18126#alg1)with batch sizemm\. Then, with probability at least

1−Nell​\(εell\)​e−porc​m=1−exp⁡\{−porc​TNell​\(εell\)\+log⁡Nell​\(εell\)\},1\-N\_\{\\mathrm\{ell\}\}\(\\varepsilon\_\{\\mathrm\{ell\}\}\)e^\{\-p\_\{\\mathrm\{orc\}\}m\}=1\-\\exp\\left\\\{\-\\frac\{p\_\{\\mathrm\{orc\}\}T\}\{N\_\{\\mathrm\{ell\}\}\(\\varepsilon\_\{\\mathrm\{ell\}\}\)\}\+\\log N\_\{\\mathrm\{ell\}\}\(\\varepsilon\_\{\\mathrm\{ell\}\}\)\\right\\\},\(54\)all batched separation calls satisfy their weak\-pricing guarantees, and the returned portfolio obeys

𝔼rnd​\[Πλ,γ​\(𝑴^\)\]≥max⁡\{V¯,1ϱ​OPTλ,γ​\(𝒢\)−εell−εrnd​\(𝒦,λ\)\}\.\\mathbb\{E\}\_\{\\mathrm\{rnd\}\}\\\!\\left\[\\Pi\_\{\\lambda,\\gamma\}\(\\widehat\{\\bm\{M\}\}\)\\right\]\\geq\\max\\left\\\{\\underline\{V\},\\,\\frac\{1\}\{\\varrho\}\\mathrm\{OPT\}\_\{\\lambda,\\gamma\}\(\\mathcal\{G\}\)\-\\varepsilon\_\{\\mathrm\{ell\}\}\-\\varepsilon\_\{\\mathrm\{rnd\}\}\(\\mathcal\{K\},\\lambda\)\\right\\\}\.\(55\)Consequently,

𝔼rnd​\[Πλ,γ​\(𝑴^\)\]≥1ϱ​V¯V¯\+εell\+εrnd​\(𝒦,λ\)​OPTλ,γ​\(𝒢\)\.\\mathbb\{E\}\_\{\\mathrm\{rnd\}\}\\\!\\big\[\\Pi\_\{\\lambda,\\gamma\}\(\\widehat\{\\bm\{M\}\}\)\\big\]\\geq\\frac\{1\}\{\\varrho\}\\frac\{\\underline\{V\}\}\{\\underline\{V\}\+\\varepsilon\_\{\\mathrm\{ell\}\}\+\\varepsilon\_\{\\mathrm\{rnd\}\}\(\\mathcal\{K\},\\lambda\)\}\\mathrm\{OPT\}\_\{\\lambda,\\gamma\}\(\\mathcal\{G\}\)\.\(56\)Here, the expectation𝔼rnd\\mathbb\{E\}\_\{\\mathrm\{rnd\}\}is taken over the final randomized\-rounding step, conditional on the successful batched\-separation event\.

To make the probability of any failed batch at mostβ\\beta, it is sufficient to choose

m≥1porc​\[log⁡Nell​\(εell\)\+log⁡1β\]\.m\\geq\\frac\{1\}\{p\_\{\\mathrm\{orc\}\}\}\\left\[\\log N\_\{\\mathrm\{ell\}\}\(\\varepsilon\_\{\\mathrm\{ell\}\}\)\+\\log\\frac\{1\}\{\\beta\}\\right\]\.\(57\)The resulting primitive\-call complexity is

T=O⁡\(Nell​\(εell\)porc​log⁡Nell​\(εell\)β\)\.T=O\\\!\\left\(\\frac\{N\_\{\\mathrm\{ell\}\}\(\\varepsilon\_\{\\mathrm\{ell\}\}\)\}\{p\_\{\\mathrm\{orc\}\}\}\\log\\frac\{N\_\{\\mathrm\{ell\}\}\(\\varepsilon\_\{\\mathrm\{ell\}\}\)\}\{\\beta\}\\right\)\.\(58\)
The guarantee separates three distinct sources of loss\. The factor1/ϱ1/\\varrhocomes from replacing exhaustive search over all run sizes by the cardinality grid\. The termεell\\varepsilon\_\{\\mathrm\{ell\}\}is the error from solving the implicit LP only approximately, andεrnd​\(𝒦,λ\)\\varepsilon\_\{\\mathrm\{rnd\}\}\(\\mathcal\{K\},\\lambda\)is the loss from converting the recovered fractional allocations into integral execution portfolios\. The stochastic pricing oracle affects the confidence level, but not the value bound conditional on successful separation\.

Most importantly, the benchmark in \([55](https://arxiv.org/html/2609.18126#S6.E55)\)–\([56](https://arxiv.org/html/2609.18126#S6.E56)\) is the best portfolio over the full implicit workflow class𝒢\\mathcal\{G\}, not merely over the workflows explicitly generated during the algorithm\. The method therefore does not require enumerating𝒢\\mathcal\{G\}or discovering every workflow with positive reduced cost; it requires only enough global pricing calls to optimize the implicit dual to the prescribed tolerance\.

## 7Stochastic Workflow Outcomes

Sections[3](https://arxiv.org/html/2609.18126#S3)through[6](https://arxiv.org/html/2609.18126#S6)treatai​ga\_\{ig\}as deterministic: workflowggeither solves development taskiior it does not, and repeated executions reproduce the same outcome\. In practice, an execution of the same workflow on the same task need not return the same answer twice\. In this section, we reinterpretai​g∈\[0,1\]a\_\{ig\}\\in\[0,1\]as the probability that a single execution of workflowggis correct on taskii, with the deterministic model recovered when everyai​ga\_\{ig\}is zero or one\.

The deployment decision is unchanged\. It remains the execution\-count vector𝒎=\(mg\)g∈𝒢∈ℤ\+𝒢\\bm\{m\}=\(m\_\{g\}\)\_\{g\\in\\mathcal\{G\}\}\\in\\mathbb\{Z\}\_\{\+\}^\{\\mathcal\{G\}\}, with finite support\. What changes is that the number of correct candidate answers on a task is now random rather than determined by𝒎\\bm\{m\}\.

The stochastic model requires the following independent\-sampling assumption\.

\{assumption\}

\[IID workflow execution\] For every development taskiiand workflow typegg, there is a parameterai​g∈\[0,1\]a\_\{ig\}\\in\[0,1\]\. Each execution of workflowggon taskiihas a correctness indicator distributed asBernoulli⁡\(ai​g\)\\mathrm\{Bernoulli\}\(a\_\{ig\}\)\. Correctness indicators are mutually independent across tasks, workflow types, execution copies, and sampling stages\. Thus, for each fixed pair\(i,g\)\(i,g\), repeated executions are i\.i\.d\. Bernoulli draws with meanai​ga\_\{ig\}\.

For a feasible𝒎\\bm\{m\}, letZi​g​q∼Bernoulli⁡\(ai​g\)Z\_\{igq\}\\sim\\mathrm\{Bernoulli\}\(a\_\{ig\}\)denote the correctness indicator of copyqqof workflowggon taskii\. Under[Section7](https://arxiv.org/html/2609.18126#S7), the random number of correct outputs on taskiiis

Ri​\(𝒎\)=∑g∈𝒢∑q=1mgZi​g​q∈\{0,…,k⁡\(𝒎\)\},𝔼⁡\[Ri​\(𝒎\)\]=∑g∈𝒢mg​ai​g\.R\_\{i\}\(\\bm\{m\}\)=\\sum\_\{g\\in\\mathcal\{G\}\}\\sum\_\{q=1\}^\{m\_\{g\}\}Z\_\{igq\}\\in\\\{0,\\ldots,k\(\\bm\{m\}\)\\\},\\qquad\\mathbb\{E\}\[R\_\{i\}\(\\bm\{m\}\)\]=\\sum\_\{g\\in\\mathcal\{G\}\}m\_\{g\}a\_\{ig\}\.\(59\)The stochastic net value is

Πψ,γstoch​\(𝒎\)=𝔼⁡\[1n​∑i=1nψk⁡\(𝒎\)​\(Ri​\(𝒎\)\)\]−γ​C​\(𝒎\),Πψ,γstoch​\(𝟎\)=0\.\\Pi\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\bm\{m\}\)=\\mathbb\{E\}\\left\[\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\psi\_\{k\(\\bm\{m\}\)\}\\bigl\(R\_\{i\}\(\\bm\{m\}\)\\bigr\)\\right\]\-\\gamma C\(\\bm\{m\}\),\\qquad\\Pi\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\bm\{0\}\)=0\.\(60\)Recovery is evaluated at the realized correct count and only then averaged\. In general,𝔼⁡\[ψk​\(Ri​\(𝒎\)\)\]≠ψk​\(𝔼⁡\[Ri​\(𝒎\)\]\)\\mathbb\{E\}\[\\psi\_\{k\}\(R\_\{i\}\(\\bm\{m\}\)\)\]\\neq\\psi\_\{k\}\(\\mathbb\{E\}\[R\_\{i\}\(\\bm\{m\}\)\]\), so the deterministic objective is not recovered by substituting expected correctness for realized correctness\. For a finite workflow poolℳ⊆𝒢\\mathcal\{M\}\\subseteq\\mathcal\{G\}, define

OPTψ,γstoch​\(ℳ\)=max𝒎∈ℤ\+ℳk⁡\(𝒎\)≤Kmax⁡Πψ,γstoch​\(𝒎\)\.\\mathrm\{OPT\}\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\mathcal\{M\}\)=\\max\_\{\\begin\{subarray\}\{c\}\\bm\{m\}\\in\\mathbb\{Z\}\_\{\+\}^\{\\mathcal\{M\}\}\\\\ k\(\\bm\{m\}\)\\leq K\_\{\\max\}\\end\{subarray\}\}\\Pi\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\bm\{m\}\)\.
#### Selector\-strength bound\.

The deterministic selector\-strength bound of[Theorem4\.4](https://arxiv.org/html/2609.18126#S4.Thmtheorem4)extends to stochastic execution\. Only the reading ofa¯g\\bar\{a\}\_\{g\}changes: under \([20](https://arxiv.org/html/2609.18126#S4.E20)\) it was the fraction of development tasks workflowggsolves, whereas it is now the average probability that one execution ofggis correct\. Recall thathΛ​\(p\)=Λ​p/\{1\+\(Λ−1\)​p\}h\_\{\\Lambda\}\(p\)=\\Lambda p/\\\{1\+\(\\Lambda\-1\)p\\\}from[Lemma4\.2](https://arxiv.org/html/2609.18126#S4.Thmtheorem2)\.

###### Proposition 7\.1\(Selector odds lift under stochastic execution\)

Suppose[Section4\.1](https://arxiv.org/html/2609.18126#S4.SS1)holds andΛ⁡\(H\)≤Λ<∞\\Lambda\(H\)\\leq\\Lambda<\\infty\. Every feasible nonzero𝐦\\bm\{m\}satisfies

1n​∑i=1n𝔼⁡\[ψk⁡\(𝒎\)​\(Ri​\(𝒎\)\)\]≤hΛ​\(1k⁡\(𝒎\)​∑g∈𝒢mg​a¯g\)\.\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\mathbb\{E\}\\\!\\left\[\\psi\_\{k\(\\bm\{m\}\)\}\\bigl\(R\_\{i\}\(\\bm\{m\}\)\\bigr\)\\right\]\\leq h\_\{\\Lambda\}\\\!\\left\(\\frac\{1\}\{k\(\\bm\{m\}\)\}\\sum\_\{g\\in\\mathcal\{G\}\}m\_\{g\}\\bar\{a\}\_\{g\}\\right\)\.\(61\)For a finite poolℳ\\mathcal\{M\}, leta⋆=maxg∈ℳ⁡a¯ga^\{\\star\}=\\max\_\{g\\in\\mathcal\{M\}\}\\bar\{a\}\_\{g\},cmin=ming∈ℳ⁡cgc\_\{\\min\}=\\min\_\{g\\in\\mathcal\{M\}\}c\_\{g\}, andV1stoch=max⁡\{0,maxg∈ℳ⁡\(a¯g−γ​cg\)\}\.V\_\{1\}^\{\\mathrm\{stoch\}\}=\\max\\\{0,\\max\_\{g\\in\\mathcal\{M\}\}\(\\bar\{a\}\_\{g\}\-\\gamma c\_\{g\}\)\\\}\.Then

V1stoch≤OPTψ,γstoch​\(ℳ\)≤max⁡\{0,max1≤k≤Kmax⁡\[hΛ​\(a⋆\)−γ​k​cmin\]\}\.V\_\{1\}^\{\\mathrm\{stoch\}\}\\leq\\mathrm\{OPT\}\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\mathcal\{M\}\)\\leq\\max\\left\\\{0,\\max\_\{1\\leq k\\leq K\_\{\\max\}\}\\left\[h\_\{\\Lambda\}\(a^\{\\star\}\)\-\\gamma kc\_\{\\min\}\\right\]\\right\\\}\.\(62\)Ifγ=0\\gamma=0, then

a⋆≤OPTψ,0stoch​\(ℳ\)≤hΛ​\(a⋆\),OPTψ,0stoch​\(ℳ\)−a⋆≤Λ−1Λ\+1\.a^\{\\star\}\\leq\\mathrm\{OPT\}\_\{\\psi,0\}^\{\\mathrm\{stoch\}\}\(\\mathcal\{M\}\)\\leq h\_\{\\Lambda\}\(a^\{\\star\}\),\\qquad\\mathrm\{OPT\}\_\{\\psi,0\}^\{\\mathrm\{stoch\}\}\(\\mathcal\{M\}\)\-a^\{\\star\}\\leq\\frac\{\\sqrt\{\\Lambda\}\-1\}\{\\sqrt\{\\Lambda\}\+1\}\.\(63\)

The conclusion is the same as in the deterministic model\. Under the iid law,𝔼⁡\[Ri​\(𝒎\)\]=∑gmg​ai​g\\mathbb\{E\}\[R\_\{i\}\(\\bm\{m\}\)\]=\\sum\_\{g\}m\_\{g\}a\_\{ig\}, and Jensen’s inequality reduces the bound to the same average success probabilitiesa¯g\\bar\{a\}\_\{g\}\. Thus stochastic execution adds no new term to the selector\-strength cap; the additional statistical issue is that the probabilitiesai​ga\_\{ig\}must now be estimated\.

#### Implicit workflow generation\.

The probabilitiesai​ga\_\{ig\}are unknown\. Whenever a workflowggis first made available to the optimizer, we execute it independentlyL1L\_\{1\}times on every development task and define

a^i​g\(1\)=1L1​∑ℓ=1L1Zi​g\(1,ℓ\),Zi​g\(1,ℓ\)∼iidBernoulli⁡\(ai​g\)\.\\widehat\{a\}\_\{ig\}^\{\(1\)\}=\\frac\{1\}\{L\_\{1\}\}\\sum\_\{\\ell=1\}^\{L\_\{1\}\}Z\_\{ig\}^\{\(1,\\ell\)\},\\qquad Z\_\{ig\}^\{\(1,\\ell\)\}\\stackrel\{\{\\scriptstyle\\mathrm\{iid\}\}\}\{\{\\sim\}\}\\mathrm\{Bernoulli\}\(a\_\{ig\}\)\.\(64\)The superscript\(1\)\(1\)marks the Stage\-1 estimation sample\. For the analysis, an independentL1L\_\{1\}\-sample panel may be associated with every workflow in the finite class𝒢\\mathcal\{G\}and revealed only when that workflow is queried\. This lazy\-sampling interpretation does not require the algorithm to enumerate𝒢\\mathcal\{G\}\. For everyδ1∈\(0,1\)\\delta\_\{1\}\\in\(0,1\), Hoeffding’s inequality and a union bound give

ℙ\{maxi=1,…,ng∈𝒢\|a^i​g\(1\)−ai​g\|≤εa\(L1,δ1\)\}≥1−δ1,εa\(L1,δ1\):=log⁡\(2​n​\|𝒢\|/δ1\)2​L1\.\\mathbb\{P\}\\left\\\{\\max\_\{\\begin\{subarray\}\{c\}i=1,\\ldots,n\\\\ g\\in\\mathcal\{G\}\\end\{subarray\}\}\\left\|\\widehat\{a\}\_\{ig\}^\{\(1\)\}\-a\_\{ig\}\\right\|\\leq\\varepsilon\_\{a\}\(L\_\{1\},\\delta\_\{1\}\)\\right\\\}\\geq 1\-\\delta\_\{1\},\\qquad\\varepsilon\_\{a\}\(L\_\{1\},\\delta\_\{1\}\):=\\sqrt\{\\frac\{\\log\(2n\|\\mathcal\{G\}\|/\\delta\_\{1\}\)\}\{2L\_\{1\}\}\}\.\(65\)
Letℙ^1\\widehat\{\\mathbb\{P\}\}\_\{1\}be the product\-Bernoulli execution law obtained by replacingai​ga\_\{ig\}witha^i​g\(1\)\\widehat\{a\}\_\{ig\}^\{\(1\)\}, and let𝔼^1\\widehat\{\\mathbb\{E\}\}\_\{1\}denote expectation under this plug\-in law\. Define

Π^ψ,γ\(1\)​\(𝒎\)=𝔼^1​\[1n​∑i=1nψk⁡\(𝒎\)​\(Ri​\(𝒎\)\)\]−γ​C​\(𝒎\),Π^ψ,γ\(1\)​\(𝟎\)=0\.\\widehat\{\\Pi\}\_\{\\psi,\\gamma\}^\{\(1\)\}\(\\bm\{m\}\)=\\widehat\{\\mathbb\{E\}\}\_\{1\}\\left\[\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\psi\_\{k\(\\bm\{m\}\)\}\\bigl\(R\_\{i\}\(\\bm\{m\}\)\\bigr\)\\right\]\-\\gamma C\(\\bm\{m\}\),\\qquad\\widehat\{\\Pi\}\_\{\\psi,\\gamma\}^\{\(1\)\}\(\\bm\{0\}\)=0\.\(66\)On the concentration event in \([65](https://arxiv.org/html/2609.18126#S7.E65)\), a common\-uniform coupling of the true and plug\-in Bernoulli executions gives

sup𝒎∈ℤ\+𝒢k⁡\(𝒎\)≤Kmax\|Π^ψ,γ\(1\)​\(𝒎\)−Πψ,γstoch​\(𝒎\)\|≤εest​\(L1,δ1\):=min⁡\{1,Kmax​εa​\(L1,δ1\)\}\.\\sup\_\{\\begin\{subarray\}\{c\}\\bm\{m\}\\in\\mathbb\{Z\}\_\{\+\}^\{\\mathcal\{G\}\}\\\\ k\(\\bm\{m\}\)\\leq K\_\{\\max\}\\end\{subarray\}\}\\left\|\\widehat\{\\Pi\}\_\{\\psi,\\gamma\}^\{\(1\)\}\(\\bm\{m\}\)\-\\Pi\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\bm\{m\}\)\\right\|\\leq\\varepsilon\_\{\\mathrm\{est\}\}\(L\_\{1\},\\delta\_\{1\}\):=\\min\\left\\\{1,K\_\{\\max\}\\varepsilon\_\{a\}\(L\_\{1\},\\delta\_\{1\}\)\\right\\\}\.\(67\)
Algorithm[2](https://arxiv.org/html/2609.18126#alg2)in Appendix[F](https://arxiv.org/html/2609.18126#A6)applies the grid, ellipsoid, batched\-pricing, primal\-recovery, and rounding arguments to the plug\-in iid law\. Its workflow\-indexed constraints are separated using the plug\-in stochastic pricing score in \([131](https://arxiv.org/html/2609.18126#A6.E131)\)\. Although the iid execution law is determined by the marginal probabilities, this stochastic reduced cost is not generally the deterministic mean\-column score∑iμi​a^i​g\(1\)−γ​cg\\sum\_\{i\}\\mu\_\{i\}\\widehat\{a\}\_\{ig\}^\{\(1\)\}\-\\gamma c\_\{g\}: the marginal task price depends on the realized number of other correct candidates in the same execution scenario\. Every newly proposed workflow is therefore evaluated using its own freshL1L\_\{1\}\-sample panel before its plug\-in stochastic score is used\.

#### Fresh\-sample evaluation\.

LetℳT\\mathcal\{M\}\_\{T\}be the workflow pool produced by the generation stage\. BecauseℳT\\mathcal\{M\}\_\{T\}was assembled using the Stage\-1 estimates, those same executions are not reused to select the final portfolio\. We freezeℳT\\mathcal\{M\}\_\{T\}and collect a second, independent sample\. Specifically, for every replicationℓ=1,…,L2\\ell=1,\\ldots,L\_\{2\}, taskii, workflowg∈ℳTg\\in\\mathcal\{M\}\_\{T\}, and potential execution copyq=1,…,Kmaxq=1,\\ldots,K\_\{\\max\}, letZi​g​q\(2,ℓ\)∼Bernoulli⁡\(ai​g\)Z\_\{igq\}^\{\(2,\\ell\)\}\\sim\\mathrm\{Bernoulli\}\(a\_\{ig\}\)be mutually independent Stage\-2 draws, independent of the complete generation history\. For a candidate execution\-count vector𝒎\\bm\{m\}, define

Ri\(2,ℓ\)​\(𝒎\)=∑g∈ℳT∑q=1mgZi​g​q\(2,ℓ\)\.R\_\{i\}^\{\(2,\\ell\)\}\(\\bm\{m\}\)=\\sum\_\{g\\in\\mathcal\{M\}\_\{T\}\}\\sum\_\{q=1\}^\{m\_\{g\}\}Z\_\{igq\}^\{\(2,\\ell\)\}\.\(68\)Thus one Stage\-2 replication supplies a common fresh execution table from which every candidate portfolio overℳT\\mathcal\{M\}\_\{T\}can be evaluated\. Set

Π^L2final​\(𝒎\)=1L2​∑ℓ=1L2\[1n​∑i=1nψk⁡\(𝒎\)​\(Ri\(2,ℓ\)​\(𝒎\)\)\]−γ​C​\(𝒎\),Π^L2final​\(𝟎\)=0\.\\widehat\{\\Pi\}\_\{L\_\{2\}\}^\{\\mathrm\{final\}\}\(\\bm\{m\}\)=\\frac\{1\}\{L\_\{2\}\}\\sum\_\{\\ell=1\}^\{L\_\{2\}\}\\left\[\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\psi\_\{k\(\\bm\{m\}\)\}\\bigl\(R\_\{i\}^\{\(2,\\ell\)\}\(\\bm\{m\}\)\\bigr\)\\right\]\-\\gamma C\(\\bm\{m\}\),\\qquad\\widehat\{\\Pi\}\_\{L\_\{2\}\}^\{\\mathrm\{final\}\}\(\\bm\{0\}\)=0\.\(69\)For a cardinality grid𝒦\\mathcal\{K\}, define𝒞𝒦​\(ℳT\)=\{𝟎\}∪\{𝒎∈ℤ\+ℳT:k⁡\(𝒎\)∈𝒦\}\.\\mathcal\{C\}\_\{\\mathcal\{K\}\}\(\\mathcal\{M\}\_\{T\}\)=\\\{\\bm\{0\}\\\}\\cup\\left\\\{\\bm\{m\}\\in\\mathbb\{Z\}\_\{\+\}^\{\\mathcal\{M\}\_\{T\}\}:k\(\\bm\{m\}\)\\in\\mathcal\{K\}\\right\\\}\.LetNTN\_\{T\}denote the size of this class\. Since repeated execution is allowed,

NT=1\+∑k∈𝒦\(\|ℳT\|\+k−1k\),εsamp​\(L2,δ2\):=log⁡\(2​NT/δ2\)2​n​L2\.N\_\{T\}=1\+\\sum\_\{k\\in\\mathcal\{K\}\}\\binom\{\|\\mathcal\{M\}\_\{T\}\|\+k\-1\}\{k\},\\qquad\\varepsilon\_\{\\mathrm\{samp\}\}\(L\_\{2\},\\delta\_\{2\}\):=\\sqrt\{\\frac\{\\log\(2N\_\{T\}/\\delta\_\{2\}\)\}\{2nL\_\{2\}\}\}\.\(70\)Conditional on the frozen generated pool, Hoeffding’s inequality and a union bound imply that, with probability at least1−δ21\-\\delta\_\{2\}, every portfolio in𝒞𝒦​\(ℳT\)\\mathcal\{C\}\_\{\\mathcal\{K\}\}\(\\mathcal\{M\}\_\{T\}\)is evaluated withinεsamp​\(L2,δ2\)\\varepsilon\_\{\\mathrm\{samp\}\}\(L\_\{2\},\\delta\_\{2\}\)of its true stochastic value\.

Let

𝒎~∈arg​max𝒎∈𝒞𝒦​\(ℳT\)⁡Π^L2final​\(𝒎\)\\widetilde\{\\bm\{m\}\}\\in\\operatorname\*\{arg\\,max\}\_\{\\bm\{m\}\\in\\mathcal\{C\}\_\{\\mathcal\{K\}\}\(\\mathcal\{M\}\_\{T\}\)\}\\widehat\{\\Pi\}\_\{L\_\{2\}\}^\{\\mathrm\{final\}\}\(\\bm\{m\}\)be the Stage\-2 empirical maximizer\. Retain a feasible baseline𝒎0∈𝒞𝒦​\(ℳT\)\\bm\{m\}^\{0\}\\in\\mathcal\{C\}\_\{\\mathcal\{K\}\}\(\\mathcal\{M\}\_\{T\}\)whose true stochastic value is known to satisfyΠψ,γstoch​\(𝒎0\)≥V¯\>0\\Pi\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\bm\{m\}^\{0\}\)\\geq\\underline\{V\}\>0, and return

𝒎^=\{𝒎~,Π^L2final​\(𝒎~\)−εsamp​\(L2,δ2\)≥V¯,𝒎0,otherwise\.\\widehat\{\\bm\{m\}\}=\\begin\{cases\}\\widetilde\{\\bm\{m\}\},&\\widehat\{\\Pi\}\_\{L\_\{2\}\}^\{\\mathrm\{final\}\}\(\\widetilde\{\\bm\{m\}\}\)\-\\varepsilon\_\{\\mathrm\{samp\}\}\(L\_\{2\},\\delta\_\{2\}\)\\geq\\underline\{V\},\\\\\[2\.84526pt\] \\bm\{m\}^\{0\},&\\text\{otherwise\}\.\\end\{cases\}\(71\)The rule switches away from the baseline only when the fresh Stage\-2 value clearsV¯\\underline\{V\}by more than the uniform sampling deviation\.

The complete two\-stage procedure is stated as Algorithm[2](https://arxiv.org/html/2609.18126#alg2)in the appendix\. The following theorem records its end\-to\-end guarantee\.

###### Theorem 7\.2\(High\-probability stochastic implicit\-class guarantee\)

Suppose[Section7](https://arxiv.org/html/2609.18126#S7)holds,𝒢\\mathcal\{G\}is a finite implicitly represented workflow\-type class withcg≤cmaxc\_\{g\}\\leq c\_\{\\max\}\. Supposeψk​\(r\)=h⁡\(r/k\)\\psi\_\{k\}\(r\)=h\(r/k\)for a common increasing concave functionh:\[0,1\]→\[0,1\]h:\[0,1\]\\to\[0,1\]withh⁡\(0\)=0h\(0\)=0andh⁡\(1\)=1h\(1\)=1, and let𝒦\\mathcal\{K\}have coverage ratioϱ≥1\\varrho\\geq 1\. Retain a baseline𝐦0\\bm\{m\}^\{0\}satisfyingΠψ,γstoch​\(𝐦0\)≥V¯\>0\\Pi\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\bm\{m\}^\{0\}\)\\geq\\underline\{V\}\>0\.

Run Algorithm[2](https://arxiv.org/html/2609.18126#alg2)with sample sizesL1,L2L\_\{1\},L\_\{2\}, confidence levelsδ1,δ2∈\(0,1\)\\delta\_\{1\},\\delta\_\{2\}\\in\(0,1\), ellipsoid toleranceεell\>0\\varepsilon\_\{\\mathrm\{ell\}\}\>0, and pricing\-batch sizemm\. Suppose every primitive pricing call satisfies \([134](https://arxiv.org/html/2609.18126#A6.E134)\), conditional on the query history, with probability at leastporc\>0p\_\{\\mathrm\{orc\}\}\>0, and letNpricestoch​\(εell\)N\_\{\\mathrm\{price\}\}^\{\\mathrm\{stoch\}\}\(\\varepsilon\_\{\\mathrm\{ell\}\}\)bound the total number of pricing batches\. Defineεrnd:=maxk∈𝒦⁡cψ,ke\.\\varepsilon\_\{\\mathrm\{rnd\}\}:=\\max\_\{k\\in\\mathcal\{K\}\}\\frac\{c\_\{\\psi,k\}\}\{e\}\.Then, with probability at least1−δ1−δ2−Npricestoch​\(εell\)​e−porc​m,1\-\\delta\_\{1\}\-\\delta\_\{2\}\-N\_\{\\mathrm\{price\}\}^\{\\mathrm\{stoch\}\}\(\\varepsilon\_\{\\mathrm\{ell\}\}\)e^\{\-p\_\{\\mathrm\{orc\}\}m\},the returned execution\-count vector satisfies

Πψ,γstoch​\(𝒎^\)≥max⁡\{V¯,1ϱ​OPTψ,γstoch​\(𝒢\)−εell−εrnd−2​εest​\(L1,δ1\)−2​εsamp​\(L2,δ2\)\}\.\\Pi\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\widehat\{\\bm\{m\}\}\)\\geq\\max\\left\\\{\\underline\{V\},\\frac\{1\}\{\\varrho\}\\mathrm\{OPT\}\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\mathcal\{G\}\)\-\\varepsilon\_\{\\mathrm\{ell\}\}\-\\varepsilon\_\{\\mathrm\{rnd\}\}\-2\\varepsilon\_\{\\mathrm\{est\}\}\(L\_\{1\},\\delta\_\{1\}\)\-2\\varepsilon\_\{\\mathrm\{samp\}\}\(L\_\{2\},\\delta\_\{2\}\)\\right\\\}\.\(72\)Consequently,

Πψ,γstoch​\(𝒎^\)≥1ϱ​V¯V¯\+εell\+εrnd\+2​εest​\(L1,δ1\)\+2​εsamp​\(L2,δ2\)​OPTψ,γstoch​\(𝒢\)\.\\Pi\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\widehat\{\\bm\{m\}\}\)\\geq\\frac\{1\}\{\\varrho\}\\frac\{\\underline\{V\}\}\{\\underline\{V\}\+\\varepsilon\_\{\\mathrm\{ell\}\}\+\\varepsilon\_\{\\mathrm\{rnd\}\}\+2\\varepsilon\_\{\\mathrm\{est\}\}\(L\_\{1\},\\delta\_\{1\}\)\+2\\varepsilon\_\{\\mathrm\{samp\}\}\(L\_\{2\},\\delta\_\{2\}\)\}\\mathrm\{OPT\}\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\mathcal\{G\}\)\.\(73\)

## 8Numerical Experiments

The numerical study addresses two questions\. First, how reliably can a post\-output selector identify a correct answer from a mixed candidate set, and how does this ability vary across applications and selector models? We study selector recovery on three domains \(ABCD, Schema\-Guided Dialogue \(SGD\), and HotpotQA\) using three selector model families\. Second, does selector\-aware workflow portfolio design improve deployment performance? We answer these questions using each data set by evaluating the complete pipeline: initial workflow evaluation, selector calibration, cost\-aware portfolio optimization, dual\-guided workflow generation, selector recalibration after the workflow bank expands, and fresh held\-out deployment\.

The three domains differ in the type of decision the system must make and the information available for making it\. ABCD\([Chen et al\. 2021a](https://arxiv.org/html/2609.18126#bib.bib44)\)is a customer\-service routing task with a common operational taxonomy: after observing the first few customer turns, the system must identify the appropriate service subflow\. SGD\([Rastogi et al\. 2020](https://arxiv.org/html/2609.18126#bib.bib45)\)contains task\-oriented conversations between users and virtual assistants across many services and domains\. Each service is accompanied by a schema describing the intents it supports and the information fields relevant to those intents\. The system must therefore interpret a conversation relative to a service\-specific set of possible intents and slots, rather than a single taxonomy shared across all conversations\. HotpotQA\([Yang et al\. 2018](https://arxiv.org/html/2609.18126#bib.bib46)\)is an open\-domain question\-answering task in which answering a question typically requires combining information from multiple passages\. It therefore provides a substantively different setting in which candidate workflows must perform multistep evidence\-based reasoning\. Together, the three domains span fixed\-taxonomy service routing, schema\-dependent dialogue understanding, and multistep question answering\.

#### Common experimental protocol\.

Candidate\-generating workflows are stochastic\. For every task–workflow pair, we execute the workflow independentlyL=5L=5times usingdo\_sample=True, temperature one, and distinct random seeds\. We retain each execution’s output, correctness indicator, seed, and parse status and estimate

a^i​g=15​∑ℓ=15Zi​g​ℓ\.\\widehat\{a\}\_\{ig\}=\\frac\{1\}\{5\}\\sum\_\{\\ell=1\}^\{5\}Z\_\{ig\\ell\}\.Selectors are evaluated deterministically \(do\_sample=False\)\. Candidate order is randomized to reduce position effects\. When one task contributes multiple selector panels or orderings, confidence intervals treat the task, rather than each individual selector decision, as the independent sampling unit\.

In all three end\-to\-end experiments, we setKmax=6K\_\{\\max\}=6and searchK∈\{1,…,6\}K\\in\\\{1,\\ldots,6\\\}\. The primary formulation permits repeated execution of a workflow type, and we also report a no\-repeat benchmark imposingmg≤1m\_\{g\}\\leq 1\. We optimize the LP relaxation using a 16\-scenario sample\-average approximation \(SAA\), where each scenario is drawn from the plug\-in independent Bernoulli execution law defined by the estimated probabilitiesa^i​g\\widehat\{a\}\_\{ig\}\. We then apply 512 randomized\-rounding trials\. For every returned integral portfolio, we compute the distribution of the number of correct outputs exactly: because this count is a sum of independent Bernoulli variables with potentially different probabilities, it follows a Poisson–binomial distribution\. We average the selector recovery curve over this exact distribution rather than reporting the finite 16\-scenario average\.

A correct final decision is valued at $1, recurring compute cost is expressed using the small\-model reference\-call normalization described below, and we setγ=1\\gamma=1\. To conserve space, we report ABCD in detail, summarize the SGD and HotpotQA results in[Section8\.3](https://arxiv.org/html/2609.18126#S8.SS3), and provide their complete experiments in[SectionsG\.1](https://arxiv.org/html/2609.18126#A7.SS1)and[G\.2](https://arxiv.org/html/2609.18126#A7.SS2)in the Appendix\.

### 8\.1ABCD customer\-service routing

#### Task and data\.

ABCD is a customer\-service dialogue dataset organized around operational flows and policy\-constrained service actions\([Chen et al\. 2021a](https://arxiv.org/html/2609.18126#bib.bib44)\)\. We study an early routing task: after observing at most the customer’s first three messages and the available dialogue context, the system must identify the appropriate service subflow from 96 possible labels, including routes associated with order cancellation, refund status, subscription changes, troubleshooting, and escalation\.We preserve the official data splits and the naturally occurring distribution of labels within each split\. The experiment uses 800 development tasks to evaluate workflows and optimize the portfolio, all 1,004 available tasks from the development split to calibrate the selector, and 800 tasks from the official test split for final held\-out evaluation\. No single label accounts for more than 4\.2% of the observations in any split\. Thus, performance cannot be driven by repeatedly predicting a small number of dominant classes; the system must distinguish among a large set of possible routing decisions\.

#### Initial workflow bank and stochastic execution\.

We begin with a deliberately simple bank of 18 single\-call workflows, formed by crossing three generator models \(Qwen2\.5\-3B\-Instruct, Mistral\-7B\-Instruct\-v0\.3, and Granite\-3\.3\-8B\-Instruct\) with six prompting strategies\. The strategies differ in the reasoning process the model is asked to follow before making the routing decision\. The*direct*strategy asks for the routing label immediately\.*Evidence first*asks the model to identify the parts of the conversation that are most informative for the decision before choosing a label\.*Decomposition*asks it to consider the customer’s intent, the current service state, and the routing logic separately before combining them into a final decision\.*Verify and revise*asks the model to make an initial judgment, check it for possible errors, and revise it if needed within the same response\.*Alternatives*asks the model to compare several plausible routing labels before selecting one, while*limited context*deliberately withholds part of the available dialogue information to create a less informed workflow\. Importantly, these differences do not change the workflow topology: every initial workflow consists of a single model call, with no separate verifier, fallback call, or multistep agent interaction\. This provides a controlled starting bank against which we later evaluate whether richer multistep agentic workflows add value\.

We set the reference cost of one 7B\-model call to0\.0010\.001and scale execution cost linearly with model size and the number of model calls\.444As of August 22, 2026, Together AI prices Qwen2\.5\-7B\-Instruct\-Turbo, a 7B\-parameter model, at0\.300\.30per million input tokens and0\.300\.30per million output tokens\. At these rates, a call to this 7B model using a total of 3,333\.33 input and output tokens costs3,333\.33×0\.30106=0\.001\.3\{,\}333\.33\\times\\frac\{0\.30\}\{10^\{6\}\}=0\.001\.Thus, the0\.0010\.001reference cost corresponds directly to the market price of a 3,333\.33\-token call to a hosted 7B model\.The generated workflows instead consist of executable compositions of GPT\-4o\-mini calls connected through Python control flow\. Depending on the generated program, calls may be sequenced or run in parallel and combined through voting, critique, verification, clarification, or conditional fallback\. OpenAI describes GPT\-4o\-mini as a small model and, at its release, characterized it as belonging to roughly the same small\-model tier as models such as Llama 3 8B\. For consistency in experimental cost accounting, we assign each GPT\-4o\-mini invocation one reference small\-model call and the same normalized cost of $0\.001\.

Across the 800 development tasks, the Mistral decomposition workflow has the highest estimated one\-execution accuracy, at 46\.625%\. However, when each of the 18 initial workflows is executed once, the probability that at least one of them produces the correct answer rises to 69\.896%\. This difference shows that the workflows are complementary: many tasks missed by the best individual workflow are solved by another workflow in the pool\.

Table 1:ABCD design and selector calibration\. The primary selector uses execution slots rather than answer deduplication and aggregates cyclic candidate\-order rotations by vote\. Confidence intervals forλ\\lambdaare clustered at the task level\.
#### Precision and stability of theL=5L=5execution design\.

With only five executions for each task\-workflow pair, the individuala^i​g\\widehat\{a\}\_\{ig\}estimates can be noisy\.555For reference, estimating a single Bernoulli success probability with a worst\-case 95% Wilson interval requiresL=93L=93for a±0\.10\\pm 0\.10margin andL=381L=381for a±0\.05\\pm 0\.05margin\.Our goal, however, is not to estimate every task\-workflow success probability precisely\. Rather, we need sufficiently stable estimates of each workflow’s average performance across the 800 development tasks to support portfolio construction\. We therefore examine whether the workflow\-level results are sensitive to using only a few executions\. For each workflow, we recompute its average accuracy using only the firstL′=1,2,3L^\{\\prime\}=1,2,3executions and compare the resulting workflow ranking with the ranking based on allL=5L=5executions\. We measure agreement using Spearman’s rank correlation\([Spearman 1961](https://arxiv.org/html/2609.18126#bib.bib59)\), where values close to one indicate similar rankings\. The correlations are 0\.945, 0\.971, and 0\.958 forL′=1,2,3L^\{\\prime\}=1,2,3, respectively\. Relative to theL=5L=5estimates, the largest absolute change in any workflow’s average accuracy across these comparisons is 1\.25 percentage points, approximately 2\.7% of the best workflow’s 46\.625% average accuracy, and four of the five highest\-ranked workflows remain in the top five\. These results suggest that, although individual task\-workflow estimates remain noisy,L=5L=5provides reasonably stable workflow\-level information for portfolio construction\.

#### Bank\-specific selector calibration\.

We next estimate how well the selector can distinguish correct from incorrect outputs produced by the initial ABCD workflow bank\. For eachK∈\{2,…,6\}K\\in\\\{2,\\ldots,6\\\}and each possible number of correct candidatesr∈\{1,…,K−1\}r\\in\\\{1,\\ldots,K\-1\\\}, we construct 100 candidate panels containing exactlyrrcorrect execution outputs andK−rK\-rincorrect outputs\. The selector observes the task, dialogue context, and candidate labels, but does not observe workflow identity, generator identity, execution cost, or which candidates are correct\. The resulting calibration sample contains 1,500 panels drawn from 586 distinct tasks\.

We estimate selector strength by maximum likelihood under the Plackett–Luce recovery model\. LetSj∈\{0,1\}S\_\{j\}\\in\\\{0,1\\\}indicate whether the selector chooses a correct candidate on calibration paneljj, whererjr\_\{j\}of theKjK\_\{j\}candidates are correct\. Under the Plackett–Luce specification,

logit⁡Pr⁡\(Sj=1∣rj,Kj\)=α\+log⁡\(rjKj−rj\),λ=exp⁡\(α\)\.\\operatorname\{logit\}\\Pr\(S\_\{j\}=1\\mid r\_\{j\},K\_\{j\}\)=\\alpha\+\\log\\\!\\left\(\\frac\{r\_\{j\}\}\{K\_\{j\}\-r\_\{j\}\}\\right\),\\qquad\\lambda=\\exp\(\\alpha\)\.Thus,λ\\lambdameasures how strongly the selector favors correct candidates relative to incorrect ones, and we estimate it by fitting the logistic model above\. The resulting estimate isλ^=2\.9848\\widehat\{\\lambda\}=2\.9848, with a task\-clustered 95% confidence interval of\[2\.4930,3\.5844\]\[2\.4930,3\.5844\]\. Under the fitted model, when exactly one of two candidates is correct, the selector chooses the correct candidate with probability 74\.90%\. Averaged across all panel compositions in the calibration design, its fitted probability of selecting a correct candidate is 21\.33 percentage points higher than under uniform random selection\.

[Figure5](https://arxiv.org/html/2609.18126#S8.F5)compares the observed selector recovery rates with two benchmarks: uniform random selection and the fitted Plackett–Luce curve\. The fitted curve captures the main pattern that selector performance improves as the fraction of correct candidates increases, although the empirical recovery rates do not lie exactly on the one\-parameter model\. We thus use Plackett–Luce as a simple operational approximation for portfolio optimization rather than as an exact behavioral description of the selector\. To check whether this approximation leads to useful deployment decisions, we subsequently evaluate the optimized portfolios using the actual selector on fresh held\-out tasks\. The final deployment results therefore do not depend solely on the fitted parametric recovery model\.

Figure 5:ABCD selector recovery\. Points show task\-clustered estimates for controlled\(K,r\)\(K,r\)cells, vertical bars show 95% task\-clustered bootstrap intervals, the dashed line is uniform selection, and the solid curve is the fitted Plackett–Luce model\. Small horizontal offsets separate cells with the same correct fraction\.
#### Initial\-pool optimization\.

We next test whether the complementarity in the initial bank can be converted into deployed value after accounting for selector confusion and compute cost\. We setKmax=6K\_\{\\max\}=6, matching the largest candidate set used in selector calibration\. For everyK∈\{1,…,6\}K\\in\\\{1,\\ldots,6\\\}, we solve the repeat\-allowed stochastic LP relaxation using 16 independent execution\-outcome scenarios drawn from the plug\-in Bernoulli law\. We then apply randomized rounding 512 times and evaluate each distinct integral portfolio using the exact Poisson–binomial calculation described above\. The best portfolio from the initial workflow bank hasK=6K=6and contains one Granite verify\-and\-revise execution, one Mistral decomposition execution, two Mistral verify\-and\-revise executions, one Qwen evidence\-first execution, and one Qwen verify\-and\-revise execution\. On the development sample, its calibrated selector accuracy is 52\.523%, and its workflow\-cost\-adjusted value is 0\.520074 per task\. Thus, the optimized portfolio assigns two execution slots to the Mistral verify\-and\-revise workflow\. When repeated execution is prohibited, the best no\-repeat portfolio has a slightly lower development value of 0\.518755\.

#### Dual\-guided workflow generation\.

We use ADAS Meta Agent Search\([Hu et al\. 2025](https://arxiv.org/html/2609.18126#bib.bib15)\)to generate new workflows, with OpenAI’s GPT\-4o\-mini serving both as the meta\-agent that proposes workflow structures and as the execution model within the generated workflows\. The search is guided by task prices from the stochastic dual\. These prices identify development tasks on which an additional correct workflow output would be most valuable, allowing ADAS to focus its search on weaknesses of the current workflow bank\. To implement this idea, for eachKKwe run the ellipsoid separation procedure associated with[Algorithm2](https://arxiv.org/html/2609.18126#alg2), using a one\-scenario sample\-average approximation \(SAA\) of the stochastic projected dual\. Because the stochastic dual assigns a separate slot priceθq\\theta\_\{q\}to each labeled execution copyq=1,…,Kq=1,\\ldots,K, workflow pricing is checked separately for each copy\. At a given ellipsoid iterate, we first check the explicit dual constraints and the workflow constraints corresponding to workflows already evaluated\. ADAS is invoked only if these constraints do not separate the current iterate\. Whenever ADAS proposes a new workflow, we freeze its design and evaluate it independently using the sameL=5L=5execution protocol before using its estimated performance in subsequent pricing calculations\.

We impose a prespecified budget of 20 distinct new workflow evaluations during generation\. After the generation stage, the final finite\-pool portfolio optimization uses 16 SAA scenarios and 512 randomized\-rounding trials\. Across the fixed\-KKsearches, the procedure makes 26 ADAS pricing queries and incorporates four generated workflows, expanding the evaluated bank from 18 to 22 workflow types\. Because each fixed\-KKsearch is subject to its prespecified workflow\-evaluation budget, the resulting expanded bank should be viewed as the output of a budgeted dual\-guided search rather than an exhaustive search over the full workflow class\.

#### Expanded\-bank recalibration and reoptimization\.

After adding the generated workflows, we recalibrate the same Qwen selector using a new controlled panel constructed from outputs in the expanded workflow bank\. This step allows the estimated selector model to reflect the new kinds of candidate outputs it will encounter after workflow generation\. The estimated selector strength isλ^=2\.5229\\widehat\{\\lambda\}=2\.5229, with a 95% confidence interval of\[2\.0918,3\.0670\]\[2\.0918,3\.0670\]\. This interval overlaps the initial\-bank confidence interval, although the point estimate is lower\. The comparison illustrates why we recalibrate after expanding the workflow bank:λ\\lambdasummarizes selector performance for the candidate distribution it faces and should not be treated as a fixed intrinsic property of the selector\.

We then reoptimize the portfolio over all 22 workflow types\. The best repeat\-allowed portfolio hasK=4K=4, whereKKcounts complete workflow executions, or equivalently the four candidate outputs ultimately presented to the selector\. Two of these four slots are assigned to generated workflow G4 \(FinalABCDRouterWithLabel\) in[Table2](https://arxiv.org/html/2609.18126#S8.T2)\. A single execution of G4 makes three independently seeded GPT\-4o\-mini routing calls on the same ABCD task and returns the label receiving the most votes\. Assigning G4 to two slots therefore means running this entire three\-call workflow twice independently, producing two candidate outputs\. The remaining two slots contain one initial\-bank Mistral verify\-and\-revise execution and one Qwen evidence\-first execution\. Thus, theK=4K=4portfolio presents four candidate outputs to the final selector but uses eight underlying model calls per task: six from the two G4 executions and one from each of the two initial\-bank workflows\.

On the development sample, this portfolio has calibrated selector accuracy of 54\.190% and a workflow\-cost\-adjusted value of 0\.534410\. The no\-repeat benchmark also selectsK=4K=4, but replaces the second copy of G4 with G3 \(Final\_Single\_Routing\_Agent\), a generated router with critique and fallback\. Its workflow\-cost\-adjusted value is 0\.533396\.

#### Composition of the generated workflows\.

[Table2](https://arxiv.org/html/2609.18126#S8.T2)summarizes the four workflows generated by ADAS and added to the evaluated workflow bank\. Every model call within these workflows uses GPT\-4o\-mini\. Thus, the generated workflows differ not in the underlying model, but in how they organize multiple calls and combine their outputs, using mechanisms such as parallel sampling, voting, critique, clarification, and conditional fallback\. The repeat\-allowed optimal portfolio assigns two execution slots toFinalABCDRouterWithLabel\. Each deployed task therefore receives two independent runs of this three\-call voting workflow, together with one Mistral and one Qwen workflow from the initial bank\. In the no\-repeat benchmark, the second copy is replaced by one execution ofFinal\_Single\_Routing\_Agent\.

Table 2:ADAS\-generated workflows incorporated into the expanded ABCD bank\. All generated model calls use GPT\-4o\-mini through the ADASLLMAgentBaseinterface\. Each GPT\-4o\-mini invocation is assigned one $0\.001 reference\-call unit; mean calls and normalized cost are measured during independent workflow evaluation\.*Notes\.*G1–G4 correspond respectively to the frozen ADAS programsFinalEnhancedRoutingAgent,FinalRoutingAgent,Final\_Single\_Routing\_Agent, andFinalABCDRouterWithLabel\. Cost is recurring normalized execution cost per workflow run\.

#### Fresh held\-out deployment\.

We freeze all reported portfolios before examining the 800 held\-out tasks\. For each workflow type appearing in a reported portfolio, we collect five new stochastic executions per task\. We first evaluate each portfolio under the plug\-in Plackett–Luce model, using the estimated workflow success probabilities and selector strength\. We then evaluate actual deployment performance by using Qwen2\.5\-7B\-Instruct as the final selector\. The selector does not observe workflow identity or correctness\. To reduce sensitivity to candidate position, we present each candidate set under up to six cyclic rotations of the candidate order and use the candidate receiving the most selection votes across rotations\.

### 8\.2Results on ABCD

[Tables3](https://arxiv.org/html/2609.18126#S8.T3)and[6](https://arxiv.org/html/2609.18126#S8.F6)report deployment performance on 800 fresh held\-out ABCD tasks\. Actual selector accuracy increases from 43\.625% for the best initial singleton to 46\.750% for the optimized initial\-bank portfolio and to 50\.250% for the portfolio obtained after ellipsoid\-guided workflow generation\. Thus, optimizing the initial workflow bank contributes 3\.125 percentage points, and expanding the bank through workflow generation contributes an additional 3\.500 percentage points\. Relative to the initial singleton, the complete procedure gains 6\.625 percentage points, or 15\.2%\. The same ordering appears under the calibrated Plackett–Luce model, where predicted selector accuracy rises from 43\.550% to 47\.269% and then to 50\.648%\. Notably, the final portfolio uses only four execution slots, compared with six for the optimized initial\-bank portfolio\. The generated workflows therefore allow the system to use fewer workflow executions while achieving higher held\-out accuracy\. The no\-repeat benchmark reaches 50\.125%, only 0\.125 percentage points below the primary repeat\-allowed solution\.

The improvement does not come simply from increasing the probability that at least one candidate is correct\. That probability, which we refer to as oracle coverage, decreases modestly from 59\.517% for the optimized initial\-bank portfolio to 58\.264% for the expanded\-bank portfolio\. At the same time, the accuracy obtained by selecting uniformly at random from the realized candidate outputs rises from 40\.533% to 45\.681%\. Thus, the generated workflows produce candidate sets in which correct outputs are more prevalent, even though the chance of having at least one correct output is slightly lower\. After subtracting both workflow execution cost and the realized cost of the selector, actual total\-system net value increases from 0\.435207 for the singleton to 0\.455829 for the optimized initial\-bank portfolio and 0\.490671 for the final ADAS\-expanded portfolio\.

Table 3:Fresh held\-out ABCD deployment on 800 tasks\. PL accuracy is the plug\-in Plackett–Luce prediction; actual accuracy uses the blind deterministic Qwen selector\. Brackets report 95% Wilson intervals\. Workflow and selector costs are per task, and actual total\-system net subtracts both\.Figure 6:Fresh held\-out ABCD accuracy\. Bars report actual deterministic selector accuracy, error bars show 95% Wilson intervals over 800 tasks, and diamonds report the corresponding plug\-in Plackett–Luce predictions\.
### 8\.3Results on SGD and HotpotQA

The same end\-to\-end pipeline produces different sources of value across the other two domains\. On SGD, the initial workflow bank is already strong and complementary: optimizing it raises actual held\-out selector accuracy from 85\.250% for the best singleton to 92\.750%\. None of the 20 evaluated ADAS candidates is incorporated at the queried dual prices, so the final portfolio coincides with the optimized initial\-bank portfolio\. Thus, SGD demonstrates the value of portfolio optimization without requiring workflow generation\.

On HotpotQA, optimizing the initial bank raises actual held\-out accuracy only from 30\.250% to 31\.125%, but dual\-guided generation incorporates a three\-call evidence\-extraction, verification, and multihop\-solving workflow\. Two independent executions of this workflow attain 55\.250% actual accuracy; one execution attains 54\.875% and has slightly higher total\-system net value after selector cost is included\. Together with ABCD, these results show that portfolio optimization, workflow generation, and multiplicity are distinct sources of value whose importance varies across applications\. Complete experimental designs, calibration results, portfolio compositions, and held\-out evaluations are reported in[SectionsG\.1](https://arxiv.org/html/2609.18126#A7.SS1)and[G\.2](https://arxiv.org/html/2609.18126#A7.SS2)\.

## 9Conclusion

Agentic AI systems often approach the same task through multiple workflows and then rely on a selector to choose the final answer\. This paper shows that workflow variety creates both opportunity and risk\. While additional executions may solve cases that the current portfolio misses, they also consume compute and may introduce plausible wrong answers that make final selection more difficult\. We develop a framework for managing this tradeoff\. Our selector strength bound places a sharp limit on the value of workflow variety and identifies conditions under which a single workflow or no deployment is optimal\. We then develop optimization methods for choosing the run size and allocating execution slots across workflow types\. For large implicit workflow classes, dual prices guide a workflow generation oracle toward residual tasks on which an additional correct output would be most valuable, while the ellipsoid method provides performance guarantees without requiring the workflow class to be enumerated\.

The numerical results illustrate both the value and the limits of workflow portfolios\. Relative to the best initial singleton, the final portfolio raises actual held\-out selector accuracy by 6\.625 percentage points on ABCD, 7\.500 points on SGD, and 25\.000 points on HotpotQA\. Workflow generation adds value on ABCD and HotpotQA but not on SGD, where optimizing the already strong initial bank is sufficient\. Multiplicity likewise has heterogeneous value: it binds in ABCD and HotpotQA but not in SGD, and its incremental accuracy benefit need not justify its additional compute and selection cost\. The central message is that more inference time variety is not automatically better\. Effective deployment requires firms to coordinate workflow generation, run size, execution allocation, and compute expenditure with the recovery capabilities of the selector\.

## References

- Abdolmalekiet al\.\(2026\)M\. Abdolmaleki, I\. Duenyas, and R\. KapuscinskiPricing delayed agentic AI services: when high\-value jobs wait longer\.Note:Available at SSRN 6721378External Links:[Document](https://dx.doi.org/10.2139/ssrn.6721378),[Link](https://ssrn.com/abstract=6721378)Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px6.p1.1)\.
- Abdolmaleki and Duenyas \(2026\)M\. Abdolmaleki and I\. DuenyasHold back, team up, or take over? optimal timing of AI access\.Note:Available at SSRN 6953340External Links:[Document](https://dx.doi.org/10.2139/ssrn.6953340),[Link](https://ssrn.com/abstract=6953340)Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px6.p1.1)\.
- Ageev and Sviridenko \(2004\)A\. A\. Ageev and M\. I\. SviridenkoPipage Rounding: A New Method of Constructing Algorithms with Proven Performance Guarantee\.Journal of Combinatorial Optimization8\(3\),pp\. 307–328\.Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px5.p1.1),[§5](https://arxiv.org/html/2609.18126#S5.p2.2)\.
- Agrawalet al\.\(2014\)S\. Agrawal, Z\. Wang, and Y\. YeA dynamic near\-optimal algorithm for online linear programming\.Operations Research62\(4\),pp\. 876–890\.Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px2.p1.1)\.
- Alaeiet al\.\(2010\)S\. Alaei, A\. Makhdoumi, and A\. MalekianMaximizing sequence\-submodular functions and its application to online advertising\.arXiv preprint arXiv:1009\.4153\.Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1)\.
- Baeket al\.\(2026\)J\. Baek, Y\. Fu, W\. Ma, and T\. PengAI agents for inventory control: human\-llm\-or complementarity\.arXiv preprint arXiv:2602\.12631\.Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px6.p1.1)\.
- Barmanet al\.\(2021\)S\. Barman, O\. Fawzi, and P\. FerméTight Approximation Guarantees for Concave Coverage Problems\.In38th International Symposium on Theoretical Aspects of Computer Science \(STACS 2021\),Leibniz International Proceedings in Informatics \(LIPIcs\), Vol\.187,pp\. 9:1–9:17\.Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px5.p1.1),[§5](https://arxiv.org/html/2609.18126#S5.p2.2)\.
- Bertsimas and Mišić \(2019\)D\. Bertsimas and V\. V\. MišićExact first\-choice product line optimization\.Operations Research67\(3\),pp\. 651–670\.Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1)\.
- Calinescuet al\.\(2011\)G\. Calinescu, C\. Chekuri, M\. Pál, and J\. VondrákMaximizing a Monotone Submodular Function Subject to a Matroid Constraint\.SIAM Journal on Computing40\(6\),pp\. 1740–1766\.Cited by:[§5](https://arxiv.org/html/2609.18126#S5.p2.2)\.
- Chekuriet al\.\(2010\)C\. Chekuri, J\. Vondrák, and R\. ZenklusenDependent Randomized Rounding via Exchange Properties of Combinatorial Structures\.InProceedings of the 51st Annual IEEE Symposium on Foundations of Computer Science,pp\. 575–584\.Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px5.p1.1),[§5](https://arxiv.org/html/2609.18126#S5.p2.2)\.
- Chenet al\.\(2021a\)D\. Chen, H\. Chen, Y\. Yang, A\. Lin, and Z\. YuAction\-Based Conversations Dataset: A Corpus for Building More In\-Depth Task\-Oriented Dialogue Systems\.External Links:2104\.00783,[Link](https://arxiv.org/abs/2104.00783)Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p13.1.1),[§8\.1](https://arxiv.org/html/2609.18126#S8.SS1.SSS0.Px1.p1.1),[§8](https://arxiv.org/html/2609.18126#S8.p2.1)\.
- Chen and Chua \(2024\)D\. Chen and G\. A\. ChuaNoisy dual mirror descent: a near optimal algorithm for jointly\-dp convex resource allocation\.Advances in Neural Information Processing Systems37,pp\. 91723–91755\.Cited by:[footnote 2](https://arxiv.org/html/2609.18126#footnote2)\.
- Chenet al\.\(2025\)L\. Chen, J\. Q\. Davis, B\. Hanin, P\. Bailis, M\. Zaharia, J\. Zou, and I\. StoicaOptimizing Model Selection for Compound AI Systems\.External Links:2502\.14815,[Link](https://arxiv.org/abs/2502.14815)Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px1.p1.1)\.
- Chenet al\.\(2024\)L\. Chen, M\. Zaharia, and J\. ZouFrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance\.Transactions on Machine Learning Research\.Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px3.p1.1)\.
- Chenet al\.\(2019\)Q\. Chen, S\. Jasin, and I\. DuenyasNonparametric Self\-Adjusting Control for Joint Learning and Optimization of Multiproduct Pricing with Finite Resource Capacity\.Mathematics of Operations Research44\(2\),pp\. 601–631\.Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px2.p1.1)\.
- Chenet al\.\(2021b\)Q\. Chen, S\. Jasin, and I\. DuenyasTechnical Note—Joint Learning and Optimization of Multi\-Product Pricing with Finite Resource Capacity and Unknown Demand Parameters\.Operations Research69\(2\),pp\. 560–573\.Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px2.p1.1)\.
- Chenget al\.\(2010\)W\. Cheng, K\. Dembczyński, and E\. HüllermeierLabel Ranking Methods Based on the Plackett–Luce Model\.InProceedings of the 27th International Conference on Machine Learning,pp\. 215–222\.Cited by:[§4\.3](https://arxiv.org/html/2609.18126#S4.SS3.p1.1)\.
- Dantzig and Wolfe \(1960\)G\. B\. Dantzig and P\. WolfeDecomposition Principle for Linear Programs\.Operations Research8\(1\),pp\. 101–111\.Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1)\.
- Deanet al\.\(2026\)A\. Dean, Z\. Zhang, S\. Jasin, and Y\. LiuMulti\-LLM Query Optimization\.External Links:2603\.24617,[Link](https://arxiv.org/abs/2603.24617)Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px6.p1.1)\.
- G\. Desaulniers, J\. Desrosiers, and M\. M\. Solomon \(Eds\.\) \(2005\)G\. Desaulniers, J\. Desrosiers, and M\. M\. Solomon \(Eds\.\)Column Generation\.Springer\.Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1)\.
- Donget al\.\(2023\)J\. Dong, W\. Mo, Z\. Qi, C\. Shi, E\. X\. Fang, and V\. TarokhPasta: pessimistic assortment optimization\.InInternational Conference on Machine Learning,pp\. 8276–8295\.Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1)\.
- Echeniqueet al\.\(2025\)F\. Echenique, A\. Fallah, and M\. I\. JordanA general framework for estimating preferences using response time data\.arXiv preprint arXiv:2507\.20403\.Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px5.p1.1)\.
- Elmachtoub and Grigas \(2022\)A\. N\. Elmachtoub and P\. GrigasSmart “predict, then optimize”\.Management science68\(1\),pp\. 9–26\.Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px2.p1.1)\.
- Faradonbeh and Faradonbeh \(2023\)M\. K\. S\. Faradonbeh and M\. S\. S\. FaradonbehOnline reinforcement learning in stochastic continuous\-time systems\.InThe Thirty Sixth Annual Conference on Learning Theory,pp\. 612–656\.Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px2.p1.1)\.
- Feige \(1998\)U\. FeigeA Threshold ofln⁡n\\ln nfor Approximating Set Cover\.Journal of the ACM45\(4\),pp\. 634–652\.Cited by:[§5\.5](https://arxiv.org/html/2609.18126#S5.SS5.p1.1)\.
- Gartner \(2025a\)GartnerGartner Predicts 40% of Enterprise Apps Will Feature Task\-Specific AI Agents by 2026, Up from Less Than 5% in 2025\.Note:Gartner NewsroomAccessed: 2026\-06\-26External Links:[Link](https://www.gartner.com/en/newsroom/press-releases/2025-08-26-gartner-predicts-40-percent-of-enterprise-apps-will-feature-task-specific-ai-agents-by-2026-up-from-less-than-5-percent-in-2025)Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p1.1)\.
- Gartner \(2025b\)GartnerGartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027\.Note:Gartner NewsroomAccessed: 2026\-06\-26External Links:[Link](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027)Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p1.1)\.
- Grötschelet al\.\(1981\)M\. Grötschel, L\. Lovász, and A\. SchrijverThe ellipsoid method and its consequences in combinatorial optimization\.Combinatorica1\(2\),pp\. 169–197\.External Links:[Document](https://dx.doi.org/10.1007/BF02579273)Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px5.p1.1),[§6](https://arxiv.org/html/2609.18126#S6.p1.1),[footnote 3](https://arxiv.org/html/2609.18126#footnote3)\.
- Grötschelet al\.\(1988\)M\. Grötschel, L\. Lovász, and A\. SchrijverGeometric algorithms and combinatorial optimization\.Algorithms and Combinatorics, Vol\.2,Springer\-Verlag,Berlin\.External Links:[Document](https://dx.doi.org/10.1007/978-3-642-97881-4)Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px5.p1.1),[§6](https://arxiv.org/html/2609.18126#S6.p1.1),[footnote 3](https://arxiv.org/html/2609.18126#footnote3)\.
- Guiver and Snelson \(2009\)J\. Guiver and E\. SnelsonBayesian Inference for Plackett–Luce Ranking Models\.InProceedings of the 26th Annual International Conference on Machine Learning,pp\. 377–384\.External Links:[Document](https://dx.doi.org/10.1145/1553374.1553423)Cited by:[§4\.3](https://arxiv.org/html/2609.18126#S4.SS3.p1.1)\.
- Guoet al\.\(2026\)Y\. Guo, S\. Jasin, and C\. LinCongestion\-Aware Static LLM Cascades: Analysis of a Steady\-State Framework\.Note:Available at SSRNSSRN 6759238Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px6.p1.1)\.
- Hajeket al\.\(2014\)B\. Hajek, S\. Oh, and J\. XuMinimax\-Optimal Inference from Partial Rankings\.InAdvances in Neural Information Processing Systems,Vol\.27,pp\. 1475–1483\.Cited by:[§4\.3](https://arxiv.org/html/2609.18126#S4.SS3.p1.1)\.
- Harvard Business Review Analytic Services \(2026\)Harvard Business Review Analytic ServicesFrom the Edge to the Core: Bringing Agentic AI to the Heart of the Enterprise\.Note:Harvard Business Review Sponsored ContentAccessed: 2026\-06\-26External Links:[Link](https://hbr.org/sponsored/2026/01/from-the-edge-to-the-core)Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p1.1)\.
- Huet al\.\(2024\)Q\. J\. Hu, J\. Bieker, X\. Li, N\. Jiang, B\. Keigwin, G\. Ranganath, K\. Keutzer, and S\. UpadhyayRouterBench: A Benchmark for Multi\-LLM Routing System\.External Links:2403\.12031,[Link](https://arxiv.org/abs/2403.12031)Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px3.p1.1)\.
- Huet al\.\(2025\)S\. Hu, C\. Lu, and J\. CluneAutomated Design of Agentic Systems\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p4.1),[§1](https://arxiv.org/html/2609.18126#S1.p7.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px1.p1.1),[§8\.1](https://arxiv.org/html/2609.18126#S8.SS1.SSS0.Px6.p1.1)\.
- Huanget al\.\(2025\)K\. Huang, Y\. Shi, D\. Ding, Y\. Li, Y\. Fei, L\. V\. S\. Lakshmanan, and X\. XiaoThriftLLM: On Cost\-Effective Selection of Large Language Models for Classification Queries\.Proceedings of the VLDB Endowment18\(11\),pp\. 4410–4423\.Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px3.p1.1)\.
- Jasin \(2014\)S\. JasinReoptimization and Self\-Adjusting Price Control for Network Revenue Management\.Operations Research62\(5\),pp\. 1168–1178\.Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px2.p1.1)\.
- Jianget al\.\(2023\)D\. Jiang, X\. Ren, and B\. Y\. LinLLM\-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics,Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px4.p1.1)\.
- Keskin and Zeevi \(2014\)N\. B\. Keskin and A\. ZeeviDynamic pricing with an unknown demand model: asymptotically optimal semi\-myopic policies\.Operations research62\(5\),pp\. 1142–1167\.Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px2.p1.1)\.
- Kleywegtet al\.\(2002\)A\. J\. Kleywegt, A\. Shapiro, and T\. Homem\-de\-MelloThe Sample Average Approximation Method for Stochastic Discrete Optimization\.SIAM Journal on Optimization12\(2\),pp\. 479–502\.Cited by:[Appendix E](https://arxiv.org/html/2609.18126#A5.SS0.SSS0.Px4.p1.2)\.
- Liet al\.\(2026\)G\. Li, J\. Liang, M\. Liu, Y\. Lei, S\. Jasin, F\. Yang, and P\. BaxiAsymptotically Optimal Sequential Testing with Heterogeneous LLMs\.External Links:2604\.01086,[Link](https://arxiv.org/abs/2604.01086)Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px3.p1.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px6.p1.1)\.
- Lianget al\.\(2024\)Y\. Liang, C\. Wu, T\. Song, W\. Wu, Y\. Xia, Y\. Liu, Y\. Ou, S\. Lu, L\. Ji, S\. Mao, Y\. Wang, L\. Shou, M\. Gong, and N\. DuanTaskMatrix\.AI: completing tasks by connecting foundation models with millions of APIs\.Intelligent Computing3,pp\. 0063\.External Links:[Document](https://dx.doi.org/10.34133/icomputing.0063)Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p4.1)\.
- Liuet al\.\(2026a\)F\. Liu, Y\. Liu, Q\. Zhang, X\. Tong, and M\. YuanEoH\-S: Evolution of Heuristic Set Using LLMs for Automated Heuristic Design\.Proceedings of the AAAI Conference on Artificial Intelligence40\(43\),pp\. 37090–37098\.Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px2.p1.1)\.
- Liuet al\.\(2026b\)F\. Liu, Y\. Yao, P\. Guo, Z\. Yang, Z\. Zhao, X\. Lin, X\. Tong, M\. Yuan, Z\. Lu, Z\. Wang, and Q\. ZhangA Systematic Survey on Large Language Models for Algorithm Design\.ACM Computing Surveys\.Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px2.p1.1)\.
- Luet al\.\(2023\)P\. Lu, B\. Peng, H\. Cheng, M\. Galley, K\. Chang, Y\. N\. Wu, S\. Zhu, and J\. GaoChameleon: plug\-and\-play compositional reasoning with large language models\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 43447–43478\.Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p4.1)\.
- Luce \(1959\)R\. D\. LuceIndividual Choice Behavior: A Theoretical Analysis\.Wiley,New York\.Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px5.p1.1),[§4\.3](https://arxiv.org/html/2609.18126#S4.SS3.p1.1)\.
- McFadden \(1974\)D\. McFaddenConditional Logit Analysis of Qualitative Choice Behavior\.InFrontiers in Econometrics,P\. Zarembka \(Ed\.\),pp\. 105–142\.Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px5.p1.1)\.
- McKinsey & Company \(2025\)McKinsey & CompanyThe State of AI: Global Survey 2025\.Note:McKinsey QuantumBlackAccessed: 2026\-06\-26External Links:[Link](https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai)Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p1.1)\.
- McKinsey & Company \(2026\)McKinsey & CompanyBuilding the Foundations for Agentic AI at Scale\.Note:McKinsey TechnologyAccessed: 2026\-06\-26External Links:[Link](https://www.mckinsey.com/capabilities/mckinsey-technology/our-insights/building-the-foundations-for-agentic-ai-at-scale)Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p1.1)\.
- Nemhauseret al\.\(1978\)G\. L\. Nemhauser, L\. A\. Wolsey, and M\. L\. FisherAn Analysis of Approximations for Maximizing Submodular Set Functions—I\.Mathematical Programming14,pp\. 265–294\.Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px5.p1.1),[§5\.5](https://arxiv.org/html/2609.18126#S5.SS5.p1.1)\.
- Novikovet al\.\(2025\)A\. Novikov, N\. Vũ, M\. Eisenberger, E\. Dupont, P\. Huang, A\. Z\. Wagner, S\. Shirobokov, B\. Kozlovskii, F\. J\. R\. Ruiz, A\. Mehrabian, M\. P\. Kumar, A\. See, S\. Chaudhuri, G\. Holland, A\. Davies, S\. Nowozin, P\. Kohli, and M\. BalogAlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery\.External Links:2506\.13131,[Link](https://arxiv.org/abs/2506.13131)Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px2.p1.1)\.
- Onget al\.\(2025\)I\. Ong, A\. Almahairi, V\. Wu, W\. Chiang, T\. Wu, J\. E\. Gonzalez, M\. W\. Kadous, and I\. StoicaRouteLLM: Learning to Route LLMs with Preference Data\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px3.p1.1)\.
- Plackett \(1975\)R\. L\. PlackettThe Analysis of Permutations\.Journal of the Royal Statistical Society: Series C24\(2\),pp\. 193–202\.Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px5.p1.1),[§4\.3](https://arxiv.org/html/2609.18126#S4.SS3.p1.1)\.
- Radvandet al\.\(2026\)T\. Radvand, M\. Abdolmaleki, M\. Mostagir, and A\. TewariA training\-free method for LLM text attribution\.Note:Available at SSRN 6169527; also available as arXiv:2501\.02406External Links:[Document](https://dx.doi.org/10.2139/ssrn.6169527),[Link](https://ssrn.com/abstract=6169527)Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px6.p1.1)\.
- Ransbothamet al\.\(2025\)S\. Ransbotham, D\. Kiron, S\. Khodabandeh, S\. Iyer, and A\. DasThe Emerging Agentic Enterprise: How Leaders Must Navigate a New Age of AI\.Note:MIT Sloan Management Review and Boston Consulting GroupAccessed: 2026\-06\-26External Links:[Link](https://shop.sloanreview.mit.edu/the-emerging-agentic-enterprise-how-leaders-must-navigate-a-new-age-of-ai)Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p1.1)\.
- Rastogiet al\.\(2020\)A\. Rastogi, X\. Zang, S\. Sunkara, R\. Gupta, and P\. KhaitanTowards Scalable Multi\-Domain Conversational Agents: The Schema\-Guided Dialogue Dataset\.InProceedings of the AAAI Conference on Artificial Intelligence,Vol\.34,pp\. 8689–8696\.Cited by:[§G\.1](https://arxiv.org/html/2609.18126#A7.SS1.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.18126#S1.p13.1.1),[§8](https://arxiv.org/html/2609.18126#S8.p2.1)\.
- Saad\-Falconet al\.\(2025\)J\. Saad\-Falcon, A\. G\. Lafuente, S\. Natarajan, N\. Maru, H\. Todorov, E\. Guha, E\. K\. Buchanan, M\. Chen, N\. Guha, C\. Ré, and A\. MirhoseiniArchon: An Architecture Search Framework for Inference\-Time Techniques\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px4.p1.1)\.
- Shenet al\.\(2023\)Y\. Shen, K\. Song, X\. Tan, D\. Li, W\. Lu, and Y\. ZhuangHuggingGPT: solving AI tasks with ChatGPT and its friends in Hugging Face\.InAdvances in Neural Information Processing Systems,Vol\.36,pp\. 38154–38180\.Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p4.1)\.
- Spearman \(1961\)C\. SpearmanThe proof and measurement of association between two things\.\.Cited by:[§8\.1](https://arxiv.org/html/2609.18126#S8.SS1.SSS0.Px3.p1.1)\.
- Szörényiet al\.\(2015\)B\. Szörényi, R\. Busa\-Fekete, A\. Paul, and E\. HüllermeierOnline Rank Elicitation for Plackett–Luce: A Dueling Bandits Approach\.InAdvances in Neural Information Processing Systems,Vol\.28,pp\. 604–612\.Cited by:[§4\.3](https://arxiv.org/html/2609.18126#S4.SS3.p1.1)\.
- Talluri and van Ryzin \(2004\)K\. T\. Talluri and G\. J\. van RyzinThe Theory and Practice of Revenue Management\.Springer\.Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px5.p1.1)\.
- van Ryzin and Mahajan \(1999\)G\. van Ryzin and S\. MahajanOn the Relationship Between Inventory Costs and Variety Benefits in Retail Assortments\.Management Science45\(11\),pp\. 1496–1509\.Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px5.p1.1)\.
- Wanget al\.\(2025\)J\. Wang, J\. Wang, B\. Athiwaratkun, C\. Zhang, and J\. ZouMixture\-of\-Agents Enhances Large Language Model Capabilities\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px4.p1.1)\.
- Wanget al\.\(2023\)X\. Wang, J\. Wei, D\. Schuurmans, Q\. Le, E\. H\. Chi, S\. Narang, A\. Chowdhery, and D\. ZhouSelf\-Consistency Improves Chain of Thought Reasoning in Language Models\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px4.p1.1)\.
- Wanget al\.\(2024\)Z\. Wang, H\. Peura, and W\. WiesemannRandomized assortment optimization\.Operations Research72\(5\),pp\. 2042–2060\.Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1)\.
- Yanget al\.\(2018\)Z\. Yang, P\. Qi, S\. Zhang, Y\. Bengio, W\. W\. Cohen, R\. Salakhutdinov, and C\. D\. ManningHotpotQA: A Dataset for Diverse, Explainable Multi\-Hop Question Answering\.InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing,pp\. 2369–2380\.Cited by:[§G\.2](https://arxiv.org/html/2609.18126#A7.SS2.SSS0.Px1.p1.1),[§1](https://arxiv.org/html/2609.18126#S1.p13.1.1),[§8](https://arxiv.org/html/2609.18126#S8.p2.1)\.
- Zhanget al\.\(2025a\)G\. Zhang, L\. Niu, J\. Fang, K\. Wang, L\. Bai, and X\. WangMulti\-Agent Architecture Search via Agentic Supernet\.External Links:2502\.04180,[Link](https://arxiv.org/abs/2502.04180)Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p4.1),[§1](https://arxiv.org/html/2609.18126#S1.p7.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px1.p1.1)\.
- Zhang and Jasin \(2022\)H\. Zhang and S\. JasinOnline Learning and Optimization of \(Some\) Cyclic Pricing Policies in the Presence of Patient Customers\.Manufacturing & Service Operations Management24\(2\),pp\. 1165–1182\.Cited by:[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px2.p1.1)\.
- Zhanget al\.\(2025b\)J\. Zhang, J\. Xiang, Z\. Yu, F\. Teng, X\. Chen, J\. Chen, M\. Zhuge, X\. Cheng, S\. Hong, J\. Wang, B\. Zheng, B\. Liu, Y\. Luo, and C\. WuAFlow: Automating Agentic Workflow Generation\.InInternational Conference on Learning Representations,Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p4.1),[§1](https://arxiv.org/html/2609.18126#S1.p7.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px1.p1.1)\.
- Zhugeet al\.\(2024\)M\. Zhuge, W\. Wang, L\. Kirsch, F\. Faccio, D\. Khizbullin, and J\. SchmidhuberGPTSwarm: Language Agents as Optimizable Graphs\.InInternational Conference on Machine Learning,Cited by:[§1](https://arxiv.org/html/2609.18126#S1.p7.1),[§2](https://arxiv.org/html/2609.18126#S2.SS0.SSS0.Px1.p1.1)\.

## Appendices for “Managing Agentic AI Workflow Portfolios under Imperfect Selection and Compute Cost\.”

## Appendix AA Graph\-Based Formalization of Workflow Classes

This appendix gives one formal representation of the implicit workflow class𝒢\\mathcal\{G\}\. The main analysis requires only that each workflowggcan be executed on a task, produces a candidate answer and trace, and has an evaluated correctness column and recurring cost\. A graph grammar is useful for making the feasible class explicit, but the portfolio optimization results do not depend on this particular representation\.

Let𝒰=\{u1,…,uQ\}\\mathcal\{U\}=\\\{u\_\{1\},\\ldots,u\_\{Q\}\\\}denote a library of primitive executable modules\. A module may be a prompt\-based agent, retriever, solver, verifier, critic, tool call, aggregation rule, learned classifier, neural\-network block, or answer formatter\. A workflowggis a finite directed computation graph

g=\(Vg,Eg,ℓg,sg,tg\),g=\(V\_\{g\},E\_\{g\},\\ell\_\{g\},s\_\{g\},t\_\{g\}\),whereVgV\_\{g\}is a finite node set,EgE\_\{g\}represents information flow,ℓg:Vg→𝒰\\ell\_\{g\}:V\_\{g\}\\to\\mathcal\{U\}assigns a module to each node, andsg,tgs\_\{g\},t\_\{g\}are input and output nodes\. We impose\|Vg\|≤Lmax\|V\_\{g\}\|\\leq L\_\{\\max\}and any application\-specific restrictions on retrieval, tools, model calls, memory, context length, or termination\. The feasible workflow class𝒢\\mathcal\{G\}consists of all graphs satisfying these restrictions\.

## Appendix BOther selector technologies

The endogenous framework allows a different recovery curveψk\\psi\_\{k\}for each candidate\-set sizekk\. The selector\-strength bound in[Section4](https://arxiv.org/html/2609.18126#S4)applies whenever these curves admit a common fraction\-based envelope with finite odds\-lift\. The fixed\-size IP remains valid for any nondecreasing recovery curve, while the randomized\-rounding certificate additionally requires diminishing increments\. A selector that reaches recovery probability one at an interior correct fraction has infinite odds\-lift, so the uniform selector\-strength cap becomes vacuous even though the exact size\-conditioned IP remains applicable\.

#### Conservative clean recovery\.

Suppose each correct candidate is rejected with probabilityβ\\beta, each incorrect candidate is falsely certified with probabilityα\\alpha, and verifier outputs are conditionally independent\. The event that at least one correct candidate is certified and no incorrect candidate is certified has probability

ψα,β,kclean\(r\)=𝟏\{r≥1\}\(1−βr\)\(1−α\)k−r\.\\psi\_\{\\alpha,\\beta,k\}^\{\\mathrm\{clean\}\}\(r\)=\\mathbf\{1\}\\\{r\\geq 1\\\}\(1\-\\beta^\{r\}\)\(1\-\\alpha\)^\{k\-r\}\.\(74\)

#### Uniform selection among certified candidates\.

If the selector chooses uniformly among certified candidates, withU∼Binomial⁡\(r,1−β\)U\\sim\\mathrm\{Binomial\}\(r,1\-\\beta\)andW∼Binomial⁡\(k−r,α\)W\\sim\\mathrm\{Binomial\}\(k\-r,\\alpha\),

ψα,β,kunif​\(r\)=∑u=1r∑w=0k−ruu\+w​\(ru\)​\(1−β\)u​βr−u​\(k−rw\)​αw​\(1−α\)k−r−w\.\\psi\_\{\\alpha,\\beta,k\}^\{\\mathrm\{unif\}\}\(r\)=\\sum\_\{u=1\}^\{r\}\\sum\_\{w=0\}^\{k\-r\}\\frac\{u\}\{u\+w\}\\binom\{r\}\{u\}\(1\-\\beta\)^\{u\}\\beta^\{r\-u\}\\binom\{k\-r\}\{w\}\\alpha^\{w\}\(1\-\\alpha\)^\{k\-r\-w\}\.\(75\)

#### Voting\.

Under a conservative model in which accepted wrong workflows coordinate on the same wrong answer, letCr∼Binomial⁡\(r,1−β\)C\_\{r\}\\sim\\mathrm\{Binomial\}\(r,1\-\\beta\)andWk−r∼Binomial⁡\(k−r,α\)W\_\{k\-r\}\\sim\\mathrm\{Binomial\}\(k\-r,\\alpha\)\. A sufficient recovery event isCr\>Wk−rC\_\{r\}\>W\_\{k\-r\}, giving

ψα,β,kvote\(r\)=ℙ\{Cr\>Wk−r\}\.\\psi\_\{\\alpha,\\beta,k\}^\{\\mathrm\{vote\}\}\(r\)=\\mathbb\{P\}\\\{C\_\{r\}\>W\_\{k\-r\}\\\}\.\(76\)Voting is typically threshold\-like and need not be concave\. The size\-conditioned IP remains exact, but the concave\-coverage rounding certificate does not apply unless the estimated curve satisfies diminishing increments\.

## Appendix CA Numerical Illustration of Workflow Multiplicity

This example illustrates why it can be optimal to assign multiple execution slots to the same workflow type under an imperfect selector\. Suppose there are two deterministic workflow types,AAandBB, and three task types with the correctness patterns and population shares shown in[Table4](https://arxiv.org/html/2609.18126#A3.T4)\.

Table 4:Task types and workflow correctness\.For compactness, write execution\-count vectors in the coordinate order\(A,B\)\(A,B\), and define

𝒎A=\(1,0\),𝒎A​B=\(1,1\),𝒎A​A​B=\(2,1\)\.\\bm\{m\}^\{A\}=\(1,0\),\\qquad\\bm\{m\}^\{AB\}=\(1,1\),\\qquad\\bm\{m\}^\{AAB\}=\(2,1\)\.
WorkflowAAhas standalone accuracy0\.400\.40, whereas workflowBBhas standalone accuracy0\.200\.20\. WorkflowBBis nevertheless complementary because it solves Type 2 tasks, which workflowAAmisses\.

Suppose the selector follows the Plackett–Luce recovery curve withλ=2\\lambda=2:

ψk,2PL​\(r\)=2​r2​r\+k−r\.\\psi^\{\\mathrm\{PL\}\}\_\{k,2\}\(r\)=\\frac\{2r\}\{2r\+k\-r\}\.\(77\)For the portfolio𝒎A​B\\bm\{m\}^\{AB\}, exactly one of the two candidates is correct on Types 1 and 2, so

J2​\(𝒎A​B\)=0\.40​\(23\)\+0\.20​\(23\)=0\.40\.J\_\{2\}\(\\bm\{m\}^\{AB\}\)=0\.40\\left\(\\frac\{2\}\{3\}\\right\)\+0\.20\\left\(\\frac\{2\}\{3\}\\right\)=0\.40\.\(78\)
Now consider the portfolio𝒎A​A​B\\bm\{m\}^\{AAB\}\. On Type 1 tasks, two of the three candidates are correct, giving recovery probability4/54/5\. On Type 2 tasks, only workflowBBis correct, giving recovery probability1/21/2\. Therefore,

J2​\(𝒎A​A​B\)=0\.40​\(45\)\+0\.20​\(12\)=0\.42\.J\_\{2\}\(\\bm\{m\}^\{AAB\}\)=0\.40\\left\(\\frac\{4\}\{5\}\\right\)\+0\.20\\left\(\\frac\{1\}\{2\}\\right\)=0\.42\.\(79\)Thus,

J2​\(𝒎A​A​B\)=0\.42\>J2​\(𝒎A​B\)=J2​\(𝒎A\)=0\.40\.J\_\{2\}\(\\bm\{m\}^\{AAB\}\)=0\.42\>J\_\{2\}\(\\bm\{m\}^\{AB\}\)=J\_\{2\}\(\\bm\{m\}^\{A\}\)=0\.40\.\(80\)
The additional execution ofAAcreates no new task coverage\. Instead, it increases the representation of the stronger workflow on the more common Type 1 tasks, raising recovery from2/32/3to4/54/5, while retaining workflowBBto cover Type 2 tasks\. The gain on Type 1 tasks exceeds the corresponding loss in selector recovery on Type 2 tasks\.

The result also survives a small execution cost\. If every execution has accuracy\-equivalent costκ=0\.005\\kappa=0\.005, then

Π2​\(𝒎A\)\\displaystyle\\Pi\_\{2\}\(\\bm\{m\}^\{A\}\)=0\.400−0\.005=0\.395,\\displaystyle=0\.400\-0\.005=0\.395,Π2​\(𝒎A​B\)\\displaystyle\\Pi\_\{2\}\(\\bm\{m\}^\{AB\}\)=0\.400−0\.010=0\.390,\\displaystyle=0\.400\-0\.010=0\.390,Π2​\(𝒎A​A​B\)\\displaystyle\\Pi\_\{2\}\(\\bm\{m\}^\{AAB\}\)=0\.420−0\.015=0\.405\.\\displaystyle=0\.420\-0\.015=0\.405\.\(81\)Hence,𝒎A​A​B\\bm\{m\}^\{AAB\}remains optimal among these portfolios after accounting for recurring execution cost\. By contrast, repeatingAAwithout including another workflow type would leave its selector\-aware accuracy unchanged at0\.400\.40and would only add cost\. Multiplicity is valuable here because it adjusts the relative representation of complementary workflow types, not because pure repetition expands coverage\.

## Appendix DProofs

### D\.1Proof of Lemma[4\.2](https://arxiv.org/html/2609.18126#S4.Thmtheorem2)

Forp∈\(0,1\)p\\in\(0,1\), the conditionΛ⁡\(H\)≤Λ\\Lambda\(H\)\\leq\\Lambdagives

H⁡\(p\)1−H⁡\(p\)≤Λ​p1−p\.\\frac\{H\(p\)\}\{1\-H\(p\)\}\\leq\\Lambda\\frac\{p\}\{1\-p\}\.\(82\)Finite odds\-lift impliesH⁡\(p\)<1H\(p\)<1at every interiorpp, so all denominators are positive\. Multiplying through and collecting terms yields

H⁡\(p\)​\{1−p\+Λ​p\}≤Λ​p\.H\(p\)\\\{1\-p\+\\Lambda p\\\}\\leq\\Lambda p\.\(83\)Because1−p\+Λ​p=1\+\(Λ−1\)​p\>01\-p\+\\Lambda p=1\+\(\\Lambda\-1\)p\>0,

H⁡\(p\)≤Λ​p1\+\(Λ−1\)​p=hΛ​\(p\)\.H\(p\)\\leq\\frac\{\\Lambda p\}\{1\+\(\\Lambda\-1\)p\}=h\_\{\\Lambda\}\(p\)\.\(84\)Atp=0p=0, both sides are zero underH⁡\(0\)=0H\(0\)=0; atp=1p=1, the right\-hand side is one and the inequality follows fromH⁡\(1\)≤1H\(1\)\\leq 1\.□\\square

### D\.2Proof of Lemma[4\.3](https://arxiv.org/html/2609.18126#S4.Thmtheorem3)

Forp∈\[0,1\]p\\in\[0,1\],

hΛ′​\(p\)=Λ\{1\+\(Λ−1\)​p\}2≥0,hΛ′′​\(p\)=−2​Λ​\(Λ−1\)\{1\+\(Λ−1\)​p\}3≤0\.h\_\{\\Lambda\}^\{\\prime\}\(p\)=\\frac\{\\Lambda\}\{\\\{1\+\(\\Lambda\-1\)p\\\}^\{2\}\}\\geq 0,\\qquad h\_\{\\Lambda\}^\{\\prime\\prime\}\(p\)=\-\\frac\{2\\Lambda\(\\Lambda\-1\)\}\{\\\{1\+\(\\Lambda\-1\)p\\\}^\{3\}\}\\leq 0\.\(85\)ThushΛh\_\{\\Lambda\}is increasing and concave\.□\\square

### D\.3Proof of Lemma[4\.6](https://arxiv.org/html/2609.18126#S4.Thmtheorem6)

For continuousrr,

dd​r​λ​rk\+\(λ−1\)​r=λ​k\{k\+\(λ−1\)​r\}2≥0,\\frac\{d\}\{dr\}\\frac\{\\lambda r\}\{k\+\(\\lambda\-1\)r\}=\\frac\{\\lambda k\}\{\\\{k\+\(\\lambda\-1\)r\\\}^\{2\}\}\\geq 0,\(86\)and

d2d​r2​λ​rk\+\(λ−1\)​r=−2​λ​k​\(λ−1\)\{k\+\(λ−1\)​r\}3≤0\.\\frac\{d^\{2\}\}\{dr^\{2\}\}\\frac\{\\lambda r\}\{k\+\(\\lambda\-1\)r\}=\-\\frac\{2\\lambda k\(\\lambda\-1\)\}\{\\\{k\+\(\\lambda\-1\)r\\\}^\{3\}\}\\leq 0\.\(87\)The discrete increment follows by subtracting adjacent values and simplifying, which yields \([31](https://arxiv.org/html/2609.18126#S4.E31)\); its denominator increases inℓ\\ell\. Finally,

hλ′​\(p\)=λ\{1\+\(λ−1\)​p\}2\>0,hλ′′​\(p\)=−2​λ​\(λ−1\)\{1\+\(λ−1\)​p\}3≤0\.h\_\{\\lambda\}^\{\\prime\}\(p\)=\\frac\{\\lambda\}\{\\\{1\+\(\\lambda\-1\)p\\\}^\{2\}\}\>0,\\qquad h\_\{\\lambda\}^\{\\prime\\prime\}\(p\)=\-\\frac\{2\\lambda\(\\lambda\-1\)\}\{\\\{1\+\(\\lambda\-1\)p\\\}^\{3\}\}\\leq 0\.\(88\)□\\square

### D\.4Proof of Proposition[4\.7](https://arxiv.org/html/2609.18126#S4.Thmtheorem7)

Consider one task and zero compute costs\. Letc1,c2c\_\{1\},c\_\{2\}be workflows that are correct on the task and letw1,w2w\_\{1\},w\_\{2\}be wrong\. Nonmonotonicity follows from

Jλ​\(\{c1\}\)=1,Jλ​\(\{c1,w1\}\)=hλ​\(1/2\)<1\.J\_\{\\lambda\}\(\\\{c\_\{1\}\\\}\)=1,\\qquad J\_\{\\lambda\}\(\\\{c\_\{1\},w\_\{1\}\\\}\)=h\_\{\\lambda\}\(1/2\)<1\.\(89\)For submodularity, takeA=\{c1\}A=\\\{c\_\{1\}\\\},B=\{c1,w1\}B=\\\{c\_\{1\},w\_\{1\}\\\}, ande=c2e=c\_\{2\}\. Then

Δ⁡\(e∣A\)=0,Δ⁡\(e∣B\)=hλ​\(2/3\)−hλ​\(1/2\)\>0,\\Delta\(e\\mid A\)=0,\\qquad\\Delta\(e\\mid B\)=h\_\{\\lambda\}\(2/3\)\-h\_\{\\lambda\}\(1/2\)\>0,\(90\)which violates diminishing returns\. For supermodularity, takeA=\{w1\}A=\\\{w\_\{1\}\\\},B=\{w1,w2\}B=\\\{w\_\{1\},w\_\{2\}\\\}, ande=c1e=c\_\{1\}\. Then

Δ⁡\(e∣A\)=hλ​\(1/2\)\>hλ​\(1/3\)=Δ⁡\(e∣B\),\\Delta\(e\\mid A\)=h\_\{\\lambda\}\(1/2\)\>h\_\{\\lambda\}\(1/3\)=\\Delta\(e\\mid B\),\(91\)which violates increasing returns\. Adding or subtracting a modular function does not change the submodularity or supermodularity inequalities\.□\\square

### D\.5Proof of Theorem[5\.1](https://arxiv.org/html/2609.18126#S5.Thmtheorem1)

Fix an execution\-count vector𝒎∈ℤ\+ℳ\\bm\{m\}\\in\\mathbb\{Z\}\_\{\+\}^\{\\mathcal\{M\}\}satisfyingk⁡\(𝒎\)=kk\(\\bm\{m\}\)=k, label itskkexecution slots, and samplebblabeled slots uniformly without replacement\. Let𝑴\(b\)\\bm\{M\}^\{\(b\)\}be the resulting random execution\-count vector and letα=b/k\\alpha=b/k\. For taskii, writeri=ri​\(𝒎\)r\_\{i\}=r\_\{i\}\(\\bm\{m\}\)and let

Ri=ri​\(𝑴\(b\)\)\.R\_\{i\}=r\_\{i\}\\bigl\(\\bm\{M\}^\{\(b\)\}\\bigr\)\.Then

Ri∼Hypergeom⁡\(k,ri,b\),𝔼⁡\[Ri\]=α​ri\.R\_\{i\}\\sim\\mathrm\{Hypergeom\}\(k,r\_\{i\},b\),\\qquad\\mathbb\{E\}\[R\_\{i\}\]=\\alpha r\_\{i\}\.Ifri=0r\_\{i\}=0, then both the size\-kktask value and the expected size\-bbtask value are zero\. Supposeri\>0r\_\{i\}\>0and definepi=ri/kp\_\{i\}=r\_\{i\}/kandXi=Ri/bX\_\{i\}=R\_\{i\}/b\. Sincehhis increasing, concave, and satisfiesh⁡\(0\)=0h\(0\)=0,

h⁡\(x\)≥h⁡\(pi\)​min⁡\{x/pi,1\},x∈\[0,1\]\.h\(x\)\\geq h\(p\_\{i\}\)\\min\\\{x/p\_\{i\},1\\\},\\qquad x\\in\[0,1\]\.\(92\)Indeed, forx≤pix\\leq p\_\{i\}the inequality follows from concavity between00andpip\_\{i\}, and forx≥pix\\geq p\_\{i\}it follows from monotonicity\. Now

Xipi=Riα​ri\.\\frac\{X\_\{i\}\}\{p\_\{i\}\}=\\frac\{R\_\{i\}\}\{\\alpha r\_\{i\}\}\.\(93\)LetZi=Ri/ri∈\[0,1\]Z\_\{i\}=R\_\{i\}/r\_\{i\}\\in\[0,1\]\. Then𝔼⁡\[Zi\]=α\\mathbb\{E\}\[Z\_\{i\}\]=\\alphaand, for everyz∈\[0,1\]z\\in\[0,1\],min⁡\{z,α\}≥α​z\\min\\\{z,\\alpha\\\}\\geq\\alpha z\. Hence

𝔼⁡\[min⁡\{Xipi,1\}\]\\displaystyle\\mathbb\{E\}\\left\[\\min\\left\\\{\\frac\{X\_\{i\}\}\{p\_\{i\}\},1\\right\\\}\\right\]=1α​𝔼​\[min⁡\{Zi,α\}\]≥1α​α​𝔼​\[Zi\]=α\.\\displaystyle=\\frac\{1\}\{\\alpha\}\\mathbb\{E\}\[\\min\\\{Z\_\{i\},\\alpha\\\}\]\\geq\\frac\{1\}\{\\alpha\}\\alpha\\mathbb\{E\}\[Z\_\{i\}\]=\\alpha\.\(94\)Combining this with \([92](https://arxiv.org/html/2609.18126#A4.E92)\) gives

𝔼⁡\[h⁡\(Ri/b\)\]≥α​h​\(ri/k\)\.\\mathbb\{E\}\[h\(R\_\{i\}/b\)\]\\geq\\alpha h\(r\_\{i\}/k\)\.\(95\)Averaging across tasks yields

𝔼⁡\[Jψ​\(𝑴\(b\)\)\]≥α​Jψ​\(𝒎\),\\mathbb\{E\}\\\!\\left\[J\_\{\\psi\}\\bigl\(\\bm\{M\}^\{\(b\)\}\\bigr\)\\right\]\\geq\\alpha J\_\{\\psi\}\(\\bm\{m\}\),\(96\)Because costs are additive,

𝔼⁡\[C⁡\(𝑴\(b\)\)\]=α​C​\(𝒎\),\\mathbb\{E\}\\\!\\left\[C\\bigl\(\\bm\{M\}^\{\(b\)\}\\bigr\)\\right\]=\\alpha C\(\\bm\{m\}\),\(97\)Therefore

𝔼⁡\[Πψ,γ​\(𝑴\(b\)\)\]≥α​Πψ,γ​\(𝒎\)\.\\mathbb\{E\}\\\!\\left\[\\Pi\_\{\\psi,\\gamma\}\\bigl\(\\bm\{M\}^\{\(b\)\}\\bigr\)\\right\]\\geq\\alpha\\Pi\_\{\\psi,\\gamma\}\(\\bm\{m\}\)\.\(98\)Some realization of𝑴\(b\)\\bm\{M\}^\{\(b\)\}attains at least this expected value\. If𝒎\\bm\{m\}is an optimal size\-kkexecution\-count vector with positive value, then

OPTb​\(ℳ\)≥α​OPTk​\(ℳ\),\\mathrm\{OPT\}\_\{b\}\(\\mathcal\{M\}\)\\geq\\alpha\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\),\(99\)which proves \([34](https://arxiv.org/html/2609.18126#S5.E34)\); ifOPTk​\(ℳ\)≤0\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\)\\leq 0, the statement with positive parts is immediate\. For the grid result, choose for eachkkan anchorb∈𝒦b\\in\\mathcal\{K\}withb≤k≤ϱ​bb\\leq k\\leq\\varrho b\. Then

OPTb\+​\(ℳ\)≥bk​OPTk\+​\(ℳ\)≥1ϱ​OPTk\+​\(ℳ\)\.\\mathrm\{OPT\}\_\{b\}^\{\+\}\(\\mathcal\{M\}\)\\geq\\frac\{b\}\{k\}\\mathrm\{OPT\}\_\{k\}^\{\+\}\(\\mathcal\{M\}\)\\geq\\frac\{1\}\{\\varrho\}\\mathrm\{OPT\}\_\{k\}^\{\+\}\(\\mathcal\{M\}\)\.\(100\)Maximizing overkkproves \([35](https://arxiv.org/html/2609.18126#S5.E35)\)\. The dyadic and ratio\-1/\(1−ηgrid\)1/\(1\-\\eta\_\{\\mathrm\{grid\}\}\)claims follow from the definition of the grid coverage ratio\.□\\square

### D\.6Proof of Theorem[4\.4](https://arxiv.org/html/2609.18126#S4.Thmtheorem4)

The lower bound follows from the outside option and singleton feasibility\. With one candidate, the selector must return that candidate, so the net value of the singleton execution\-count vector𝒆g\\bm\{e\}\_\{g\}isa¯g−γ​cg\\bar\{a\}\_\{g\}\-\\gamma c\_\{g\}\.

Fix a nonzero execution\-count vector𝒎∈ℤ\+ℳ\\bm\{m\}\\in\\mathbb\{Z\}\_\{\+\}^\{\\mathcal\{M\}\}satisfyingk⁡\(𝒎\)=kk\(\\bm\{m\}\)=k, and define

pi​\(𝒎\)=ri​\(𝒎\)k\.p\_\{i\}\(\\bm\{m\}\)=\\frac\{r\_\{i\}\(\\bm\{m\}\)\}\{k\}\.By[Sections4\.1](https://arxiv.org/html/2609.18126#S4.SS1)and[4\.2](https://arxiv.org/html/2609.18126#S4.Thmtheorem2),

Jψ​\(𝒎\)\\displaystyle J\_\{\\psi\}\(\\bm\{m\}\)=1n​∑i=1nψk​\(ri​\(𝒎\)\)\\displaystyle=\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\psi\_\{k\}\\bigl\(r\_\{i\}\(\\bm\{m\}\)\\bigr\)\(101\)≤1n​∑i=1nH⁡\(pi​\(𝒎\)\)\\displaystyle\\leq\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}H\\bigl\(p\_\{i\}\(\\bm\{m\}\)\\bigr\)\(102\)≤1n​∑i=1nhΛ​\(pi​\(𝒎\)\)\.\\displaystyle\\leq\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}h\_\{\\Lambda\}\\bigl\(p\_\{i\}\(\\bm\{m\}\)\\bigr\)\.\(103\)By[Lemma4\.3](https://arxiv.org/html/2609.18126#S4.Thmtheorem3)and Jensen’s inequality,

1n​∑i=1nhΛ​\(pi​\(𝒎\)\)\\displaystyle\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}h\_\{\\Lambda\}\\bigl\(p\_\{i\}\(\\bm\{m\}\)\\bigr\)≤hΛ​\(1n​∑i=1npi​\(𝒎\)\)\\displaystyle\\leq h\_\{\\Lambda\}\\left\(\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}p\_\{i\}\(\\bm\{m\}\)\\right\)\(104\)=hΛ​\(1k​∑g∈ℳa¯g​mg\)\\displaystyle=h\_\{\\Lambda\}\\left\(\\frac\{1\}\{k\}\\sum\_\{g\\in\\mathcal\{M\}\}\\bar\{a\}\_\{g\}m\_\{g\}\\right\)\(105\)≤hΛ​\(Ak\)\.\\displaystyle\\leq h\_\{\\Lambda\}\(A\_\{k\}\)\.\(106\)AlsoC⁡\(𝒎\)≥CkC\(\\bm\{m\}\)\\geq C\_\{k\}\. Hence

Πψ,γ​\(𝒎\)≤hΛ​\(Ak\)−γ​Ck\.\\Pi\_\{\\psi,\\gamma\}\(\\bm\{m\}\)\\leq h\_\{\\Lambda\}\(A\_\{k\}\)\-\\gamma C\_\{k\}\.Maximizing overkkand including the outside option proves \([22](https://arxiv.org/html/2609.18126#S4.E22)\)\. Whenγ=0\\gamma=0,Ak≤a⋆A\_\{k\}\\leq a^\{\\star\}for allkk, giving \([23](https://arxiv.org/html/2609.18126#S4.E23)\)\.

For tightness over the selector class, take recoveryH=hΛH=h\_\{\\Lambda\}\. Supposea⋆=r/ka^\{\\star\}=r/kandk≤Kmaxk\\leq K\_\{\\max\}\. Constructkktasks andkkworkflows using a cyclic incidence matrix in which each workflow solves exactlyrrtasks and each task is solved by exactlyrrworkflows\. Let𝒎⋆\\bm\{m\}^\{\\star\}assign one execution slot to each of thesekkworkflows\. Thenpi​\(𝒎⋆\)=a⋆p\_\{i\}\(\\bm\{m\}^\{\\star\}\)=a^\{\\star\}for every task, and thereforeJψ​\(𝒎⋆\)=hΛ​\(a⋆\)J\_\{\\psi\}\(\\bm\{m\}^\{\\star\}\)=h\_\{\\Lambda\}\(a^\{\\star\}\)\. Rational values can be represented exactly after replication, and arbitrary values can be approximated\.□\\square

### D\.7Proof of Corollary[4\.5](https://arxiv.org/html/2609.18126#S4.Thmtheorem5)

The first inequality follows from \([23](https://arxiv.org/html/2609.18126#S4.E23)\)\. Let

GΛ​\(a\)=hΛ​\(a\)−a=\(Λ−1\)​a​\(1−a\)1\+\(Λ−1\)​a\.G\_\{\\Lambda\}\(a\)=h\_\{\\Lambda\}\(a\)\-a=\\frac\{\(\\Lambda\-1\)a\(1\-a\)\}\{1\+\(\\Lambda\-1\)a\}\.\(107\)ForΛ\>1\\Lambda\>1, differentiation gives the unique maximizera=1/\(1\+Λ\)a=1/\(1\+\\sqrt\{\\Lambda\}\)and maximum\(Λ−1\)/\(Λ\+1\)\(\\sqrt\{\\Lambda\}\-1\)/\(\\sqrt\{\\Lambda\}\+1\)\. AtΛ=1\\Lambda=1, both sides are zero\.

If all workflows costcc, any execution\-count vector𝒎\\bm\{m\}satisfyingk⁡\(𝒎\)=k≥2k\(\\bm\{m\}\)=k\\geq 2obeys

Πψ,γ​\(𝒎\)≤hΛ​\(a⋆\)−k​γ​c\.\\Pi\_\{\\psi,\\gamma\}\(\\bm\{m\}\)\\leq h\_\{\\Lambda\}\(a^\{\\star\}\)\-k\\gamma c\.Under \([26](https://arxiv.org/html/2609.18126#S4.E26)\),

hΛ​\(a⋆\)−k​γ​c≤a⋆−γ​c,h\_\{\\Lambda\}\(a^\{\\star\}\)\-k\\gamma c\\leq a^\{\\star\}\-\\gamma c,\(108\)so no multi\-workflow portfolio beats the best singleton\. The outside option dominates only when the singleton net value is negative\.□\\square

### D\.8Proof of Proposition[5\.2](https://arxiv.org/html/2609.18126#S5.Thmtheorem2)

Fix an execution\-count vector𝒎∈ℤ\+ℳ\\bm\{m\}\\in\\mathbb\{Z\}\_\{\+\}^\{\\mathcal\{M\}\}satisfyingk⁡\(𝒎\)=kk\(\\bm\{m\}\)=k, and set

zg=mg,g∈ℳ\.z\_\{g\}=m\_\{g\},\\qquad g\\in\\mathcal\{M\}\.Then

∑g∈ℳai​gzg=ri\(𝒎\),i=1,…,n\.\\sum\_\{g\\in\\mathcal\{M\}\}a\_\{ig\}z\_\{g\}=r\_\{i\}\(\\bm\{m\}\),\\qquad i=1,\\ldots,n\.For this fixedzz, the prefix constraints imply that the active tier variables for taskiiform a prefix, while the capacity constraint implies that the length of this prefix is at mostri​\(𝒎\)r\_\{i\}\(\\bm\{m\}\)\. Sincedℓ,k≥0d\_\{\\ell,k\}\\geq 0, there exists an optimal choice

yi​ℓ=𝟏\{ℓ≤ri\(𝒎\)\},ℓ=1,…,k\.y\_\{i\\ell\}=\\mathbf\{1\}\\\{\\ell\\leq r\_\{i\}\(\\bm\{m\}\)\\\},\\qquad\\ell=1,\\ldots,k\.Its task\-iicontribution is

∑ℓ=1kdℓ,k​yi​ℓ=∑ℓ=1ri​\(𝒎\)dℓ,k=ψk​\(ri​\(𝒎\)\),\\sum\_\{\\ell=1\}^\{k\}d\_\{\\ell,k\}y\_\{i\\ell\}=\\sum\_\{\\ell=1\}^\{r\_\{i\}\(\\bm\{m\}\)\}d\_\{\\ell,k\}=\\psi\_\{k\}\\bigl\(r\_\{i\}\(\\bm\{m\}\)\\bigr\),where the last equality usesψk​\(0\)=0\\psi\_\{k\}\(0\)=0\. Moreover, the cost term is

−γ∑g∈ℳcgzg=−γC\(𝒎\)\.\-\\gamma\\sum\_\{g\\in\\mathcal\{M\}\}c\_\{g\}z\_\{g\}=\-\\gamma C\(\\bm\{m\}\)\.Thus every feasible size\-kkexecution\-count vector induces an integer\-program solution with the same objective value, and hence

Ik​\(ℳ\)≥OPTk​\(ℳ\)\.I\_\{k\}\(\\mathcal\{M\}\)\\geq\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\)\.
Conversely, let\(y,z\)\(y,z\)be any integral feasible solution of \([39](https://arxiv.org/html/2609.18126#S5.E39)\), and define the execution\-count vector𝒎∈ℤ\+ℳ\\bm\{m\}\\in\\mathbb\{Z\}\_\{\+\}^\{\\mathcal\{M\}\}bymg=zgm\_\{g\}=z\_\{g\}\. Thenk⁡\(𝒎\)=kk\(\\bm\{m\}\)=k\. For each taskii, let

ti=∑ℓ=1kyi​ℓ\.t\_\{i\}=\\sum\_\{\\ell=1\}^\{k\}y\_\{i\\ell\}\.Becauseyyis binary and satisfies the prefix constraints, its active entries form a prefix of lengthtit\_\{i\}\. Feasibility gives

ti≤∑g∈ℳai​g​zg=ri​\(𝒎\)\.t\_\{i\}\\leq\\sum\_\{g\\in\\mathcal\{M\}\}a\_\{ig\}z\_\{g\}=r\_\{i\}\(\\bm\{m\}\)\.Therefore, using the nonnegativity of the increments,

∑ℓ=1kdℓ,k​yi​ℓ=∑ℓ=1tidℓ,k=ψk​\(ti\)≤ψk​\(ri​\(𝒎\)\)\.\\sum\_\{\\ell=1\}^\{k\}d\_\{\\ell,k\}y\_\{i\\ell\}=\\sum\_\{\\ell=1\}^\{t\_\{i\}\}d\_\{\\ell,k\}=\\psi\_\{k\}\(t\_\{i\}\)\\leq\\psi\_\{k\}\\bigl\(r\_\{i\}\(\\bm\{m\}\)\\bigr\)\.The integer\-program objective is consequently at mostΠψ,γ​\(𝒎\)\\Pi\_\{\\psi,\\gamma\}\(\\bm\{m\}\), and hence at mostOPTk​\(ℳ\)\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\)\. Thus

Ik​\(ℳ\)≤OPTk​\(ℳ\)\.I\_\{k\}\(\\mathcal\{M\}\)\\leq\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\)\.Combining the two inequalities givesIk​\(ℳ\)=OPTk​\(ℳ\)I\_\{k\}\(\\mathcal\{M\}\)=\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\)\. Finally, combining this equality with the size decomposition \([33](https://arxiv.org/html/2609.18126#S5.E33)\) proves \([40](https://arxiv.org/html/2609.18126#S5.E40)\)\.□\\square

### D\.9Proof of Theorem[5\.3](https://arxiv.org/html/2609.18126#S5.Thmtheorem3)

Writedℓ=dℓ,kd\_\{\\ell\}=d\_\{\\ell,k\}\. For a feasible repeated\-execution LP vectorzz, define

xi=∑gai​g​zg,qg=zg/k\.x\_\{i\}=\\sum\_\{g\}a\_\{ig\}z\_\{g\},\\qquad q\_\{g\}=z\_\{g\}/k\.DrawG1,…,GkG\_\{1\},\\ldots,G\_\{k\}independently fromqq, and letRi=∑j=1kai​GjR\_\{i\}=\\sum\_\{j=1\}^\{k\}a\_\{iG\_\{j\}\}\. Then

Ri∼Binomial⁡\(k,xi/k\),𝔼⁡\[Ri\]=xi\.R\_\{i\}\\sim\\mathrm\{Binomial\}\(k,x\_\{i\}/k\),\\qquad\\mathbb\{E\}\[R\_\{i\}\]=x\_\{i\}\.For every integert≥1t\\geq 1andr≥0r\\geq 0,

min⁡\{r,t\}≥t⁡\{1−\(1−1/t\)r\}\.\\min\\\{r,t\\\}\\geq t\\left\\\{1\-\(1\-1/t\)^\{r\}\\right\\\}\.Therefore

𝔼⁡\[min⁡\{Ri,t\}\]\\displaystyle\\mathbb\{E\}\[\\min\\\{R\_\{i\},t\\\}\]≥t⁡\[1−𝔼⁡\{\(1−1/t\)Ri\}\]\\displaystyle\\geq t\\left\[1\-\\mathbb\{E\}\\\{\(1\-1/t\)^\{R\_\{i\}\}\\\}\\right\]=t⁡\[1−\(1−xik​t\)k\]\\displaystyle=t\\left\[1\-\\left\(1\-\\frac\{x\_\{i\}\}\{kt\}\\right\)^\{k\}\\right\]≥t\(1−e−xi/t\)\\displaystyle\\geq t\\left\(1\-e^\{\-x\_\{i\}/t\}\\right\)≥\(1−1e\)​min⁡\{xi,t\}\.\\displaystyle\\geq\\left\(1\-\\frac\{1\}\{e\}\\right\)\\min\\\{x\_\{i\},t\\\}\.The last inequality follows from1−e−u≥\(1−1/e\)​min⁡\{u,1\}1\-e^\{\-u\}\\geq\(1\-1/e\)\\min\\\{u,1\\\}foru≥0u\\geq 0\. Using the decomposition

ψ~k​\(r\)=dk,k​r\+∑t=1k−1\(dt,k−dt\+1,k\)​min⁡\{r,t\},\\widetilde\{\\psi\}\_\{k\}\(r\)=d\_\{k,k\}r\+\\sum\_\{t=1\}^\{k\-1\}\(d\_\{t,k\}\-d\_\{t\+1,k\}\)\\min\\\{r,t\\\},and averaging across tasks gives

𝔼⁡\[Jψ​\(𝑴kRR​\(z\)\)\]≥\(1−1e\)​Qk​\(z\)\+dk,ke​Rk​\(z\)\.\\mathbb\{E\}\\\!\\left\[J\_\{\\psi\}\\bigl\(\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\(z\)\\bigr\)\\right\]\\geq\\left\(1\-\\frac\{1\}\{e\}\\right\)Q\_\{k\}\(z\)\+\\frac\{d\_\{k,k\}\}\{e\}R\_\{k\}\(z\)\.Moreover, ifNgN\_\{g\}is the sampled multiplicity of typegg, then

𝔼⁡\[C⁡\(𝑴kRR​\(z\)\)\]=∑gcg​𝔼​\[Mk,gRR​\(z\)\]=∑gcg​zg=Ck​\(z\)\.\\mathbb\{E\}\\\!\\left\[C\\bigl\(\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\(z\)\\bigr\)\\right\]=\\sum\_\{g\}c\_\{g\}\\mathbb\{E\}\\\!\\left\[M\_\{k,g\}^\{\\mathrm\{RR\}\}\(z\)\\right\]=\\sum\_\{g\}c\_\{g\}z\_\{g\}=C\_\{k\}\(z\)\.Subtracting cost proves \([46](https://arxiv.org/html/2609.18126#S5.E46)\) and \([47](https://arxiv.org/html/2609.18126#S5.E47)\)\. Ifz=zk⋆z=z^\{k\\star\}is LP\-optimal, thenOPTk\(ℳ\)≤Lk\(ℳ\)=Fk\(zk⋆\)\\mathrm\{OPT\}\_\{k\}\(\\mathcal\{M\}\)\\leq L\_\{k\}\(\\mathcal\{M\}\)=F\_\{k\}\(z^\{k\\star\}\), which gives the first inequality in \([48](https://arxiv.org/html/2609.18126#S5.E48)\)\. Finally, concavity impliesQk​\(z\)≤d1,k​Rk​\(z\)Q\_\{k\}\(z\)\\leq d\_\{1,k\}R\_\{k\}\(z\), so

dk,k​Rk​\(z\)=\(1−cψ,k\)​d1,k​Rk​\(z\)≥\(1−cψ,k\)​Qk​\(z\)\.d\_\{k,k\}R\_\{k\}\(z\)=\(1\-c\_\{\\psi,k\}\)d\_\{1,k\}R\_\{k\}\(z\)\\geq\(1\-c\_\{\\psi,k\}\)Q\_\{k\}\(z\)\.This proves the curvature bound\.□\\square

### D\.10Proof of Corollary[5\.4](https://arxiv.org/html/2609.18126#S5.Thmtheorem4)

By[Theorem5\.3](https://arxiv.org/html/2609.18126#S5.Thmtheorem3),

𝔼⁡\[Πλ,γ​\(𝑴kRR\)\]≥Lk​\(ℳ\)−ck,λPLe​Qk⋆\.\\mathbb\{E\}\\\!\\left\[\\Pi\_\{\\lambda,\\gamma\}\(\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\)\\right\]\\geq L\_\{k\}\(\\mathcal\{M\}\)\-\\frac\{c\_\{k,\\lambda\}^\{\\mathrm\{PL\}\}\}\{e\}Q\_\{k\}^\{\\star\}\.Since the LP relaxation has0≤yi​ℓ≤10\\leq y\_\{i\\ell\}\\leq 1and∑ℓ=1kdℓ,k=1\\sum\_\{\\ell=1\}^\{k\}d\_\{\\ell,k\}=1, we have0≤Qk⋆≤10\\leq Q\_\{k\}^\{\\star\}\\leq 1\. Hence

𝔼⁡\[Πλ,γ​\(𝑴kRR\)\]≥Lk​\(ℳ\)−ck,λPLe,\\mathbb\{E\}\\\!\\left\[\\Pi\_\{\\lambda,\\gamma\}\(\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\)\\right\]\\geq L\_\{k\}\(\\mathcal\{M\}\)\-\\frac\{c\_\{k,\\lambda\}^\{\\mathrm\{PL\}\}\}\{e\},This establishes the fixed\-size additive net\-value bound used below\.

For the accuracy\-only bound,[Theorem5\.3](https://arxiv.org/html/2609.18126#S5.Thmtheorem3)gives

𝔼⁡\[Jλ​\(𝑴kRR\)\]≥\(1−1e\)​Qk⋆\+dk,ke​Rk⋆\.\\mathbb\{E\}\\\!\\left\[J\_\{\\lambda\}\(\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\)\\right\]\\geq\\left\(1\-\\frac\{1\}\{e\}\\right\)Q\_\{k\}^\{\\star\}\+\\frac\{d\_\{k,k\}\}\{e\}R\_\{k\}^\{\\star\}\.By the definition of curvature,

dk,k=\(1−ck,λPL\)​d1,k\.d\_\{k,k\}=\\left\(1\-c\_\{k,\\lambda\}^\{\\mathrm\{PL\}\}\\right\)d\_\{1,k\}\.Concavity impliesQk⋆≤d1,k​Rk⋆Q\_\{k\}^\{\\star\}\\leq d\_\{1,k\}R\_\{k\}^\{\\star\}, so

dk,k​Rk⋆≥\(1−ck,λPL\)​Qk⋆\.d\_\{k,k\}R\_\{k\}^\{\\star\}\\geq\\left\(1\-c\_\{k,\\lambda\}^\{\\mathrm\{PL\}\}\\right\)Q\_\{k\}^\{\\star\}\.Substituting this inequality yields

𝔼⁡\[Jλ​\(𝑴kRR\)\]≥\(1−1e\)​Qk⋆\+1−ck,λPLe​Qk⋆=\(1−ck,λPLe\)​Qk⋆,\\mathbb\{E\}\\\!\\left\[J\_\{\\lambda\}\(\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\)\\right\]\\geq\\left\(1\-\\frac\{1\}\{e\}\\right\)Q\_\{k\}^\{\\star\}\+\\frac\{1\-c\_\{k,\\lambda\}^\{\\mathrm\{PL\}\}\}\{e\}Q\_\{k\}^\{\\star\}=\\left\(1\-\\frac\{c\_\{k,\\lambda\}^\{\\mathrm\{PL\}\}\}\{e\}\\right\)Q\_\{k\}^\{\\star\},This establishes an auxiliary accuracy\-only bound\.

It remains to prove the grid\-level multiplicative bound\. Let

U=ULP​\(𝒦,ℳ\),a=1e​maxk∈𝒦​ck,λPL\.U=U\_\{\\mathrm\{LP\}\}\(\\mathcal\{K\};\\mathcal\{M\}\),\\qquad a=\\frac\{1\}\{e\}\\max\_\{k\\in\\mathcal\{K\}\}c\_\{k,\\lambda\}^\{\\mathrm\{PL\}\}\.IfU=0U=0, the result is immediate\. Otherwise, choosek⋆∈\\argmaxk∈𝒦​Lk​\(ℳ\)k^\{\\star\}\\in\\argmax\_\{k\\in\\mathcal\{K\}\}L\_\{k\}\(\\mathcal\{M\}\), so thatLk⋆​\(ℳ\)=UL\_\{k^\{\\star\}\}\(\\mathcal\{M\}\)=U\. By the fixed\-size additive net\-value bound derived above,

𝔼⁡\[Πλ,γ​\(Sk⋆RR\)\]≥U−a\.\\mathbb\{E\}\\\!\\left\[\\Pi\_\{\\lambda,\\gamma\}\(S\_\{k^\{\\star\}\}^\{\\mathrm\{RR\}\}\)\\right\]\\geq U\-a\.Because𝑴^\\widehat\{\\bm\{M\}\}is chosen after comparing the rounded portfolios with the baseline𝒎0\\bm\{m\}^\{0\}and the outside option, Jensen’s inequality for the convex maximum function gives

𝔼⁡\[Πλ,γ​\(𝑴^\)\]≥max⁡\{V¯,U−a\}\.\\mathbb\{E\}\\\!\\left\[\\Pi\_\{\\lambda,\\gamma\}\(\\widehat\{\\bm\{M\}\}\)\\right\]\\geq\\max\\\{\\underline\{V\},U\-a\\\}\.For everyU≥0U\\geq 0,a≥0a\\geq 0, andV¯\>0\\underline\{V\}\>0,

max⁡\{V¯,U−a\}≥V¯V¯\+a​U\.\\max\\\{\\underline\{V\},U\-a\\\}\\geq\\frac\{\\underline\{V\}\}\{\\underline\{V\}\+a\}U\.Indeed, ifU≤V¯\+aU\\leq\\underline\{V\}\+a, then

max⁡\{V¯,U−a\}≥V¯≥V¯V¯\+a​U\.\\max\\\{\\underline\{V\},U\-a\\\}\\geq\\underline\{V\}\\geq\\frac\{\\underline\{V\}\}\{\\underline\{V\}\+a\}U\.IfU\>V¯\+aU\>\\underline\{V\}\+a, then

max⁡\{V¯,U−a\}≥U−a≥V¯V¯\+a​U\.\\max\\\{\\underline\{V\},U\-a\\\}\\geq U\-a\\geq\\frac\{\\underline\{V\}\}\{\\underline\{V\}\+a\}U\.Combining the last two displays proves \([51](https://arxiv.org/html/2609.18126#S5.E51)\)\. Finally, exact integer\-program solves overℳ\\mathcal\{M\}dominate the rounded feasible portfolios, so the same lower bound also applies when the final restricted problems are solved exactly\.□\\square

### D\.11Projected dual, weak separation, and proof details

Forx∈\[0,k\]x\\in\[0,k\], let

φk​\(x\)=1n​ψ~k​\(x\),\\varphi\_\{k\}\(x\)=\\frac\{1\}\{n\}\\widetilde\{\\psi\}\_\{k\}\(x\),and define

χk​\(u\)=maxr=0,…,k⁡\{ψk​\(r\)n−u​r\},0≤u≤d1,kn\.\\chi\_\{k\}\(u\)=\\max\_\{r=0,\\ldots,k\}\\left\\\{\\frac\{\\psi\_\{k\}\(r\)\}\{n\}\-ur\\right\\\},\\qquad 0\\leq u\\leq\\frac\{d\_\{1,k\}\}\{n\}\.\(109\)Becauseφk\\varphi\_\{k\}is concave and piecewise linear with slopes in\[0,d1,k/n\]\[0,d\_\{1,k\}/n\],

φk​\(x\)=min0≤u≤d1,k/n⁡\{u​x\+χk​\(u\)\}\.\\varphi\_\{k\}\(x\)=\\min\_\{0\\leq u\\leq d\_\{1,k\}/n\}\\\{ux\+\\chi\_\{k\}\(u\)\\\}\.\(110\)
###### Lemma D\.1\(Projected dual for repeated executions\)

For every nonempty finite workflow\-type class𝒢\\mathcal\{G\}and fixed sizekk,

Lk​\(𝒢\)=min0≤μi≤d1,k/n⁡\{∑i=1nχk​\(μi\)\+k​maxg∈𝒢⁡\(∑i=1nμi​ai​g−γ​cg\)\}\.L\_\{k\}\(\\mathcal\{G\}\)=\\min\_\{0\\leq\\mu\_\{i\}\\leq d\_\{1,k\}/n\}\\left\\\{\\sum\_\{i=1\}^\{n\}\\chi\_\{k\}\(\\mu\_\{i\}\)\+k\\max\_\{g\\in\\mathcal\{G\}\}\\left\(\\sum\_\{i=1\}^\{n\}\\mu\_\{i\}a\_\{ig\}\-\\gamma c\_\{g\}\\right\)\\right\\\}\.\(111\)Equivalently, introducingθ\\thetaand epigraph variablesξi\\xi\_\{i\}yields \([52](https://arxiv.org/html/2609.18126#S6.E52)\)\.

###### Proof D\.2

After optimizing the tier variables, the repeated\-execution LP is

maxz≥0:∑gzg=k\{∑i=1nφk\(∑gai​gzg\)−γ∑gcgzg\}\.\\max\_\{z\\geq 0:\\,\\sum\_\{g\}z\_\{g\}=k\}\\left\\\{\\sum\_\{i=1\}^\{n\}\\varphi\_\{k\}\\left\(\\sum\_\{g\}a\_\{ig\}z\_\{g\}\\right\)\-\\gamma\\sum\_\{g\}c\_\{g\}z\_\{g\}\\right\\\}\.Substitute \([110](https://arxiv.org/html/2609.18126#A4.E110)\) and apply finite\- dimensional minimax to interchange the minimization overμ\\muand the maximization over the scaled simplex\. For fixedμ\\mu, linear optimization over that simplex gives

maxz≥0:∑gzg=k∑gzg\(∑iμiai​g−γcg\)=kmaxg∈𝒢\(∑iμiai​g−γcg\)\.\\max\_\{z\\geq 0:\\,\\sum\_\{g\}z\_\{g\}=k\}\\sum\_\{g\}z\_\{g\}\\left\(\\sum\_\{i\}\\mu\_\{i\}a\_\{ig\}\-\\gamma c\_\{g\}\\right\)=k\\max\_\{g\\in\\mathcal\{G\}\}\\left\(\\sum\_\{i\}\\mu\_\{i\}a\_\{ig\}\-\\gamma c\_\{g\}\\right\)\.This proves \([111](https://arxiv.org/html/2609.18126#A4.E111)\)\. The epigraph representation ofχk\\chi\_\{k\}is

ξi\+rμi≥ψk​\(r\)n,r=0,…,k,\\xi\_\{i\}\+r\\mu\_\{i\}\\geq\\frac\{\\psi\_\{k\}\(r\)\}\{n\},\\qquad r=0,\\ldots,k,which gives \([52](https://arxiv.org/html/2609.18126#S6.E52)\)\.□\\square

###### Lemma D\.3\(Batched pricing gives high\-confidence weak separation\)

Fix a candidate dual point\(μ,ξ,θ\)\(\\mu,\\xi,\\theta\)and a toleranceδ\>0\\delta\>0\. Suppose each primitive pricing call made at task pricesμ\\mureturns aδ\\delta\-accurate workflow, meaning a workflowg~\\widetilde\{g\}satisfying

sg~​\(μ\)≥maxg∈𝒢⁡sg​\(μ\)−δ,sg​\(μ\)=∑iμi​ai​g−γ​cg,s\_\{\\widetilde\{g\}\}\(\\mu\)\\geq\\max\_\{g\\in\\mathcal\{G\}\}s\_\{g\}\(\\mu\)\-\\delta,\\qquad s\_\{g\}\(\\mu\)=\\sum\_\{i\}\\mu\_\{i\}a\_\{ig\}\-\\gamma c\_\{g\},with conditional probability at leastporc\>0p\_\{\\mathrm\{orc\}\}\>0\. Runmmprimitive calls at the same prices and keep the highest\-scoring returned workflow\. Then, conditional on the history before the batch, the batch isδ\\delta\-accurate with probability at least

1−e−porc​m\.1\-e^\{\-p\_\{\\mathrm\{orc\}\}m\}\.On this event, if the best returned workflow has score aboveθ\\theta, its workflow inequality is a valid violated constraint\. If its score is at mostθ\\theta, then all workflow inequalities are satisfied after replacingθ\\thetabyθ\+δ\\theta\+\\delta, increasing the dual objective by at mostk​δk\\delta\.

###### Proof D\.4

LetYjY\_\{j\}be the indicator that primitive calljjin the batch isδ\\delta\-accurate\. The calls may be adaptive to previous failed attempts inside the batch, but by assumption

ℙ⁡\{Yj=1∣past within the batch\}≥porc\.\\mathbb\{P\}\\\{Y\_\{j\}=1\\mid\\text\{past within the batch\}\\\}\\geq p\_\{\\mathrm\{orc\}\}\.Therefore

ℙ⁡\{Y1=⋯=Ym=0∣history\}≤\(1−porc\)m≤e−porc​m\.\\mathbb\{P\}\\\{Y\_\{1\}=\\cdots=Y\_\{m\}=0\\mid\\text\{history\}\\\}\\leq\(1\-p\_\{\\mathrm\{orc\}\}\)^\{m\}\\leq e^\{\-p\_\{\\mathrm\{orc\}\}m\}\.Thus, with probability at least1−e−porc​m1\-e^\{\-p\_\{\\mathrm\{orc\}\}m\}, at least one primitive call isδ\\delta\-accurate\. Since the batch keeps the highest\-scoring returned workflow, the batch output is alsoδ\\delta\-accurate\.

If the batch output satisfiessg~​\(μ\)\>θs\_\{\\widetilde\{g\}\}\(\\mu\)\>\\theta, then the constraint

sg~​\(μ\)≤θs\_\{\\widetilde\{g\}\}\(\\mu\)\\leq\\thetais violated and supplies a valid separating hyperplane\. If insteadsg~​\(μ\)≤θs\_\{\\widetilde\{g\}\}\(\\mu\)\\leq\\theta, thenδ\\delta\-accuracy gives

maxg∈𝒢⁡sg​\(μ\)≤sg~​\(μ\)\+δ≤θ\+δ\.\\max\_\{g\\in\\mathcal\{G\}\}s\_\{g\}\(\\mu\)\\leq s\_\{\\widetilde\{g\}\}\(\\mu\)\+\\delta\\leq\\theta\+\\delta\.Thussg​\(μ\)≤θ\+δs\_\{g\}\(\\mu\)\\leq\\theta\+\\deltafor everyg∈𝒢g\\in\\mathcal\{G\}\. Since the coefficient ofθ\\thetain \([52](https://arxiv.org/html/2609.18126#S6.E52)\) iskk, this relaxation increases the dual objective by at mostk​δk\\delta\.□\\square

### D\.12Proof of Theorem[6\.1](https://arxiv.org/html/2609.18126#S6.Thmtheorem1)

For each fixedkk, \([52](https://arxiv.org/html/2609.18126#S6.E52)\) is a rational LP in2​n\+12n\+1variables, with polynomially many explicit recovery and box constraints and an implicit family of workflow constraints\. By[LemmaD\.3](https://arxiv.org/html/2609.18126#A4.Thmtheorem3), a successful batch supplies weak separation for the implicit workflow constraints\. The weak optimization–separation equivalence for the ellipsoid method therefore returns anεell/2\\varepsilon\_\{\\mathrm\{ell\}\}/2\-optimal dual solution in a number of separation queries polynomial in

n,k,B,log⁡\(1/εell\)\.n,\\quad k,\\quad B,\\quad\\log\(1/\\varepsilon\_\{\\mathrm\{ell\}\}\)\.Withτk=εell/\(2​k\)\\tau\_\{k\}=\\varepsilon\_\{\\mathrm\{ell\}\}/\(2k\), the possible slot\-price relaxation in[LemmaD\.3](https://arxiv.org/html/2609.18126#A4.Thmtheorem3)contributes at mostk​τk=εell/2k\\tau\_\{k\}=\\varepsilon\_\{\\mathrm\{ell\}\}/2to the dual objective\. Hence, on the event that all batches used by the ellipsoid routine are successful, standard primal recovery from the workflow inequalities generated by the routine gives, for everyk∈𝒦k\\in\\mathcal\{K\}, a feasible fractional execution vectorzkz^\{k\}supported only on oracle\-returned workflows and satisfying

Fk​\(zk\)≥Lk​\(𝒢\)−εell\.F\_\{k\}\(z^\{k\}\)\\geq L\_\{k\}\(\\mathcal\{G\}\)\-\\varepsilon\_\{\\mathrm\{ell\}\}\.\(112\)
It remains to bound the probability that all batches are successful\. LetNell​\(εell\)N\_\{\\mathrm\{ell\}\}\(\\varepsilon\_\{\\mathrm\{ell\}\}\)be the deterministic upper bound on the total number of ellipsoid separation queries across all grid sizes\. Index these separation queries byq=1,…,Nellq=1,\\ldots,N\_\{\\mathrm\{ell\}\}\. Letℬq\\mathcal\{B\}\_\{q\}be the event that batchqqcontains at least oneτk\\tau\_\{k\}\-accurate primitive pricing call, wherekkis the grid size being solved at that query\. By[LemmaD\.3](https://arxiv.org/html/2609.18126#A4.Thmtheorem3),

ℙ⁡\(ℬqc∣history before batch​q\)≤e−porc​m\.\\mathbb\{P\}\(\\mathcal\{B\}\_\{q\}^\{c\}\\mid\\text\{history before batch \}q\)\\leq e^\{\-p\_\{\\mathrm\{orc\}\}m\}\.Therefore,

ℙ⁡\(⋃q=1Nellℬqc\)\\displaystyle\\mathbb\{P\}\\left\(\\bigcup\_\{q=1\}^\{N\_\{\\mathrm\{ell\}\}\}\\mathcal\{B\}\_\{q\}^\{c\}\\right\)≤∑q=1Nell𝔼⁡\[ℙ⁡\(ℬqc∣history before batch​q\)\]\\displaystyle\\leq\\sum\_\{q=1\}^\{N\_\{\\mathrm\{ell\}\}\}\\mathbb\{E\}\\\!\\left\[\\mathbb\{P\}\(\\mathcal\{B\}\_\{q\}^\{c\}\\mid\\text\{history before batch \}q\)\\right\]\(113\)≤Nell​\(εell\)​e−porc​m\.\\displaystyle\\leq N\_\{\\mathrm\{ell\}\}\(\\varepsilon\_\{\\mathrm\{ell\}\}\)e^\{\-p\_\{\\mathrm\{orc\}\}m\}\.\(114\)Equivalently, sinceT=m​Nell​\(εell\)T=mN\_\{\\mathrm\{ell\}\}\(\\varepsilon\_\{\\mathrm\{ell\}\}\),

Nell​\(εell\)​e−porc​m=exp⁡\{−porc​TNell​\(εell\)\+log⁡Nell​\(εell\)\}\.N\_\{\\mathrm\{ell\}\}\(\\varepsilon\_\{\\mathrm\{ell\}\}\)e^\{\-p\_\{\\mathrm\{orc\}\}m\}=\\exp\\left\\\{\-\\frac\{p\_\{\\mathrm\{orc\}\}T\}\{N\_\{\\mathrm\{ell\}\}\(\\varepsilon\_\{\\mathrm\{ell\}\}\)\}\+\\log N\_\{\\mathrm\{ell\}\}\(\\varepsilon\_\{\\mathrm\{ell\}\}\)\\right\\\}\.This proves the probability statement in \([54](https://arxiv.org/html/2609.18126#S6.E54)\)\.

On the event that all batches are successful, apply[Theorem5\.3](https://arxiv.org/html/2609.18126#S5.Thmtheorem3)to each recovered fractional solutionzkz^\{k\}\. SinceQk​\(zk\)≤1Q\_\{k\}\(z^\{k\}\)\\leq 1, for everyk∈𝒦k\\in\\mathcal\{K\},

𝔼rnd​\[Πλ,γ​\(𝑴kRR\)\]≥Fk​\(zk\)−ck,λPLe≥Lk​\(𝒢\)−εell−ck,λPLe\.\\mathbb\{E\}\_\{\\mathrm\{rnd\}\}\\\!\\left\[\\Pi\_\{\\lambda,\\gamma\}\(\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\)\\right\]\\geq F\_\{k\}\(z^\{k\}\)\-\\frac\{c\_\{k,\\lambda\}^\{\\mathrm\{PL\}\}\}\{e\}\\geq L\_\{k\}\(\\mathcal\{G\}\)\-\\varepsilon\_\{\\mathrm\{ell\}\}\-\\frac\{c\_\{k,\\lambda\}^\{\\mathrm\{PL\}\}\}\{e\}\.Taking the best grid point gives

maxk∈𝒦⁡𝔼rnd​\[Πλ,γ​\(𝑴kRR\)\]≥ULP​\(𝒦,𝒢\)−εell−εrnd​\(𝒦,λ\)\.\\max\_\{k\\in\\mathcal\{K\}\}\\mathbb\{E\}\_\{\\mathrm\{rnd\}\}\\\!\\left\[\\Pi\_\{\\lambda,\\gamma\}\(\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\)\\right\]\\geq U\_\{\\mathrm\{LP\}\}\(\\mathcal\{K\};\\mathcal\{G\}\)\-\\varepsilon\_\{\\mathrm\{ell\}\}\-\\varepsilon\_\{\\mathrm\{rnd\}\}\(\\mathcal\{K\},\\lambda\)\.By[Theorem5\.1](https://arxiv.org/html/2609.18126#S5.Thmtheorem1),

ULP​\(𝒦,𝒢\)≥1ϱ​OPTλ,γ​\(𝒢\)\.U\_\{\\mathrm\{LP\}\}\(\\mathcal\{K\};\\mathcal\{G\}\)\\geq\\frac\{1\}\{\\varrho\}\\mathrm\{OPT\}\_\{\\lambda,\\gamma\}\(\\mathcal\{G\}\)\.Comparing the rounded grid candidates with the retained baseline𝒎0\\bm\{m\}^\{0\}proves \([55](https://arxiv.org/html/2609.18126#S6.E55)\)\.

Finally, for everyA≥0A\\geq 0,e≥0e\\geq 0, andV¯\>0\\underline\{V\}\>0,

max⁡\{V¯,A−e\}≥V¯V¯\+e​A\.\\max\\\{\\underline\{V\},A\-e\\\}\\geq\\frac\{\\underline\{V\}\}\{\\underline\{V\}\+e\}A\.Apply this inequality with

A=1ϱ​OPTλ,γ​\(𝒢\),e=εell\+εrnd​\(𝒦,λ\),A=\\frac\{1\}\{\\varrho\}\\mathrm\{OPT\}\_\{\\lambda,\\gamma\}\(\\mathcal\{G\}\),\\qquad e=\\varepsilon\_\{\\mathrm\{ell\}\}\+\\varepsilon\_\{\\mathrm\{rnd\}\}\(\\mathcal\{K\},\\lambda\),to obtain \([56](https://arxiv.org/html/2609.18126#S6.E56)\)\.□\\square

## Appendix EManagerial Implications and Extensions

#### Accuracy value and the price of compute\.

The parameterγ\\gammais not an arbitrary regularizer\. If a correct decision is worthvvdollars relative to an incorrect one, maximizingv​Jψ​\(𝒎\)−C⁡\(𝒎\)vJ\_\{\\psi\}\(\\bm\{m\}\)\-C\(\\bm\{m\}\)is equivalent to maximizingJψ​\(𝒎\)−γ​C​\(𝒎\)J\_\{\\psi\}\(\\bm\{m\}\)\-\\gamma C\(\\bm\{m\}\)withγ=1/v\\gamma=1/v\. Varyingγ\\gammatraces an accuracy–compute frontier\. The screening bound in[Theorem4\.4](https://arxiv.org/html/2609.18126#S4.Thmtheorem4)can eliminate run sizes that cannot be efficient before solving any IP\.

#### Selector investment versus workflow investment\.

The selector\-strength cap \([25](https://arxiv.org/html/2609.18126#S4.E25)\) maps an odds\-lift boundΛ\\Lambdainto the largest possible gross return from workflow complementarity\. WhenΛ\\Lambdais close to one, spending on additional workflows has little possible upside; improving the selector may dominate expanding the portfolio\. WhenΛ\\Lambdais large, workflow variety can become more valuable, subject to cost\. Under Plackett–Luce,Λ=λ\\Lambda=\\lambda, so selector\-model investment can be evaluated through the fitted discrimination parameter\.

#### Shared modules, latency, and nonadditive costs\.

The additive costC⁡\(𝒎\)=∑g∈𝒢cg​mgC\(\\bm\{m\}\)=\\sum\_\{g\\in\\mathcal\{G\}\}c\_\{g\}m\_\{g\}is appropriate for token or dollar expenditure when all assigned workflow executions are run\. Workflow graphs may share retrieval results, cached model calls, or verification modules, and parallel execution makes latency depend on a critical path rather than a sum\. These features can be represented by fixed\-charge module variables or a set\-dependent costC⁡\(S\)C\(S\)\. The structural accuracy bound remains valid, while the finite\-pool optimization becomes a richer mixed\-integer problem\.

#### Stochastic workflow outputs\.

If correctness is stochastic, the exact selector\-aware contribution of taskiiunder execution\-count vector𝒎\\bm\{m\}is

𝔼⁡\[ψk⁡\(𝒎\)​\(Ri​\(𝒎\)\)\]\.\\mathbb\{E\}\\left\[\\psi\_\{k\(\\bm\{m\}\)\}\\bigl\(R\_\{i\}\(\\bm\{m\}\)\\bigr\)\\right\]\.where the expectation is over the joint distribution of workflow outcomes\. In general this is not equal to applyingψ\\psito the sum of marginal success probabilities\. A practical approach is scenario\-based sample\-average approximation: repeated workflow runs create deterministic correctness scenarios, and the portfolio is optimized against their average selector value\([Kleywegt et al\. 2002](https://arxiv.org/html/2609.18126#bib.bib8)\)\.

#### Task\-dependent run sets\.

The paper chooses one persistent portfolio for a task population\. A pre\-output router could select a task\-specific subset before execution, yielding a joint routing and post\-output selection problem\. The current model provides the value and cost of each candidate run set; an outer policy can then allocate run sets by task features or congestion state\.

#### Voting and nonconcave selectors\.

Majority voting and threshold verification can have increasing marginal value near a decision threshold\. The general selector\-strength bound still applies if a finite common odds\-lift envelope exists, even when the raw recovery curve is nonconcave\. If recovery reaches one at an interior correct fraction, however, the odds\-lift index is infinite and that uniform cap is vacuous\. For each fixed size, the exact IP remains valid with arbitrary nondecreasing increments, but the concave\-coverage rounding certificate may fail\. Endogenous size still matters because adding a wrong vote can move the system away from the threshold and incurs compute\. Selector\-specific approximation or robust envelope methods are natural extensions\.

#### Creation cost versus execution cost\.

Execution costcgc\_\{g\}recurs on every deployed task\. Workflow creation and evaluation costs are one\-time investments\. The generation algorithm can impose a separate budget on oracle calls, candidate evaluation, or human engineering\. A longer planning horizon increases the relative importance of recurring execution cost, while a short\-lived application may place more weight on creation cost\.

#### Robust selector calibration\.

The fittedλ\\lambdamay vary across task types or drift when portfolios are optimized\. For the structural cap, a robust analysis can construct an upper confidence envelopeH\+H^\{\+\}for recovery and useΛ⁡\(H\+\)\\Lambda\(H^\{\+\}\), or use an upper confidence bound forλ\\lambdaunder Plackett–Luce\. For portfolio optimization, one can optimize worst\-case net value over a family of recovery curves or impose a distribution over task\-specific selector strengths\. The cardinality\-grid and IP architecture remains unchanged; only the recovery coefficientsdℓ,kd\_\{\\ell,k\}vary\. If the recovery curves do not share a common concave fraction\-based representation, the exact fixed\-size IPs remain valid but the geometric cardinality approximation should be treated as a Plackett–Luce or common\-fraction specialization rather than a universal guarantee\.

## Appendix FProofs for Stochastic Workflow Outcomes

This appendix gives the complete two\-stage stochastic procedure and its proof\. Stage 1 usesL1L\_\{1\}independent executions to estimate each task–workflow success probability and constructs a plug\-in iid execution law\. Workflow generation, stochastic pricing, and primal recovery are performed under that plug\-in law\. After the generated workflow pool is frozen, Stage 2 discards the Stage\-1 outcomes for selection purposes and usesL2L\_\{2\}fresh execution panels to choose the final portfolio\.

For the analysis, associate with every task–workflow pair\(i,g\)\(i,g\)a potential Stage\-1 pilot panel

\(Zi​g\(1,ℓ\):ℓ=1,…,L1\),Zi​g\(1,ℓ\)∼iidBernoulli\(ai​g\),\\left\(Z\_\{ig\}^\{\(1,\\ell\)\}:\\ell=1,\\ldots,L\_\{1\}\\right\),\\qquad Z\_\{ig\}^\{\(1,\\ell\)\}\\stackrel\{\{\\scriptstyle\\mathrm\{iid\}\}\}\{\{\\sim\}\}\\mathrm\{Bernoulli\}\(a\_\{ig\}\),\(115\)with all panels mutually independent\. The algorithm reveals a panel only when the corresponding workflow is first evaluated\. This lazy revelation is equivalent to drawing the entire finite table in advance and does not require enumerating𝒢\\mathcal\{G\}computationally\.

###### Lemma F\.1\(Uniform concentration of theL1L\_\{1\}workflow estimates\)

Under[Section7](https://arxiv.org/html/2609.18126#S7), for everyδ1∈\(0,1\)\\delta\_\{1\}\\in\(0,1\),

ℙ\{maxi=1,…,ng∈𝒢\|a^i​g\(1\)−ai​g\|≤εa\(L1,δ1\)\}≥1−δ1,\\mathbb\{P\}\\left\\\{\\max\_\{\\begin\{subarray\}\{c\}i=1,\\ldots,n\\\\ g\\in\\mathcal\{G\}\\end\{subarray\}\}\\left\|\\widehat\{a\}\_\{ig\}^\{\(1\)\}\-a\_\{ig\}\\right\|\\leq\\varepsilon\_\{a\}\(L\_\{1\},\\delta\_\{1\}\)\\right\\\}\\geq 1\-\\delta\_\{1\},\(116\)whereεa​\(L1,δ1\)\\varepsilon\_\{a\}\(L\_\{1\},\\delta\_\{1\}\)is defined in \([65](https://arxiv.org/html/2609.18126#S7.E65)\)\.

###### Proof F\.2

For each fixed\(i,g\)\(i,g\), Hoeffding’s inequality gives

ℙ\{\|a^i​g\(1\)−ai​g\|\>t\}≤2e−2​L1​t2\.\\mathbb\{P\}\\left\\\{\\left\|\\widehat\{a\}\_\{ig\}^\{\(1\)\}\-a\_\{ig\}\\right\|\>t\\right\\\}\\leq 2e^\{\-2L\_\{1\}t^\{2\}\}\.A union bound over then​\|𝒢\|n\|\\mathcal\{G\}\|potential pilot panels and the choicet=εa​\(L1,δ1\)t=\\varepsilon\_\{a\}\(L\_\{1\},\\delta\_\{1\}\)prove the result\. Since the complete pilot table can be regarded as drawn before the generation process begins, revealing its entries adaptively does not alter the bound\.□\\square

###### Lemma F\.3\(Perturbation of iid stochastic portfolio values\)

Letb=\(bi​g\)b=\(b\_\{ig\}\)andb′=\(bi​g′\)b^\{\\prime\}=\(b^\{\\prime\}\_\{ig\}\)be two success\-probability matrices such thatmaxi,g⁡\|bi​g−bi​g′\|≤η\\max\_\{i,g\}\|b\_\{ig\}\-b^\{\\prime\}\_\{ig\}\|\\leq\\eta\. LetΠψ,γb\\Pi\_\{\\psi,\\gamma\}^\{b\}andΠψ,γb′\\Pi\_\{\\psi,\\gamma\}^\{b^\{\\prime\}\}denote portfolio values under the corresponding product\-Bernoulli execution laws\. Then

sup𝒎∈ℤ\+𝒢k⁡\(𝒎\)≤Kmax\|Πψ,γb​\(𝒎\)−Πψ,γb′​\(𝒎\)\|≤min⁡\{1,Kmax​η\}\.\\sup\_\{\\begin\{subarray\}\{c\}\\bm\{m\}\\in\\mathbb\{Z\}\_\{\+\}^\{\\mathcal\{G\}\}\\\\ k\(\\bm\{m\}\)\\leq K\_\{\\max\}\\end\{subarray\}\}\\left\|\\Pi\_\{\\psi,\\gamma\}^\{b\}\(\\bm\{m\}\)\-\\Pi\_\{\\psi,\\gamma\}^\{b^\{\\prime\}\}\(\\bm\{m\}\)\\right\|\\leq\\min\\\{1,K\_\{\\max\}\\eta\\\}\.\(117\)Consequently, on the event in \([116](https://arxiv.org/html/2609.18126#A6.E116)\),

sup𝒎∈ℤ\+𝒢k⁡\(𝒎\)≤Kmax\|Π^ψ,γ\(1\)​\(𝒎\)−Πψ,γstoch​\(𝒎\)\|≤εest​\(L1,δ1\)\.\\sup\_\{\\begin\{subarray\}\{c\}\\bm\{m\}\\in\\mathbb\{Z\}\_\{\+\}^\{\\mathcal\{G\}\}\\\\ k\(\\bm\{m\}\)\\leq K\_\{\\max\}\\end\{subarray\}\}\\left\|\\widehat\{\\Pi\}\_\{\\psi,\\gamma\}^\{\(1\)\}\(\\bm\{m\}\)\-\\Pi\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\bm\{m\}\)\\right\|\\leq\\varepsilon\_\{\\mathrm\{est\}\}\(L\_\{1\},\\delta\_\{1\}\)\.\(118\)

###### Proof F\.4

Fix𝐦\\bm\{m\}and couple every Bernoulli execution underbbandb′b^\{\\prime\}using the same independent uniform random variable\. On taskii, one execution of workflowggdiffers under the two laws with probability\|bi​g−bi​g′\|\|b\_\{ig\}\-b^\{\\prime\}\_\{ig\}\|\. Hence the probability that any of thek⁡\(𝐦\)k\(\\bm\{m\}\)coupled execution outcomes differs is at most

∑g∈𝒢mg​\|bi​g−bi​g′\|≤k⁡\(𝒎\)​η\.\\sum\_\{g\\in\\mathcal\{G\}\}m\_\{g\}\|b\_\{ig\}\-b^\{\\prime\}\_\{ig\}\|\\leq k\(\\bm\{m\}\)\\eta\.If no coupled outcome differs, the two correct counts and therefore the two recovery values coincide\. Since recovery lies in\[0,1\]\[0,1\], the difference in expected task value is at mostmin⁡\{1,k⁡\(𝐦\)​η\}\\min\\\{1,k\(\\bm\{m\}\)\\eta\\\}\. Averaging over tasks and observing that the cost term is identical under the two laws prove \([117](https://arxiv.org/html/2609.18126#A6.E117)\)\. The second claim follows from[LemmaF\.1](https://arxiv.org/html/2609.18126#A6.Thmtheorem1)\.□\\square

For a fixed run sizekk, label the execution copies byq=1,…,kq=1,\\ldots,k\. The Stage\-1 estimates are integer multiples of1/L11/L\_\{1\}\. This permits a common finite scenario representation of the entire plug\-in law that does not change when a new workflow is revealed\. Let

Ωk=\{1,…,L1\}n×k,pω=L1−n​k,\\Omega\_\{k\}=\\\{1,\\ldots,L\_\{1\}\\\}^\{n\\times k\},\\qquad p\_\{\\omega\}=L\_\{1\}^\{\-nk\},\(119\)and writeUi​q​\(ω\)∈\{1,…,L1\}U\_\{iq\}\(\\omega\)\\in\\\{1,\\ldots,L\_\{1\}\\\}for coordinate\(i,q\)\(i,q\)of scenarioω\\omega\. For every workflowgg, define its potential plug\-in correctness outcome in scenarioω\\omegaby

A^i​g​q\(1\)\(ω\)=𝟏\{Ui​q\(ω\)≤L1a^i​g\(1\)\}\.\\widehat\{A\}\_\{igq\}^\{\(1\)\}\(\\omega\)=\\mathbf\{1\}\\left\\\{U\_\{iq\}\(\\omega\)\\leq L\_\{1\}\\widehat\{a\}\_\{ig\}^\{\(1\)\}\\right\\\}\.\(120\)For each fixed integral assignment of workflows to execution copies, the variables in \([120](https://arxiv.org/html/2609.18126#A6.E120)\) are independent across tasks and copies and have the required Bernoulli meansa^i​g\(1\)\\widehat\{a\}\_\{ig\}^\{\(1\)\}\. The same uniforms may be used for different workflow choices within one copy because an integral assignment selects only one workflow for that copy\.

For a finite workflow classℳ\\mathcal\{M\}, define

𝒞k​\(ℳ\)=\{𝒎∈ℤ\+ℳ:∑g∈ℳmg=k\},\|𝒞k​\(ℳ\)\|=\(\|ℳ\|\+k−1k\),\\mathcal\{C\}\_\{k\}\(\\mathcal\{M\}\)=\\left\\\{\\bm\{m\}\\in\\mathbb\{Z\}\_\{\+\}^\{\\mathcal\{M\}\}:\\sum\_\{g\\in\\mathcal\{M\}\}m\_\{g\}=k\\right\\\},\\qquad\|\\mathcal\{C\}\_\{k\}\(\\mathcal\{M\}\)\|=\\binom\{\|\\mathcal\{M\}\|\+k\-1\}\{k\},\(121\)and let

OPT^k\(1\)​\(ℳ\)=max𝒎∈𝒞k​\(ℳ\)⁡Π^ψ,γ\(1\)​\(𝒎\)\.\\widehat\{\\mathrm\{OPT\}\}\_\{k\}^\{\(1\)\}\(\\mathcal\{M\}\)=\\max\_\{\\bm\{m\}\\in\\mathcal\{C\}\_\{k\}\(\\mathcal\{M\}\)\}\\widehat\{\\Pi\}\_\{\\psi,\\gamma\}^\{\(1\)\}\(\\bm\{m\}\)\.\(122\)
###### Lemma F\.5\(Plug\-in stochastic LP relaxation and projected dual\)

Fixkkand supposed1,k≥⋯≥dk,k≥0d\_\{1,k\}\\geq\\cdots\\geq d\_\{k,k\}\\geq 0\. Letψ~k\\widetilde\{\\psi\}\_\{k\}be defined by \([38](https://arxiv.org/html/2609.18126#S5.E38)\), and set

𝒳k\(𝒢\)=\{x≥0:∑g∈𝒢xg​q=1,q=1,…,k\}\.\\mathcal\{X\}\_\{k\}\(\\mathcal\{G\}\)=\\left\\\{x\\geq 0:\\sum\_\{g\\in\\mathcal\{G\}\}x\_\{gq\}=1,\\ q=1,\\ldots,k\\right\\\}\.\(123\)Forx∈𝒳k​\(𝒢\)x\\in\\mathcal\{X\}\_\{k\}\(\\mathcal\{G\}\), define

s^i​ω\(1\)​\(x\)\\displaystyle\\widehat\{s\}\_\{i\\omega\}^\{\(1\)\}\(x\):=∑q=1k∑g∈𝒢A^i​g​q\(1\)​\(ω\)​xg​q,\\displaystyle:=\{\}\\sum\_\{q=1\}^\{k\}\\sum\_\{g\\in\\mathcal\{G\}\}\\widehat\{A\}\_\{igq\}^\{\(1\)\}\(\\omega\)x\_\{gq\},\(124\)ℱ^k\(1\)​\(x\)\\displaystyle\\widehat\{\\mathcal\{F\}\}\_\{k\}^\{\(1\)\}\(x\):=∑ω∈Ωkpω​1n​∑i=1nψ~k​\(s^i​ω\(1\)​\(x\)\)−γ​∑q=1k∑g∈𝒢cg​xg​q,\\displaystyle:=\{\}\\sum\_\{\\omega\\in\\Omega\_\{k\}\}p\_\{\\omega\}\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\widetilde\{\\psi\}\_\{k\}\\bigl\(\\widehat\{s\}\_\{i\\omega\}^\{\(1\)\}\(x\)\\bigr\)\-\\gamma\\sum\_\{q=1\}^\{k\}\\sum\_\{g\\in\\mathcal\{G\}\}c\_\{g\}x\_\{gq\},\(125\)ℒ^k\(1\)​\(𝒢\)\\displaystyle\\widehat\{\\mathcal\{L\}\}\_\{k\}^\{\(1\)\}\(\\mathcal\{G\}\):=maxx∈𝒳k​\(𝒢\)⁡ℱ^k\(1\)​\(x\)\.\\displaystyle:=\{\}\\max\_\{x\\in\\mathcal\{X\}\_\{k\}\(\\mathcal\{G\}\)\}\\widehat\{\\mathcal\{F\}\}\_\{k\}^\{\(1\)\}\(x\)\.\(126\)Then

ℒ^k\(1\)​\(𝒢\)≥OPT^k\(1\)​\(𝒢\)\.\\widehat\{\\mathcal\{L\}\}\_\{k\}^\{\(1\)\}\(\\mathcal\{G\}\)\\geq\\widehat\{\\mathrm\{OPT\}\}\_\{k\}^\{\(1\)\}\(\\mathcal\{G\}\)\.\(127\)Define

χk​\(u\)=maxr=0,…,k⁡\{ψk​\(r\)n−u​r\},0≤u≤d1,kn\.\\chi\_\{k\}\(u\)=\\max\_\{r=0,\\ldots,k\}\\left\\\{\\frac\{\\psi\_\{k\}\(r\)\}\{n\}\-ur\\right\\\},\\qquad 0\\leq u\\leq\\frac\{d\_\{1,k\}\}\{n\}\.\(128\)Strong duality gives

ℒ^k\(1\)​\(𝒢\)=minμ,θ\\displaystyle\\widehat\{\\mathcal\{L\}\}\_\{k\}^\{\(1\)\}\(\\mathcal\{G\}\)=\\min\_\{\\mu,\\theta\}\\quad∑ω∈Ωkpω​∑i=1nχk​\(μi​ω\)\+∑q=1kθq\\displaystyle\\sum\_\{\\omega\\in\\Omega\_\{k\}\}p\_\{\\omega\}\\sum\_\{i=1\}^\{n\}\\chi\_\{k\}\(\\mu\_\{i\\omega\}\)\+\\sum\_\{q=1\}^\{k\}\\theta\_\{q\}\(129\)s\.t\.∑ω∈Ωkpω​∑i=1nμi​ω​A^i​g​q\(1\)​\(ω\)−γ​cg≤θq,\\displaystyle\\sum\_\{\\omega\\in\\Omega\_\{k\}\}p\_\{\\omega\}\\sum\_\{i=1\}^\{n\}\\mu\_\{i\\omega\}\\widehat\{A\}\_\{igq\}^\{\(1\)\}\(\\omega\)\-\\gamma c\_\{g\}\\leq\\theta\_\{q\},g∈𝒢,q=1,…,k,\\displaystyle g\\in\\mathcal\{G\},\\ q=1,\\ldots,k,0≤μi​ω≤d1,kn,\\displaystyle 0\\leq\\mu\_\{i\\omega\}\\leq\\frac\{d\_\{1,k\}\}\{n\},i=1,…,n,ω∈Ωk\.\\displaystyle i=1,\\ldots,n,\\ \\omega\\in\\Omega\_\{k\}\.Equivalently, introducing epigraph variablesξi​ω\\xi\_\{i\\omega\}gives the rational LP

ℒ^k\(1\)​\(𝒢\)=minμ,ξ,θ\\displaystyle\\widehat\{\\mathcal\{L\}\}\_\{k\}^\{\(1\)\}\(\\mathcal\{G\}\)=\\min\_\{\\mu,\\xi,\\theta\}\\quad∑ω∈Ωkpω​∑i=1nξi​ω\+∑q=1kθq\\displaystyle\\sum\_\{\\omega\\in\\Omega\_\{k\}\}p\_\{\\omega\}\\sum\_\{i=1\}^\{n\}\\xi\_\{i\\omega\}\+\\sum\_\{q=1\}^\{k\}\\theta\_\{q\}\(130\)s\.t\.ξi​ω\+r​μi​ω≥ψk​\(r\)n,\\displaystyle\\xi\_\{i\\omega\}\+r\\mu\_\{i\\omega\}\\geq\\frac\{\\psi\_\{k\}\(r\)\}\{n\},i=1,…,n,ω∈Ωk,r=0,…,k,\\displaystyle i=1,\\ldots,n,\\ \\omega\\in\\Omega\_\{k\},\\ r=0,\\ldots,k,∑ω∈Ωkpω​∑i=1nμi​ω​A^i​g​q\(1\)​\(ω\)−γ​cg≤θq,\\displaystyle\\sum\_\{\\omega\\in\\Omega\_\{k\}\}p\_\{\\omega\}\\sum\_\{i=1\}^\{n\}\\mu\_\{i\\omega\}\\widehat\{A\}\_\{igq\}^\{\(1\)\}\(\\omega\)\-\\gamma c\_\{g\}\\leq\\theta\_\{q\},g∈𝒢,q=1,…,k,\\displaystyle g\\in\\mathcal\{G\},\\ q=1,\\ldots,k,0≤μi​ω≤d1,kn,0≤ξi​ω≤1n,\\displaystyle 0\\leq\\mu\_\{i\\omega\}\\leq\\frac\{d\_\{1,k\}\}\{n\},\\qquad 0\\leq\\xi\_\{i\\omega\}\\leq\\frac\{1\}\{n\},i=1,…,n,ω∈Ωk,\\displaystyle i=1,\\ldots,n,\\ \\omega\\in\\Omega\_\{k\},−γ​cmax≤θq≤d1,k,\\displaystyle\-\\gamma c\_\{\\max\}\\leq\\theta\_\{q\}\\leq d\_\{1,k\},q=1,…,k\.\\displaystyle q=1,\\ldots,k\.The plug\-in stochastic pricing score is therefore

𝒮^k​q\(1\)​\(g,μ\)=∑ω∈Ωkpω​∑i=1nμi​ω​A^i​g​q\(1\)​\(ω\)−γ​cg\.\\widehat\{\\mathcal\{S\}\}\_\{kq\}^\{\(1\)\}\(g;\\mu\)=\\sum\_\{\\omega\\in\\Omega\_\{k\}\}p\_\{\\omega\}\\sum\_\{i=1\}^\{n\}\\mu\_\{i\\omega\}\\widehat\{A\}\_\{igq\}^\{\(1\)\}\(\\omega\)\-\\gamma c\_\{g\}\.\(131\)

###### Proof F\.6

An integralxxassigns one workflow to every execution copyqqand induces countsmg=∑qxg​qm\_\{g\}=\\sum\_\{q\}x\_\{gq\}\. Its scenario\-wise correct counts are integral, soψ~k=ψk\\widetilde\{\\psi\}\_\{k\}=\\psi\_\{k\}at those counts\. By the construction in \([120](https://arxiv.org/html/2609.18126#A6.E120)\), its scenario distribution is exactly the Stage\-1 plug\-in iid law\. Thus its objective isΠ^ψ,γ\(1\)​\(𝐦\)\\widehat\{\\Pi\}\_\{\\psi,\\gamma\}^\{\(1\)\}\(\\bm\{m\}\), and relaxing integrality proves \([127](https://arxiv.org/html/2609.18126#A6.E127)\)\.

Letφk​\(t\)=ψ~k​\(t\)/n\\varphi\_\{k\}\(t\)=\\widetilde\{\\psi\}\_\{k\}\(t\)/n\. Concavity gives

φk​\(t\)=min0≤u≤d1,k/n⁡\{u​t\+χk​\(u\)\},0≤t≤k\.\\varphi\_\{k\}\(t\)=\\min\_\{0\\leq u\\leq d\_\{1,k\}/n\}\\\{ut\+\\chi\_\{k\}\(u\)\\\},\\qquad 0\\leq t\\leq k\.\(132\)Substitute \([132](https://arxiv.org/html/2609.18126#A6.E132)\) into \([125](https://arxiv.org/html/2609.18126#A6.E125)\) and apply finite\-dimensional minimax to interchange maximization overxxand minimization overμ\\mu\. For fixedμ\\mu, maximization separates by execution copy:

maxx∈𝒳k​\(𝒢\)∑q=1k∑g∈𝒢xg​q𝒮^k​q\(1\)\(g;μ\)\\displaystyle\\max\_\{x\\in\\mathcal\{X\}\_\{k\}\(\\mathcal\{G\}\)\}\\sum\_\{q=1\}^\{k\}\\sum\_\{g\\in\\mathcal\{G\}\}x\_\{gq\}\\widehat\{\\mathcal\{S\}\}\_\{kq\}^\{\(1\)\}\(g;\\mu\)=∑q=1kmaxg∈𝒢⁡𝒮^k​q\(1\)​\(g,μ\)\.\\displaystyle\\hskip 85\.35826pt=\\sum\_\{q=1\}^\{k\}\\max\_\{g\\in\\mathcal\{G\}\}\\widehat\{\\mathcal\{S\}\}\_\{kq\}^\{\(1\)\}\(g;\\mu\)\.\(133\)Introducing the variablesθq\\theta\_\{q\}yields \([129](https://arxiv.org/html/2609.18126#A6.E129)\); replacing eachχk\\chi\_\{k\}by its finite epigraph representation yields \([130](https://arxiv.org/html/2609.18126#A6.E130)\)\.□\\square

Even under iid execution, the score in \([131](https://arxiv.org/html/2609.18126#A6.E131)\) is not generally equal to the deterministic mean\-column score∑iμ¯i​a^i​g\(1\)−γ​cg\\sum\_\{i\}\\bar\{\\mu\}\_\{i\}\\widehat\{a\}\_\{ig\}^\{\(1\)\}\-\\gamma c\_\{g\}\. The scenario\-dependent marginal priceμi​ω\\mu\_\{i\\omega\}and the candidate’s correctness indicator are evaluated in the same execution scenario, so averaging them separately need not preserve their product\.

Algorithm 2IID Stochastic Dual\-Guided Workflow Optimization1:Input: initial workflow pool

ℳ0\\mathcal\{M\}\_\{0\}; cardinality grid

𝒦\\mathcal\{K\}; recovery curves

\{ψk\}\\\{\\psi\_\{k\}\\\}; cost price

γ\\gamma; Stage\-1 sample size

L1L\_\{1\}and confidence level

δ1\\delta\_\{1\}; Stage\-2 sample size

L2L\_\{2\}and confidence level

δ2\\delta\_\{2\}; ellipsoid tolerance

εell\>0\\varepsilon\_\{\\mathrm\{ell\}\}\>0; pricing\-batch size

mm; primitive stochastic pricing oracle; baseline

𝒎0\\bm\{m\}^\{0\}with certified value

V¯\>0\\underline\{V\}\>0and

supp⁡\(𝒎0\)⊆ℳ0\\operatorname\{supp\}\(\\bm\{m\}^\{0\}\)\\subseteq\\mathcal\{M\}\_\{0\}\.

2:For every

g∈ℳ0g\\in\\mathcal\{M\}\_\{0\}and task

ii, collect

L1L\_\{1\}independent executions and compute

a^i​g\(1\)\\widehat\{a\}\_\{ig\}^\{\(1\)\}using \([64](https://arxiv.org/html/2609.18126#S7.E64)\)\. Set

ℳrec←ℳ0\\mathcal\{M\}\_\{\\mathrm\{rec\}\}\\leftarrow\\mathcal\{M\}\_\{0\}\.

3:for

k∈𝒦k\\in\\mathcal\{K\}do

4:Set

τk←εell/\(2​k\)\\tau\_\{k\}\\leftarrow\\varepsilon\_\{\\mathrm\{ell\}\}/\(2k\)\.

5:Run the ellipsoid method on the plug\-in epigraph dual \([130](https://arxiv.org/html/2609.18126#A6.E130)\) with optimization tolerance

εell/2\\varepsilon\_\{\\mathrm\{ell\}\}/2\.

6:whilethe ellipsoid routine has not terminateddo

7:At the current point

\(μt,ξt,θt\)\(\\mu^\{t\},\\xi^\{t\},\\theta^\{t\}\), separate violated explicit recovery and box constraints directly\.

8:ifno explicit constraint is violatedthen

9:for

q=1,…,kq=1,\\ldots,kdo

10:Make

mmprimitive pricing calls at

\(k,q,μt,τk\)\(k,q,\\mu^\{t\},\\tau\_\{k\}\)\. Whenever a call proposes a workflow

ggthat has not been evaluated, collect its independent

L1L\_\{1\}\-sample panel, compute

\(a^1​g\(1\),…,a^n​g\(1\)\)\(\\widehat\{a\}\_\{1g\}^\{\(1\)\},\\ldots,\\widehat\{a\}\_\{ng\}^\{\(1\)\}\), and set

ℳrec←ℳrec∪\{g\}\\mathcal\{M\}\_\{\\mathrm\{rec\}\}\\leftarrow\\mathcal\{M\}\_\{\\mathrm\{rec\}\}\\cup\\\{g\\\}before scoring it\.

11:Retain the proposed workflow

gt​qg^\{tq\}with the largest score

𝒮^k​q\(1\)​\(g,μt\)\\widehat\{\\mathcal\{S\}\}\_\{kq\}^\{\(1\)\}\(g;\\mu^\{t\}\)\.

12:if

𝒮^k​q\(1\)​\(gt​q,μt\)\>θqt\\widehat\{\\mathcal\{S\}\}\_\{kq\}^\{\(1\)\}\(g^\{tq\};\\mu^\{t\}\)\>\\theta\_\{q\}^\{t\}then

13:Add its violated workflow constraint to the ellipsoid system\.

14:else

15:Use the weak\-separation certificate

maxg∈𝒢⁡𝒮^k​q\(1\)​\(g,μt\)≤θqt\+τk\\max\_\{g\\in\\mathcal\{G\}\}\\widehat\{\\mathcal\{S\}\}\_\{kq\}^\{\(1\)\}\(g;\\mu^\{t\}\)\\leq\\theta\_\{q\}^\{t\}\+\\tau\_\{k\}\.

16:endif

17:endfor

18:endif

19:endwhile

20:By primal recovery, construct

xk∈𝒳k​\(𝒢\)x^\{k\}\\in\\mathcal\{X\}\_\{k\}\(\\mathcal\{G\}\)supported on

ℳrec\\mathcal\{M\}\_\{\\mathrm\{rec\}\}such that

ℱ^k\(1\)​\(xk\)≥ℒ^k\(1\)​\(𝒢\)−εell\\widehat\{\\mathcal\{F\}\}\_\{k\}^\{\(1\)\}\(x^\{k\}\)\\geq\\widehat\{\\mathcal\{L\}\}\_\{k\}^\{\(1\)\}\(\\mathcal\{G\}\)\-\\varepsilon\_\{\\mathrm\{ell\}\}\.

21:endfor

22:Freeze the generated pool

ℳT←ℳrec\\mathcal\{M\}\_\{T\}\\leftarrow\\mathcal\{M\}\_\{\\mathrm\{rec\}\}\. Do not reuse Stage\-1 execution outcomes for final portfolio selection\.

23:Draw the

L2L\_\{2\}independent Stage\-2 execution panels

\{Zi​g​q\(2,ℓ\)\}\\\{Z\_\{igq\}^\{\(2,\\ell\)\}\\\}, compute

Π^L2final\\widehat\{\\Pi\}\_\{L\_\{2\}\}^\{\\mathrm\{final\}\}for every

𝒎∈𝒞𝒦​\(ℳT\)\\bm\{m\}\\in\\mathcal\{C\}\_\{\\mathcal\{K\}\}\(\\mathcal\{M\}\_\{T\}\), and choose the empirical maximizer

𝒎~\\widetilde\{\\bm\{m\}\}\.

24:Return

𝒎^\\widehat\{\\bm\{m\}\}according to the conservative rule \([71](https://arxiv.org/html/2609.18126#S7.E71)\)\.

At a query\(k,q,μ\)\(k,q,\\mu\), a primitive plug\-in pricing call isτk\\tau\_\{k\}\-accurate if it returnsg~\\widetilde\{g\}satisfying

𝒮^k​q\(1\)​\(g~,μ\)≥maxg∈𝒢⁡𝒮^k​q\(1\)​\(g,μ\)−τk\.\\widehat\{\\mathcal\{S\}\}\_\{kq\}^\{\(1\)\}\(\\widetilde\{g\};\\mu\)\\geq\\max\_\{g\\in\\mathcal\{G\}\}\\widehat\{\\mathcal\{S\}\}\_\{kq\}^\{\(1\)\}\(g;\\mu\)\-\\tau\_\{k\}\.\(134\)The guarantee is defined relative to the complete potential Stage\-1 pilot table in \([115](https://arxiv.org/html/2609.18126#A6.E115)\); only the panels of workflows actually proposed by the oracle need to be materialized\.

###### Lemma F\.7\(Batched plug\-in stochastic pricing\)

Suppose each primitive pricing call satisfies \([134](https://arxiv.org/html/2609.18126#A6.E134)\), conditional on the complete query history, with probability at leastporc\>0p\_\{\\mathrm\{orc\}\}\>0\. If each pricing query usesmmprimitive calls and retains the highest\-scoring returned workflow, letℰorc\\mathcal\{E\}\_\{\\mathrm\{orc\}\}be the event that every resulting batch isτk\\tau\_\{k\}\-accurate\. If at mostNpricestochN\_\{\\mathrm\{price\}\}^\{\\mathrm\{stoch\}\}batches are used, then

ℙ⁡\(ℰorc\)≥1−Npricestoch​e−porc​m\.\\mathbb\{P\}\(\\mathcal\{E\}\_\{\\mathrm\{orc\}\}\)\\geq 1\-N\_\{\\mathrm\{price\}\}^\{\\mathrm\{stoch\}\}e^\{\-p\_\{\\mathrm\{orc\}\}m\}\.\(135\)

###### Proof F\.8

A batch fails only if allmmprimitive calls fail, which has conditional probability at most\(1−porc\)m≤e−porc​m\(1\-p\_\{\\mathrm\{orc\}\}\)^\{m\}\\leq e^\{\-p\_\{\\mathrm\{orc\}\}m\}\. A union bound over the pricing batches proves the claim\.□\\square

###### Lemma F\.9\(Ellipsoid recovery under plug\-in iid pricing\)

For eachk∈𝒦k\\in\\mathcal\{K\}, run the plug\-in stochastic dual with ellipsoid optimization toleranceεell,k/2\\varepsilon\_\{\\mathrm\{ell\},k\}/2and pricing toleranceτk=εell,k/\(2​k\)\\tau\_\{k\}=\\varepsilon\_\{\\mathrm\{ell\},k\}/\(2k\)\. Onℰorc\\mathcal\{E\}\_\{\\mathrm\{orc\}\}, primal recovery returns anxk∈𝒳k​\(𝒢\)x^\{k\}\\in\\mathcal\{X\}\_\{k\}\(\\mathcal\{G\}\)supported on the returned workflow poolℳT\\mathcal\{M\}\_\{T\}such that

ℱ^k\(1\)​\(xk\)≥ℒ^k\(1\)​\(𝒢\)−εell,k\.\\widehat\{\\mathcal\{F\}\}\_\{k\}^\{\(1\)\}\(x^\{k\}\)\\geq\\widehat\{\\mathcal\{L\}\}\_\{k\}^\{\(1\)\}\(\\mathcal\{G\}\)\-\\varepsilon\_\{\\mathrm\{ell\},k\}\.\(136\)

###### Proof F\.10

Successful plug\-in pricing supplies weak separation\. If the best returned score for copyqqis at mostθq\\theta\_\{q\}, then \([134](https://arxiv.org/html/2609.18126#A6.E134)\) implies that all copy\-qqworkflow constraints hold after replacingθq\\theta\_\{q\}byθq\+τk\\theta\_\{q\}\+\\tau\_\{k\}\. Across allkkcopies, this increases the dual objective by at mostk​τk=εell,k/2k\\tau\_\{k\}=\\varepsilon\_\{\\mathrm\{ell\},k\}/2\. Ellipsoid optimization contributes the remaining half\. Standard optimization–separation and primal recovery give \([136](https://arxiv.org/html/2609.18126#A6.E136)\)\.□\\square

###### Lemma F\.11\(Rounding the plug\-in stochastic LP\)

Fixkkandx∈𝒳k​\(𝒢\)x\\in\\mathcal\{X\}\_\{k\}\(\\mathcal\{G\}\)\. Independently for every copyqq, drawGqG\_\{q\}withℙ\{Gq=g\}=xg​q\\mathbb\{P\}\\\{G\_\{q\}=g\\\}=x\_\{gq\}and define

Mk,gRR\(x\)=∑q=1k𝟏\{Gq=g\},𝑴kRR\(x\)=\(Mk,gRR\(x\):g∈𝒢\)\.M\_\{k,g\}^\{\\mathrm\{RR\}\}\(x\)=\\sum\_\{q=1\}^\{k\}\\mathbf\{1\}\\\{G\_\{q\}=g\\\},\\qquad\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\(x\)=\\bigl\(M\_\{k,g\}^\{\\mathrm\{RR\}\}\(x\):g\\in\\mathcal\{G\}\\bigr\)\.Let

Q^k\(1\)​\(x\)\\displaystyle\\widehat\{Q\}\_\{k\}^\{\(1\)\}\(x\):=∑ω∈Ωkpω​1n​∑i=1nψ~k​\(s^i​ω\(1\)​\(x\)\),\\displaystyle:=\{\}\\sum\_\{\\omega\\in\\Omega\_\{k\}\}p\_\{\\omega\}\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\widetilde\{\\psi\}\_\{k\}\\bigl\(\\widehat\{s\}\_\{i\\omega\}^\{\(1\)\}\(x\)\\bigr\),\(137\)R^k\(1\)​\(x\)\\displaystyle\\widehat\{R\}\_\{k\}^\{\(1\)\}\(x\):=∑ω∈Ωkpω​1n​∑i=1ns^i​ω\(1\)​\(x\),Ck​\(x\):=∑q=1k∑g∈𝒢cg​xg​q\.\\displaystyle:=\{\}\\sum\_\{\\omega\\in\\Omega\_\{k\}\}p\_\{\\omega\}\\frac\{1\}\{n\}\\sum\_\{i=1\}^\{n\}\\widehat\{s\}\_\{i\\omega\}^\{\(1\)\}\(x\),\\qquad C\_\{k\}\(x\):=\{\}\\sum\_\{q=1\}^\{k\}\\sum\_\{g\\in\\mathcal\{G\}\}c\_\{g\}x\_\{gq\}\.\(138\)Then

𝔼rnd​\[Π^ψ,γ\(1\)​\(𝑴kRR​\(x\)\)\]\\displaystyle\\mathbb\{E\}\_\{\\mathrm\{rnd\}\}\\\!\\left\[\\widehat\{\\Pi\}\_\{\\psi,\\gamma\}^\{\(1\)\}\\bigl\(\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\(x\)\\bigr\)\\right\]≥\(1−1e\)​Q^k\(1\)​\(x\)\+dk,ke​R^k\(1\)​\(x\)−γ​Ck​\(x\),\\displaystyle\\geq\\left\(1\-\\frac\{1\}\{e\}\\right\)\\widehat\{Q\}\_\{k\}^\{\(1\)\}\(x\)\+\\frac\{d\_\{k,k\}\}\{e\}\\widehat\{R\}\_\{k\}^\{\(1\)\}\(x\)\-\\gamma C\_\{k\}\(x\),\(139\)ℱ^k\(1\)​\(x\)−𝔼rnd​\[Π^ψ,γ\(1\)​\(𝑴kRR​\(x\)\)\]\\displaystyle\\widehat\{\\mathcal\{F\}\}\_\{k\}^\{\(1\)\}\(x\)\-\\mathbb\{E\}\_\{\\mathrm\{rnd\}\}\\\!\\left\[\\widehat\{\\Pi\}\_\{\\psi,\\gamma\}^\{\(1\)\}\\bigl\(\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\(x\)\\bigr\)\\right\]≤1e​\{Q^k\(1\)​\(x\)−dk,k​R^k\(1\)​\(x\)\}≤cψ,ke​Q^k\(1\)​\(x\)\.\\displaystyle\\leq\\frac\{1\}\{e\}\\left\\\{\\widehat\{Q\}\_\{k\}^\{\(1\)\}\(x\)\-d\_\{k,k\}\\widehat\{R\}\_\{k\}^\{\(1\)\}\(x\)\\right\\\}\\leq\\frac\{c\_\{\\psi,k\}\}\{e\}\\widehat\{Q\}\_\{k\}^\{\(1\)\}\(x\)\.\(140\)SinceQ^k\(1\)​\(x\)≤1\\widehat\{Q\}\_\{k\}^\{\(1\)\}\(x\)\\leq 1, a uniform rounding loss is

εrnd,k=cψ,ke\.\\varepsilon\_\{\\mathrm\{rnd\},k\}=\\frac\{c\_\{\\psi,k\}\}\{e\}\.\(141\)

###### Proof F\.12

Fix\(i,ω\)\(i,\\omega\)and setBq=A^i,Gq,q\(1\)​\(ω\)B\_\{q\}=\\widehat\{A\}\_\{i,G\_\{q\},q\}^\{\(1\)\}\(\\omega\)andpq=∑gxg​q​A^i​g​q\(1\)​\(ω\)p\_\{q\}=\\sum\_\{g\}x\_\{gq\}\\widehat\{A\}\_\{igq\}^\{\(1\)\}\(\\omega\)\. Conditional onω\\omega, theBqB\_\{q\}are independent under the rounding andN=∑qBqN=\\sum\_\{q\}B\_\{q\}has means^i​ω\(1\)​\(x\)\\widehat\{s\}\_\{i\\omega\}^\{\(1\)\}\(x\)\. For every integert≥1t\\geq 1,

min⁡\{r,t\}≥t⁡\{1−\(1−1/t\)r\}\.\\min\\\{r,t\\\}\\geq t\\\{1\-\(1\-1/t\)^\{r\}\\\}\.\(142\)Hence

𝔼rnd​\[min⁡\{N,t\}∣ω\]\\displaystyle\\mathbb\{E\}\_\{\\mathrm\{rnd\}\}\[\\min\\\{N,t\\\}\\mid\\omega\]≥t⁡\[1−∏q=1k\(1−pqt\)\]\\displaystyle\\geq t\\left\[1\-\\prod\_\{q=1\}^\{k\}\\left\(1\-\\frac\{p\_\{q\}\}\{t\}\\right\)\\right\]≥t\(1−e−s^i​ω\(1\)\(x\)/t\)\\displaystyle\\geq t\\left\(1\-e^\{\-\\widehat\{s\}\_\{i\\omega\}^\{\(1\)\}\(x\)/t\}\\right\)≥\(1−1e\)​min⁡\{s^i​ω\(1\)​\(x\),t\}\.\\displaystyle\\geq\\left\(1\-\\frac\{1\}\{e\}\\right\)\\min\\\{\\widehat\{s\}\_\{i\\omega\}^\{\(1\)\}\(x\),t\\\}\.\(143\)Using

ψ~k​\(r\)=dk,k​r\+∑t=1k−1\(dt,k−dt\+1,k\)​min⁡\{r,t\},\\widetilde\{\\psi\}\_\{k\}\(r\)=d\_\{k,k\}r\+\\sum\_\{t=1\}^\{k\-1\}\(d\_\{t,k\}\-d\_\{t\+1,k\}\)\\min\\\{r,t\\\},\(144\)and averaging over tasks and scenarios prove \([139](https://arxiv.org/html/2609.18126#A6.E139)\)\. Rounding also preserves execution cost in expectation\. Subtraction gives the first inequality in \([140](https://arxiv.org/html/2609.18126#A6.E140)\)\. Finally,ψ~k​\(r\)≤d1,k​r\\widetilde\{\\psi\}\_\{k\}\(r\)\\leq d\_\{1,k\}randdk,k=\(1−cψ,k\)​d1,kd\_\{k,k\}=\(1\-c\_\{\\psi,k\}\)d\_\{1,k\}give the curvature bound, while0≤ψ~k​\(r\)≤10\\leq\\widetilde\{\\psi\}\_\{k\}\(r\)\\leq 1givesQ^k\(1\)​\(x\)≤1\\widehat\{Q\}\_\{k\}^\{\(1\)\}\(x\)\\leq 1\.□\\square

###### Lemma F\.13\(Stochastic cardinality grid under the plug\-in iid law\)

Supposeψk​\(r\)=h⁡\(r/k\)\\psi\_\{k\}\(r\)=h\(r/k\)for a common increasing concave functionh:\[0,1\]→\[0,1\]h:\[0,1\]\\to\[0,1\]withh⁡\(0\)=0h\(0\)=0\. For every workflow classℳ\\mathcal\{M\}and1≤b≤k≤Kmax1\\leq b\\leq k\\leq K\_\{\\max\},

\{OPT^b\(1\)​\(ℳ\)\}\+≥bk​\{OPT^k\(1\)​\(ℳ\)\}\+\.\\left\\\{\\widehat\{\\mathrm\{OPT\}\}\_\{b\}^\{\(1\)\}\(\\mathcal\{M\}\)\\right\\\}^\{\+\}\\geq\\frac\{b\}\{k\}\\left\\\{\\widehat\{\\mathrm\{OPT\}\}\_\{k\}^\{\(1\)\}\(\\mathcal\{M\}\)\\right\\\}^\{\+\}\.\(145\)Consequently, if𝒦\\mathcal\{K\}has coverage ratioϱ\\varrho,

max⁡\{0,maxb∈𝒦⁡OPT^b\(1\)​\(ℳ\)\}≥1ϱ​OPT^ψ,γ\(1\)​\(ℳ\),\\max\\left\\\{0,\\max\_\{b\\in\\mathcal\{K\}\}\\widehat\{\\mathrm\{OPT\}\}\_\{b\}^\{\(1\)\}\(\\mathcal\{M\}\)\\right\\\}\\geq\\frac\{1\}\{\\varrho\}\\widehat\{\\mathrm\{OPT\}\}\_\{\\psi,\\gamma\}^\{\(1\)\}\(\\mathcal\{M\}\),\(146\)where

OPT^ψ,γ\(1\)​\(ℳ\)=max𝒎∈ℤ\+ℳk⁡\(𝒎\)≤Kmax⁡Π^ψ,γ\(1\)​\(𝒎\)\.\\widehat\{\\mathrm\{OPT\}\}\_\{\\psi,\\gamma\}^\{\(1\)\}\(\\mathcal\{M\}\)=\\max\_\{\\begin\{subarray\}\{c\}\\bm\{m\}\\in\\mathbb\{Z\}\_\{\+\}^\{\\mathcal\{M\}\}\\\\ k\(\\bm\{m\}\)\\leq K\_\{\\max\}\\end\{subarray\}\}\\widehat\{\\Pi\}\_\{\\psi,\\gamma\}^\{\(1\)\}\(\\bm\{m\}\)\.

###### Proof F\.14

Fix𝐦∈𝒞k​\(ℳ\)\\bm\{m\}\\in\\mathcal\{C\}\_\{k\}\(\\mathcal\{M\}\), label itskkexecution copies, retainbbof them uniformly without replacement, and let𝐌\(b\)\\bm\{M\}^\{\(b\)\}be the retained count vector\. Setα=b/k\\alpha=b/k\. Conditional on a realized plug\-in size\-kkexecution, if taskiihasrir\_\{i\}correct copies andRiR\_\{i\}of them are retained, thenRi∼Hypergeom⁡\(k,ri,b\)R\_\{i\}\\sim\\mathrm\{Hypergeom\}\(k,r\_\{i\},b\)\. Forri\>0r\_\{i\}\>0, setpi=ri/kp\_\{i\}=r\_\{i\}/kandXi=Ri/bX\_\{i\}=R\_\{i\}/b\. Concavity and monotonicity give

h⁡\(x\)≥h⁡\(pi\)​min⁡\{x/pi,1\},x∈\[0,1\]\.h\(x\)\\geq h\(p\_\{i\}\)\\min\\\{x/p\_\{i\},1\\\},\\qquad x\\in\[0,1\]\.\(147\)WritingZi=Ri/riZ\_\{i\}=R\_\{i\}/r\_\{i\}, we haveXi/pi=Zi/αX\_\{i\}/p\_\{i\}=Z\_\{i\}/\\alpha,𝔼⁡\[Zi\]=α\\mathbb\{E\}\[Z\_\{i\}\]=\\alpha, andmin⁡\{z,α\}≥α​z\\min\\\{z,\\alpha\\\}\\geq\\alpha zon\[0,1\]\[0,1\]\. Hence𝔼⁡\[h⁡\(Ri/b\)\]≥α​h​\(ri/k\)\\mathbb\{E\}\[h\(R\_\{i\}/b\)\]\\geq\\alpha h\(r\_\{i\}/k\)\. The caseri=0r\_\{i\}=0is immediate\.

By the iid plug\-in execution law, any retained set of labeled copies has the same distribution as a direct run of the retained execution\-count vector, and𝔼⁡\[C⁡\(𝐌\(b\)\)\]=α​C​\(𝐦\)\\mathbb\{E\}\[C\(\\bm\{M\}^\{\(b\)\}\)\]=\\alpha C\(\\bm\{m\}\)\. Therefore

𝔼⁡\[Π^ψ,γ\(1\)​\(𝑴\(b\)\)\]≥α​Π^ψ,γ\(1\)​\(𝒎\)\.\\mathbb\{E\}\\left\[\\widehat\{\\Pi\}\_\{\\psi,\\gamma\}^\{\(1\)\}\(\\bm\{M\}^\{\(b\)\}\)\\right\]\\geq\\alpha\\widehat\{\\Pi\}\_\{\\psi,\\gamma\}^\{\(1\)\}\(\\bm\{m\}\)\.Some retained multiset attains at least this expected value\. Applying the result to an optimal positive\-valued size\-kkvector proves \([145](https://arxiv.org/html/2609.18126#A6.E145)\); choosing a grid anchorb≤k≤ϱ​bb\\leq k\\leq\\varrho bproves \([146](https://arxiv.org/html/2609.18126#A6.E146)\)\.□\\square

###### Lemma F\.15\(Plug\-in generated\-pool certificate\)

Onℰorc\\mathcal\{E\}\_\{\\mathrm\{orc\}\}, define

εell=maxk∈𝒦⁡εell,k,εrnd=maxk∈𝒦⁡εrnd,k,\\varepsilon\_\{\\mathrm\{ell\}\}=\\max\_\{k\\in\\mathcal\{K\}\}\\varepsilon\_\{\\mathrm\{ell\},k\},\\qquad\\varepsilon\_\{\\mathrm\{rnd\}\}=\\max\_\{k\\in\\mathcal\{K\}\}\\varepsilon\_\{\\mathrm\{rnd\},k\},and let

V^T\(1\)=max𝒎∈𝒞𝒦​\(ℳT\)⁡Π^ψ,γ\(1\)​\(𝒎\)\.\\widehat\{V\}\_\{T\}^\{\(1\)\}=\\max\_\{\\bm\{m\}\\in\\mathcal\{C\}\_\{\\mathcal\{K\}\}\(\\mathcal\{M\}\_\{T\}\)\}\\widehat\{\\Pi\}\_\{\\psi,\\gamma\}^\{\(1\)\}\(\\bm\{m\}\)\.Then

V^T\(1\)≥1ϱ​OPT^ψ,γ\(1\)​\(𝒢\)−εell−εrnd\.\\widehat\{V\}\_\{T\}^\{\(1\)\}\\geq\\frac\{1\}\{\\varrho\}\\widehat\{\\mathrm\{OPT\}\}\_\{\\psi,\\gamma\}^\{\(1\)\}\(\\mathcal\{G\}\)\-\\varepsilon\_\{\\mathrm\{ell\}\}\-\\varepsilon\_\{\\mathrm\{rnd\}\}\.\(148\)

###### Proof F\.16

For everyk∈𝒦k\\in\\mathcal\{K\}, the recoveredxkx^\{k\}is supported onℳT\\mathcal\{M\}\_\{T\}\. By[LemmasF\.9](https://arxiv.org/html/2609.18126#A6.Thmtheorem9)and[F\.11](https://arxiv.org/html/2609.18126#A6.Thmtheorem11),

OPT^k\(1\)​\(ℳT\)\\displaystyle\\widehat\{\\mathrm\{OPT\}\}\_\{k\}^\{\(1\)\}\(\\mathcal\{M\}\_\{T\}\)≥𝔼rnd​\[Π^ψ,γ\(1\)​\(𝑴kRR​\(xk\)\)\]\\displaystyle\\geq\\mathbb\{E\}\_\{\\mathrm\{rnd\}\}\\\!\\left\[\\widehat\{\\Pi\}\_\{\\psi,\\gamma\}^\{\(1\)\}\\bigl\(\\bm\{M\}\_\{k\}^\{\\mathrm\{RR\}\}\(x^\{k\}\)\\bigr\)\\right\]≥ℒ^k\(1\)​\(𝒢\)−εell,k−εrnd,k\\displaystyle\\geq\\widehat\{\\mathcal\{L\}\}\_\{k\}^\{\(1\)\}\(\\mathcal\{G\}\)\-\\varepsilon\_\{\\mathrm\{ell\},k\}\-\\varepsilon\_\{\\mathrm\{rnd\},k\}≥OPT^k\(1\)​\(𝒢\)−εell,k−εrnd,k\.\\displaystyle\\geq\\widehat\{\\mathrm\{OPT\}\}\_\{k\}^\{\(1\)\}\(\\mathcal\{G\}\)\-\\varepsilon\_\{\\mathrm\{ell\},k\}\-\\varepsilon\_\{\\mathrm\{rnd\},k\}\.\(149\)By[LemmaF\.13](https://arxiv.org/html/2609.18126#A6.Thmtheorem13), somek∈𝒦k\\in\\mathcal\{K\}satisfies

OPT^k\(1\)​\(𝒢\)≥1ϱ​OPT^ψ,γ\(1\)​\(𝒢\)\.\\widehat\{\\mathrm\{OPT\}\}\_\{k\}^\{\(1\)\}\(\\mathcal\{G\}\)\\geq\\frac\{1\}\{\\varrho\}\\widehat\{\\mathrm\{OPT\}\}\_\{\\psi,\\gamma\}^\{\(1\)\}\(\\mathcal\{G\}\)\.Combining the two displays proves \([148](https://arxiv.org/html/2609.18126#A6.E148)\)\.□\\square

###### Lemma F\.17\(Transfer of the generated\-pool certificate to the true iid law\)

On the intersection of the Stage\-1 concentration event \([116](https://arxiv.org/html/2609.18126#A6.E116)\) andℰorc\\mathcal\{E\}\_\{\\mathrm\{orc\}\},

max𝒎∈𝒞𝒦​\(ℳT\)⁡Πψ,γstoch​\(𝒎\)≥1ϱ​OPTψ,γstoch​\(𝒢\)−εell−εrnd−\(1\+1ϱ\)​εest​\(L1,δ1\)\.\\max\_\{\\bm\{m\}\\in\\mathcal\{C\}\_\{\\mathcal\{K\}\}\(\\mathcal\{M\}\_\{T\}\)\}\\Pi\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\bm\{m\}\)\\geq\\frac\{1\}\{\\varrho\}\\mathrm\{OPT\}\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\mathcal\{G\}\)\-\\varepsilon\_\{\\mathrm\{ell\}\}\-\\varepsilon\_\{\\mathrm\{rnd\}\}\-\\left\(1\+\\frac\{1\}\{\\varrho\}\\right\)\\varepsilon\_\{\\mathrm\{est\}\}\(L\_\{1\},\\delta\_\{1\}\)\.\(150\)In particular, the final term is at most2​εest​\(L1,δ1\)2\\varepsilon\_\{\\mathrm\{est\}\}\(L\_\{1\},\\delta\_\{1\}\)\.

###### Proof F\.18

By[LemmaF\.3](https://arxiv.org/html/2609.18126#A6.Thmtheorem3), the true and plug\-in values of every feasible portfolio differ by at mostεest​\(L1,δ1\)\\varepsilon\_\{\\mathrm\{est\}\}\(L\_\{1\},\\delta\_\{1\}\)\. Therefore

max𝒎∈𝒞𝒦​\(ℳT\)⁡Πψ,γstoch​\(𝒎\)\\displaystyle\\max\_\{\\bm\{m\}\\in\\mathcal\{C\}\_\{\\mathcal\{K\}\}\(\\mathcal\{M\}\_\{T\}\)\}\\Pi\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\bm\{m\}\)≥V^T\(1\)−εest​\(L1,δ1\)\\displaystyle\\geq\\widehat\{V\}\_\{T\}^\{\(1\)\}\-\\varepsilon\_\{\\mathrm\{est\}\}\(L\_\{1\},\\delta\_\{1\}\)≥1ϱ​OPT^ψ,γ\(1\)​\(𝒢\)−εell−εrnd−εest​\(L1,δ1\)\\displaystyle\\geq\\frac\{1\}\{\\varrho\}\\widehat\{\\mathrm\{OPT\}\}\_\{\\psi,\\gamma\}^\{\(1\)\}\(\\mathcal\{G\}\)\-\\varepsilon\_\{\\mathrm\{ell\}\}\-\\varepsilon\_\{\\mathrm\{rnd\}\}\-\\varepsilon\_\{\\mathrm\{est\}\}\(L\_\{1\},\\delta\_\{1\}\)≥1ϱ​OPTψ,γstoch​\(𝒢\)−εell−εrnd−\(1\+1ϱ\)​εest​\(L1,δ1\),\\displaystyle\\geq\\frac\{1\}\{\\varrho\}\\mathrm\{OPT\}\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\mathcal\{G\}\)\-\\varepsilon\_\{\\mathrm\{ell\}\}\-\\varepsilon\_\{\\mathrm\{rnd\}\}\-\\left\(1\+\\frac\{1\}\{\\varrho\}\\right\)\\varepsilon\_\{\\mathrm\{est\}\}\(L\_\{1\},\\delta\_\{1\}\),which proves the claim\.□\\square

LetℋT\\mathcal\{H\}\_\{T\}denote the sigma\-field generated by the completed workflow generation stage, including all Stage\-1 pilot outcomes and oracle randomness\.

###### Lemma F\.19\(Conditional fresh\-L2L\_\{2\}deviation\)

Let𝒜T\\mathcal\{A\}\_\{T\}be a finite,ℋT\\mathcal\{H\}\_\{T\}\-measurable class of execution\-count vectors\. Conditional onℋT\\mathcal\{H\}\_\{T\}, with probability at least1−δ21\-\\delta\_\{2\},

sup𝒎∈𝒜T\|Π^L2final​\(𝒎\)−Πψ,γstoch​\(𝒎\)\|≤log⁡\(2​\|𝒜T\|/δ2\)2​n​L2\.\\sup\_\{\\bm\{m\}\\in\\mathcal\{A\}\_\{T\}\}\\left\|\\widehat\{\\Pi\}\_\{L\_\{2\}\}^\{\\mathrm\{final\}\}\(\\bm\{m\}\)\-\\Pi\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\bm\{m\}\)\\right\|\\leq\\sqrt\{\\frac\{\\log\(2\|\\mathcal\{A\}\_\{T\}\|/\\delta\_\{2\}\)\}\{2nL\_\{2\}\}\}\.\(151\)Consequently, an exact empirical maximizer over𝒜T\\mathcal\{A\}\_\{T\}has true value at least the best true value in𝒜T\\mathcal\{A\}\_\{T\}minus twice the right\-hand side\.

###### Proof F\.20

For fixed𝐦\\bm\{m\}, then​L2nL\_\{2\}random variables

ψk⁡\(𝒎\)​\(Ri\(2,ℓ\)​\(𝒎\)\),i=1,…,n,ℓ=1,…,L2,\\psi\_\{k\(\\bm\{m\}\)\}\\bigl\(R\_\{i\}^\{\(2,\\ell\)\}\(\\bm\{m\}\)\\bigr\),\\qquad i=1,\\ldots,n,\\quad\\ell=1,\\ldots,L\_\{2\},are conditionally independent under[Section7](https://arxiv.org/html/2609.18126#S7), lie in\[0,1\]\[0,1\], and have average mean equal to the stochastic accuracy component of𝐦\\bm\{m\}\. Hoeffding’s inequality and a union bound over𝒜T\\mathcal\{A\}\_\{T\}prove \([151](https://arxiv.org/html/2609.18126#A6.E151)\)\. Applying the uniform bound once to the empirical maximizer and once to a true maximizer gives the final statement\.□\\square

#### Proof of Proposition[7\.1](https://arxiv.org/html/2609.18126#S7.Thmtheorem1)\.

Fix a nonzero𝒎\\bm\{m\}and writek=k⁡\(𝒎\)k=k\(\\bm\{m\}\)\. For every taskii,[Sections4\.1](https://arxiv.org/html/2609.18126#S4.SS1)and[4\.2](https://arxiv.org/html/2609.18126#S4.Thmtheorem2)and concavity ofhΛh\_\{\\Lambda\}give

𝔼⁡\[ψk​\(Ri​\(𝒎\)\)\]≤𝔼⁡\[hΛ​\(Ri​\(𝒎\)k\)\]≤hΛ​\(𝔼​\[Ri​\(𝒎\)\]k\)\.\\mathbb\{E\}\[\\psi\_\{k\}\(R\_\{i\}\(\\bm\{m\}\)\)\]\\leq\\mathbb\{E\}\\\!\\left\[h\_\{\\Lambda\}\\\!\\left\(\\frac\{R\_\{i\}\(\\bm\{m\}\)\}\{k\}\\right\)\\right\]\\leq h\_\{\\Lambda\}\\\!\\left\(\\frac\{\\mathbb\{E\}\[R\_\{i\}\(\\bm\{m\}\)\]\}\{k\}\\right\)\.\(152\)Using \([59](https://arxiv.org/html/2609.18126#S7.E59)\), averaging over tasks, and applying Jensen once more yield

1n​∑i𝔼⁡\[ψk​\(Ri​\(𝒎\)\)\]≤hΛ​\(1k​∑gmg​a¯g\),\\frac\{1\}\{n\}\\sum\_\{i\}\\mathbb\{E\}\[\\psi\_\{k\}\(R\_\{i\}\(\\bm\{m\}\)\)\]\\leq h\_\{\\Lambda\}\\\!\\left\(\\frac\{1\}\{k\}\\sum\_\{g\}m\_\{g\}\\bar\{a\}\_\{g\}\\right\),which proves \([61](https://arxiv.org/html/2609.18126#S7.E61)\)\. For a size\-kkvector overℳ\\mathcal\{M\}, the argument ofhΛh\_\{\\Lambda\}is at mosta⋆a^\{\\star\}andC⁡\(𝒎\)≥k​cminC\(\\bm\{m\}\)\\geq kc\_\{\\min\}, proving \([62](https://arxiv.org/html/2609.18126#S7.E62)\)\. Singleton feasibility gives the lower bound\. MaximizinghΛ​\(a\)−ah\_\{\\Lambda\}\(a\)\-aovera∈\[0,1\]a\\in\[0,1\]proves \([63](https://arxiv.org/html/2609.18126#S7.E63)\)\.□\\square

#### Proof of Theorem[7\.2](https://arxiv.org/html/2609.18126#S7.Thmtheorem2)\.

Letℰest\\mathcal\{E\}\_\{\\mathrm\{est\}\}be the Stage\-1 concentration event in \([116](https://arxiv.org/html/2609.18126#A6.E116)\), letℰorc\\mathcal\{E\}\_\{\\mathrm\{orc\}\}be the event that every pricing batch isτk\\tau\_\{k\}\-accurate, and letℰsamp\\mathcal\{E\}\_\{\\mathrm\{samp\}\}be the Stage\-2 uniform\-deviation event in \([151](https://arxiv.org/html/2609.18126#A6.E151)\) for𝒞𝒦​\(ℳT\)\\mathcal\{C\}\_\{\\mathcal\{K\}\}\(\\mathcal\{M\}\_\{T\}\)\. By[LemmasF\.1](https://arxiv.org/html/2609.18126#A6.Thmtheorem1),[F\.7](https://arxiv.org/html/2609.18126#A6.Thmtheorem7)and[F\.19](https://arxiv.org/html/2609.18126#A6.Thmtheorem19),

ℙ⁡\(ℰest∩ℰorc∩ℰsamp\)≥1−δ1−δ2−Npricestoch​\(εell\)​e−porc​m\.\\mathbb\{P\}\\left\(\\mathcal\{E\}\_\{\\mathrm\{est\}\}\\cap\\mathcal\{E\}\_\{\\mathrm\{orc\}\}\\cap\\mathcal\{E\}\_\{\\mathrm\{samp\}\}\\right\)\\geq 1\-\\delta\_\{1\}\-\\delta\_\{2\}\-N\_\{\\mathrm\{price\}\}^\{\\mathrm\{stoch\}\}\(\\varepsilon\_\{\\mathrm\{ell\}\}\)e^\{\-p\_\{\\mathrm\{orc\}\}m\}\.\(153\)
Work on this intersection and write

VT⋆=max𝒎∈𝒞𝒦​\(ℳT\)⁡Πψ,γstoch​\(𝒎\)\.V\_\{T\}^\{\\star\}=\\max\_\{\\bm\{m\}\\in\\mathcal\{C\}\_\{\\mathcal\{K\}\}\(\\mathcal\{M\}\_\{T\}\)\}\\Pi\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\bm\{m\}\)\.By[LemmaF\.17](https://arxiv.org/html/2609.18126#A6.Thmtheorem17),

VT⋆≥1ϱ​OPTψ,γstoch​\(𝒢\)−εell−εrnd−2​εest​\(L1,δ1\)\.V\_\{T\}^\{\\star\}\\geq\\frac\{1\}\{\\varrho\}\\mathrm\{OPT\}\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\mathcal\{G\}\)\-\\varepsilon\_\{\\mathrm\{ell\}\}\-\\varepsilon\_\{\\mathrm\{rnd\}\}\-2\\varepsilon\_\{\\mathrm\{est\}\}\(L\_\{1\},\\delta\_\{1\}\)\.\(154\)Letεs=εsamp​\(L2,δ2\)\\varepsilon\_\{s\}=\\varepsilon\_\{\\mathrm\{samp\}\}\(L\_\{2\},\\delta\_\{2\}\)and let𝒎T⋆\\bm\{m\}\_\{T\}^\{\\star\}maximize true value over𝒞𝒦​\(ℳT\)\\mathcal\{C\}\_\{\\mathcal\{K\}\}\(\\mathcal\{M\}\_\{T\}\)\.

Suppose first that the conservative rule returns𝒎~\\widetilde\{\\bm\{m\}\}\. By the rule and uniform deviation,

Πψ,γstoch​\(𝒎~\)≥Π^L2final​\(𝒎~\)−εs≥V¯\.\\Pi\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\widetilde\{\\bm\{m\}\}\)\\geq\\widehat\{\\Pi\}\_\{L\_\{2\}\}^\{\\mathrm\{final\}\}\(\\widetilde\{\\bm\{m\}\}\)\-\\varepsilon\_\{s\}\\geq\\underline\{V\}\.Empirical optimality and uniform deviation also give

Πψ,γstoch​\(𝒎~\)\\displaystyle\\Pi\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\widetilde\{\\bm\{m\}\}\)≥Π^L2final​\(𝒎~\)−εs\\displaystyle\\geq\\widehat\{\\Pi\}\_\{L\_\{2\}\}^\{\\mathrm\{final\}\}\(\\widetilde\{\\bm\{m\}\}\)\-\\varepsilon\_\{s\}≥Π^L2final​\(𝒎T⋆\)−εs\\displaystyle\\geq\\widehat\{\\Pi\}\_\{L\_\{2\}\}^\{\\mathrm\{final\}\}\(\\bm\{m\}\_\{T\}^\{\\star\}\)\-\\varepsilon\_\{s\}≥VT⋆−2​εs\.\\displaystyle\\geq V\_\{T\}^\{\\star\}\-2\\varepsilon\_\{s\}\.Hence this case yields

Πψ,γstoch​\(𝒎^\)≥max⁡\{V¯,VT⋆−2​εs\}\.\\Pi\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\widehat\{\\bm\{m\}\}\)\\geq\\max\\\{\\underline\{V\},V\_\{T\}^\{\\star\}\-2\\varepsilon\_\{s\}\\\}\.
Suppose instead that the rule returns the baseline𝒎0\\bm\{m\}^\{0\}\. Then

Π^L2final​\(𝒎~\)−εs<V¯\.\\widehat\{\\Pi\}\_\{L\_\{2\}\}^\{\\mathrm\{final\}\}\(\\widetilde\{\\bm\{m\}\}\)\-\\varepsilon\_\{s\}<\\underline\{V\}\.Empirical optimality and uniform deviation imply

Π^L2final​\(𝒎~\)−εs≥VT⋆−2​εs,\\widehat\{\\Pi\}\_\{L\_\{2\}\}^\{\\mathrm\{final\}\}\(\\widetilde\{\\bm\{m\}\}\)\-\\varepsilon\_\{s\}\\geq V\_\{T\}^\{\\star\}\-2\\varepsilon\_\{s\},soVT⋆−2​εs<V¯V\_\{T\}^\{\\star\}\-2\\varepsilon\_\{s\}<\\underline\{V\}\. Since the baseline has true value at leastV¯\\underline\{V\}, this case gives the same bound\. Combining the two cases with \([154](https://arxiv.org/html/2609.18126#A6.E154)\) proves \([72](https://arxiv.org/html/2609.18126#S7.E72)\)\.

Finally, for everyA,e≥0A,e\\geq 0andV¯\>0\\underline\{V\}\>0,

max⁡\{V¯,A−e\}≥V¯V¯\+e​A\.\\max\\\{\\underline\{V\},A\-e\\\}\\geq\\frac\{\\underline\{V\}\}\{\\underline\{V\}\+e\}A\.Apply this inequality with

A=1ϱ​OPTψ,γstoch​\(𝒢\),e=εell\+εrnd\+2​εest​\(L1,δ1\)\+2​εsamp​\(L2,δ2\),A=\\frac\{1\}\{\\varrho\}\\mathrm\{OPT\}\_\{\\psi,\\gamma\}^\{\\mathrm\{stoch\}\}\(\\mathcal\{G\}\),\\qquad e=\\varepsilon\_\{\\mathrm\{ell\}\}\+\\varepsilon\_\{\\mathrm\{rnd\}\}\+2\\varepsilon\_\{\\mathrm\{est\}\}\(L\_\{1\},\\delta\_\{1\}\)\+2\\varepsilon\_\{\\mathrm\{samp\}\}\(L\_\{2\},\\delta\_\{2\}\),to obtain \([73](https://arxiv.org/html/2609.18126#S7.E73)\)\.□\\square

## Appendix GAdditional End\-to\-End Experiments

The main text reports the ABCD experiment in detail\. This appendix applies the same stochastic workflow\-evaluation, selector\-calibration, portfolio\- optimization, and dual\-guided generation pipeline to the two other domains in the study\. We organize each domain in the same order: task construction and data, stochastic performance of the initial workflow bank, selector calibration, finite\-pool optimization, budgeted workflow generation, and fresh held\-out deployment\. The SGD and HotpotQA experiments are reported below\.

### G\.1Schema\-Guided Dialogue

#### Task construction and data\.

Schema\-Guided Dialogue \(SGD\) contains task\-oriented conversations spanning multiple services, each accompanied by a natural\-language schema that defines the service, its available intents, and its slots\([Rastogi et al\. 2020](https://arxiv.org/html/2609.18126#bib.bib45)\)\. We form a schema\-conditioned active\-intent classification task\. At a user turn, the workflow observes the service description, the descriptions of all intents available for that service, required and optional slots, slot descriptions, and the dialogue prefix through the current user message\. It must return exactly one intent from that service’s allowed intent set\.

To reduce nearly repeated observations within a dialogue, we retain the first user turn at which each distinct service–active\-intent pair appears\. We exclude frames whose active intent isNONEand services with fewer than two available intents\. We preserve the official train, development, and test split structure and sample 800 development tasks for workflow evaluation and portfolio construction, 1,004 development\-split tasks for selector calibration, and 800 official test tasks for held\-out deployment\. The sampled development, calibration, and held\-out sets contain 33, 24, and 30 observed active\-intent labels, respectively; no single intent accounts for more than 6\.13%, 9\.27%, and 7\.38% of the corresponding samples\.

#### Initial workflow bank and stochastic execution\.

We use the same initial bank as in the ABCD experiment: 18 single\-call workflows obtained by crossing Qwen2\.5\-3B\-Instruct, Mistral\-7B\-Instruct\-v0\.3, and Granite\-3\.3\-8B\-Instruct with the direct, evidence\-first, decomposition, verify\-and\-revise, alternatives, and limited\-context prompting strategies\. Every workflow–task pair is executed independentlyL=5L=5times at temperature one\. The best estimated standalone workflow on the development sample is Granite alternatives, with one\-execution accuracy 82\.175%\. If every workflow in the initial bank is run once independently, expected oracle coverage is 97\.467%, indicating that the bank retains meaningful complementarity despite its strong best singleton\.

The execution parser succeeds on 95\.207% of development executions, and the five draws produce 1\.433 distinct intent predictions per task–workflow cell on average, confirming that repeated executions are not simply identical copies of a deterministic output\. Workflow\-level estimates are also stable at smallLL\. Relative to the fullL=5L=5estimates, the workflow rankings based on the first one, two, and three draws have Spearman correlations 0\.897, 0\.989, and 0\.994; the largest corresponding absolute change in a workflow’s average accuracy is 1\.975, 0\.700, and 0\.492 percentage points\.

Table 5:SGD design, selector calibration, and workflow\-generation summary\.
#### Selector calibration\.

We construct 100 feasible candidate panels for everyK∈\{2,…,6\}K\\in\\\{2,\\ldots,6\\\}and every interior correct countr∈\{1,…,K−1\}r\\in\\\{1,\\ldots,K\-1\\\}\. The selector observes the service schema, dialogue prefix, and candidate intent labels but not workflow identity, cost, or correctness\. The primary rotation\-vote calibration contains 1,500 panels from 702 distinct tasks and has a 100% parse rate\. Maximum\-likelihood estimation of the Plackett–Luce recovery model givesλ^=7\.2145\\widehat\{\\lambda\}=7\.2145with a 95% confidence interval of\[5\.7747,9\.0901\]\[5\.7747,9\.0901\], obtained by resampling tasks\. This estimate corresponds to an 87\.83% fitted probability of selecting the correct candidate when exactly one of two candidates is correct and an average 34\.33 percentage\-point advantage over uniform selection on the calibration design\.

Figure 7:SGD selector recovery\. Points report observed recovery on controlled\(K,r\)\(K,r\)panels, vertical bars report 95% bootstrap intervals obtained by resampling tasks, the dotted line is uniform selection, and the solid curve is the fitted Plackett–Luce model\. Small horizontal offsets separate observations having the same correct fraction\.
#### Initial\-pool optimization\.

For eachK∈\{1,…,6\}K\\in\\\{1,\\ldots,6\\\}, we apply the common 16\-scenario SAA and 512\-trial randomized\-rounding procedure described in the main text and then evaluate each distinct integral portfolio using the corresponding exact plug\-in calculation\. Net development value increases throughout the searched range, and the selected portfolio hasK=6K=6\. It contains Granite alternatives, Granite direct, Granite limited context, Mistral evidence first, Qwen decomposition, and Qwen limited context, one execution each\. Its calibrated development accuracy is 92\.612%, and its workflow\-cost\-adjusted value is 0\.920758\. Thus, although repetition is allowed, the selected SGD portfolio uses six distinct workflow types\. The separately optimized no\-repeat solution has essentially the same development value\.

Figure 8:Development performance of the initial SGD workflow bank by run size\. Bars report exact plug\-in selector\-aware accuracy and the same value after subtracting recurring workflow cost\. The primary search selectsK=6K=6\.
#### Budgeted dual\-guided workflow generation\.

We apply the same practical stochastic ellipsoid\-generation procedure as in the ABCD experiment\. For eachKK, workflow generation uses one common\-uniform SAA scenario, checks the explicit and previously materialized workflow constraints before invoking ADAS, and evaluates every distinct new workflow with five independent executions on each development task\. The prespecified paper budget permits 20 distinct new workflow evaluations, allocated across the six run sizes\. The SGD run makes 26 ADAS pricing queries and evaluates all 20 candidate workflows\. None satisfies the stochastic pricing criterion required for incorporation at the queried dual points, so the optimization bank remains at its original 18 workflow types\. Each fixed\-KKsearch reaches its candidate\-evaluation budget; this result is therefore evidence from a budgeted search, not a claim that no useful workflow exists anywhere in the implicit class\.

Because the workflow bank does not change, the pipeline’s second calibration is a repeated calibration of the same 18\-workflow candidate distribution rather than an expanded\-bank calibration\. It givesλ^=8\.1506\\widehat\{\\lambda\}=8\.1506with 95% confidence interval\[6\.5437,10\.4949\]\[6\.5437,10\.4949\], overlapping the initial estimate\. We use this second estimate for the final model\-based predictions below; the held\-out deployment results additionally run the selector directly and do not rely solely on the parametric recovery model\.

#### Fresh held\-out deployment\.

We freeze the reported portfolios, collect five new stochastic executions per selected workflow on each of 800 official test tasks, and run the deterministic Qwen selector using cyclic candidate\-order rotations and rotation vote\.[Tables6](https://arxiv.org/html/2609.18126#A7.T6)and[9](https://arxiv.org/html/2609.18126#A7.F9)report the results\. Actual selector accuracy rises from 85\.250% for the best initial singleton to 92\.750% for the optimized initial\-bank portfolio, an increase of 7\.500 percentage points or 8\.8% relative to the singleton\. Since the budgeted ADAS search incorporates no new workflow, the final repeat\-allowed portfolio is the same optimized initial\-bank portfolio and has the same held\-out accuracy\. The no\-repeat benchmark reaches 92\.375%\.

The calibrated model predicts the same ordering, with selector\-aware accuracy 84\.700% for the singleton, 94\.708% for the optimized repeat\-allowed portfolio, and 94\.543% for the no\-repeat benchmark\. For the optimized portfolio, uniform selection among the realized candidate slots would attain 83\.150%, while oracle coverage is 97\.904%\. After subtracting both workflow cost and realized selector cost, actual total\-system net value rises from 0\.851357 for the singleton to 0\.915629 for the optimized portfolio\.

Table 6:Fresh held\-out SGD deployment on 800 tasks\. PL accuracy is the plug\-in Plackett–Luce prediction; actual accuracy uses the blind deterministic Qwen selector\. Brackets report 95% Wilson intervals\. Workflow and selector costs are per task, and actual total\-system net subtracts both\.Figure 9:Fresh held\-out SGD accuracy\. Bars report actual deterministic selector accuracy, error bars show 95% Wilson intervals over 800 tasks, and diamonds report the corresponding plug\-in Plackett–Luce predictions\.
#### Multiplicity\.

Within the selector\-calibrated rangeK≤6K\\leq 6, the repeat\-allowed solution uses only distinct workflow types, and the repeat and no\-repeat values are effectively identical\. Repetition first appears atK=7K=7in the extended model\-based sensitivity analysis and produces only small gains\. Thus, SGD supports the value of portfolio diversity and post\-output selection but, unlike ABCD, does not provide meaningful evidence that multiplicity is valuable at the endogenous deployment solution\.

### G\.2HotpotQA

#### Task construction, data, and answer scoring\.

HotpotQA is an open\-domain question\-answering benchmark designed to require reasoning across multiple pieces of evidence\([Yang et al\. 2018](https://arxiv.org/html/2609.18126#bib.bib46)\)\. We use the distractor configuration\. Each workflow receives the question together with all supplied Wikipedia passages and must return a concise answer span or short phrase\. Supporting\-fact annotations are retained only as audit metadata and are not shown to candidate workflows or to the selector\.

The public distractor test labels are not distributed\. We therefore sample 800 workflow\-development questions from the official training split and form disjoint selector\-calibration and held\-out samples of 1,004 and 800 questions from the labeled distractor validation split\. These samples contain 733, 913, and 741 distinct normalized answers, respectively\. The largest answer share is 2\.63% in the development sample, 3\.98% in the calibration sample, and 2\.63% in the held\-out sample\. Answers are evaluated using the normalized exact\-match rule implemented in the notebooks: text is lowercased, articles and nonalphanumeric punctuation are removed, and whitespace is collapsed before comparison with the reference answer\.

#### Initial workflow bank and stochastic execution\.

We use the same 18\-workflow initial bank as in the other domains, obtained by crossing Qwen2\.5\-3B\-Instruct, Mistral\-7B\-Instruct\-v0\.3, and Granite\-3\.3\-8B\-Instruct with the direct, evidence\-first, decomposition, verify\-and\-revise, alternatives, and limited\-context prompting strategies\. Every workflow–question pair is executed independentlyL=5L=5times at temperature one\.

The best estimated standalone workflow on the development sample is Granite evidence first, with one\-execution accuracy 44\.950%\. If every workflow in the initial bank is executed once independently, expected oracle coverage is 77\.578%, showing substantial complementarity among the candidate generators\. The answer parser succeeds on 88\.468% of executions, and the five stochastic draws produce 2\.908 distinct normalized answers per workflow–question cell on average, confirming substantial within\-cell stochastic variation\. Workflow\-level estimates are stable despite noisy task\-level probabilities\. Relative to the fullL=5L=5ranking, rankings based on the first one, two, and three executions have Spearman correlations 0\.939, 0\.993, and 0\.990; all five of the highest\-ranked workflows remain in the top five\. The largest absolute change in a workflow’s average accuracy is 3\.000, 0\.875, and 0\.633 percentage points, respectively\.

Table 7:HotpotQA design, selector calibration, and workflow\-generation summary\.
#### Selector calibration\.

For eachK∈\{2,…,6\}K\\in\\\{2,\\ldots,6\\\}and each interior correct countr∈\{1,…,K−1\}r\\in\\\{1,\\ldots,K\-1\\\}, we construct 100 feasible panels containing exactlyrrcorrect candidate slots andK−rK\-rincorrect slots\. The selector observes the question, all supplied passages, and the candidate answers, but not workflow identity, cost, or correctness\. The primary rotation\-vote calibration contains 1,500 panels from 633 distinct questions and has a 100% base\-panel parse rate\.

Maximum\-likelihood estimation of the Plackett–Luce recovery model givesλ^=1\.7062\\widehat\{\\lambda\}=1\.7062with a 95% confidence interval of\[1\.4557,2\.0068\]\[1\.4557,2\.0068\], obtained by resampling tasks\. The fitted selector therefore chooses the correct answer with probability 63\.05% when exactly one of two candidates is correct\. Averaged over the calibration design, selector recovery exceeds uniform selection by 10\.87 percentage points\. Selector recovery is positive but substantially weaker than in the two service\-operation domains, consistent with the greater difficulty of recognizing a correct free\-form multi\-hop answer among plausible alternatives\.

Figure 10:HotpotQA selector recovery\. Points report observed recovery on controlled\(K,r\)\(K,r\)panels, vertical bars report 95% bootstrap intervals obtained by resampling tasks, the dotted line is uniform selection, and the solid curve is the fitted Plackett–Luce model\. Small horizontal offsets separate observations having the same correct fraction\.
#### Initial\-pool optimization\.

For eachK∈\{1,…,6\}K\\in\\\{1,\\ldots,6\\\}, we apply the common 16\-scenario SAA and 512\-trial randomized\-rounding procedure described in the main text and then evaluate each distinct integral portfolio using the corresponding exact plug\-in calculation\. The best initial\-bank solution hasK=6K=6and assigns five execution slots to Granite evidence first and one slot to Granite decomposition\. Its calibrated development accuracy is 48\.270%, compared with 44\.950% for the best singleton, and its workflow\-cost\-adjusted value is 0\.475840\.

Multiplicity is already valuable in the initial bank\. When repeated workflow types are prohibited, the best no\-repeat solution is attained atK=3K=3by Granite alternatives, Granite decomposition, and Granite evidence first and has workflow\-cost\-adjusted value 0\.466358\. Thus, allowing repetition increases the initial\-bank endogenous optimum by 0\.009482 in accuracy\-equivalent net\-value units\.

#### Budgeted dual\-guided workflow generation\.

We apply the same practical stochastic ellipsoid\-generation procedure used in the ABCD and SGD experiments\. For eachKK, generation uses one common\-uniform SAA scenario, checks explicit and already materialized workflow constraints before invoking ADAS, and evaluates every distinct new workflow using five independent executions on each development question\. The prespecified paper budget permits 20 distinct workflow evaluations across the six run sizes\.

The search makes 26 ADAS pricing queries, evaluates all 20 distinct candidates, and incorporates one workflow during theK=2K=2search, expanding the final optimization bank from 18 to 19 workflow types\. We label this workflow G1; its frozen ADAS name isFinalizedMultiHopQASolver\. Each fixed\-KKsearch reaches its prespecified candidate\-evaluation budget, so the result should be interpreted as a budgeted ellipsoid search rather than as a claim that the complete implicit workflow class has been exhausted\.

Table 8:Executable structure of the ADAS\-generated HotpotQA workflow incorporated into the final bank\. All three model calls use GPT\-4o\-mini through the ADASLLMAgentBaseinterface\.WorkflowExecutable structureMean callsCostG1A first GPT\-4o\-mini call extracts evidence from the supplied passages\. A second call verifies the relevance and accuracy of that evidence\. A third call receives the question, passages, and verified evidence and derives the final answer\. If either evidence stage returns no usable output, the workflow returns a fixed insufficient\-evidence response\.3\.00\.003
*Notes\.*G1 is the frozen ADAS programadas\_ell\_5860177799ff/ FinalizedMultiHopQASolver\. Cost is recurring normalized execution cost per complete workflow run\.

After adding G1, we recalibrate the deterministic Qwen selector on the expanded candidate distribution\. The estimate increases modestly toλ^=1\.7529\\widehat\{\\lambda\}=1\.7529with 95% confidence interval\[1\.5028,2\.0499\]\[1\.5028,2\.0499\]\.

#### Expanded\-bank reoptimization and multiplicity\.

G1 changes the portfolio substantially\. As a standalone workflow, it has estimated development accuracy 66\.100%, compared with 44\.950% for the best initial workflow\. Under the primary objective, which subtracts recurring workflow cost but not selector cost, the final repeat\-allowed optimum hasK=2K=2and assigns both execution slots to G1\. Its calibrated development accuracy is 66\.863%, and its workflow\-cost\-adjusted value is 0\.662631\. The endogenous no\-repeat optimum is one execution of G1, with calibrated accuracy 66\.100% and workflow\-cost\-adjusted value 0\.658000\.

At fixedK=2K=2, multiplicity is valuable under the primary workflow\-cost objective\. When repetition is prohibited, the bestK=2K=2portfolio combines G1 with Granite evidence first and has value 0\.596823\. Allowing two independent executions of G1 raises that fixed\-size value by 0\.065808\. After optimizing run size separately under each formulation, the repeat\-allowed optimum remains 0\.004631 above the no\-repeat optimum\.

Figure 11:HotpotQA development value before and after workflow generation\. Lines compare repeat\-allowed and no\-repeat optimization over the initial and expanded workflow banks\. Workflow generation shifts the selected run size fromK=6K=6toK=2K=2and makes repeated execution of G1 optimal under the primary workflow\-cost objective\.Table 9:Key HotpotQA development portfolios\. Accuracy is the exact plug\-in Plackett–Luce value of the reported integral portfolio\. Net value subtracts recurring workflow cost but not selector cost, matching the primary optimization objective\.
#### Fresh held\-out deployment\.

We freeze the reported portfolios, collect five new stochastic executions per selected workflow on each of 800 held\-out questions, and then deploy the Qwen2\.5\-7B\-Instruct selector deterministically using cyclic candidate\-order rotations and rotation vote\.[Tables10](https://arxiv.org/html/2609.18126#A7.T10)and[12](https://arxiv.org/html/2609.18126#A7.F12)report both the plug\-in Plackett–Luce prediction and the realized selector performance\.

Actual selector accuracy is 30\.250% for the best initial singleton and 31\.125% for the optimized initial\-bank portfolio\. After workflow generation, accuracy rises to 55\.250% for the two\-copy G1 portfolio and 54\.875% for the single\-G1 no\-repeat benchmark\. Relative to the best initial singleton, the full repeat\-allowed procedure gains 25\.000 percentage points, or 82\.6%\. Workflow generation contributes 24\.125 percentage points relative to the optimized initial portfolio\. The fitted recovery model is especially accurate for the final repeated portfolio, predicting 55\.017% compared with the realized 55\.250%\.

The held\-out comparison also clarifies the value and cost of multiplicity\. A second independent G1 execution raises actual accuracy by 0\.375 percentage points relative to one G1 execution\. Under the primary objective, which charges workflow cost but not selector cost, the two\-copy portfolio has slightly higher net value: 0\.546500 versus 0\.545750\. Once realized selector cost is also subtracted, however, the one\-copy benchmark has slightly higher total\-system net value, 0\.545750 versus 0\.544329\. Thus, HotpotQA provides evidence that multiplicity can improve accuracy while also illustrating that the incremental gain may not justify the additional post\-output selection cost\.

Table 10:Fresh held\-out HotpotQA deployment on 800 questions\. PL accuracy is the plug\-in Plackett–Luce prediction; actual accuracy uses the blind deterministic Qwen selector\. Brackets report 95% Wilson intervals\. Workflow and selector costs are per question, and actual total\-system net subtracts both\.Figure 12:Fresh held\-out HotpotQA accuracy\. Bars report actual deterministic\-selector accuracy, error bars show 95% Wilson intervals over 800 questions, and diamonds report the corresponding plug\-in Plackett–Luce predictions\.

Similar Articles

Architectural Implications of Agentic AI Workflows

arXiv cs.AI

This paper presents the first architectural characterization of agentic AI workflows, revealing fragmented, heterogeneous execution patterns that mismatch conventional server designs, and introduces a prototype server called Agora to improve CPU/GPU utilization and throughput.

Improving the speed and energy-efficiency of AI agents

MIT News — Artificial Intelligence

Researchers from MIT and Microsoft developed an intelligent system that automatically optimizes agentic workflows, reducing computational resources and energy usage while maintaining performance.

Learning to Construct Practical Agentic Systems

arXiv cs.LG

This paper proposes principled approaches for designing and optimizing practical agentic LLM systems, introducing a framework with pseudo-tools and fixed workflows to improve modularity, cost-efficiency, and accuracy across diverse tasks.