MetaRoute-Bench: Evaluating Meta-Decision Policies for Agentic Workflow Routing
Summary
This paper introduces MetaRoute-Bench, an open benchmark for evaluating meta-decision policies in agentic workflows, comparing routing policies on success, cost, and latency under a shared execution model.
View Cached Full Text
Cached at: 08/04/26, 07:36 AM
# MetaRoute-Bench: Evaluating Meta-Decision Policies for Agentic Workflows
Source: [https://arxiv.org/html/2608.00107](https://arxiv.org/html/2608.00107)
\(2026\)
###### Abstract\.
Agentic systems must repeatedly decide whether to answer directly, decompose a task, invoke a tool, execute code, delegate to a specialist, verify an intermediate result, or recover from failure\. These meta\-decisions affect not only task success but also operating cost and latency, yet they are often embedded inside an orchestration framework and evaluated only through aggregate task accuracy\. We present MetaRoute\-Bench, an open, inspectable framework for comparing meta\-decision policies under a shared execution model\. The initial benchmark contains 180 synthetic task profiles spanning data analysis, research, and document processing, eight routing policies, and 30 paired random seeds\. Across 43,200 traces, a task\-aware compositional policy achieves 79\.4% success compared with 76\.7% for a strong workload\-specific static policy, 67\.4% for one\-shot task routing, and 52\.9% for direct answering\. Relative to the static policy, this is a 2\.7 percentage\-point improvement \(paired 95% CI:±\\pm2\.0 points\) at 4\.7% higher mean cost and 6\.4% higher latency\. Ablations show the largest losses when route composition is restricted to one operation and when verification is removed\. These results are generated by a seeded offline execution model rather than a live deployment; accordingly, our primary contribution is a reproducible evaluation method and an analysis of routing\-policy tradeoffs, not evidence of production effectiveness\. We release task generation, policies, traces, tests, and analysis artifacts to support live\-system validation\.
agentic systems, routing, orchestration, evaluation, tool use, reproducibility
††copyright:none††journalyear:2026††conference:8th International Conference on Distributed Artificial Intelligence; November 29–December 2, 2026; Hong Kong††booktitle:8th International Conference on Distributed Artificial Intelligence \(DAI ’26\), November 29–December 2, 2026, Hong Kong††ccs:Computing methodologies Multi\-agent systems††ccs:Software and its engineering Software performance††ccs:General and reference Evaluation## 1\.Introduction
An agentic application is more than a language model call\. It is a distributed decision process involving models, retrieval systems, code executors, specialist agents, verification stages, retry logic, and human escalation\. Before any component can help, the system must decide which component to invoke and when\. A poor meta\-decision can waste time on unnecessary decomposition, call an irrelevant tool, execute unsafe or unhelpful code, or stop before a result has been checked\.
Existing research has established the value of interleaving reasoning and action\(Yaoet al\.,[2023](https://arxiv.org/html/2608.00107#bib.bib1)\), learning API use\(Schicket al\.,[2023](https://arxiv.org/html/2608.00107#bib.bib2); Qinet al\.,[2024](https://arxiv.org/html/2608.00107#bib.bib3)\), routing among models\(Chenet al\.,[2024](https://arxiv.org/html/2608.00107#bib.bib7); Onget al\.,[2025](https://arxiv.org/html/2608.00107#bib.bib8)\), and optimizing function\-call plans\(Kimet al\.,[2024](https://arxiv.org/html/2608.00107#bib.bib6)\)\. Agent benchmarks reveal persistent failures in long\-horizon reasoning and decision making\(Liuet al\.,[2024](https://arxiv.org/html/2608.00107#bib.bib4)\), while domain benchmarks increasingly emphasize executable outcomes\(Jimenezet al\.,[2024](https://arxiv.org/html/2608.00107#bib.bib5); Yaoet al\.,[2024](https://arxiv.org/html/2608.00107#bib.bib10)\)\. However, system builders still lack a small, inspectable protocol for isolating the policy that chooses among reasoning modes and measuring its success, cost, and latency tradeoffs\.
This paper introduces MetaRoute\-Bench, an offline benchmark and trace format for that meta\-decision layer\. We ask:*under controlled execution assumptions, does task\-aware adaptive routing improve operational outcomes over fixed and one\-shot policies?*Our contributions are:
- •a typed framework separating task profiles, routing policies, execution, traces, and evaluation;
- •a balanced suite of 180 synthetic profiles across three operational workload families;
- •a paired comparison of eight policies over 30 seeds and 43,200 traces; and
- •ablations and workload\-level analyses identifying when composition, verification, decomposition, and recovery affect outcomes\.
The benchmark is intentionally transparent, but currently simulated\. This boundary is central: the reported values validate the framework’s reproducibility and diagnostic behavior, while live models, tools, and organizational workloads are required to establish external validity\.
## 2\.Related Work
Adaptive routing and efficient inference\.Prior routers establish that computational strategy should vary with the request rather than remain fixed\. Adaptive\-RAG classifies question complexity and selects among no retrieval, single\-step retrieval, and iterative retrieval, improving the accuracy–efficiency balance\(Jeonget al\.,[2024](https://arxiv.org/html/2608.00107#bib.bib11)\)\. CP\-Router similarly uses uncertainty to choose between a standard language model and a longer\-reasoning model\(Suet al\.,[2026](https://arxiv.org/html/2608.00107#bib.bib16)\)\. FrugalGPT and RouteLLM route requests among models to balance quality and cost\(Chenet al\.,[2024](https://arxiv.org/html/2608.00107#bib.bib7); Onget al\.,[2025](https://arxiv.org/html/2608.00107#bib.bib8)\), while ZeroRouter explicitly optimizes accuracy, cost, and latency and supports onboarding unseen models\(Yanet al\.,[2026](https://arxiv.org/html/2608.00107#bib.bib17)\)\. These systems motivate our task\-difficulty thresholds and separate reporting of success, cost, and latency\. MetaRoute\-Bench differs by routing among workflow operations rather than only retrieval modes or model endpoints; its current threshold policy is transparent and hand specified rather than learned or uncertainty calibrated\.
Planning, decomposition, and verification\.ReAct interleaves reasoning and environment actions so plans can evolve from observations\(Yaoet al\.,[2023](https://arxiv.org/html/2608.00107#bib.bib1)\)\. ACPBench uses formal planning domains to synthesize scalable tasks with provably correct answers and shows that current models retain uneven planning abilities\(Kokelet al\.,[2025](https://arxiv.org/html/2608.00107#bib.bib15)\)\. SPIRAL assigns proposing, simulation, and critique to specialized agents inside grounded reflective search\(Zhanget al\.,[2026](https://arxiv.org/html/2608.00107#bib.bib18)\)\. Together, these works support treating decomposition and verification as separable operations and motivate our ablations of each\. Our implementation is intentionally less ambitious: it composes an annotated route once and permits an executor\-level retry, but does not perform search or observation\-conditioned replanning\.
Tool selection and execution\.Toolformer learns whether, when, and how to invoke APIs\(Schicket al\.,[2023](https://arxiv.org/html/2608.00107#bib.bib2)\), and ToolLLM scales tool learning to thousands of real APIs\(Qinet al\.,[2024](https://arxiv.org/html/2608.00107#bib.bib3)\)\. API\-Bank separates planning, API retrieval, and API calling in a runnable benchmark\(Liet al\.,[2023](https://arxiv.org/html/2608.00107#bib.bib12)\); RESTful\-Llama demonstrates an industry\-oriented path from natural\-language requests and API documentation to REST calls\(Xuet al\.,[2024](https://arxiv.org/html/2608.00107#bib.bib13)\)\. AnyTool adds hierarchical retrieval and reflection\(Duet al\.,[2024](https://arxiv.org/html/2608.00107#bib.bib9)\), while LLMCompiler optimizes function\-call plans for latency and cost\(Kimet al\.,[2024](https://arxiv.org/html/2608.00107#bib.bib6)\)\. This literature justifies representing tool use, code execution, delegation, and verification as explicit trace events with independent costs and failures\. Unlike API\-Bank or RESTful\-Llama, the present study simulates those events and therefore cannot establish live tool\-call robustness\.
Agent evaluation and trajectory diagnosis\.AgentBench evaluates agents across interactive environments and identifies long\-horizon reasoning and decision\-making failures\(Liuet al\.,[2024](https://arxiv.org/html/2608.00107#bib.bib4)\)\. SWE\-bench grounds coding\-agent evaluation in repository issues\(Jimenezet al\.,[2024](https://arxiv.org/html/2608.00107#bib.bib5)\), andτ\\tau\-bench evaluates conversational tool agents under domain policies\(Yaoet al\.,[2024](https://arxiv.org/html/2608.00107#bib.bib10)\)\. AgentDiagnose argues that final success alone obscures decomposition, observation reading, verification, and backtracking behavior, and instead analyzes full trajectories\(Ouet al\.,[2025](https://arxiv.org/html/2608.00107#bib.bib14)\)\. These findings directly motivate our typed traces, failure taxonomy, action\-frequency analysis, and workload slices\. MetaRoute\-Bench contributes a controlled policy\-comparison layer; it complements rather than replaces benchmarks with executable tasks and real environments\.
## 3\.Framework
### 3\.1\.Design Requirements
MetaRoute\-Bench is designed around four requirements derived from operational agent evaluation\.*Policy isolation*requires routing logic to be interchangeable without changing tasks or the executor\.*Paired evaluation*requires every policy to encounter the same profiles and seed schedule\.*Trace completeness*, consistent with trajectory\-level diagnosis\(Ouet al\.,[2025](https://arxiv.org/html/2608.00107#bib.bib14)\), requires the artifact to retain actions, failures, retries, cost, and latency rather than only final correctness\. Finally,*assumption visibility*requires simulator parameters to remain inspectable and configurable\. These requirements make the artifact useful for controlled diagnostics today and for later shadow evaluation against live systems\.
The unit of evaluation is a route, not an isolated model response\. This distinction matters because two systems can use the same base model yet produce different operational behavior through decomposition, tool access, verification, and retry policies\. Conversely, a more capable model can be operationally inferior if its controller invokes costly components unnecessarily\. MetaRoute\-Bench therefore treats the controller as an independently testable system component\.
### 3\.2\.Meta\-Decision Process
Each taskxxhas a workload, difficulty, ambiguity, operation\-need annotations, and cost and latency budgets\. The present policies map task metadata to a complete routerr:
\(1\)r=π\(x\),r⊆\{D,T,C,G,V,A\},r=\\pi\(x\),\\quad r\\subseteq\\\{D,T,C,G,V,A\\\},
whereDDdenotes decomposition,TTtool use,CCcode execution,GGdelegation,VVverification, andAAfinal answering\. Every route must end inAA; operations are unique in the current implementation\. A trace records actions, success, expected success, cost, latency, retries, budget compliance, confidence, and failure mode\. The framework can be extended to policies over execution history, but the evaluated policy adapts only through one executor\-level retry after a failed operation; it does not replan from arbitrary intermediate outputs\.
We report success, cost, and latency separately\. For ranking in the command\-line summary only, we define a secondary utility
\(2\)U=S−0\.025K−0\.0015L,U=S\-0\.025K\-0\.0015L,
whereSSis success rate,KKmean normalized cost, andLLmean latency in seconds\. Conclusions do not depend solely on this weighting\.
### 3\.3\.System Architecture
The implementation separates five interfaces: \(1\) a deterministic workload generator; \(2\) routing policies; \(3\) a seeded offline executor; \(4\) typed trace export; and \(5\) aggregate and paired evaluation\. This separation permits replacement of the simulator with live model and tool adapters without changing policy or analysis interfaces\.
Task profile→\\rightarrowrouting policy→\\rightarrowroute plan↓\\downarrowseeded executor→\\rightarrowtyped trace→\\rightarrowevaluator
Figure 1\.MetaRoute\-Bench separates task generation, route selection, execution, and evaluation\. Live adapters can replace the seeded executor while preserving trace and metric interfaces\.A pipeline begins with a task profile, passes through a routing policy to form a route plan, executes the plan, records a typed trace, and sends the trace to an evaluator\.The executor makes its assumptions explicit\. Each operation has a normalized cost, latency, benefit scaled by task need, and—for external tools, code, and delegation—a failure probability\. Difficulty and ambiguity reduce base success\. Unnecessary operations impose overhead\. Adaptive recovery makes one retry after an execution failure\. All randomness is keyed by seed, task identifier, and policy name\.
### 3\.4\.Policies
We evaluate four fixed policies \(direct, always decompose, always tool, and always code\), a random route policy, a workload\-specific static rule table, a one\-shot router selecting the highest annotated operation need, and the proposed task\-aware compositional policy, labeled*adaptive*in the artifacts\. It selects up to three operations above a difficulty\-dependent threshold, orders decomposition first and verification last, and enables one recovery attempt\. Both one\-shot and adaptive policies receive the same task\-level need annotations; their comparison therefore isolates route composition and recovery rather than feature prediction\.
The adaptive policy uses a threshold of \.66 for difficulty levels one and two and \.56 for levels three and four\. This coarse complexity conditioning follows the same design principle as Adaptive\-RAG\(Jeonget al\.,[2024](https://arxiv.org/html/2608.00107#bib.bib11)\), although our thresholds are fixed rather than classifier learned\. If no operation clears the threshold, the policy selects the highest\-scoring operation\. A maximum of three support operations prevents unbounded orchestration overhead\. Decomposition is moved to the beginning of a route because it affects downstream work allocation, while verification is moved to the end because it evaluates the assembled result\. The current policy is deliberately transparent: every decision can be reconstructed from exported task annotations and constants\. This favors auditability over model flexibility and provides a reproducible baseline for future learned or uncertainty\-aware routers\(Suet al\.,[2026](https://arxiv.org/html/2608.00107#bib.bib16); Yanet al\.,[2026](https://arxiv.org/html/2608.00107#bib.bib17)\)\.
The static workload baseline is intentionally strong\. Data\-analysis profiles execute code and verify, research profiles decompose and use a tool, and document\-processing profiles use a tool and verify\. It represents the kind of rule table an engineering team might deploy before investing in a learned or task\-aware controller\. The one\-shot policy receives richer task annotations but can choose only one support operation\. Comparing these policies distinguishes three questions: whether orchestration helps at all, whether workload rules are sufficient, and whether composing multiple task\-specific operations adds value\.
## 4\.Evaluation
### 4\.1\.Workloads and Protocol
The suite contains 60 task profiles for each of data analysis, research, and document processing\. Profiles cover four difficulty levels and use seeded variation around workload\-specific needs\. Data analysis emphasizes code and verification, research emphasizes decomposition and tools, and document processing emphasizes tools and verification\. These profiles represent workflow characteristics rather than natural\-language task instances\.
Table 1\.Mean annotations for the 180 task profiles\. Dcmp\. denotes decomposition and Verif\. denotes verification\.Table[1](https://arxiv.org/html/2608.00107#S4.T1)summarizes the resulting suite\. The profiles are balanced by workload and difficulty, but intentionally heterogeneous within each workload: every need receives uniform jitter of up to \.25 before clipping\. This prevents the workload label from fully determining the best route and creates cases where a workload\-level rule is unnecessarily expensive or omits a useful operation\. Scalable synthetic evaluation has precedent in planning benchmarks such as ACPBench\(Kokelet al\.,[2025](https://arxiv.org/html/2608.00107#bib.bib15)\); however, our annotations are not backed by formal semantics or provably correct plans\. They should be viewed as controlled routing signals rather than a realistic task\-understanding stage\.
We execute every policy on every profile for 30 paired seeds\. This produces 5,400 traces per policy and 43,200 main\-experiment traces\. We report mean success and a 95% confidence interval computed over seed\-level success rates, plus mean normalized cost, latency, and cost per success\. Paired differences use the same seed\-level aggregation\. The complete trace file is exported as CSV\.
### 4\.2\.Main Results
Table 2\.Main benchmark results\. CI is the 95% interval half\-width for success over 30 seeds\.Table[2](https://arxiv.org/html/2608.00107#S4.T2)shows a clear operational tradeoff\. Adaptive routing has the highest success rate, but direct answering is least expensive and fastest\. Against the strongest baseline, static workload routing, adaptive routing improves success by 2\.74 points \(paired 95% CI±\\pm1\.96\), while increasing cost by 0\.13 units and latency by 1\.22 seconds\. Compared with one\-shot routing, adaptive routing improves success by 12\.02 points \(±\\pm1\.87\), with 1\.09 additional cost units and 8\.00 additional seconds\.
Figure 2\.Success versus mean normalized cost \(left\) and mean latency \(right\)\. Error bars show 95% confidence intervals over seed\-level success\. No policy dominates all three dimensions: direct answering is cheapest and fastest, while adaptive routing has the highest success\.Two scatter plots compare eight routing policies\. In both plots, adaptive routing has the highest success and highest cost or latency, static workload routing is second in success and slightly cheaper, and direct answering is cheapest and fastest but has the lowest success\.Figure[2](https://arxiv.org/html/2608.00107#S4.F2)makes the absence of a single universally best policy explicit\. Direct answering and always\-decompose are attractive when latency or cost dominates, while adaptive and static routing occupy the high\-success region\. The adaptive policy’s cost per successful task is 3\.67 units, compared with 3\.63 for static routing and 2\.71 for one\-shot routing\. Thus, its success advantage over static routing does not translate into lower cost per success under the current weights\. A deployment should select a policy from this frontier using service\-level objectives rather than ranking by success alone\.
Workload\-level adaptive success is \.802 for data analysis, \.835 for research, and \.746 for document processing\. The static policy is particularly competitive for data analysis \(\.791\), but trails more on research \(\.786 versus \.835\)\. This suggests that task\-level composition has the most value where decomposition and information access interact\.
### 4\.3\.Ablations
Table 3\.Adaptive\-policy ablations over 5,400 traces each\.Restricting routes to one operation produces the largest reduction: 11\.30 points \(±\\pm1\.85\)\. Removing verification reduces success by 4\.70 points \(±\\pm1\.68\), and removing decomposition reduces it by 2\.57 points \(±\\pm1\.68\)\. The 1\.13\-point recovery difference has an interval of±\\pm1\.70 and therefore does not support a confident recovery benefit in this experiment\. This is a useful negative finding: recovery is operationally plausible, but the current failure frequency and sample design do not establish its effect\.
### 4\.4\.Routing and Failure Analysis
The adaptive policy verifies 68\.3% of routes, uses a tool in 50\.6%, executes code in 33\.3%, decomposes 28\.3%, and delegates 13\.3%\. In contrast, the static policy applies exactly two support operations to every task\. The adaptive policy therefore does not improve by simply constructing longer routes; it reallocates operations according to profile\-level signals and uses three operations only where multiple needs clear the threshold\.
Failure traces separate unrecovered execution failures from ordinary task failures\. Adaptive routing records 109 unrecovered tool, code, or delegation failures among 5,400 traces \(2\.0%\), compared with 408 tool or code failures for static routing \(7\.6%\)\. Ordinary task failures are similar: 1,060 for adaptive and 1,050 for static\. The lower unrecovered\-execution count is consistent with retry behavior, but the no\-recovery ablation’s confidence interval includes zero\. We therefore treat this pattern as diagnostic evidence about the trace model, not proof that recovery improves overall success\.
Document processing is the weakest adaptive workload at \.746 success\. Its high verification and tool needs make routes vulnerable to external\-operation failure while offering less benefit from decomposition\. Research shows the largest advantage over static routing: 4\.89 points\. These differences illustrate why aggregate scores should be accompanied by workload slices before an orchestration policy is deployed broadly\.
## 5\.Industry Application and Lessons
First, a strong static policy deserves inclusion in agent evaluations\. It captures much of the benefit of adaptive routing and is simpler to inspect and operate\. Second, success improvements must be reported alongside cost and latency: the adaptive policy is neither cheapest nor fastest\. Third, route composition matters more than any single fixed action in this model\. Fourth, structured traces make policy behavior auditable; aggregate accuracy alone cannot reveal unnecessary calls, retry behavior, or workload\-specific regressions\.
### 5\.1\.Integration Pattern
For deployment, the framework should sit above existing model and tool adapters\. A task\-intake service would provide observable metadata and available capabilities; the router would return a structured route plan; an executor would enforce permissions and budgets; and the trace service would record decisions, outcomes, and stop reasons\. This architecture does not require storage of private chain\-of\-thought\. Short structured rationales, confidence values, tool inputs and outputs, timing, and failure codes are sufficient for policy analysis\.
The separation between router and executor is operationally important\. The router may recommend code execution or delegation, but the executor remains responsible for sandboxing, access control, timeouts, data residency, and allowlists\. Hard controls should not depend on the router’s language\-model judgment\. Likewise, verification should use an independent check where possible rather than asking the same component to approve its own output\.
### 5\.2\.Deployment Protocol
A production evaluation should proceed in four stages\. First, replay historical, consented traces offline and compare proposed routes with the incumbent policy\. Second, run in shadow mode: generate route decisions without executing them, then estimate disagreement, expected cost, and policy coverage\. Third, execute only low\-risk tasks under hard cost and latency limits, retaining a static fallback and human escalation\. Finally, use a randomized or stepped\-wedge comparison where organizational constraints permit, measuring end\-to\-end task success rather than proxy judgments alone\.
Operational acceptance criteria should be specified before testing\. Examples include a lower bound on task success, a maximum p95 latency, a cost\-per\-success ceiling, and a maximum unrecovered tool\-failure rate\. Results should also be sliced by workload, difficulty, data sensitivity, and route type\. A global average can hide a policy that is beneficial for research tasks but harmful for document workflows\.
### 5\.3\.Practical Lessons
The study yields four immediate lessons\. First, strong static policies are credible production baselines, not strawmen\. Second, route composition can improve success, but each additional operation consumes budget and expands the failure surface\. Third, verification appears valuable in the simulator and should be isolated in live ablations rather than assumed beneficial\. Fourth, complete traces turn routing into an observable engineering problem: teams can inspect unnecessary calls, failure recovery, policy disagreement, and workload regressions instead of debugging from final responses alone\.
## 6\.Limitations, Ethics, and Next Steps
The primary limitation is external validity\. Outcomes are sampled from a hand\-specified execution model whose operation benefits depend on the same need dimensions supplied to task\-aware policies\. The adaptive policy should therefore be interpreted as an annotated upper\-bound controller, not a learned router\. Cost units and latency values are normalized assumptions, not invoices or wall\-clock measurements\. Task profiles do not contain natural\-language inputs, live APIs, concurrent agents, security constraints, or human judgments\.
Construct validity is also limited\. “Success” is a Bernoulli outcome generated from the simulator rather than an independently graded artifact, and the confidence intervals quantify sampling variation under the fixed model, not uncertainty about the model assumptions themselves\. The utility weights are illustrative and can change policy rankings\. Internal validity is stronger because policies share task profiles, seeds, budgets, and executor logic, but policy\-specific random streams mean traces are paired at the seed aggregate rather than by identical random draws for every operation\.
The generated workload profiles encode our expectations about which operations help each task family\. This makes the benchmark appropriate for checking whether evaluation machinery behaves coherently, but it also creates a form of evaluator\-policy alignment\. A live study must derive routing signals from task text or operational metadata without revealing ground\-truth operation utilities\. It should also test distribution shift, missing tools, correlated failures, concurrent execution, and adversarial or malformed inputs\.
These limitations prevent claims of deployment readiness\. The next evaluation must replace task annotations with predictions derived from task text, replay real anonymized traces, and then run controlled live tools under fixed budgets\. A learned router, model\-routing baselines, calibration analysis, and human escalation should also be added\. We publish all simulator constants and raw traces so these assumptions can be challenged rather than hidden\.
Agent routing can create privacy, security, and accountability risks when tasks are delegated to external services or code is executed\. Practical deployments need data\-minimization rules, sandboxed execution, allowlisted tools, access control, audit logs, and human review for high\-impact actions\. The current benchmark performs no real external action and contains no personal or proprietary data\.
## 7\.Conclusion
MetaRoute\-Bench isolates the orchestration policy governing how an agentic system decomposes, routes, executes, verifies, and recovers\. In a reproducible offline study, task\-aware adaptive composition improves success over fixed and one\-shot policies while incurring measurable cost and latency\. The framework and traces provide a foundation for the more consequential next step: validation on live, production\-representative workflows\.
###### Acknowledgements\.
This work was conducted with organizational support from Anote AI\. We thank the Anote AI research and engineering team for research infrastructure, technical feedback, and support in developing the reproducible evaluation artifact\. No human participants or proprietary data were used in the reported offline benchmark\.
## Generative AI Disclosure
OpenAI Codex materially assisted with software implementation, experiment scripting, literature organization, and manuscript drafting\. The authors are responsible for reviewing the code, verifying the generated results against the released artifacts, checking all citations, and approving the final claims and text\.
## References
- L\. Chen, M\. Zaharia, and J\. Zou \(2024\)FrugalGPT: how to use large language models while reducing cost and improving performance\.Transactions on Machine Learning Research\.External Links:[Link](https://openreview.net/forum?id=cSimKw5p6R)Cited by:[§1](https://arxiv.org/html/2608.00107#S1.p2.1),[§2](https://arxiv.org/html/2608.00107#S2.p1.1)\.
- Y\. Du, F\. Wei, and H\. Zhang \(2024\)AnyTool: self\-reflective, hierarchical agents for large\-scale api calls\.InProceedings of the 41st International Conference on Machine Learning,External Links:[Link](https://openreview.net/forum?id=qFILbkTQWw)Cited by:[§2](https://arxiv.org/html/2608.00107#S2.p3.1)\.
- S\. Jeong, J\. Baek, S\. Cho, S\. J\. Hwang, and J\. Park \(2024\)Adaptive\-RAG: learning to adapt retrieval\-augmented large language models through question complexity\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,pp\. 7036–7050\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.naacl-long.389),[Link](https://aclanthology.org/2024.naacl-long.389/)Cited by:[§2](https://arxiv.org/html/2608.00107#S2.p1.1),[§3\.4](https://arxiv.org/html/2608.00107#S3.SS4.p2.1)\.
- C\. E\. Jimenez, J\. Yang, A\. Wettig, S\. Yao, K\. Pei, O\. Press, and K\. Narasimhan \(2024\)SWE\-bench: can language models resolve real\-world github issues?\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=VTF8yNQM66)Cited by:[§1](https://arxiv.org/html/2608.00107#S1.p2.1),[§2](https://arxiv.org/html/2608.00107#S2.p4.1)\.
- S\. Kim, S\. Moon, R\. Tabrizi, N\. Lee, M\. W\. Mahoney, K\. Keutzer, and A\. Gholami \(2024\)An llm compiler for parallel function calling\.InProceedings of the 41st International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.235,pp\. 24370–24391\.External Links:[Link](https://proceedings.mlr.press/v235/kim24y.html)Cited by:[§1](https://arxiv.org/html/2608.00107#S1.p2.1),[§2](https://arxiv.org/html/2608.00107#S2.p3.1)\.
- H\. Kokel, M\. Katz, K\. Srinivas, and S\. Sohrabi \(2025\)ACPBench: reasoning about action, change, and planning\.Proceedings of the AAAI Conference on Artificial Intelligence39\(25\),pp\. 26559–26568\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v39i25.34857),[Link](https://ojs.aaai.org/index.php/AAAI/article/view/34857)Cited by:[§2](https://arxiv.org/html/2608.00107#S2.p2.1),[§4\.1](https://arxiv.org/html/2608.00107#S4.SS1.p2.1)\.
- M\. Li, Y\. Zhao, B\. Yu, F\. Song, H\. Li, H\. Yu, Z\. Li, F\. Huang, and Y\. Li \(2023\)API\-bank: a comprehensive benchmark for tool\-augmented LLMs\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 3102–3116\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.emnlp-main.187),[Link](https://aclanthology.org/2023.emnlp-main.187/)Cited by:[§2](https://arxiv.org/html/2608.00107#S2.p3.1)\.
- X\. Liu, H\. Yu, H\. Zhang, Y\. Xu, X\. Lei, H\. Lai, Y\. Gu, H\. Ding, K\. Men, K\. Yang,et al\.\(2024\)AgentBench: evaluating llms as agents\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=zAdUB0aCTQ)Cited by:[§1](https://arxiv.org/html/2608.00107#S1.p2.1),[§2](https://arxiv.org/html/2608.00107#S2.p4.1)\.
- I\. Ong, A\. Almahairi, V\. Wu, W\. Chiang, T\. Wu, J\. E\. Gonzalez, M\. W\. Kadous, and I\. Stoica \(2025\)RouteLLM: learning to route llms with preference data\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=8sSqNntaMr)Cited by:[§1](https://arxiv.org/html/2608.00107#S1.p2.1),[§2](https://arxiv.org/html/2608.00107#S2.p1.1)\.
- T\. Ou, W\. Guo, A\. Gandhi, G\. Neubig, and X\. Yue \(2025\)AgentDiagnose: an open toolkit for diagnosing LLM agent trajectories\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing: System Demonstrations,pp\. 207–215\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.emnlp-demos.15),[Link](https://aclanthology.org/2025.emnlp-demos.15/)Cited by:[§2](https://arxiv.org/html/2608.00107#S2.p4.1),[§3\.1](https://arxiv.org/html/2608.00107#S3.SS1.p1.1)\.
- Y\. Qin, S\. Liang, Y\. Ye, K\. Zhu, L\. Yan, Y\. Lu, Y\. Lin, X\. Cong, X\. Tang, B\. Qian,et al\.\(2024\)ToolLLM: facilitating large language models to master 16000\+ real\-world apis\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=dHng2O0Jjr)Cited by:[§1](https://arxiv.org/html/2608.00107#S1.p2.1),[§2](https://arxiv.org/html/2608.00107#S2.p3.1)\.
- T\. Schick, J\. Dwivedi\-Yu, R\. Dessi, R\. Raileanu, M\. Lomeli, E\. Hambro, L\. Zettlemoyer, N\. Cancedda, and T\. Scialom \(2023\)Toolformer: language models can teach themselves to use tools\.InAdvances in Neural Information Processing Systems,Vol\.36\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/d842425e4bf79ba039352da0f658a906-Abstract-Conference.html)Cited by:[§1](https://arxiv.org/html/2608.00107#S1.p2.1),[§2](https://arxiv.org/html/2608.00107#S2.p3.1)\.
- J\. Su, F\. Lin, Z\. Feng, H\. Zheng, T\. Wang, Z\. Xiao, X\. Zhao, Z\. Liu, L\. Cheng, and H\. Wang \(2026\)CP\-router: an uncertainty\-aware router between LLM and LRM\.Proceedings of the AAAI Conference on Artificial Intelligence40\(39\),pp\. 33065–33073\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v40i39.40589),[Link](https://ojs.aaai.org/index.php/AAAI/article/view/40589)Cited by:[§2](https://arxiv.org/html/2608.00107#S2.p1.1),[§3\.4](https://arxiv.org/html/2608.00107#S3.SS4.p2.1)\.
- H\. Xu, R\. Zhao, J\. Wang, and H\. Chen \(2024\)RESTful\-llama: connecting user queries to RESTful APIs\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track,pp\. 1433–1443\.External Links:[Document](https://dx.doi.org/10.18653/v1/2024.emnlp-industry.105),[Link](https://aclanthology.org/2024.emnlp-industry.105/)Cited by:[§2](https://arxiv.org/html/2608.00107#S2.p3.1)\.
- C\. Yan, W\. Zhang, Z\. Ning, F\. Xu, Z\. Tao, L\. Zhang, B\. Yin, and Y\. Zhang \(2026\)Breaking model lock\-in: cost\-efficient zero\-shot LLM routing via a universal latent space\.Proceedings of the AAAI Conference on Artificial Intelligence40\(43\),pp\. 36483–36490\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v40i43.40970),[Link](https://ojs.aaai.org/index.php/AAAI/article/view/40970)Cited by:[§2](https://arxiv.org/html/2608.00107#S2.p1.1),[§3\.4](https://arxiv.org/html/2608.00107#S3.SS4.p2.1)\.
- S\. Yao, N\. Shinn, P\. Razavi, and K\. Narasimhan \(2024\)Tau\-bench: a benchmark for tool\-agent\-user interaction in real\-world domains\.arXiv preprint arXiv:2406\.12045\.External Links:[Link](https://arxiv.org/abs/2406.12045)Cited by:[§1](https://arxiv.org/html/2608.00107#S1.p2.1),[§2](https://arxiv.org/html/2608.00107#S2.p4.1)\.
- S\. Yao, J\. Zhao, D\. Yu, N\. Du, I\. Shafran, K\. Narasimhan, and Y\. Cao \(2023\)ReAct: synergizing reasoning and acting in language models\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=WE_vluYUL-X)Cited by:[§1](https://arxiv.org/html/2608.00107#S1.p2.1),[§2](https://arxiv.org/html/2608.00107#S2.p2.1)\.
- Y\. Zhang, G\. Ganapavarapu, S\. Jayaraman, B\. Agrawal, D\. Patel, and A\. Fokoue \(2026\)SPIRAL: symbolic LLM planning via grounded and reflective search\.Proceedings of the AAAI Conference on Artificial Intelligence40\(43\),pp\. 36527–36535\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v40i43.40975),[Link](https://ojs.aaai.org/index.php/AAAI/article/view/40975)Cited by:[§2](https://arxiv.org/html/2608.00107#S2.p2.1)\.
## Appendix AExecution Model
The offline executor uses explicit per\-operation assumptions, listed in Table[4](https://arxiv.org/html/2608.00107#A1.T4)\. Difficulty multiplies cost and latency by1\+0\.08\(d−1\)1\+0\.08\(d\-1\)\. Base success is0\.69−0\.075\(d−1\)−0\.13a0\.69\-0\.075\(d\-1\)\-0\.13a, whereddis difficulty andaais ambiguity\. Each selected operation adds a benefit proportional to its annotated task need and a small penalty when unnecessary\. Routes longer than two support operations incur coordination overhead\. External operations can fail; the adaptive condition permits one retry with additional cost and latency\. Probabilities are clipped to\[0\.02,0\.98\]\[0\.02,0\.98\]\. The implementation insrc/metarouter/simulator\.pyis authoritative\.
Table 4\.Offline executor parameters\. Failure probability is zero where omitted\.
## Appendix BReproducibility
The artifact requires Python 3\.10 or later\. The following commands regenerate the main and ablation results and verify that the headline manuscript values match the exported CSV files\.
```
python -m venv .venv
source .venv/bin/activate
pip install -e ".[dev]"
metarouter-benchmark --seeds 30 \
--output results/meta-routing/dai2026/main
python experiments/meta-routing/dai2026/run_ablations.py
python experiments/meta-routing/dai2026/check_paper_results.py
pytest -q
```
The artifact exports task profiles, raw traces, policy summaries, paired comparisons, workload\-level success, action counts, and failure distributions\. The random stream is keyed by seed, task identifier, and policy name\.Similar Articles
Learning Compositional Meta-Routing for Agentic Workflows: An Executable Benchmark
This paper introduces an executable benchmark for learning compositional meta-routing in agentic workflows, where a budget-aware controller composes operations like retrieval, code execution, and verification. The learned policy outperforms static workflows on held-out tests but shows lower lexical generalization on challenge splits.
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows
DecisionBench introduces a standardized benchmark for evaluating emergent delegation in long-horizon multi-agent workflows, providing a substrate with task suites, peer models, and multi-axis metrics to isolate orchestration capabilities.
AgentRouter: Heterogeneous Model Routing for Cost-Optimal Multi-Step Agentic Workflows
AgentRouter is a lightweight classifier for routing steps in agentic workflows to different model tiers, achieving 72% cost reduction with minimal quality degradation compared to frontier-only models.
@tomas_hk: Today we’re releasing our methodology for evaluating model routing with interactive benchmarks, which represent agent c…
Releasing a methodology for evaluating model routing with interactive benchmarks that achieve Pareto-dominance over leading benchmarks, offering higher quality at lower cost.
Task- and Session-Level Model Routing: A Common-Interface Hybrid Evaluation of Four Open-Source Routers Across Four Benchmarks
This paper evaluates four open-source LLM routers using a common protocol across four benchmarks, finding that performance gains are more closely tied to model tier composition than task-specific targeting.