MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale
Summary
MERA introduces a multi-cycle adaptation approach for LLM agents that improves small models by distilling execution-verified demonstrations into a SkillBook and fine-tuning LoRA adapters, with router-based deployment and verifier-backed fallback. Experiments show Qwen2.5-Coder-1.5B improves from 28.7% to 49.7% pass on HumanEval++MBPP while retaining most quality at lower cost.
View Cached Full Text
Cached at: 08/12/26, 08:28 AM
# MERA: Model Evolution and Routing with Skill Adaptation for Agentic Systems at Scale
Source: [https://arxiv.org/html/2608.10333](https://arxiv.org/html/2608.10333)
\\correspondence
yuhangyao8@gmail\.com, tianyu@gradient\.network\\sourcecodehttps://github\.com/yh\-yao/MERA\-Evolve
Yuhang Yao4,†\\daggerZeyu Wang6Wanyi Chen2Tongyun Yang3Yuhang Han5 Jie Xiao1Chengke Bao1Tianyi Zhao1Lynn Ai1Eric Yang1Tianyu Shi1,†\\dagger1Gradient2Soochow University3Independent Researcher4Carnegie Mellon University 5Shanghai Jiao Tong University6University of California, Los Angeles †\\daggerCorresponding author
###### Abstract
LLM agents execute heterogeneous sequences of model calls within a single task: some invocations require careful reasoning, while others are structured steps such as formatting or tool\-argument construction\. Prior routing methods exploit this asymmetry by assigning easy invocations to a cheaper small model and difficult ones to a large model\. Such policies reduce inference cost, but they leave the small model’s capability unchanged, so attainable savings remain bounded by the work the student can already solve\. MERA instead improves the small model itself, using a single model invocation as the unit of adaptation\. In each cycle, MERA replays failed student invocations to obtain execution\-verified teacher demonstrations, distills recurring procedures into an iteratively updated SkillBook, and fine\-tunes a student LoRA adapter via supervised learning and optional GRPO\. Routing serves as supporting machinery for deployment: the improved student is served behind a cost\-calibrated router with verifier\-backed fallback, and a candidate SkillBook, adapter, or router is admitted only when joint replay preserves task quality\. Empirically, four\-cycle adaptation raises Qwen2\.5\-Coder\-1\.5B from 28\.7% to 49\.7% pass on held\-out HumanEval\+\+MBPP\. Under verifier\-backed fallback, the deployed policy retains 88\.3% pass at 60\.8% of always\-Luna cost\. On TAU\-2, a fine\-tuned Qwen3\.5\-2B improves from 14/35 to 18/35 and matches an unadapted 4B model\. These results indicate that verifier\-backed multi\-cycle adaptation can increase small\-model capability, rather than only routing around a fixed student\.
## 1Introduction
Language\-model agents execute heterogeneous chains of model calls\. A single task may contain hard reasoning or policy\-sensitive decisions, but it also contains many structured steps: extraction, formatting, tool\-argument construction, post\-processing, clarification, and templated summaries\. Running every invocation on the strongest model is reliable but expensive\. Conversely, routing only once at the user\-task level is too coarse, because the easy and hard parts of a workflow often appear side by side\. Recent work has improved agents through prompting, tool use, planning, search, and agent optimization\(react2023;toolformer2023;lats2024;toolllm2023;agentoptimizer2024;gptswarm2024\), but production systems also need mechanisms for adapting the models, routes, and reusable skills inside the workflow\.
This creates a practical tension for deployed agent systems\. The traces needed for adaptation are naturally produced online, but directly changing the serving policy from raw traces is risky: a cheaper model may fail silently, a learned adapter may regress on rare cases, and a reusable template may only be safe under a narrow prompt signature\. A useful evolution loop therefore needs to separate observation from admission\. It should collect evidence at the granularity of individual invocations, update several candidate components, and admit the resulting runtime state only when replay shows that the combined policy still satisfies verification constraints\.
MERA addresses this problem by treating a single model invocation as the unit of adaptation\. At runtime, an input\-only router chooses among strong, cheap, and specialized models; a skill layer can dispatch stable templates for recurring local structure; and a verifier protects quality through fallback\. This design keeps serving simple: the router does not depend on hidden agent state, and unsafe down\-routing is corrected by verification\. The heavier logic is moved offline, where traces are replayed to decide which prompts are easy, which failures should become training examples, and which recurrent patterns are stable enough to become skills\.
Agent workloadheterogeneousmulti\-step tasksAgent executionrouting, skills,verificationOnline tracesprompts, outputs,verifier outcomesSkillBooksignature successstatisticsLLM updatehard\-exampleadaptationLearned routercheap vs\.strong routesJoint replay gateadmit nextruntime statenext round
Figure 1:Overview of MERA\. Online traces drive scheduled SkillBook, LLM\-update, and router tracks; their combined state is admitted through joint replay evaluation\.Figure[1](https://arxiv.org/html/2608.10333#S1.F1)shows the resulting feedback loop\. Online traces drive update rounds in which SkillBook statistics, router training, and LLM adaptation share evidence\. Within a cycle we use the dependency\-respecting Skill→\\rightarrowLLM→\\rightarrowRouter schedule: the adapter is trained with the current SkillBook procedure, and the router is fit last so its labels reflect current\-cycle student outcomes\. The combined state is admitted only after joint replay evaluation, so routing, skill promotion, and adapter updates are judged by their end\-to\-end effect rather than by isolated component metrics\.
Our contribution is a verifier\-backed multi\-cycle protocol for improving the small model from shared executable traces; routing and SkillBook are supporting machinery for safe deployment\. On 582 held\-out HumanEval\+\+MBPP tasks \(3 seeds\), four\-cycle adaptation raises Qwen2\.5\-Coder\-1\.5B from 28\.7% to 44\.2% under SFT and to 49\.7% with matched SFT\+GRPO\. With executable verification and GPT\-5\.6 Luna fallback, the deployed policy retains 88\.3% pass at 60\.8% of always\-Luna cost\. On TAU\-2, an adapted Qwen3\.5\-2B improves from 14/35 to 18/35 and matches an unadapted 4B endpoint, though the comparison is underpowered\. A finance break\-even analysis further connects adaptation cost to serving savings\. MERA is therefore an evolution protocol first: it raises small\-model capability rather than only routing around a fixed student\.
## 2Related Work
Runtime loopInvocationserialized promptlocal contextRouterstrong / cheap /specialized studentSkill selectoroptional templatedispatchExecutionLLM calltool useVerifierschemasuccessfallbackfallbackIterative update loopTrace storepromptsoutputsverifier logsStep slicescanonicalizedinvocation statesSkillBooksignature statsmotifsLLM updatehard examplesSFT / GRPORouter updateexecutable labelsthresholdJoint replayadmission /next roundSkill→\\rightarrowLLMLLM→\\rightarrowRouteradmitIterative dependency\.Each round updates SkillBook, the LLM adapter, then the router over shared traces \(Skill→\\rightarrowLLM→\\rightarrowRouter\)\. Admission depends on joint replay evaluation\.
Figure 2:Detailed method view of MERA\. Runtime routing remains input\-only and verifier\-protected; update tracks share traces and are admitted through joint replay\.LLM routing\.Prior work on model routing studies how to dispatch requests across model pools under cost\-quality trade\-offs\. FrugalGPT\-style cascades, learned routers, and preference\- or uncertainty\-based selectors use prompt features, predicted quality, or explicit routing objectives to decide when a cheaper model is sufficient and when a stronger model is needed\(frugalgpt2024;routellm2024;dekoninck2025routing;bestroute2025\)\. Recent systems further consider upgraded model tiers, multi\-round routing, and answer aggregation rather than one\-shot dispatch\(llmat2025;routerr12025\)\. This literature establishes routing as a practical mechanism for reducing inference cost, but most settings route whole user requests or single\-turn examples\. The unit of decision is usually the prompt submitted to a model, not an internal invocation within a longer agent execution\.
Agent optimization\.Agent research has focused on making agents more capable through tool use, reasoning\-action loops, execution\-time search, and automated optimization of agent graphs or functions\(react2023;toolformer2023;lats2024;toolllm2023;agentoptimizer2024;gptswarm2024\)\. Recent post\-training and environment\-based approaches also train agents from interaction traces or simulated feedback\(autopdl2025;agentgymrl2026\)\. These methods improve the policy, planner, or tool\-use behavior of an agent\. They typically treat the model choice, router, and reusable procedures as fixed engineering choices, or evaluate a single improved agent policy\. They do not directly address how a deployed agent should use its own execution traces to update the model mixture and routing policy over time\.
Model specialization and reusable skills\.Distillation, step\-by\-step supervision, small\-model adaptation, reflective self\-training, and explicit skill\-library construction all aim to reuse behavior learned from stronger models or successful trajectories\(distillingstepbystep2023;agentr2025;skillx2026\)\. These methods show that smaller models and externalized skills can capture useful structure\. However, their success is often reported as a standalone training or benchmark result\. In deployed agent systems, specialization is conditional: a student or skill is useful only on slices where it remains reliable, and failures may need to trigger fallback rather than silently enter production\.
Evaluation for tool\-using agents\.Benchmarks for grounded web interaction, tool use, multi\-hop tools, agent\-user interaction, computer control, and software engineering provide increasingly realistic environments for measuring agent behavior\(webshop2022;toolllm2023;toolhop2025;taubench2025;osworld2024;swebench2024\)\. These benchmarks are valuable because many failures can be checked by tools, task assertions, execution results, or natural\-language judges\. At the same time, they expose a deployment difficulty: as tool ecosystems grow and workflows become longer, a single user request contains many heterogeneous steps, only some of which are safe for cheaper execution or specialization\.
The limitations across these methods are not a lack of routers, skills, or post\-training methods in isolation\. The missing systems layer is a conservative mechanism that connects them inside a running agent: traces should identify repeated invocation types, train or update students on hard examples, update routers at the same granularity, and admit all changes only through replay with verifier\-backed fallback\. MERA addresses this gap by treating execution traces as shared supervision for SkillBook statistics, invocation\-level routing, and model adaptation\. Rather than only routing around a fixed small model, MERA expands cheap execution where replay shows that the combined router, skill state, and adapter preserve task quality\.
## 3Method
### 3\.1Runtime Routing
MERA separates the serving path from the update path\. At runtime, the router observes only the serialized prompt for the current invocation and selects among a strong model, a cheap model, and optionally a specialized student\. A skill selector may dispatch stable templates for recurring local structure, such as formatting, extraction, or single\-tool argument construction\. The selected model produces an output, and a verifier checks schema validity, tool\-call legality, executable tests, or downstream success\. If verification fails, the invocation falls back to a stronger model and the full event is logged\.
This input\-only router is intentionally restrictive\. It avoids coupling routing decisions to hidden agent state or implementation\-specific tool traces, making the interface easier to deploy across agent harnesses\. Reliability is instead protected after execution by the verifier and fallback path\. The consequence is that MERA can start conservatively: early routers may prioritize low unsafe down\-routing, while replay and SkillBook evidence gradually identify regions that can be served cheaply\.
### 3\.2Trace Products
The update loop operates on complete traces collected from runtime execution or replay\. MERA canonicalizes each trace into step slices containing the prompt, local context, tool schemas, generated output, verifier result, retry count, fallback metadata, and any skill assignment\. These slices produce three evidence streams\. SkillBook records success and failure statistics for recurring prompt signatures\. The learned router trains on prompt text with cheap/strong labels derived from executable small\-model outcomes\. The LLM adapter trains on selected hard examples where cheaper execution fails or where the router and SkillBook disagree\.
The same slice can therefore support different updates without forcing the components to share the same supervision format\. Router examples can remain input\-only, while LLM examples preserve the context needed to execute the step\. Skill examples are grouped by repeated local structure rather than by label alone\.
Algorithm 1MERA runtime and iterative update loop1:router
RR, model registry
ℳ\\mathcal\{M\}, SkillBook
𝒦\\mathcal\{K\}, verifier
VV
2:foreach agent invocation
xtx\_\{t\}do
3:choose model
mt←R\(xt\)m\_\{t\}\\leftarrow R\(x\_\{t\}\)and optional skill
kt∈𝒦k\_\{t\}\\in\\mathcal\{K\}
4:execute
yt←mt\(xt,kt\)y\_\{t\}\\leftarrow m\_\{t\}\(x\_\{t\},k\_\{t\}\)
5:if
V\(xt,yt\)V\(x\_\{t\},y\_\{t\}\)failsthen
6:fallback to a stronger model
7:endif
8:log
\(xt,yt,mt,kt,V\(xt,yt\)\)\(x\_\{t\},y\_\{t\},m\_\{t\},k\_\{t\},V\(x\_\{t\},y\_\{t\}\)\)
9:endfor
10:foreach update rounddo
11:canonicalize traces into step slices
12:update SkillBook, router, and LLM adapter under chosen schedule
13:run joint replay; admit only if quality is preserved
14:endfor
### 3\.3Scheduled Updates
Algorithm[1](https://arxiv.org/html/2608.10333#alg1)summarizes the loop\. Each update round refreshes the three tracks over shared traces, and the order respects a data dependency rather than being free\. The LLM adapter is trained and queried with the SkillBook procedure prepended to its prompt, so the SkillBook update must precede LLM adaptation within a cycle \(Skill→\\rightarrowLLM\); training the two in parallel would fit the adapter on a stale procedure\. The router is placed last so its labels reflect the current\-cycle skill state and small\-model outcomes\. This yields the canonical Skill→\\rightarrowLLM→\\rightarrowRouter schedule we use throughout; because the loop iterates, any residual staleness is absorbed by the next cycle\.
This design lets MERA distinguish scientific effects from systems effects\. Figure[2](https://arxiv.org/html/2608.10333#S2.F2)makes the dependency structure explicit: runtime execution produces trace slices, the three update tracks consume different evidence products from those slices, and replay admits only the combined state\. The evaluation therefore reports direct small\-model quality separately from pre\-routing, verification and fallback, end\-task quality, and normalized cost\. This prevents a high cascade pass rate from being mistaken for either strong standalone model capability or accurate pre\-routing\.
### 3\.4Operational Component Definitions
In the code\-generation implementation, SkillBook is external procedural prompt memory rather than a response cache or adapter weight\. The signature function maps tasks to two coarse dataset\-level keys,humanevalandmbpp\. Each rendered entry combines static task\-format instructions with bounded successful exemplars and exemplar\-grounded recurring pitfalls and patterns\. It is prepended to a new task and never returns a stored answer for an identical prompt\.
Router supervision is also operationally grounded\. A task is cheap\-eligible when the current small\-model rollout passes the benchmark verifier and requires escalation otherwise\. The code\-generation router uses frozen Qwen3\-Embedding\-0\.6B prompt features and a class\-balanced logistic\-regression head\. Repeated observations of a task are kept in the same cross\-validation group, threshold calibration uses a disjoint shard, and policy results are reported on held\-out task identifiers\. This construction replaces the earlier weak\-supervision and BERT\-tiny description\.
Verification is benchmark\-specific\. For HumanEval and MBPP, generated code is executed in an isolated Python subprocess against the benchmark\-provided test program\. For TAU\-2, the official evaluator checks completion using environment state, required tool actions, and configured natural\-language assertions\. The router does not call an LLM judge at inference\.
### 3\.5Joint Admission
MERA admits updates through joint replay rather than isolated component metrics\. SkillBook, router, and verifier evidence define easy, hard, and uncertain regions\. Easy regions are candidates for cheap serving or future student admission; hard examples feed the LLM update; uncertain regions remain protected by fallback\. A new router, skill state, or adapter is promoted only if replay preserves quality while reducing cost or fallback risk\.
This yields a conservative deployment strategy\. The runtime system can continue using strong\-model fallback while update rounds search for cheaper safe regions\. When an update does not improve the joint replay result, it remains an experimental artifact rather than entering the serving registry\. Replay itself is only a pre\-deployment gate: a production deployment still requires shadow or canary validation, drift monitoring, and rollback because an updated policy can change the traffic distribution\.
## 4Experiments
### 4\.1Experimental Setup
We evaluate three questions\. First, does multi\-cycle adaptation improve the small model beyond a matched SFT\-only control? Second, can executable verification and fallback turn that stronger SLM into a favorable deployed cost–quality operating point? Third, does the model\-update result transfer to a multi\-turn tool\-use setting under a controlled adapter\-only comparison?
The code\-generation study merges HumanEval and MBPP into 546 training tasks and 582 held\-out evaluation tasks with disjoint task identifiers and zero exact prompt overlap\. The small model is Qwen2\.5\-Coder\-1\.5B\-Instruct and the large teacher/fallback model is GPT\-5\.6 Luna\. We compare four\-cycle SkillBook\+SFT and SkillBook\+SFT\+GRPO schedules using three independent training seeds\. Teacher outputs are cached so that matched runs share the same supervision\. Direct\-SLM evaluation disables both routing and fallback\. For deployed policies, cost is normalized to always using the large model, with a small:large cost ratio of1:101\{:\}10\. Generated code is executed in an isolated Python subprocess against the benchmark\-provided test program; these are not self\-generated tests\.
The SkillBook uses the documented dataset\-level signatureshumanevalandmbpp\. Router targets are generated from executable small\-model outcomes\. Frozen Qwen3\-Embedding\-0\.6B features feed a class\-balanced logistic\-regression head; repetitions are grouped during cross\-validation, and threshold calibration is disjoint from held\-out policy evaluation\. Thus no result below uses the previously described UncommonRoute weak\-label set\.
### 4\.2Multi\-Cycle Small\-Model Adaptation
Table 1:Three\-seed means on the 582 held\-out HumanEval\+MBPP tasks\. Direct SLM pass disables routing and fallback; normalized cost is relative to always using GPT\-5\.6 Luna\.PolicyPass \(%\)Cost \(%\)Base SLM \(Qwen2\.5\-Coder\-1\.5B\)28\.710\.0Multi\-cycle SFT44\.210\.0Multi\-cycle SFT\+GRPO49\.710\.0Always Luna86\.9100\.0
Table 2:Final\-cycle routing and fallback comparison on the matched SFT\+GRPO artifacts \(three\-seed mean\)\. Thresholds are chosen on a calibration shard only; RouteLLM\-/FrugalGPT\-style rows use matched adaptations rather than official checkpoints\.PolicyPass \(%\)Cost \(%\)Always small / exact\-cache\+small49\.710\.0Always Luna86\.9100\.0RouteLLM\-style \+ fallback87\.097\.4FrugalGPT\-style response cascade85\.5106\.7MERA router \+ verifier fallback88\.360\.8
Table 3:Strict TAU\-2 comparison on the fixed 35\-task split\. Evaluation uses the official environment outcome and disables SkillBook, routing, and fallback\.PolicyAgentAirlineRetailTelecomOverallBase SLMQwen3\.5\-2B5/97/182/814/35 \(40\.0%\)Larger local endpointQwen3\.5\-4B \(unadapted\)5/99/183/817/35 \(48\.6%\)Trained SLMQwen3\.5\-2B \+ VERL GRPO5/910/183/818/35 \(51\.4%\)
Table 4:Finance priority\-data cost planning\. Break\-even divides one\-time training cost by estimated daily serving savings\. The final column applies a conservative 70% savings realization factor\.Train rowsTrain cost \($\)Deploy cost \($/hr\)Savings \($/hr\)Break\-even70% break\-even300168\.750\.350\.4615\.29 days31\.96 days500281\.250\.351\.358\.68 days12\.47 days
Table[1](https://arxiv.org/html/2608.10333#S4.T1)reports the rebuilt multi\-cycle result \(final\-cycle, three\-seed mean\)\. Multi\-cycle fine\-tuning lifts the SLM from 28\.7% to 44\.2% \(SFT\) and 49\.7% \(SFT\+GRPO\)\. Matched GRPO exceeds SFT by 5\.5–6\.7 points \(95% paired\-ttintervals exclude zero\)\. Cascade quality is already near 88%; the gain is a stronger directly usable SLM with lower fallback demand, not a cascade lift\.
### 4\.3Verifier\-Backed Fallback
Table[2](https://arxiv.org/html/2608.10333#S4.T2)separates direct model improvement from the deployed operating point\. Exact\-cache coverage is zero on held\-out prompts, so it equals always\-small\. With verifier fallback, MERA matches near\-Luna quality at 60\.8% cost, while the RouteLLM\- and FrugalGPT\-style baselines remain near always\-Luna cost\. The system win is the evolved SLM plus verifier fallback; the learned pre\-router is weak, and most quality preservation comes from verification rather than routing alone\.
### 4\.4TAU\-2 Adapter Ablation
Table[3](https://arxiv.org/html/2608.10333#S4.T3)is a clean base\-versus\-GRPO comparison under the same parser, split, user simulator, and no\-fallback policy\. The trained Qwen3\.5\-2B improves from 14/35 to 18/35 and can reach the performance of the unadapted Qwen3\.5\-4B endpoint\. The paired test has seven wins, three losses, and 25 ties \(one\-sided McNemarp=0\.171875p=0\.171875\), so the comparison is underpowered and supports feasibility on tool use rather than broad cross\-domain generality\.
### 4\.5Finance Deployment Planning
Finance is retained as a deployment\-planning case rather than a public benchmark\. Under the stated workload and serving\-cost assumptions, the 500\-row adaptation setting has the larger one\-time training cost but the shorter nominal break\-even time: 8\.68 days, or 12\.47 days when only 70% of estimated savings are realized\. This analysis connects the paper’s primary small\-model improvement objective to an operational decision about when adaptation pays for itself\.
### 4\.6Discussion
The rebuilt results place small\-model evolution at the center of the paper\. The strongest evidence is the three\-seed SLM lift under matched SFT and SFT\+GRPO\. Verifier fallback turns that improvement into a near\-Luna cost–quality point at 60\.8% cost, while query/response routers stay near always\-Luna cost\. TAU\-2 supplies a clean but underpowered adapter\-only check in which a fine\-tuned 2B agent matches an unadapted 4B endpoint, and finance illustrates when adaptation pays for itself\. A statistically established TAU\-2 gain remains outside the evidence\.
## 5Conclusion
MERA uses shared executable traces primarily to improve the small model, with SkillBook, routing, and verifier\-backed admission as supporting machinery\. On held\-out HumanEval\+\+MBPP, four\-cycle adaptation raises Qwen2\.5\-Coder\-1\.5B from 28\.7% to 49\.7% direct pass, and verifier\-backed deployment retains 88\.3% pass at 60\.8% of always\-Luna cost\. On TAU\-2, an adapted Qwen3\.5\-2B improves from 14/35 to 18/35 and matches an unadapted 4B endpoint, though the comparison is underpowered\. A finance break\-even analysis further connects one\-time adaptation cost to serving savings\. Overall, MERA offers a conservative protocol for raising small\-model capability from verifier\-grounded traces rather than only routing around a fixed student; broader verified workloads and multi\-seed agentic evaluation remain necessary\.
## References
## 6Additional Experiment Details
### 6\.1Replay and Router Label Construction
MERA treats replay as the common interface between routing, SkillBook updates, and model adaptation\. For each logged invocation, we retain the serialized input, available tool schema, model output, verifier result, retry and fallback metadata, and any skill identifier\. Replay then executes candidate runtime states against the same slice representation\. A slice is considered eligible for cheaper execution only when the candidate output satisfies the same verifier used by the strong\-model path\. If the cheap model, adapter, or skill state fails verification, the slice remains assigned to the stronger model or the fallback region\.
Router labels are therefore conservative\. Positive cheap\-model labels come from slices where cheaper execution preserves the verifier outcome, while strong\-model labels come from failures, verifier uncertainty, or examples outside the observed support of the current SkillBook statistics\. This is stricter than training a router from model confidence alone: the router is not asked to predict whether a cheap model sounds plausible, but whether the current runtime state can safely handle the invocation under replay\.
### 6\.2Admission Metrics
The normalized cost metric reports the estimated serving cost of the admitted policy relative to always using the large model\. Letcsc\_\{s\}andclc\_\{l\}be the small\- and large\-model costs for an invocation, and letffdenote whether fallback is triggered\. The per\-invocation replay cost is counted ascs\+fclc\_\{s\}\+fc\_\{l\}when a cheap path is attempted and asclc\_\{l\}when the large model is selected directly\. We report the aggregate ratio over the held\-out replay set\. Fallback rate is the fraction of invocations for which the cheap path fails verification and requires escalation\.
An update is a candidate for admission only when it improves the cost\-quality operating point without introducing additional unverified failures\. In the experiments, this rule is applied to the joint state rather than to isolated component metrics: a router that looks accurate in isolation is not sufficient if the resulting skill or adapter path fails replay\. Conversely, an adapter improvement is reported both as a direct\-SLM gain and as a deployed cascade with verifier fallback, so a weak pre\-router is not mistaken for a weak student\.
### 6\.3Component Update Details
SkillBook updates aggregate repeated prompt signatures and local execution motifs\. The goal is not to memorize full trajectories, but to identify compact invocation types that have stable verifier outcomes across traces\. The LLM update track uses hard examples from slices where the cheaper model fails or where router and SkillBook evidence disagree\. The router update track consumes prompt\-level labels derived from executable outcomes and can be scheduled after the other tracks when it should observe current\-cycle adapter or SkillBook evidence\.
Because the LLM adapter consumes the SkillBook procedure, we run the dependency\-respecting Skill→\\rightarrowLLM→\\rightarrowRouter schedule\. The SFT\-only and SFT\+GRPO arms share the same split, cached teacher outputs, verifier, and cost accounting\. Direct\-SLM evaluation disables routing and fallback and is the primary scientific readout of student improvement; the routing comparison reuses the matched final\-cycle artifacts as supporting deployment evidence\.
### 6\.4TAU\-2 Controlled Protocol
The reportable TAU\-2 comparison uses the same 35 tasks for all agents, a Qwen3\.5\-4B user simulator, non\-thinking templates, a 1024\-token cap per agent turn, and no SkillBook, router, or fallback during evaluation\. Final pass/fail comes from the official environment outcome\. The base and GRPO rows therefore isolate the adapter under one inference protocol\. The unadapted Qwen3\.5\-4B endpoint is reported as a descriptive reference under the same split \(17/35\); it is not part of the paired base\-versus\-GRPO claim\. Older runs that changed the parser, SFT recipe, user simulator, token cap, or fallback policy are excluded rather than pooled with this comparison\.
The paired base\-versus\-GRPO outcome contains seven wins, three losses, and 25 ties \(14/35→\\to18/35\)\. We report the one\-sided exact McNemar value,p=0\.171875p=0\.171875, and treat the observed four\-task increase as underpowered evidence of feasibility rather than a statistically established gain\.
### 6\.5Finance Break\-Even Calculation
The finance table is a deployment\-planning calculation rather than a public benchmark\. Nominal break\-even is
break\-even=one\-timetrainingcostdailyservingsavings\.\\mathrm\{break\\mbox\{\-\}even\}=\\frac\{\\mathrm\{one\\mbox\{\-\}time\\ training\\ cost\}\}\{\\mathrm\{daily\\ serving\\ savings\}\}\.The conservative column applies a 70% realization factor to estimated savings before computing break\-even\. These assumptions keep the finance result separate from the open\-source HumanEval\+MBPP and TAU\-2 benchmark evidence\.
## 7Limitations
MERA is designed around replay, verification, and conservative staged deployment, and these choices introduce corresponding limitations\. First, the quality of both routing labels and student admission decisions depends on verifier coverage\. If the verifier fails to capture an important semantic failure mode, replay may overestimate the safety of down\-routing or skill promotion\. Second, our simple\-step\-first strategy is intentionally biased toward narrow and highly checkable slices\. This makes early deployment safer, but it also means that the framework may realize its gains gradually and may leave a substantial fraction of difficult long\-horizon reasoning on the strongest model for a long time\.
Third, the current evaluation protocol is trace\-centric rather than fully online\. Replay is attractive because it enables controlled counterfactual evaluation of routing, student models, and skills under a shared verifier, but replay cannot perfectly capture distribution shift induced by changing the runtime policy itself\. Fourth, skill promotion assumes that repeated local subgraphs of agent behavior can be identified and canonicalized into stable templates\. Some tasks may remain too heterogeneous for this representation to pay off\. Fifth, the learned pre\-router remains weak in our rebuilt study: most deployed quality preservation comes from executable verification and fallback rather than accurate upfront routing\. Finally, the strongest multi\-seed evidence is code generation; the TAU\-2 adapter check is clean but underpowered, and gains from specialization remain limited if the workload is highly non\-stationary or contains few reusable step types\.
## 8Broader Impact
MERA aims to reduce the cost of agentic systems without giving up end\-to\-end reliability\. A positive consequence is that stronger agent workflows may become deployable at lower serving cost, which could make high\-quality automation more accessible\. The same trace\-centric design may also improve operational safety by requiring replay, verification, and fallback before new routing or specialization decisions reach production\.
The same capabilities also carry risks\. More efficient agents may accelerate large\-scale automation in settings where reliability, privacy, or oversight matter\. If verification is incomplete, down\-routing or skill promotion could create subtle failures that are cheap to execute but costly to detect\. The framework could also be used to lower the operating cost of agents in sensitive domains without adequate human review\. These risks suggest that deployment should remain conservative, with explicit verifier coverage, staged admission, audit logs, and scope restrictions on production use\.
## 9LLM Usage
Large language models are a core methodological component of this work\. They are used both as runtime policies within the agent and as the objects being routed, specialized, and evaluated\. The method further relies on replay across multiple candidate models to generate routing labels and admission decisions\. Our use of LLMs is therefore part of the scientific contribution itself rather than a writing\-only aid\.Similar Articles
Skill is Not One-Size-Fits-All: Model-Aware Skill Alignment for LLM Agents
This paper proposes MASA, a framework that adapts skills to each LLM backbone without modifying weights, using hierarchical evolution and a model-conditioned rewriter, achieving gains of up to 25.8 points over baselines.
CERA-MoA: Co-Evolving Routing Mechanisms with Continually Learning LLM Agents
CERA-MoA introduces a co-evolving framework for mixture-of-agents systems that uses reinforcement learning to dynamically route queries and adapt agent capabilities, enhancing task performance and efficiency.
MASkills: Continual Skills Optimization for Multi-Agent LLM Systems
MASkills presents a continual learning framework that optimizes multi-agent LLM systems through agent skills, using skill-conditioned credit assignment and hierarchical aggregation to improve performance on tasks like HotpotQA and GAIA.
@dair_ai: // Evolving Meta-Skill for Multi-Agent Systems // Can a multi-agent system get better at orchestration without touching…
Skill-MAS introduces a method for evolving meta-skills in multi-agent systems to improve orchestration without modifying model weights, achieving transferable performance gains across tasks and LLMs.
ERSkill: Evolving for Skill-Guided Adaptive Memory Retrieval
Introduces ERSkill, a retrieval-centric framework for self-evolving, skill-guided adaptive memory access in LLM agents. It co-evolves retrieval skills and a routing policy, substantially outperforming strong baselines across agent memory benchmarks.