ChipMEM: Verification-Grounded Memory for EDA Agents
Summary
ChipMEM introduces a verification-grounded memory layer for EDA agents, enhancing RTL design performance by distilling reusable skills from verified tool interactions and using Bayesian statistical guidance for transfer to unseen tasks.
View Cached Full Text
Cached at: 09/24/26, 09:36 AM
# ChipMEM: Verification-Grounded Memory for EDA Agents
Source: [https://arxiv.org/html/2609.27067](https://arxiv.org/html/2609.27067)
\\workshoptitle
AI for Chip Design
Abdulrahman AlRabah1Joshua Mabry2Dilek Hakkani\-Tür1Abdussalam Alawini1Hamid Shojaei3Kartik Hegde3Sandesh Adhikary31University of Illinois Urbana\-Champaign2NVIDIA3Cadence††thanks:Corresponding author:alrabah2@illinois\.edu
###### Abstract
Large language model \(LLM\)\-based agents use Electronic Design Automation \(EDA\) tools to generate and revise register\-transfer\-level \(RTL\) designs under synthesis and verification feedback\. Recent methods learn from this feedback by distilling reusable skills from execution traces or by training on rewards derived from EDA\-tools\. Both methods are typically evaluated on the tasks that produced the experience\. Repeated access to benchmark feedback on the same task can reward task\-specific revision rather than creating reusable knowledge that transfers\. We introduce ChipMEM, a verification\-grounded memory layer for EDA agents\. It combines cross\-task procedural memory with within\-trajectory statistical guidance\. Its procedural component distills and stores a skill only after it passes synthesis, simulation, or formal checks, rather than relying on model self\-assessments\. A Bayesian component maintains hierarchical Beta estimates over tool\-call outcomes and ranks recovery strategies that succeeded under comparable errors\. A common adapter applies the same memory interface to RTL optimization and testbench\-generation agents while preserving each domain’s tools and acceptance criteria\. We measure performance on training tasks and evaluate whether learned skills transfer to unseen tasks\. On RTLRewriter\-Bench, under matched model and tool settings, ChipMEM produces equivalence\-passing outputs on 39/54 scored designs versus 35/54 without memory; on the 49\-design short suite, mean area improvement is 8\.69% versus 5\.66%\. On held\-out CVDP tasks, ChipMEM with a frozen procedural library achieves 20/20 accepted outcomes versus 18/20 without memory in a single evaluation per setting\.
Figure 1:ChipMEM’s seed data come from EDA\-agent sessions evaluated by domain tools\. The verification gate stores a reusable procedural skill only after a passing final artifact, while statistical memory records every completed tool\-call outcome to support retry and recovery decisions in later actions and sessions\. Figure[2](https://arxiv.org/html/2609.27067#S1.F2)shows the full system architecture, including skill retrieval and use\.## 1INTRODUCTION
Large language models now support several stages of chip design and verification\. They can generate register transfer level \(RTL\) code, revise hardware descriptions, and use feedback from design tools to correct their outputs\([Liu et al\., 2023](https://arxiv.org/html/2609.27067#bib.bib1);[Lu et al\., 2024](https://arxiv.org/html/2609.27067#bib.bib2);[Pinckney et al\., 2025](https://arxiv.org/html/2609.27067#bib.bib14);[Yu et al\., 2026](https://arxiv.org/html/2609.27067#bib.bib23)\)\. Recent systems also equip agents with reusable skills distilled from prior tool guided sessions\([Du and Pinckney, 2026](https://arxiv.org/html/2609.27067#bib.bib4);[Fang et al\., 2026](https://arxiv.org/html/2609.27067#bib.bib6)\)\. This approach reduces the need to encode every lesson in a prompt\. Continual learning then depends on whether the agent can discover reliable skills from verified work and reuse them on later tasks\. This problem matters in chip development because early RTL decisions shape later design quality\. Choices about bit width, resource sharing, and state encoding affect power, performance, and area, yet many useful changes require engineers to restructure the RTL source instead of relying on synthesis alone\([Lu et al\., 2026](https://arxiv.org/html/2609.27067#bib.bib7);[Yao et al\., 2024](https://arxiv.org/html/2609.27067#bib.bib12)\)\. The revised RTL must then pass simulation and equivalence checks before it can move forward\. These iterations are slow, and the lessons learned often remain with individual engineers instead of becoming reusable knowledge\. Automation is therefore essential\.
These challenges have motivated a rapidly growing body of research on large language models for hardware design\. Such models can reason over hardware descriptions, propose candidate rewrites in seconds, and iteratively refine them under feedback from EDA tools\. Existing efforts have explored both parametric approaches — adapting model weights through supervised fine\-tuning or reinforcement learning\([Chen et al\., 2026](https://arxiv.org/html/2609.27067#bib.bib3);[Shi et al\., 2025](https://arxiv.org/html/2609.27067#bib.bib22);[Zhou et al\., 2026](https://arxiv.org/html/2609.27067#bib.bib16)\)— and nonparametric approaches that equip frozen models with curated skills, search strategies, or tool\-in\-the\-loop feedback\([Arnold et al\., 2026](https://arxiv.org/html/2609.27067#bib.bib5);[Ping et al\., 2026](https://arxiv.org/html/2609.27067#bib.bib17)\)\.
Prior work suggests that agents can self\-improve by retaining skills, reflections, and strategies from earlier trajectories\([Wang et al\., 2023](https://arxiv.org/html/2609.27067#bib.bib8);[Shinn et al\., 2023](https://arxiv.org/html/2609.27067#bib.bib9);[Ouyang et al\., 2026](https://arxiv.org/html/2609.27067#bib.bib24)\)\. Recent EDA systems have begun incorporating similar forms of tool\-grounded experience\. However, these methods are typically evaluated on the same task or task family that produced the memory, leaving unresolved whether stored skills remain relevantacrosstasks andunseendesigns\. Such memory is useful if the agent is faced with the same task again, but the agent must once again start from scratch on brand new tasks\. Moreover, naively exposing the agent to memories cultivated from dissimilar tasks has limited utility and may even be counterproductive; the optimal skills for one design may actively diminish performance on a different design\. This gap motivates a memory layer whose contents are certified by EDA tools and whose value is measured through transfer beyond the tasks that produced them\.
Figure 2:ChipMEM workflow\. \(a\) Procedural memory retrieves skills by task similarity, supplies them to the EDA agent, and stores a distilled skill only after a verifiedpass\. \(b\) Statistical memory updates from each completed tool call, estimates retry success, and invokes a recovery advisor and bounded LLM plan when the estimate falls belowτ=0\.20\\tau=0\.20\.To address this gap, we propose ChipMEM111Code:[https://github\.com/aalrabah/ChipMEM\.git](https://github.com/aalrabah/ChipMEM.git), a verification\-grounded memory layer for cross\-task transfer that combinesprocedural memoryandstatistical memory\. Theprocedural memorycaptures complete agent trajectories and distills them into reusable skills to be retrieved for subsequent tasks\. These skills are collected and updated according to outcomes from verifiable evaluators for specific tasks, e\.g\. EDA tools for synthesis and verification\. A successful strategy may create a new skill or refine an existing one, while a verified failure may be retained as guidance against repeating an ineffective action\. Moreover, the skills are indexed by an embedding derived from the design RTL\. Given a new task, similarities with respect to these embeddings are used to retrieve the skills most likely to be related to and beneficial for the current task; the central hypothesis being that similar designs will likely have transferable skills\. Thestatistical memorycomplements these trajectory\-derived lessons with a Bayesian model for continual learning through current and historical successes and failures, estimating the probability that a tool strategy will succeed in the current execution context and advising the agent when an unchanged retry is unlikely to be productive\. Instead of solely relying on the agent’s intrinsic reasoning to identify the right tools, we equip it with a “data analyst” that produces advice grounded in statistics collected from prior rollouts\. With statistical memory, the agent is thus able to begin a session with an informed prior over success likelihoods of tools, and can update this prior with each tool execution\.
ChipMEM combines the procedural and statistical memories into a single memory layer through a common agent adapter that records the task representation, retrieved context, execution trajectory, returned artifacts, and domain\-specific acceptance criterion, allowing both memory components to support EDA agents without requiring those agents to share the same tools or scoring procedure\. This design allows procedural memory to evolve across a sequence of tasks or remain frozen for evaluation on unseen designs, while statistical memory summarizes execution outcomes and provides probabilistic guidance during tool use\.
We summarize our primary contributions as follows:
1. 1\.ChipMEM combines procedural memory that extracts reusable skills from tool\-scored executions with Bayesian statistical memory that accumulates tool\-level successes and failures to guide retries and recovery\.
2. 2\.ChipMEM provides an agent\-agnostic adapter between an EDA agent, its tools, and its evaluation harness\. The adapter supplies retrieved skills and recovery guidance, records execution trajectories, and returns artifacts for domain\-specific evaluation without modifying the underlying agent or tools\.
3. 3\.We evaluate ChipMEM against memory\-off baselines across multiple EDA tasks and benchmarks, examining verified outcomes, interaction cost, and the effects of each memory component\.
4. 4\.We separately evaluate transfer to unseen tasks using a fixed, read\-only procedural library, testing whether prior experience remains useful without adding new skills during evaluation\.
We evaluate ChipMEM on the RTL\-OPT\([Lu et al\., 2026](https://arxiv.org/html/2609.27067#bib.bib7)\)and RTLRewriter\([Yao et al\., 2024](https://arxiv.org/html/2609.27067#bib.bib12)\)benchmarks for PPA optimization, and on CVDP tasks for testbench generation\([Pinckney et al\., 2025](https://arxiv.org/html/2609.27067#bib.bib14)\)\. The evaluation tests whether skills distilled from verified trajectories transfer to later tasks and unseen designs, and whether accumulated tool outcomes improve decisions during execution\.
## 2RELATED WORK
##### Skill memory and design\-aware retrieval\.
There is a growing body of work that equips frozen agents with procedural memory, storing experience from past episodes as reusable artifacts and retrieving it to guide later tasks\. Voyager\([Wang et al\., 2023](https://arxiv.org/html/2609.27067#bib.bib8)\)builds an executable skill library from successful episodes, Reflexion\([Shinn et al\., 2023](https://arxiv.org/html/2609.27067#bib.bib9)\)turns execution feedback into verbal episodic memory, ReasoningBank\([Ouyang et al\., 2025](https://arxiv.org/html/2609.27067#bib.bib10)\)distills transferable strategies from both successful and failed trajectories, and SkillOpt\([Yang et al\., 2026](https://arxiv.org/html/2609.27067#bib.bib11)\)refines a skill document over a frozen model without weight updates\. In chip design, however, the relevance of a stored strategy depends on the target specification, RTL structure, tool context, and failure mode, so a skill that helps one design may be ineffective or harmful on another\. This makes retrieval as important as skill extraction and motivates evaluating memory the task that produced it\. ChipMEM extends this direction by integrating trajectory\-distilled procedural memory for cross\-task transfer with Bayesian statistical memory for within\-session guidance, while grounding both in outcomes produced by EDA tools\.
##### Tool\-grounded RTL optimization\.
RTL\-OPT\([Lu et al\., 2026](https://arxiv.org/html/2609.27067#bib.bib7)\)and RTLRewriter\([Yao et al\., 2024](https://arxiv.org/html/2609.27067#bib.bib12)\)evaluate LLM\-driven RTL optimization under both implementation\-quality and functional\-equivalence criteria\. Recent methods then fold EDA feedback into optimization in different ways\. Reinforcement learning drives a gated toolchain reward into model weights\([Chen et al\., 2026](https://arxiv.org/html/2609.27067#bib.bib3)\)\. Evolutionary and agentic search explore candidate rewrites under synthesis feedback\([Ping et al\., 2026](https://arxiv.org/html/2609.27067#bib.bib17);[Arnold et al\., 2026](https://arxiv.org/html/2609.27067#bib.bib5);[Hsin et al\., 2026](https://arxiv.org/html/2609.27067#bib.bib13)\)\. Frozen agents instead evolve reusable skills from execution traces\([Du and Pinckney, 2026](https://arxiv.org/html/2609.27067#bib.bib4);[Fang et al\., 2026](https://arxiv.org/html/2609.27067#bib.bib6);[Wang et al\., 2026](https://arxiv.org/html/2609.27067#bib.bib15)\), and test\-time training adapts weights to a single design and discards them once it changes\([Zhou et al\., 2026](https://arxiv.org/html/2609.27067#bib.bib16)\)\. These methods improve how an agent searches or adapts during optimization, whereas ChipMEM turns verified outcomes into reusable procedural and statistical memory that transfers to later tasks and different EDA agents\.
##### LLMs for RTL design and evaluation\.
VerilogEval\([Liu et al\., 2023](https://arxiv.org/html/2609.27067#bib.bib1)\), RTLLM\([Lu et al\., 2024](https://arxiv.org/html/2609.27067#bib.bib2)\), and CVDP\([Pinckney et al\., 2025](https://arxiv.org/html/2609.27067#bib.bib14)\)provide established benchmarks for RTL generation, completion, and verification\. Building on these efforts, ChipMEM extends the evaluation beyond isolated task success to examine whether tool\-verified experience remains useful across later tasks and unseen designs, and whether historical outcomes can improve decisions during tool\-guided execution\.
## 3ChipMEM
ChipMEM is an agent\-agnostic layer that converts tool\-verified execution traces into two complementary memories:procedural memorypreserves reusable strategies across tasks, whilestatistical memorylearns action\-level success and recovery patterns within and across sessions\. Both components augment the agent context without changing the underlying EDA agent, tools, or domain\-specific acceptance gate\. Figure[2](https://arxiv.org/html/2609.27067#S1.F2)summarizes this dual\-memory workflow\.
##### Procedural memory\.
We represent the procedural memory library as key\-value mapℳproc=\{ki↦\(vi,mi,hi\)\}i=1M\\mathcal\{M\}^\{\\mathrm\{proc\}\}=\\\{k\_\{i\}\\mapsto\(v\_\{i\},m\_\{i\},h\_\{i\}\)\\\}\_\{i=1\}^\{M\}\. For a given task, the value tuple\(vi,mi,hi\)\(v\_\{i\},m\_\{i\},h\_\{i\}\)consists of the the complete skill fileviv\_\{i\}, metadatamim\_\{i\}for its domain, task, provenance, and verifier outcomeshih\_\{i\}\. For theii\-th task, the keyki=ϕ\(zi\)∈ℝdk\_\{i\}=\\phi\(z\_\{i\}\)\\in\\mathbb\{R\}^\{d\}is the task embedding obtained by applying an embedding modelϕ\\phito a pre\-defined task\-artifactziz\_\{i\}\. For instance, in our experiments,ziz\_\{i\}corresponds to the top RTL file for RTL optimization and the task\-definitionPROMPT\.mdfile for CVDP tasks\. This task definition excludes agent instructions, retrieved skills, tool history, generated files, tests, and scoring artifacts\.
The library starts empty and grows acrossNseedN\_\{\\mathrm\{seed\}\}tasks run forRRrounds\. For thenn\-th task, we retrieve procedural\-memory as value\-tuples with keys exhibiting the largest similarity to the current task’s embeddingqnq\_\{n\}, i\.e\.,
ℛn=\{\(vi,mi,hi\)\|i∈TopK\[sim\(qn,kj\);τproc\]\}\.\\mathcal\{R\}\_\{n\}=\\left\\\{\(v\_\{i\},m\_\{i\},h\_\{i\}\)\\;\\middle\|\\;i\\in\\operatorname\{TopK\}\\left\[\\mathrm\{sim\}\(q\_\{n\},k\_\{j\}\);~\\tau\_\{\\mathrm\{proc\}\}\\right\]\\right\\\}\.Here,TopK\[sim\(⋅\);τproc\]\\operatorname\{TopK\}\[\\texttt\{sim\}\(\\cdot\);\\tau\_\{\\mathrm\{proc\}\}\]is the list of top\-k matches with respect to a similarity functionsimand a thresholdτproc\\tau\_\{\\mathrm\{proc\}\}; only matches with similarities higher thanτproc\\tau\_\{\\mathrm\{proc\}\}are returned\. We use cosine\-similarity over embeddings as our similarity functionsim\(qn,kj\)\\texttt\{sim\}\(q\_\{n\},k\_\{j\}\)\. Thusℛt\\mathcal\{R\}\_\{t\}contains at mostKKcomplete memory items with the highest task similarity above a thresholdτproc\\tau\_\{\\mathrm\{proc\}\}\.
Once thenn\-th task is processed, ChipMEM may expand the procedural memory library by adding a skill for the source task if none exist\. Skills are added to the library only if the corresponding task execution obtains a verifiedpassfrom a verifiable domain\-specific evaluation harness, not an LLM judge\. We can apply the procedural memory library under two modes: evolving and frozen\. While the evolving\-mode carries verified additions forward to later tasks, the frozen\-mode keeps its initial library read\-only throughout\.
##### Statistical memory\.
While procedural memory allows EDA agents to extract high\-level strategies to aid their tasks, agents can also benefit from the more structured and narrower\-scoped skill of tool\-failure\-recovery\. EDA agents often repeat failed tool calls without reusing recoveries that have succeeded in earlier trajectories\. Since EDA\-tools can have particularly high latency, repeated tool\-call failures can lead to significant delays and wasted token\-consumption\. Thus, we propose statistical memory that maintains a prior belief over tool\-call success based on statistics collected from prior runs, updates the prior based on the current trajectory, and provides advice on tool\-retry\-success and failure\-recovery strategies by learning on a tool\-step granularity\. The statistical memory component is comprised of two binomial distribution models – the retry predictor and the recovery predictor – encoding the success likelihood of tool\-calls and their corresponding failure\-recovery actions\. As illustrated in Figure[2](https://arxiv.org/html/2609.27067#S1.F2), the retry predictor informs the agent whether its proposed tool\-call is likely to succeed, and the recovery predictor provides the most\-successful strategy to recover from that failure\.
At thett\-th turn in the execution trajectory of an agent, letctc\_\{t\}be tool\-call submitted by the agent; further assume that this tool has previously been calledrtr\_\{t\}times in the same session with the last call resulting in errorete\_\{t\}\. If the previous call forctc\_\{t\}succeeded,ete\_\{t\}is set toNone; otherwise it is set to the specific error class encountered\. For example, in our experiments with the PPA agent, such error classes includesyntax\-error,unknown\-module,liberty\-missing,timing\-fail, andrc\-nonzero\. This provides us with statistics on the likelihood of a tool’s success\. Following a tool\-call error, we also record the actionata\_\{t\}taken by the agent, and whether theata\_\{t\}resulted in a successful recovery\. For each error classete\_\{t\}, we thus collect a table of associated recovery actions and their corresponding likelihoods of success\.
We define the context of the retry\-model via the tuplext=\(ct,et,bt\)x\_\{t\}=\(c\_\{t\},e\_\{t\},b\_\{t\}\), wherebt=min\(rt,3\)b\_\{t\}=\\min\(r\_\{t\},3\)\. We setbt=min\(rt,3\)b\_\{t\}=\\min\(r\_\{t\},3\), so its values are00,11,22, and33, with33grouping three or more earlier calls\. This saturation prevents sparse evidence from being split across many large retry counts\. Starting from the global Beta\(1,1\)\(1,1\)posteriorp0p\_\{0\}, the model adds one conditioning variable at each of three levels\. Levelp1p\_\{1\}conditions on command type,p2p\_\{2\}adds11\-step history of the previous error class, andp3p\_\{3\}adds prior history:p0→p1\(ct\)→p2\(ct,et\)→p3\(ct,et,bt\)p\_\{0\}\\rightarrow p\_\{1\}\(c\_\{t\}\)\\rightarrow p\_\{2\}\(c\_\{t\},e\_\{t\}\)\\rightarrow p\_\{3\}\(c\_\{t\},e\_\{t\},b\_\{t\}\), with
pj=sj\+Kpj−1sj\+fj\+K,j∈\{1,2,3\},p\_\{j\}=\\frac\{s\_\{j\}\+Kp\_\{j\-1\}\}\{s\_\{j\}\+f\_\{j\}\+K\},\\hskip 20\.00003ptj\\in\\\{1,2,3\\\},\(1\)wheresjs\_\{j\}andfjf\_\{j\}are prior successful and failed calls at leveljj, and the parameterKKessentially controls the weight placed on prior history; a lower value ofKKmakes the model more myopic\. In our experiments, we useK=8K=8\. The final estimate isPretry\(success∣ct,et,bt\)=p3P\_\{\\mathrm\{retry\}\}\(\\mathrm\{success\}\\mid c\_\{t\},e\_\{t\},b\_\{t\}\)=p\_\{3\}\. After a failed same\-type call, a value belowτstat=0\.20\\tau\_\{\\text\{stat\}\}=0\.20activates the recovery model\.
For a candidate recovery strategyaaand the current tool stagegtg\_\{t\}\. The recovery hierarchy is
q0→q1\(a\)→q2\(a,et\)→q3\(a,et,gt\)\.q\_\{0\}\\rightarrow q\_\{1\}\(a\)\\rightarrow q\_\{2\}\\left\(a,e\_\{t\}\\right\)\\rightarrow q\_\{3\}\\left\(a,e\_\{t\},g\_\{t\}\\right\)\.For example,q1\(inspect\_logs\)q\_\{1\}\(\\texttt\{inspect\\\_logs\}\)estimates how often inspecting logs has recovered a failure across all errors and stages\. At each level,
qj=uj\+Kqj−1uj\+vj\+K,j∈\{1,2,3\},q\_\{j\}=\\frac\{u\_\{j\}\+Kq\_\{j\-1\}\}\{u\_\{j\}\+v\_\{j\}\+K\},\\hskip 20\.00003ptj\\in\\\{1,2,3\\\},\(2\)whereuju\_\{j\}andvjv\_\{j\}are historical successful and failed recoveries,qj−1q\_\{j\-1\}is the probability inherited from the broader level, andK=8K=8\. HencePrecovery\(a∣et,gt\)=q3P\_\{\\mathrm\{recovery\}\}\(a\\mid e\_\{t\},g\_\{t\}\)=q\_\{3\}, and candidate strategies are ranked as
π\(et,gt\)=sortaPrecovery\(a∣et,gt\)\.\\pi\(e\_\{t\},g\_\{t\}\)=\\operatorname\{sort\}\_\{a\}\\,P\_\{\\mathrm\{recovery\}\}\(a\\mid e\_\{t\},g\_\{t\}\)\.\(3\)When activated, the recovery model passes its ranked strategies to an LLM recovery planner, which returns a bounded advisory plan to the agent\. The ranking is advisory; each completed tool call atomically persists retry and recovery evidence for the next decision, including the next decision in the same session\.
##### Algorithm\.
Algorithm[1](https://arxiv.org/html/2609.27067#alg1)summarizes ChipMEM across ordered sessions\. HereDnD\_\{n\}is the task artifact,𝒜\\mathcal\{A\}is the EDA agent,𝒯\\mathcal\{T\}is its tool set,𝒱\\mathcal\{V\}is the deterministic domain evaluation harness,qnq\_\{n\}is the task embedding,ℛn\\mathcal\{R\}\_\{n\}is the retrieved skill set,HnH\_\{n\}is the execution trajectory, andyny\_\{n\}is the final harness verdict\.statistical memoryupdates after every tool call\. Only a passing trajectory can add a skill toprocedural memory\. The algorithm shows evolving mode\. Frozen mode starts from a fixed library and skips the skill storage step\.
Algorithm 1procedural memoryandstatistical memoryacross sessions1:Tasks
\{Dn\}n=1N\\\{D\_\{n\}\\\}\_\{n=1\}^\{N\}, agent
𝒜\\mathcal\{A\}, EDA tools
𝒯\\mathcal\{T\}, domain evaluation harness
𝒱\\mathcal\{V\}, embedder
ϕ\\phi
2:
ℳproc←∅\\mathcal\{M\}^\{\\mathrm\{proc\}\}\\leftarrow\\emptyset; initialize
ℳstat\\mathcal\{M\}^\{\\mathrm\{stat\}\}with
Beta\(1,1\)\\mathrm\{Beta\}\(1,1\)priors⊳\\trianglerightprocedural library; Bayesian state
3:for
n=1,…,Nn=1,\\ldots,Ndo⊳\\trianglerightnn: session;NN: total sessions
4:
qn←EmbedTask\(ϕ,Dn\)q\_\{n\}\\leftarrow\\textsc\{EmbedTask\}\(\\phi,D\_\{n\}\)⊳\\trianglerightqnq\_\{n\}: task embedding
5:
ℛn←RetrieveSkills\(ℳproc,qn,τproc\)\\mathcal\{R\}\_\{n\}\\leftarrow\\textsc\{RetrieveSkills\}\(\\mathcal\{M\}^\{\\mathrm\{proc\}\},q\_\{n\},\\tau\_\{\\mathrm\{proc\}\}\)⊳\\trianglerightℛn\\mathcal\{R\}\_\{n\}: topKKskill payloads;τproc\\tau\_\{\\mathrm\{proc\}\}: threshold
6:Give
\(Dn,ℛn\)\(D\_\{n\},\\mathcal\{R\}\_\{n\}\)to
𝒜\\mathcal\{A\}; initialize
Hn←∅H\_\{n\}\\leftarrow\\emptyset⊳\\trianglerightHnH\_\{n\}: trajectory
7:whilesession
nnis activedo
8:
𝒜\\mathcal\{A\}selects an action and calls a tool in
𝒯\\mathcal\{T\}⊳\\trianglerightAct
9:Append the action and returned outcome to
HnH\_\{n\}
10:
ℳstat←BayesUpdate\(ℳstat,latest tool outcome\)\\mathcal\{M\}^\{\\mathrm\{stat\}\}\\leftarrow\\textsc\{BayesUpdate\}\(\\mathcal\{M\}^\{\\mathrm\{stat\}\},\\text\{latest tool outcome\}\)⊳\\trianglerightxt=\(et,ct,bt\)x\_\{t\}=\(e\_\{t\},c\_\{t\},b\_\{t\}\); retry model
11:ifa repeated failure has
Pretry<τBP\_\{\\mathrm\{retry\}\}<\\tau\_\{\\mathrm\{B\}\}then⊳\\trianglerightretry probability; nudge threshold
12:Advise
𝒜\\mathcal\{A\}using
π\(et,gt\)\\pi\(e\_\{t\},g\_\{t\}\)⊳\\trianglerightrecovery model;ete\_\{t\}: error;gtg\_\{t\}: stage
13:endif
14:endwhile
15:
yn←RunDomainHarness\(𝒱,Dn,Hn\)y\_\{n\}\\leftarrow\\textsc\{RunDomainHarness\}\(\\mathcal\{V\},D\_\{n\},H\_\{n\}\)⊳\\trianglerightyny\_\{n\}: final harness verdict
16:if
yn=passy\_\{n\}=\\textsc\{pass\}then
17:
vn←DistillProcedures\(Hn\)v\_\{n\}\\leftarrow\\textsc\{DistillProcedures\}\(H\_\{n\}\)⊳\\trianglerightprocedural distiller;vnv\_\{n\}: skill
18:
ℳproc←StoreSkill\(ℳproc,qn,vn\)\\mathcal\{M\}^\{\\mathrm\{proc\}\}\\leftarrow\\textsc\{StoreSkill\}\(\\mathcal\{M\}^\{\\mathrm\{proc\}\},q\_\{n\},v\_\{n\}\)⊳\\trianglerightqnq\_\{n\}: index;vnv\_\{n\}: skill
19:else
20:Leave
ℳproc\\mathcal\{M\}^\{\\mathrm\{proc\}\}unchanged
21:endif
22:endfor
##### Agent adapter and verification\.
ChipMEM connects to each EDA domain through a common adapter\. For taskDnD\_\{n\}, the adapter gives𝒜\\mathcal\{A\}the original task, retrieved skillsℛn\\mathcal\{R\}\_\{n\}, and any Bayesian recovery advice\. It exposes the domain tools, records each tool action and result inHnH\_\{n\}, and returns the final artifact to𝒱\\mathcal\{V\}\. Here𝒱\\mathcal\{V\}is a deterministic domain evaluation harness, not an LLM judge\. PPA acceptance requires successful synthesis, functional equivalence, and a positive audited improvement\. CVDP testbench generation requires the hidden simulation, coverage, and mutation checks to pass\. In evolving mode, a verifiedpassmay add a procedural skill\. Frozen mode leavesprocedural memoryunchanged\. Afailorinvalidverdict creates no skill, whilestatistical memorystill records completed tool outcomes\.
## 4Experimental Setup
##### Tasks
The CVDP study uses 40 testbench tasks, with 20 training tasks and 20 held\-out tasks split evenly between CID012 stimulus generation and CID013 checker generation\([Pinckney et al\., 2025](https://arxiv.org/html/2609.27067#bib.bib14)\)\. RTL\-OPT contains 38 design\-level RTL optimization tasks\([Lu et al\., 2026](https://arxiv.org/html/2609.27067#bib.bib7)\), while RTLRewriter\-Bench contains 59 available cases, 54 short and five long\([Yao et al\., 2024](https://arxiv.org/html/2609.27067#bib.bib12)\)\. OpenTitan evaluates ten IP blocks over five attempts per mode\([Meza et al\., 2023](https://arxiv.org/html/2609.27067#bib.bib18)\)\. Additionally, we also performed evaluations on a custom set of2020open\-sourced designs \(see Appendix for details\)\.
##### Implementation details\.
We use GPT\-5\.5\([OpenAI, 2026](https://arxiv.org/html/2609.27067#bib.bib19)\)for RTL\-OPT, RTLRewriter\-Bench, the custom\-design PPA study, and CVDP CID012 and CID013\. We use Qwen3\.8\-27B\([Qwen Team, 2026](https://arxiv.org/html/2609.27067#bib.bib20)\)for OpenTitan\.Procedural memoryretrieval usestext\-embedding\-3\-largefor the first group andQwen/Qwen3\-Embedding\-0\.6Bfor OpenTitan\. Reasoning is set to medium for all experiments\. OpenTitan uses at most 40 tool steps, a 600 second model timeout, and a 7200 second session limit\. Retrieval returns at most two skills at similarity threshold0\.60\.6\. The open model and its embedding model run on an NVIDIA A100 80GB GPU\([NVIDIA Corporation, 2020](https://arxiv.org/html/2609.27067#bib.bib21)\)\. Appendix[A\.5](https://arxiv.org/html/2609.27067#A1.SS5)reports the detailed resource measurements\.
##### Evaluation metrics\.
Equivalence pass rate and strict gate pass rate are binary measures of functional correctness and complete task acceptance\. For PPA, area, power, and timing gains are measured relative to the original design after synthesis and equivalence\. Failed, invalid, unproven, unchanged, and nonimproving outputs receive zero gain\. RTLRewriter additionally reports geometric mean cell and wire ratios relative to the baseline, where values below one indicate reduction\. For CID012 and CID013, task pass rate is the fraction of generated testbenches that pass every hidden Xcelium/IMC test\. For ordered passes, cumulative Pass@kkis the fraction of tasks solved by passkkand is reported as a sequential curve rather than an independent sampling estimate\.
An EDA agent works on one design over many turns\. It reads and edits files, calls design tools, inspects failures, and revises its output before the final evaluation\. The PPA agent can call synthesis, simulation, timing, and equivalence tools\. The simulation agent can compile and run testbenches\. The CVDP simulation set contains 20 training tasks and 20 unseen tasks\. CID012 stimulus generation and CID013 checker generation each contribute 10 training tasks and 10 unseen tasks\. The unseen tasks are disjoint from training\. The agent sees the task files and design context, while hidden tests and reference artifacts remain outside its workspace\.
Figure 3:Audited PPA outcomes across the custom\-design PPA suite, OpenTitan, RTL\-OPT, and RTLRewriter\-Bench\. Vertical bar graphs compare baseline EDA\-Agent with no memory \(purple\) against ChipMEM \(red\)\. Horizontal bars show the per\-task direction of change under ChipMEM as improved, unchanged, or hurt\. RTLRewriter means use 54 scored cases after excluding five zero\-cell cases\. For, RTL\-OPT we select the best valid result per design across three result sets and reports the strongest demonstrated capability\.
##### PPA metric scope\.
The PPA studies use two evaluation scopes\. We began with RTL\-OPT and RTLRewriter\-Bench as first\-pass area\-optimization studies using Yosys for synthesis and equivalence and ABC for technology mapping\([Wolf and Glaser, 2013](https://arxiv.org/html/2609.27067#bib.bib25);[Brayton and Mishchenko, 2010](https://arxiv.org/html/2609.27067#bib.bib26)\)\. This fast, reproducible loop let us verify that the agent workflow, equivalence gate, and measurement methodology operated correctly before moving to the broader Custom Design Set and OpenTitan flows\. The first\-pass runs measure mapped area, and RTLRewriter\-Bench also records cell and wire counts\. Because we did not run a matched characterized downstream power or timing flow for these two benchmarks, we do not report unmeasured power or performance values\. The later Custom Design Set and OpenTitan studies report full PPA metrics, including area, power, and timing\. Figure[3](https://arxiv.org/html/2609.27067#S4.F3)presents the first\-pass area\-optimization and full\-PPA results together under their respective measurement scopes\.
##### Baselines and ablations\.
Figure[3](https://arxiv.org/html/2609.27067#S4.F3)and Appendix Table[2](https://arxiv.org/html/2609.27067#S5.T2)compareEDA agent222The EDA Agent is Cadence ChipStack 2\.0, a proprietary system actively used in production by leading chip\-design companies; all GPT\-5\.5 experiments were run on this agent\.withEDA agent \+ ChipMEM\.EDA agentis the same domain agent with memory disabled\. It receives no retrieved skills, writes no skills, and receives no Bayesian guidance\.EDA agent \+ ChipMEMuses the same model, task, prompt, tools, and budget with memory enabled\. In evolving mode,EDA agent \+ ChipMEMupdates the skill library after verified passing tasks\. In frozen mode,EDA agent \+ ChipMEMreads a fixed library during unseen task evaluation\. These settings provide the whole\-system comparison between Memory OFF and full ChipMEM\. Separately, on five RTL\-OPT designs, we compare Memory OFF, procedural\-only, and Bayesian\-only under matched settings, while the procedural\-plus\-Bayesian loaded\-state replay reports combined system capability\.
## 5Results and Discussion
##### Cross benchmark interpretation\.
Figure[3](https://arxiv.org/html/2609.27067#S4.F3)shows that, on the two public benchmarks, ChipMEM increases mean area improvement from 6\.18% to 8\.79% on RTL\-OPT and from 5\.95% to 8\.88% on RTLRewriter\-Bench\. It also raises RTLRewriter\-Bench cell and wire reductions from 5\.9%/6\.0% to 8\.3%/8\.0%\. The per\-design results include 12 area improvements on RTL\-OPT and 10 on RTLRewriter\-Bench, showing that the gains extend across distinct public RTL optimization tasks\. In the full\-PPA studies, the Custom Design Set increases mean area and power improvement from 0\.14%/0\.39% to 0\.31%/0\.52%\. OpenTitan increases mean area and power improvement from 0\.05%/0\.00% to 0\.24%/2\.54%, while timing remains stable\. Table[3](https://arxiv.org/html/2609.27067#A1.T3)further shows that OpenTitan full\-PPA gate passes rise from 1/10 to 4/10 with equivalence maintained at 6/10 in both modes\. Together, these results show that ChipMEM improves both public area\-optimization benchmarks and the broader full\-PPA design flows\.
Figure 4:OpenTitan interaction cost over 50 sessions per mode, comprising ten designs and five attempts per design\. Bars and labels report mean±\\pmstandard error for LLM calls, tool calls, input and total tokens, and tool time\.
##### OpenTitan interaction cost\.
Figure[4](https://arxiv.org/html/2609.27067#S5.F4)reports lower means for ChipMEM on numerous resource measures\. Mean LLM calls fall from 29\.4 to 26\.1, mean tool calls from 27\.4 to 23\.3, and mean tool time from 420\.2 to 369\.8 seconds, corresponding to reductions of 11\.2%, 15\.0%, and 12\.0%\. Input tokens decrease by 6\.8%, and total tokens decrease by 3\.4%\. Across the 50 sessions, ChipMEM has lower average LLM calls, tool calls, input and total tokens, and tool time\.
##### Memory ablation\.
We evaluate four memory configurations on the same RTL\-OPT designs \(n=5n=5\) with matched model, task order, RTL inputs, tools, seed, and execution limits\. Table[1](https://arxiv.org/html/2609.27067#S5.T1)reports strict PASS counts, audited area gain, and interaction cost\. Procedural and Bayesian each produce four strict passes, compared with three for Memory OFF\. The combined state produces five strict passes, the highest mean audited gain, and the lowest tokens per accepted design\. It records the capability of the full system after both memories accumulated prior experience on the same design sequence\.
Table 1:Five\-design RTL\-OPT memory study \(n=5n=5\)\. Tokens report aggregate input/output tokens\. Tool/LLM reports tool calls and successful agent LLM turns\. Time is summed session wall time\. Best values are bold\.
##### Held\-out PPA skill retrieval\.
Figure[5](https://arxiv.org/html/2609.27067#S5.F5)shows representative summaries of skills retrieved during one\-shot PPA evaluations of 36 held\-out designs with evolving procedural memory\. The 50 skill–design retrievals comprised 25RTL Optimization, 21Flow & Validation, and 4Design Setupretrievals, indicating that the agent most often received concrete RTL lessons and procedures for obtaining trustworthy synthesis results\. Cohort A produced 38 skill–design retrievals, compared with 12 in cohort B \(Appendix[A\.2\.2](https://arxiv.org/html/2609.27067#A1.SS2.SSS2)\)\. Beyond its larger size, cohort A contains more FIFOs, arbiters, and peripheral controllers similar to designs represented in the skill library, whereas cohort B contains more heterogeneous processor, memory, and high\-speed interconnect blocks\.
Flow & ValidationGeneral flow knowledge for establishing a valid synthesis baseline and trustworthy PPA evidence\.skill\-name:sv\-cdc\-fifo\-master skill\-summary: Use the intended FIFO top and full hierarchy with every clock and valid CDC constraint before interpreting PPA\.Design SetupDesign\-specific knowledge needed to compile and elaborate a particular IP correctly\.skill\-name:sram\-ctrl skill\-summary: Preserve SRAM controller package and source order, include paths, defines, and assertion macros; local shims can leave hierarchy incomplete\.RTL OptimizationRTL transformation knowledge, including what may change and which behaviors must be preserved\.skill\-name:priority\-encoder skill\-summary: Enumerated case or threshold rewrites increased mapped area or power\. Retain the compact priority or comparator form favored by synthesis\.
Figure 5:Procedural skill types and representative stored skill summaries\.Flow & Validationcaptures general synthesis and evidence\-quality procedures,Design Setuppreserves design\-specific compilation knowledge, andRTL Optimizationrecords concrete transformations and their guardrails; dashed boxes show one stored example for each type\.Table 2:Held\-out transfer on ten unseen tasks per category, evaluated once per setting with frozen, read\-only memory\.
##### Held\-out CVDP transfer\.
Table[2](https://arxiv.org/html/2609.27067#S5.T2)evaluates frozen memory on unseen tasks\. CID012 preserves 10/10 performance in both settings\. CID013 increases from 8/10 to 10/10 and raises the overall result from 18/20 to 20/20\. Because the library is read\-only, the two additional passes reflect transfer from previously learned procedural memory\. The result shows that stored experience improves checker generation while preserving perfect stimulus generation\.
##### What the memory learns\.
The held\-out retrievals show that ChipMEM preserves cross\-design RTL knowledge, not only instructions for operating EDA tools\. This distinction matters: completing the synthesis flow establishes that a result can be evaluated, whereas improving PPA requires transferable design reasoning\. Future work should isolate the contribution of individual retrieved skills to flow completion and PPA improvement through paired, skill\-level ablations\.
## 6Limitations
ChipMEM uses a fixed top two retrieval setting at threshold0\.60\.6\. The component and frozen transfer studies cover five RTL\-OPT designs and ten unseen CVDP tasks per category\. Future work should test broader retrieval settings, additional EDA tasks, and controlled runtime environments\.
## 7Conclusion
ChipMEM converts tool verified EDA experience into reusable procedural and statistical memory\. It improves accepted outcomes without changing the underlying agent\. In the OpenTitan evaluation, it used fewer LLM and tool calls, fewer input and total tokens, and less tool time on average\. The SWE\-bench Pro\([Deng and others, 2025](https://arxiv.org/html/2609.27067#bib.bib27)\)pilot in Appendix[A\.7](https://arxiv.org/html/2609.27067#A1.SS7)extends the evaluation to software engineering and motivates testing ChipMEM across additional domains\. Future work should test larger memories and broader cross\-domain transfer\.
## References
- F\. Arnold, R\. Amaudruz, D\. Tsaras, R\. Andri, and L\. CavigelliRTLScout: joint agentic code and synthesis optimization for efficient digital circuits\.arXiv preprint arXiv:2606\.06530\.Cited by:[§1](https://arxiv.org/html/2609.27067#S1.p2.1),[§2](https://arxiv.org/html/2609.27067#S2.SS0.SSS0.Px2.p1.1)\.
- Brayton and Mishchenko \(2010\)R\. Brayton and A\. MishchenkoABC: an academic industrial\-strength verification tool\.InComputer Aided Verification,pp\.24–40\.External Links:[Document](https://dx.doi.org/10.1007/978-3-642-14295-6%5F5)Cited by:[§4](https://arxiv.org/html/2609.27067#S4.SS0.SSS0.Px4.p1.1)\.
- Chenet al\.\(2026\)Z\. Chen, K\. Chang, Z\. Li, C\. Li, X\. He, C\. Chen, M\. Wang, H\. Xu, Y\. Han, H\. Li,et al\.ChipSeek: optimizing verilog generation via eda\-integrated reinforcement learning\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\.25180–25201\.Cited by:[§1](https://arxiv.org/html/2609.27067#S1.p2.1),[§2](https://arxiv.org/html/2609.27067#S2.SS0.SSS0.Px2.p1.1)\.
- Denget al\.\(2025\)X\. Denget al\.SWE\-Bench Pro: can AI agents solve long\-horizon software engineering tasks?\.arXiv preprint arXiv:2509\.16941\.Cited by:[§7](https://arxiv.org/html/2609.27067#S7.p1.1)\.
- Du and Pinckney \(2026\)Z\. Du and N\. PinckneyTrace2Skill: verifier\-guided skill evolution for long\-context eda agents\.arXiv preprint arXiv:2605\.21810\.Cited by:[§1](https://arxiv.org/html/2609.27067#S1.p1.1),[§2](https://arxiv.org/html/2609.27067#S2.SS0.SSS0.Px2.p1.1)\.
- Fanget al\.\(2026\)W\. Fang, Y\. Lu, S\. Liu, J\. Wang, Z\. Guo, J\. He, F\. Tu, and Z\. XieDr\. rtl: autonomous agentic rtl optimization through tool\-grounded self\-improvement\.arXiv preprint arXiv:2604\.14989\.Cited by:[§1](https://arxiv.org/html/2609.27067#S1.p1.1),[§2](https://arxiv.org/html/2609.27067#S2.SS0.SSS0.Px2.p1.1)\.
- Hsinet al\.\(2026\)W\. Hsin, R\. Deng, Y\. Hsieh, E\. Huang, and S\. HungEvolVE: evolutionary search for llm\-based verilog generation and optimization\.arXiv preprint arXiv:2601\.18067\.Cited by:[§2](https://arxiv.org/html/2609.27067#S2.SS0.SSS0.Px2.p1.1)\.
- Liuet al\.\(2023\)M\. Liu, N\. Pinckney, B\. Khailany, and H\. RenVerilogeval: evaluating large language models for verilog code generation\.In2023 IEEE/ACM International Conference on Computer Aided Design \(ICCAD\),pp\.1–8\.Cited by:[§1](https://arxiv.org/html/2609.27067#S1.p1.1),[§2](https://arxiv.org/html/2609.27067#S2.SS0.SSS0.Px3.p1.1)\.
- Luet al\.\(2024\)Y\. Lu, S\. Liu, Q\. Zhang, and Z\. XieRtllm: an open\-source benchmark for design rtl generation with large language model\.In2024 29th Asia and South Pacific Design Automation Conference \(ASP\-DAC\),pp\.722–727\.Cited by:[§1](https://arxiv.org/html/2609.27067#S1.p1.1),[§2](https://arxiv.org/html/2609.27067#S2.SS0.SSS0.Px3.p1.1)\.
- Luet al\.\(2026\)Y\. Lu, S\. Liu, H\. Zhou, W\. Fang, Q\. Zhang, and Z\. XieA new benchmark for the appropriate evaluation of rtl code optimization\.arXiv preprint arXiv:2601\.01765\.Cited by:[Table 3](https://arxiv.org/html/2609.27067#A1.T3),[§1](https://arxiv.org/html/2609.27067#S1.p1.1),[§1](https://arxiv.org/html/2609.27067#S1.p7.1),[§2](https://arxiv.org/html/2609.27067#S2.SS0.SSS0.Px2.p1.1),[§4](https://arxiv.org/html/2609.27067#S4.SS0.SSS0.Px1.p1.1)\.
- Mezaet al\.\(2023\)A\. Meza, F\. Restuccia, J\. Oberg, D\. Rizzo, and R\. KastnerSecurity verification of the opentitan hardware root of trust\.IEEE Security & Privacy21\(3\),pp\.27–36\.External Links:[Document](https://dx.doi.org/10.1109/MSEC.2023.3251954)Cited by:[§4](https://arxiv.org/html/2609.27067#S4.SS0.SSS0.Px1.p1.1)\.
- NVIDIA Corporation \(2020\)NVIDIA CorporationNVIDIA a100 tensor core gpu\.Note:[https://www\.nvidia\.com/en\-us/data\-center/a100/](https://www.nvidia.com/en-us/data-center/a100/)Cited by:[§4](https://arxiv.org/html/2609.27067#S4.SS0.SSS0.Px2.p1.1)\.
- OpenAI \(2026\)OpenAIGPT\-5\.5 model\.Note:[https://developers\.openai\.com/api/docs/models/gpt\-5\.5](https://developers.openai.com/api/docs/models/gpt-5.5)Cited by:[§4](https://arxiv.org/html/2609.27067#S4.SS0.SSS0.Px2.p1.1)\.
- Ouyanget al\.\(2026\)S\. Ouyang, J\. Yan, I\. Hsu, Y\. Chen, K\. Jiang, Z\. Wang, R\. Han, L\. Le, S\. Daruki, X\. Tang,et al\.Reasoningbank: scaling agent self\-evolving with reasoning memory\.InInternational Conference on Learning Representations,Vol\.2026,pp\.94327–94354\.Cited by:[§1](https://arxiv.org/html/2609.27067#S1.p3.1)\.
- Ouyanget al\.\(2025\)S\. Ouyang, J\. Yan, I\. Hsu, Y\. Chen, K\. Jiang, Z\. Wang, R\. Han, L\. T\. Le, S\. Daruki, X\. Tang,et al\.Reasoningbank: scaling agent self\-evolving with reasoning memory\.arXiv preprint arXiv:2509\.25140\.Cited by:[§2](https://arxiv.org/html/2609.27067#S2.SS0.SSS0.Px1.p1.1)\.
- Pinckneyet al\.\(2025\)N\. Pinckney, C\. Deng, C\. Ho, Y\. Tsai, M\. Liu, W\. Zhou, B\. Khailany, and H\. RenComprehensive verilog design problems: a next\-generation benchmark dataset for evaluating large language models and agents on rtl design and verification\.arXiv preprint arXiv:2506\.14074\.Cited by:[§1](https://arxiv.org/html/2609.27067#S1.p1.1),[§1](https://arxiv.org/html/2609.27067#S1.p7.1),[§2](https://arxiv.org/html/2609.27067#S2.SS0.SSS0.Px3.p1.1),[§4](https://arxiv.org/html/2609.27067#S4.SS0.SSS0.Px1.p1.1)\.
- Pinget al\.\(2026\)H\. Ping, P\. Zhang, Z\. Wang, S\. Li, A\. Cheng, W\. Yang, P\. Bogdan, and S\. NazarianPOET: power\-oriented evolutionary tuning for llm\-based rtl ppa optimization\.arXiv preprint arXiv:2603\.19333\.Cited by:[§1](https://arxiv.org/html/2609.27067#S1.p2.1),[§2](https://arxiv.org/html/2609.27067#S2.SS0.SSS0.Px2.p1.1)\.
- Qwen Team \(2026\)Qwen TeamQwen3\.8\-27b\.Note:[https://huggingface\.co/Qwen/Qwen3\.8\-27B](https://huggingface.co/Qwen/Qwen3.8-27B)Cited by:[§4](https://arxiv.org/html/2609.27067#S4.SS0.SSS0.Px2.p1.1)\.
- Shiet al\.\(2025\)J\. Shi, Z\. Gao, C\. Ko, and D\. BoningEarl: entropy\-aware rl alignment of llms for reliable rtl code generation\.arXiv preprint arXiv:2511\.12033\.Cited by:[§1](https://arxiv.org/html/2609.27067#S1.p2.1)\.
- Shinnet al\.\(2023\)N\. Shinn, F\. Cassano, A\. Gopinath, K\. Narasimhan, and S\. YaoReflexion: language agents with verbal reinforcement learning\.Advances in neural information processing systems36,pp\.8634–8652\.Cited by:[§1](https://arxiv.org/html/2609.27067#S1.p3.1),[§2](https://arxiv.org/html/2609.27067#S2.SS0.SSS0.Px1.p1.1)\.
- Wanget al\.\(2023\)G\. Wang, Y\. Xie, Y\. Jiang, A\. Mandlekar, C\. Xiao, Y\. Zhu, L\. Fan, and A\. AnandkumarVoyager: an open\-ended embodied agent with large language models\.arXiv preprint arXiv:2305\.16291\.Cited by:[§1](https://arxiv.org/html/2609.27067#S1.p3.1),[§2](https://arxiv.org/html/2609.27067#S2.SS0.SSS0.Px1.p1.1)\.
- Wanget al\.\(2026\)Y\. Wang, Q\. Shi, S\. Li, Q\. Hu, X\. Yin, B\. Guo, X\. Han, M\. Sun, and J\. SuVeriAgent: a tool\-integrated multi\-agent system with evolving memory for ppa\-aware rtl code generation\.arXiv preprint arXiv:2603\.17613\.Cited by:[§2](https://arxiv.org/html/2609.27067#S2.SS0.SSS0.Px2.p1.1)\.
- Wolf and Glaser \(2013\)C\. Wolf and J\. GlaserYosys – a free verilog synthesis suite\.InProceedings of the 21st Austrian Workshop on Microelectronics \(Austrochip\),External Links:[Link](https://yosyshq.net/yosys/files/yosys-austrochip2013.pdf)Cited by:[§4](https://arxiv.org/html/2609.27067#S4.SS0.SSS0.Px4.p1.1)\.
- Yanget al\.\(2026\)Y\. Yang, Z\. Gong, W\. Huang, Q\. Yang, Z\. Zhou, Z\. Huang, Y\. Li, X\. Gao, Q\. Dai, B\. Liu,et al\.Skillopt: executive strategy for self\-evolving agent skills\.arXiv preprint arXiv:2605\.23904\.Cited by:[§2](https://arxiv.org/html/2609.27067#S2.SS0.SSS0.Px1.p1.1)\.
- Yaoet al\.\(2024\)X\. Yao, Y\. Wang, X\. Li, Y\. Lian, R\. Chen, L\. Chen, M\. Yuan, H\. Xu, and B\. YuRtlrewriter: methodologies for large models aided rtl code optimization\.InProceedings of the 43rd IEEE/ACM International Conference on Computer\-Aided Design,pp\.1–7\.Cited by:[Table 3](https://arxiv.org/html/2609.27067#A1.T3),[§1](https://arxiv.org/html/2609.27067#S1.p1.1),[§1](https://arxiv.org/html/2609.27067#S1.p7.1),[§2](https://arxiv.org/html/2609.27067#S2.SS0.SSS0.Px2.p1.1),[§4](https://arxiv.org/html/2609.27067#S4.SS0.SSS0.Px1.p1.1)\.
- Yuet al\.\(2026\)C\. Yu, C\. Deng, N\. Pinckney, and B\. KhailanyAgentic hardware design as repository\-level code evolution\.arXiv preprint arXiv:2606\.28279\.Cited by:[§1](https://arxiv.org/html/2609.27067#S1.p1.1)\.
- Zhouet al\.\(2026\)P\. Zhou, Z\. Chen, C\. Li, H\. Gao, K\. Chang, Z\. Qu, and Y\. WangAlpha\-rtl: test\-time training for rtl hardware optimization\.arXiv preprint arXiv:2606\.05253\.Cited by:[§1](https://arxiv.org/html/2609.27067#S1.p2.1),[§2](https://arxiv.org/html/2609.27067#S2.SS0.SSS0.Px2.p1.1)\.
ChipMEM: Verification\-Grounded Memory for EDA Agents
Supplementary Material
Appendix
Table of Contents
A\.1 Per\-Design PPA Results\.[A\.1](https://arxiv.org/html/2609.27067#A1.SS1)
A\.2 Held\-Out A and B: Per\-Design PPA Results\.[A\.2](https://arxiv.org/html/2609.27067#A1.SS2)
A\.3 OpenTitan Ten\-Design PPA Audit\.[A\.3](https://arxiv.org/html/2609.27067#A1.SS3)
A\.4 Memory ON vs OFF: Cost and Consistency\.[A\.4](https://arxiv.org/html/2609.27067#A1.SS4)
A\.5 Execution and Resource Metrics\.[A\.5](https://arxiv.org/html/2609.27067#A1.SS5)
A\.6 Experiment Grid\.[A\.6](https://arxiv.org/html/2609.27067#A1.SS6)
A\.7 SWE\-bench Pro Pilot\.[A\.7](https://arxiv.org/html/2609.27067#A1.SS7)
A\.8 Prompts\.[A\.8](https://arxiv.org/html/2609.27067#A1.SS8)
## Appendix AAppendix
Table 3:Consolidated audited results on RTL\-OPT\[[Lu et al\., 2026](https://arxiv.org/html/2609.27067#bib.bib7)\], RTLRewriter\-Bench\[[Yao et al\., 2024](https://arxiv.org/html/2609.27067#bib.bib12)\], the custom\-design PPA suite, and OpenTitan\. Equiv\. passed denotes equivalence\-passing outputs; for RTLRewriter it also includes unchanged outputs accepted as trivially equivalent, while unproven outputs are excluded\. RTL\-OPT area wins and golden comparisons are reported among equivalence\-passing outputs; its mean area gain averages all 38 designs with zero for invalid or non\-improving outputs\. RTLRewriter reports area wins over scored cases and geometric\-mean cell/wire ratios\. custom design suite gains retain only positive equivalence\-accepted improvements\. OpenTitan reports best\-of\-five strict\-gate outcomes; accepted gains come from each design’s highest\-scoring full\-gate attempt, and WNS/TNS changes are in ns\. RTL\-OPT ChipMEM selects the best valid result per design across three result sets\. Best values are bold\.
Table 4:Five\-pass CVDP training aggregates with 50 executions per setting\.### A\.1Per\-Design PPA Results
Table 5:Per\-design five\-attempt PPA results for GPT\-5\.5 underEDA agentandEDA agent \+ ChipMEM\. Pass@kkis cumulative full\-synthesis completion \(synth\_status=1\.0\) within attempts 1–kk\. Area, power, and timing are positive design\-level gains admitted by the documented equivalence check; all other outcomes contribute zero\. Equivalence is reported as a binary PASS/FAIL\.Per\-design five\-attempt PPA evaluation \(n=20n=20\)DesignModelPass@1Pass@2Pass@3Pass@4Pass@5Area GainPower GainTiming GainSynthesisEquiv\. pass02\_fifo\_syncEDA agent001110\.0000%0\.0000%0\.0000%1/5FAILEDA agent \+ ChipMEM000000\.0000%0\.0000%0\.0000%0/5FAIL03\_rv32i\_aluEDA agent111110\.0000%0\.0000%0\.0000%2/5FAILEDA agent \+ ChipMEM111110\.0000%0\.0000%0\.0000%1/5FAIL09\_apb\_gpioEDA agent011110\.0567%0\.3937%0\.0000%1/5PASSEDA agent \+ ChipMEM011110\.0000%0\.3937%0\.0000%4/5PASS11\_arbiterEDA agent011110\.0000%1\.4529%9\.5161%2/5PASSEDA agent \+ ChipMEM001110\.0000%0\.0000%0\.0000%3/5FAIL14\_axi\_adapter\_rdEDA agent000010\.0000%0\.0000%0\.0000%1/5PASSEDA agent \+ ChipMEM011110\.0000%0\.0000%0\.0000%2/5PASS18\_uart8receiverEDA agent111110\.0000%0\.0000%0\.0000%2/5FAILEDA agent \+ ChipMEM000110\.0000%0\.0000%0\.0000%2/5FAILahb3lite\_apb\_bridgeEDA agent000110\.0000%0\.0000%0\.0000%2/5FAILEDA agent \+ ChipMEM000012\.9366%4\.4379%0\.0000%1/5PASSapb\_spiEDA agent011110\.0000%0\.0000%0\.0000%2/5PASSEDA agent \+ ChipMEM011110\.0000%0\.0000%0\.0000%2/5FAILaxi\_lite\_interfaceEDA agent111110\.0000%0\.0000%0\.0000%2/5FAILEDA agent \+ ChipMEM111110\.0000%0\.0000%0\.0084%3/5PASSbsg\_ready\_to\_credit\_flow\_converterEDA agent001112\.6174%5\.4759%2\.2307%2/5PASSEDA agent \+ ChipMEM000002\.6174%5\.4759%2\.2307%0/5PASSbsg\_serial\_in\_parallel\_out\_dynamicEDA agent001110\.0000%0\.0000%0\.0000%3/5PASSEDA agent \+ ChipMEM001110\.0027%0\.0943%0\.0000%3/5PASScdc\_fifo\_master\_hierEDA agent111110\.0000%0\.0000%0\.0000%4/5FAILEDA agent \+ ChipMEM000110\.0000%0\.0000%0\.0000%2/5FAILgpioEDA agent111110\.0000%0\.0000%0\.0000%1/5PASSEDA agent \+ ChipMEM000010\.0000%0\.0000%0\.0000%1/5FAILhmacEDA agent000000\.0000%0\.0000%0\.0000%0/5FAILEDA agent \+ ChipMEM000010\.6454%0\.0194%0\.0000%1/5PASSi2cEDA agent000000\.0000%0\.0000%0\.0000%0/5FAILEDA agent \+ ChipMEM000110\.0000%0\.0000%0\.0000%1/5FAILnand\_flash\_controllerEDA agent011110\.0000%0\.0000%0\.0000%2/5FAILEDA agent \+ ChipMEM000000\.0000%0\.0000%0\.0000%0/5PASSriscv\_fetchEDA agent000110\.0000%0\.0000%0\.0000%1/5FAILEDA agent \+ ChipMEM111110\.0000%0\.0000%0\.0000%4/5PASSsdram\_axiEDA agent000110\.1264%0\.4348%0\.0000%1/5PASSEDA agent \+ ChipMEM111110\.0000%0\.0000%0\.0000%3/5PASSsram\_ctrlEDA agent000000\.0000%0\.0000%0\.0000%0/5FAILEDA agent \+ ChipMEM011110\.0000%0\.0000%0\.0000%3/5FAILsv\_cdc\_fifo\_masterEDA agent111110\.0000%0\.0000%0\.0000%3/5FAILEDA agent \+ ChipMEM000000\.0000%0\.0000%0\.0000%0/5FAIL
### A\.2Held\-Out PPA Cohorts
We evaluate two disjoint cohorts of previously unseen RTL designs: cohort A contains 20 designs and cohort B contains 16 designs\. Each design is evaluated once in cohort order\. The evaluation begins with procedural skills distilled from earlier tool\-evaluated PPA sessions, while library updates remain enabled\. Consequently, a reused skill may originate from the initial library or from an earlier run in the ordered held\-out evaluation\. In the following tables,*Reused*denotes use of existing skill context, whereas*Mined*denotes creation of a skill from the current trajectory\.
#### A\.2\.1Per\-design outcomes
Table[6](https://arxiv.org/html/2609.27067#A1.T6)reports the outcome of each held\-out evaluation\. The per\-design view distinguishes successful synthesis with available PPA measurements from runs that completed synthesis but did not produce sufficient metrics for PPA comparison\.
Table 6:Per\-design outcomes from the one\-shot ChipMEM evaluation of held\-out cohorts A and B using GPT\-5\.5\. Synthesis scores indicate full \(1\.00\) or partial \(0\.33\) completion\. PPA entries are raw relative changes; em dashes indicate unavailable metrics\.
#### A\.2\.2Skill inventory
Skills are grouped into three types\.Flow & Validationcovers tool setup, source ingestion, synthesis, constraints, and report collection\.Design Setupcaptures compilation and synthesizability requirements specific to a design family\.RTL Optimizationrecords concrete RTL transformations and the conditions under which they produced positive, neutral, or negative results\.
The retrieval columns report the number of designs in each cohort for which the agent received a skill, either as the selected skill or as additional context\. A skill is counted at most once per design\. Retrieval frequency therefore measures exposure to a skill, not whether it was applied or caused the resulting PPA outcome\.
Table 7:Skill inventory and held\-out retrieval frequency\.SkillSkill summaryCohort AretrievalsCohort BretrievalsFlow & Validationsv\-cdc\-fifo\-masterRequire the intended CDC FIFO top, complete hierarchy, all clocks, and valid CDC constraints to synthesize before comparing area, power, or slack\.7002\-fifo\-syncFix the source list and elaborate the intended FIFO top before changing RTL\. Use the baseline only after synthesis produces timing, area, and power reports\.4018\-uart8receiverCount the UART receiver run only after the intended top elaborates, maps, and produces timing, area, and power reports\.30ahb3lite\-apb\-bridgeA frontend warning does not necessarily mean the bridge failed\. Check that the AHB\-to\-APB top and its hierarchy elaborate before stopping the run\.20nand\-flash\-controllerA missing\-top error is a source\-ingestion problem, not a PPA result\. Fix the flash\-controller filelist and elaboration before changing RTL\.20streaming\-datapathUse a complete source list and fixed constraints, then write timing, area, and power reports to known files before comparing datapath revisions\.02dma\-engineResolve the DMA top, complete source list, constraints, and report paths before modifying datapath or control RTL\.0103\-rv32i\-aluReading the HDL is not enough\. The intended processor top must elaborate and synthesize before the ALU has usable QoR data\.00apb\-spiIf the APB\-SPI top is missing after the HDL read, check the exact module name, source list, and child\-module coverage before editing RTL\.00memory\-controllerUse the controller’s PPA results only after the top and full hierarchy synthesize under a defined constraint set and the reports contain numeric results\.00serial\-interface\-controllerWhen the serial\-controller top is missing after the HDL read, use the parser diagnostics to fix the source list, search path, or language mode\.00Design Setupbsg\-ready\-to\-credit\-flow\-converterKeep the original source order, include paths, defines, and synthesis guards\. Changing the preprocessor context can hide the converter top or select the wrong RTL branch\.10gpioSupply the assertion macros or synthesis\-only stubs expected by the GPIO sources\. Do not change functional RTL to work around missing verification headers\.10i2cUse the intended assertion headers, defines, include directories, and synthesis macro branch\. Unresolved macros can prevent the I²C top from parsing\.10sram\-ctrlPreserve package and source order, include directories, defines, and assertion macros in the SRAM controller filelist\. Local shims can change preprocessing or leave the hierarchy incomplete\.10fsm\-controllerCode asynchronous reset and synchronous initialization as separate conditions\. Only the true asynchronous reset belongs in the reset branch\.00RTL Optimizationriscv\-fetchDelete only state proven constant or unobservable\. Preserve redirects, stalls, flushes, skid\-buffer behavior, PC updates, reset behavior, and cycle latency\.23cdc\-fifo\-master\-hierSynthesize the complete FIFO hierarchy with no black boxes\. Pointer or register\-enable rewrites increased mapped cost and must preserve the CDC, flag, and memory behavior\.3111\-arbiterRecoding the round\-robin state\-update logic increased mapped cost\. Limit changes to local common\-term or decode cleanup, then compare post\-synthesis QoR\.30axi\-lite\-interfaceExplicit register enables, hold\-mux rewrites, and idle\-data gating were neutral or worse after mapping and can change AXI\-Lite\-visible behavior\.30dma\-interfaceDerive load enables for wide DMA registers from existing transfer or state conditions\. Preserve reset values and interface latency, then measure switching, area, and slack after mapping\.0309\-apb\-gpioVectorizing per\-bit GPIO logic or factoring APB address decode can map worse\. Preserve the register map, read/write behavior, interrupts, and cycle timing\.2014\-axi\-adapter\-rdParameter\-controlled tie\-offs on disabled AXI read paths produced no QoR gain, likely because synthesis had already removed the inactive logic\.20bsg\-serial\-in\-parallel\-out\-dynamicA one\-entry FIFO controller can be recoded only if depth, valid/occupancy behavior, handshake timing, latency, throughput, and parameter behavior stay unchanged\.10hazard\-detectorFactor repeated opcode and register\-decode terms and use direct equality comparisons in combinational hazard logic\. Prove equivalence and compare mapped QoR\.01priority\-encoderReplacing a compact priority or comparator cone with enumerated case or threshold logic increased mapped area or power\. Keep the RTL form the mapper handles better\.01bus\-bridgeForcing combinational bridge signals to zero or adding speculative enables increased logic and hurt slack\. Keep gating only when mapped power, area, and timing improve\.00hmacBuild only the digest or message word selected in the current cycle instead of a full\-width combinational bus\. Recheck timing because the narrower mux cone may have different depth\.00protocol\-bridgeRemove identical\-arm ternaries, redundant self\-assignments, combinational self\-feedback, and arithmetic identities in the bridge control logic\. Resynthesize each change separately\.00sdram\-axiOperand or payload gating on AXI/SDRAM data paths added muxing and could raise area or dynamic power\. Keep it only when mapped QoR improves under the same activity assumptions\.00
### A\.3OpenTitan Ten\-Design PPA Audit
Table 8:Complete per\-design cumulative strict\-gate audit for the OpenTitan experiment\. Attempts 1–5 are complete for all ten designs in both modes\. All runs use Qwen3\.8\-27B and the same frozen harness\. A pass requires the strict equivalence and OpenROAD gate to succeed\.
### A\.4Memory ON vs OFF—Full Cost, Success, and Consistency Data
Figure[4](https://arxiv.org/html/2609.27067#S5.F4)summarizes the interaction\-cost comparison; this section reports the complete task\- and attempt\-level evidence\. The completed legacy run mined a skill after every memory\-on session, including all 45 failed sessions, so these results are descriptive of this run rather than a clean evaluation of the current PASS\-only learning rule\.
#### A\.4\.1—Per\-design success
Table 9:Per\-design success\. EDA Agent \+ ChipMEM produces passing, equivalence\-verified edits on 4/10 designs versus 1/10 for EDA Agent, including three designs \(pattgen,spi\_host, andkeymgr\) on which the EDA Agent fails every attempt\.
#### A\.4\.2—Overall interaction cost
Table 10:Overall interaction cost across 50 sessions per mode\. Output tokens and session wall time favor EDA Agent; the increase is consistent with memory writes, reflection, and longer\-running successful sessions\. Wall time is cumulative session duration, not parallel\-run makespan\.*Medians\.*LLM calls: OFF 40 \(the budget cap\), ON 27\. Tool calls: OFF 36\.5, ON 18\. Total tokens: OFF 373\.6k, ON 379\.2k\.
#### A\.4\.3—Attempt\-level paired comparison
Table 11:Pairwise deltas over 50 design\-and\-repeat matched attempts\. Efficiency deltas are directionally favorable to EDA Agent \+ ChipMEM, but clustered confidence intervals cross zero except for output tokens, which increase with memory\. Primary full\-gate outcome summaries appear in Table[9](https://arxiv.org/html/2609.27067#A1.T9)\.
#### A\.4\.4—Per\-design consistency
Table 12:Direction of the per\-design mean change across the ten OpenTitan designs\.*Note\.*The LLM\-call ties arekeymgrandaes; both reach the 40\-call cap in both modes\.
#### A\.4\.6—Matched case study:aon\_timer
Table 13:On the only design with at least one full\-gate pass in both modes, EDA Agent \+ ChipMEM uses roughly half the token and call cost and one\-third less session wall time\. Values are per\-attempt means\.
### A\.5Execution and Resource Metrics
Table 14:Execution audit for the 100 Qwen3\.8\-27B canonical sessions underlying Table[3](https://arxiv.org/html/2609.27067#A1.T3)and Appendix Table[8](https://arxiv.org/html/2609.27067#A1.T8): ten designs, five attempts, and two modes\. Continuous entries are mean±\\pmsample standard deviation\. Across the 50 sessions per mode,EDA agent/EDA agent \+ ChipMEMused 15\.357M/14\.834M total tokens, 34\.14/39\.22 cumulative session\-hours, 1468/1303 LLM calls, and 1372/1166 tool calls\. All canonical rows and transcripts were read from the completed Brev snapshots\. Baseline/candidate OpenROAD times are reported only where gate metrics were recovered; their availablennis shown\. Four source\-migrated attempt\-1 transcripts do not reconstruct the retained canonical tool\-time field exactly, so the table uses the preserved canonical values\. The OpenTitan canonical schema did not preserve a retry counter; aborted\-session counts were recovered from the exact transcript flags\. Cached\-token and reasoning\-token counts were unavailable and are not reported as zero\.OpenTitan gate coverage and memory\-control activityModeAttemptBaseline ORFS \(min\)Candidate ORFS \(min\)GatennEquiv\. passPass/FailAbortedBayes updatesNudgesRetrieved/MinedEDA agent128\.9±49\.928\.9\\pm 49\.922\.8±33\.922\.8\\pm 33\.98/105/101/96/100\.0±0\.00\.0\\pm 0\.00\.0±0\.00\.0\\pm 0\.00\.0±0\.00\.0\\pm 0\.0/ 0EDA agent \+ ChipMEM138\.0±56\.438\.0\\pm 56\.414\.9±13\.414\.9\\pm 13\.46/103/100/107/1025\.0±15\.125\.0\\pm 15\.11\.6±1\.41\.6\\pm 1\.40\.0±0\.00\.0\\pm 0\.0/ 10EDA agent232\.3±48\.832\.3\\pm 48\.813\.6±12\.913\.6\\pm 12\.98/105/100/106/100\.0±0\.00\.0\\pm 0\.00\.0±0\.00\.0\\pm 0\.00\.0±0\.00\.0\\pm 0\.0/ 0EDA agent \+ ChipMEM232\.3±48\.832\.3\\pm 48\.813\.5±12\.713\.5\\pm 12\.78/103/100/106/1024\.4±14\.924\.4\\pm 14\.92\.7±0\.72\.7\\pm 0\.70\.1±0\.30\.1\\pm 0\.3/ 10EDA agent329\.3±49\.929\.3\\pm 49\.911\.3±13\.111\.3\\pm 13\.18/106/101/97/100\.0±0\.00\.0\\pm 0\.00\.0±0\.00\.0\\pm 0\.00\.0±0\.00\.0\\pm 0\.0/ 0EDA agent \+ ChipMEM329\.0±46\.729\.0\\pm 46\.712\.1±14\.012\.1\\pm 14\.09/106/103/74/1023\.4±14\.423\.4\\pm 14\.42\.7±0\.72\.7\\pm 0\.70\.2±0\.40\.2\\pm 0\.4/ 10EDA agent429\.0±46\.729\.0\\pm 46\.712\.0±14\.212\.0\\pm 14\.29/106/100/106/100\.0±0\.00\.0\\pm 0\.00\.0±0\.00\.0\\pm 0\.00\.0±0\.00\.0\\pm 0\.0/ 0EDA agent \+ ChipMEM432\.3±48\.832\.3\\pm 48\.812\.4±12\.512\.4\\pm 12\.58/104/101/96/1022\.4±13\.422\.4\\pm 13\.42\.8±0\.62\.8\\pm 0\.60\.2±0\.40\.2\\pm 0\.4/ 10EDA agent532\.3±48\.832\.3\\pm 48\.811\.8±13\.411\.8\\pm 13\.48/105/100/107/100\.0±0\.00\.0\\pm 0\.00\.0±0\.00\.0\\pm 0\.00\.0±0\.00\.0\\pm 0\.0/ 0EDA agent \+ ChipMEM532\.3±48\.832\.3\\pm 48\.810\.9±12\.810\.9\\pm 12\.88/103/101/94/1021\.4±13\.121\.4\\pm 13\.12\.5±0\.82\.5\\pm 0\.80\.3±0\.50\.3\\pm 0\.5/ 10
### A\.6Experiment Grid
The main controlled grids contain 200 proprietary PPA executions, 100 OpenTitan executions, and 240 CVDP simulation executions\.
### A\.7SWE\-bench Pro Pilot
Figure 6:SWE\-bench Pro resolved rate in a 100\-task pilot per mode\. Error bars show Wilson 95% confidence intervals\. SWE\-agent \+ ChipMEM resolves 43/100 tasks compared with 32/100 for the SWE\-agent baseline, an absolute difference of 11 percentage points\.Table 15:Aggregate SWE\-bench Pro outcomes for the 100\-task pilot\. The preserved pilot summary does not include matched per\-task outcomes, so we do not report a paired significance test\.
### A\.8Prompts
The prompts are grouped by the model and reported experiment in which they were used\. Task\-specific values and session data were inserted at runtime\.
##### EDA agent with GPT\-5\.5\.
The following distillation prompts produced procedural skills for the reported GPT\-5\.5 EDA\-agent experiments\.
Youdistilloneagentrunintoareusablelearnedskill\.
Useonlythesuppliedsessionandmeasuredreward\.ReturnstrictJSON:
\{
”name”:”generic\-lowercase\-skill\-name”,
”description”:”one\-linedescription”,
”body”:”Markdownwith\#\#DOand\#\#AVOIDsections”,
”metadata”:\{”failure\_modes”:”shortsummary”\}
\}
Keeprulesgenericandevidence\-based\.Afailedorregressedactionbelongs
inAVOID\.SuccessfulmeasuredactionsbelonginDO\.Donotinclude
instance\-specificpathsorsecrets\.
AGENTTYPE:\{\{agent\_type\}\}
INITIALREQUEST:\{\{initial\_request\}\}
REWARD:score=\{\{reward\_score\}\},success=\{\{reward\_success\}\},
components=\{\{reward\_components\}\},failure\_class=\{\{failure\_class\}\}
\{%ifexisting\_skill\_body%\}
EXISTINGSKILLTOUPDATE:
\{\{existing\_skill\_body\}\}
\{%endif%\}
SESSION:
\{\{session\_transcript\}\}
##### EDA agent with Qwen3\.8\-27B on OpenTitan\.
The following agent and distillation prompts were used only for the reported OpenTitan experiment\.
Youareanautonomousengineer\.Solvethetaskusingtheavailabletools\.
Whendone,replywithyourfinalsummaryinsteadofatoolcall\.
OptimizeOpenTitanIP\{\{design\}\}forlowerphysicalareaandpowerwithout
changingbehaviororports\.Youmayeditonlythefilesreturnedby
list\_files\.FirstinspecttheRTLandrunorfs\_synthforabaseline\.Make
conservativechanges,rerunorfs\_synthaftereachchange,andkeeponly
measuredimprovements\.AstrictequivalencecheckandfullORFSflowwill
gradetheresult\.
DistillexactlyonereusableflatPPAskillfromthisscoredOpenTitanPASS
trajectory\.ReturnonlyconciseDOandAVOIDlines\.
The measured session trajectory was supplied to the distiller together with this fixed instruction\. Memory OFF and Memory ON used the same agent prompt, model settings, tools, and execution budget\. Memory ON additionally received the retrieved procedural skills\.Similar Articles
MemGuard: Persisting Verifier Signals for LLM-Agent Memory Governance
MemGuard introduces a system that persists verifier signals as metadata to govern memory in LLM agents, improving reliability and performance across multiple benchmarks like SWE-Bench and WebArena.
PROJECTMEM: A Local-First, Event-Sourced Memory and Judgment Layer for AI Coding Agents
PROJECTMEM is an open-source, local-first memory and judgment layer for AI coding agents that records development events and provides deterministic warnings before repeating failed actions, reducing token waste and improving reproducibility.
@zxlzr: Introducing MemTrace: Making LLM Memory Systems Finally Debuggable Memory is becoming a core component of AI agents. Bu…
MemTrace is a new tool that makes LLM memory systems debuggable by tracing memory operations across multiple turns, addressing the black-box nature of current memory-augmented agents.
Memory-Augmented Reinforcement Learning Agent for CAD Generation
This paper proposes a memory-augmented reinforcement learning framework for CAD generation agents that integrates geometric kernel toolchains, dual-track memory, and dynamic utility retrieval to handle complex CAD models with long operation sequences and geometric constraints, achieving improved success rate and geometric consistency.
PLACEMEM: Toward a Compute-Aware Memory Plane for Lifelong Agents
PlaceMem proposes a compute-aware memory plane for lifelong agents, using versioned memory capsules that unify semantic content and reusable runtime state to enable correction-aware reuse and avoid redundant computation.