GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure
Summary
GoGoTB is an agentic framework for end-to-end RTL verification that achieves specification-grounded coverage closure, reaching high coverage on 8 designs without human intervention.
View Cached Full Text
Cached at: 07/31/26, 04:00 AM
# GoGoTB: Agentic RTL Verification with Specification-Grounded Coverage Closure
Source: [https://arxiv.org/html/2607.26181](https://arxiv.org/html/2607.26181)
Xin Xin1,†, Jincheng Lou2,†, Junhui Li3, Jinglin Yan1, Panda Xiao1, Di Wu1, Haixiao Li1 Weicong Lu1, Weijian Fan2, Xinyu Qu2, Yuxiang Zhao2, Min Yu2, Zhixiong Di3, Yibo Lin2,4,51Tencent2School of Integrated Circuits, Peking University, Beijing 3Southwest Jiaotong University4Institute of Electronic Design Automation, Peking University, Wuxi 5Beijing Advanced Innovation Center for Integrated Circuits \{wukongxin,lyndayan,pandaxiao,woodywdwu,stevehxli,lukewclu\}@tencent\.com \{jinchenglou,wjfan25,xyqu25,yuxiangzhao\}@stu\.pku\.edu\.cn ssfortynine@gmail\.comyum@pku\.edu\.cndizhixiong2@126\.comyibolin@pku\.edu\.cn
\(2026\)
###### Abstract\.
Functional verification dominates integrated circuit \(IC\) front\-end engineering effort, and a single missed bug that escapes to silicon can trigger a costly respin\. Recent large language models \(LLMs\) offer new opportunities to automate this process, yet existing LLM\-based approaches generate each component through independent single\-turn calls with no shared context, leaving interface mismatches undetected and reported coverage disconnected from specification requirements\. To address these challenges, we presentGoGoTB, an agentic framework that achieves end\-to\-end verification closure through three subsystems: an agentic execution control layer, an evolvable knowledge system, and specification\-grounded coverage closure\. The execution control layer separates deterministic enforcement from LLM reasoning at every tool and stage boundary\. The knowledge system dispatches methodology and design\-specific expertise on demand\. The coverage framework anchors every bin to a named specification behavior so that each residual gap has a diagnosable root cause and a targeted remedy\. Tested on 8 register transfer level \(RTL\) designs without any human intervention,GoGoTBachieves 100% environment generation success and averages 98\.4% line, 97\.2% branch, 97\.0% toggle, and 83\.2% functional coverage\. No prior work successfully generates a complete verification environment or achieves meaningful coverage on the same benchmarks\.
Digital IC Verification, Large Language Models, Testbench Generation, Coverage\-Driven Verification, Agentic Workflows
††copyright:none††journalyear:2026††conference:IEEE/ACM International Conference on Computer\-Aided Design; November 8–12, 2026; San Jose, CA, USA††booktitle:Proceedings of the IEEE/ACM International Conference on Computer\-Aided Design \(ICCAD ’26\), November 8–12, 2026, San Jose, CA, USA††ccs:Hardware Electronic design automation††ccs:Hardware Hardware verification††ccs:Computing methodologies Natural language processing††footnotetext:Xin Xin and Jincheng Lou contributed equally to this work\.## 1\.Introduction
Figure 1\.Four\-level verification coverage hierarchy\. Prior work covers L1–L2\.GoGoTBdelivers L3–L4\.Functional verification consumes the majority of front\-end engineering effort in IC development\(IEE,[2017](https://arxiv.org/html/2607.26181#bib.bib2);Foster,[2015](https://arxiv.org/html/2607.26181#bib.bib3)\), and assembling a complete UVM environment demands weeks of specialized effort per design\(Bergeron,[2003](https://arxiv.org/html/2607.26181#bib.bib4);Spear,[2008](https://arxiv.org/html/2607.26181#bib.bib5)\)\. Recent LLMs have motivated a wave of work on automated testbench generation\(Qiu et al\.,[2024](https://arxiv.org/html/2607.26181#bib.bib6),[2025](https://arxiv.org/html/2607.26181#bib.bib7);Ye et al\.,[2025](https://arxiv.org/html/2607.26181#bib.bib8)\)and formal assertion synthesis\(Yan et al\.,[2025](https://arxiv.org/html/2607.26181#bib.bib9);Wu et al\.,[2025](https://arxiv.org/html/2607.26181#bib.bib10);Orenes\-Vera et al\.,[2023](https://arxiv.org/html/2607.26181#bib.bib11)\); see\(Liu et al\.,[2026](https://arxiv.org/html/2607.26181#bib.bib12)\)for a comprehensive survey\.
Verification is not a best\-effort task\. An undetected bug that escapes to silicon costs orders of magnitude more to fix than one caught before tape\-out\(Foster,[2015](https://arxiv.org/html/2607.26181#bib.bib3)\)\. Existing LLM\-based approaches fall short of the quality bar this demands\. Environment generation success rates remain low and reported coverage does not reflect specification requirements\. No existing framework delivers end\-to\-end verification closure without human intervention, with coverage tied to specification behaviors\.
The root cause is structural\. Existing approaches\(Ye et al\.,[2025](https://arxiv.org/html/2607.26181#bib.bib8);Qiu et al\.,[2024](https://arxiv.org/html/2607.26181#bib.bib6),[2025](https://arxiv.org/html/2607.26181#bib.bib7)\)generate each component through an independent single\-turn LLM call with no shared context, so interface mismatches go undetected and cross\-component errors cannot be resolved\. Coverage models are built from RTL signal ranges rather than specification behaviors, so two designs with the same port list produce the same model regardless of their actual behavior\. These two structural gaps give rise to three unresolved challenges\.
Challenge 1: Execution reliability\.Assembling a verification environment needs iterative compilation, simulation, and repair\. An agentic loop provides feedback closure, but agents without guardrails invoke the wrong tools, propagate poor artifacts across stages, and repeat incorrect repairs without progress\(Yao et al\.,[2023](https://arxiv.org/html/2607.26181#bib.bib13);Wang et al\.,[2024](https://arxiv.org/html/2607.26181#bib.bib14);Shinn et al\.,[2023](https://arxiv.org/html/2607.26181#bib.bib15)\)\. Prior work\(Ye et al\.,[2025](https://arxiv.org/html/2607.26181#bib.bib8)\)does not address these failure modes\. LLM reasoning and deterministic enforcement must therefore be architecturally separated, with each failure mode handled by a dedicated layer\.
Challenge 2: On\-demand domain knowledge injection\.No existing LLM\-based verification framework provides a structured domain knowledge base\. LLMs rely on general training knowledge alone, which is insufficient for complex modules\. Injecting the full knowledge base inflates the context window, and injecting irrelevant knowledge actively misleads the model\. Knowledge injection is therefore an on\-demand dispatch problem\. Each task requires different methodology knowledge and each design category requires different domain knowledge, both of which can be dispatched deterministically\.
Challenge 3: Specification\-traceable coverage closure\.Coverage models built from RTL signals measure stimulus breadth rather than specification compliance\. When gaps remain, supplement strategies add stimulus without diagnosing why bins were missed, so the loop stalls without convergence\. Tying every coverage bin to a named specification behavior solves both problems\. Each gap then has a root cause and a targeted remedy, making convergence achievable\. As shown in Fig\.[1](https://arxiv.org/html/2607.26181#S1.F1), prior work\(Qiu et al\.,[2024](https://arxiv.org/html/2607.26181#bib.bib6),[2025](https://arxiv.org/html/2607.26181#bib.bib7)\)reaches at most L1 with code coverage only, while UVM2\(Ye et al\.,[2025](https://arxiv.org/html/2607.26181#bib.bib8)\)advances to L2 with signal\-range functional coverage\.GoGoTBtargets L3–L4 by grounding coverage bins in specification behaviors\.
We presentGoGoTB, an agentic framework that achieves end\-to\-end verification closure with coverage tied to specification behaviors\. Given RTL source, a natural\-language specification, and an optional reference model,GoGoTBautonomously constructs and refines a verification environment until coverage targets are met\. We summarize our key contributions as follows\.
- •We develop a three\-layer execution control architecture that separates deterministic enforcement from LLM reasoning at tool\-call, stage, and failure\-recovery boundaries, achieving 100% environment generation success across 8 designs and 7 backbone LLMs \(Section[3\.2](https://arxiv.org/html/2607.26181#S3.SS2)\)\.
- •We design a two\-tier knowledge system that dispatches methodology modules by deterministic stage\-level mapping and historical environments by similarity gating, injecting domain knowledge on demand while suppressing irrelevant context \(Section[3\.3](https://arxiv.org/html/2607.26181#S3.SS3)\)\.
- •We propose a specification\-grounded coverage closure framework that derives testpoints across seven behavioral dimensions, classifies each residual gap into a seven\-category root\-cause taxonomy, and drives targeted remediation with a bounded stall\-aware loop \(Section[3\.4](https://arxiv.org/html/2607.26181#S3.SS4)\)\.
- •The completeGoGoTBframework and its evaluation benchmark of 8 open\-source RTL designs will be open\-sourced at a later date\.
The rest of this paper is organized as follows\. Section[2](https://arxiv.org/html/2607.26181#S2)introduces preliminaries\. Section[3](https://arxiv.org/html/2607.26181#S3)presents the methodology\. Section[4](https://arxiv.org/html/2607.26181#S4)reports experimental results\. Section[5](https://arxiv.org/html/2607.26181#S5)concludes the paper\.
Figure 2\.GoGoTBis an eight\-stage agentic framework built on three subsystems\. It includes agentic execution control, an evolvable knowledge system, and specification\-grounded coverage closure\.
## 2\.Preliminaries
This section introduces the concepts underlyingGoGoTBand surfaces the observations that motivate each architectural choice in Section[3](https://arxiv.org/html/2607.26181#S3)\.
### 2\.1\.Digital IC Functional Verification
IC design flow and verification\.Functional verification confirms that an RTL implementation satisfies its specification before synthesis sign\-off\. Because an undetected bug at this stage propagates into silicon and costs orders of magnitude more to fix after fabrication\(Foster,[2015](https://arxiv.org/html/2607.26181#bib.bib3)\), verification constitutes the dominant fraction of front\-end engineering effort\(IEE,[2017](https://arxiv.org/html/2607.26181#bib.bib2);Foster,[2015](https://arxiv.org/html/2607.26181#bib.bib3)\)\. A complete UVM environment comprises a driver, monitor, scoreboard, functional coverage model, and reference model\(IEE,[2017](https://arxiv.org/html/2607.26181#bib.bib2)\), all traditionally hand\-authored\(Bergeron,[2003](https://arxiv.org/html/2607.26181#bib.bib4)\)and consuming weeks per design\.
Testpoint decomposition\.A verification campaign begins with*testpoint decomposition*, enumerating every verifiable behavior in the specification\(Wile et al\.,[2005](https://arxiv.org/html/2607.26181#bib.bib16)\)\. A testpoint is a tuple⟨s,r,c,d⟩\\langle s,r,c,d\\ranglewheressis the stimulus,rris the expected response,ccis the check method, andddis the verification dimension\. Testpoint completeness therefore sets a hard ceiling on verification quality\. Any behavior not enumerated at this stage cannot be verified downstream\.
Coverage metrics\.Verification coverage falls into two categories\.*Code coverage*tracks structural exercise through line, branch, and toggle metrics\.*Functional coverage*\(Piziali,[2008](https://arxiv.org/html/2607.26181#bib.bib17)\)measures which specification\-defined scenarios have been exercised using coverpoints and bins\.
Coverage model quality is determined at authoring time, not at measurement time\. A model built from RTL signal ranges can reach 100% bin coverage while leaving critical specification behaviors unverified\. Two designs with the same port list produce identical models despite doing entirely different things\. Coverage must therefore be grounded in the specification before simulation begins\. Once every bin is anchored to a specification behavior, gaps become diagnosable and closure becomes principled\.
Simulation stack\.The dominant commercial stack pairs Synopsys VCS\(Synopsys, Inc\.,[2023](https://arxiv.org/html/2607.26181#bib.bib18)\)or Cadence Xcelium\(Cadence Design Systems, Inc\.,[2023](https://arxiv.org/html/2607.26181#bib.bib19)\)with SystemVerilog/UVM testbenches\. Prior LLM\-based approaches\(Ye et al\.,[2025](https://arxiv.org/html/2607.26181#bib.bib8)\)follow the same path and generate SystemVerilog components directly\. LLMs produce substantially lower first\-pass correct code in SystemVerilog than in Python, making Python\-based cocotb\(cocotb contributors,[2025](https://arxiv.org/html/2607.26181#bib.bib20)\)a better fit for LLM\-driven generation\. Verilator\(The Verilator Project,[2026](https://arxiv.org/html/2607.26181#bib.bib21)\)emits structured, machine\-parseable diagnostics that enable deterministic error classification without LLM involvement\.
### 2\.2\.From Language Models to Agentic Systems
LLMs and their structural limitation\.An LLM autoregressively samples an output sequence given a fixed prompt:
\(1\)P\(y1:m∣x1:n\)=∏t=1mP\(yt∣x1:n,y1:t−1\)\.P\(y\_\{1:m\}\\mid x\_\{1:n\}\)=\\prod\_\{t=1\}^\{m\}P\(y\_\{t\}\\mid x\_\{1:n\},y\_\{1:t\-1\}\)\.Herex1:nx\_\{1:n\}is the input prompt,y1:my\_\{1:m\}is the generated output sequence, and each tokenyty\_\{t\}is sampled conditioned on the prompt and all previously generated tokens\. This is a single stateless mapping with no memory across turns and no mechanism to act on the world\. All prior LLM\-based verification works\(Ye et al\.,[2025](https://arxiv.org/html/2607.26181#bib.bib8);Hu et al\.,[2025](https://arxiv.org/html/2607.26181#bib.bib22);Qiu et al\.,[2024](https://arxiv.org/html/2607.26181#bib.bib6),[2025](https://arxiv.org/html/2607.26181#bib.bib7);Zhang et al\.,[2025](https://arxiv.org/html/2607.26181#bib.bib23);Liu et al\.,[2026](https://arxiv.org/html/2607.26181#bib.bib12)\)operate within this paradigm, which is structurally insufficient for verification environment construction\.
The LLM agent abstraction\.An agent augments an LLM with persistent state, tool access, and a reasoning\-action\-observation loop\(Yao et al\.,[2023](https://arxiv.org/html/2607.26181#bib.bib13);Wang et al\.,[2024](https://arxiv.org/html/2607.26181#bib.bib14)\)\. At each step the agent produces an action conditioned on accumulated context and updates its state from the observed result\. This feedback closure enables the iterative repair and failure\-adaptive recovery that verification construction requires\.
The agent abstraction alone is insufficient\. Unconstrained agentic loops corrupt protected design files, disable checking logic, and loop indefinitely without progress\. These are control failures from the absence of an enforcement layer\. LLM reasoning and deterministic enforcement must therefore be architecturally separated so that reliability depends on the control architecture rather than on any particular model\.
Knowledge injection and the case for a domain knowledge base\.Industrial verification demands two types of expertise\. The first is methodology knowledge covering how to structure testbenches, apply constrained\-random patterns, analyze RTL, and close coverage\. The second is design\-specific knowledge covering protocol timing, FSM sequencing, cryptographic patterns, and simulator constraints for a particular design category\. Neither type is reliably captured in LLM training data and no existing LLM\-based verification framework provides a structured domain knowledge base\.
Injecting the full knowledge base at every stage inflates the context window and dilutes the relevant signal, so injection must be on demand\. These two types of knowledge also require different dispatch mechanisms\. Methodology knowledge is predictable from the pipeline stage definition and can be dispatched deterministically\. Design\-specific knowledge depends on how closely the current design matches past ones\. A low\-similarity reference causes the LLM to apply the wrong template with high confidence, which is worse than injecting nothing\. Design\-specific knowledge must therefore be controlled by similarity gating and suppressed when confidence is insufficient\.
## 3\.Methodology
### 3\.1\.GoGoTBOverview
GoGoTBtakes RTL source, a natural\-language specification, and a reference model as input\. Fig\.[2](https://arxiv.org/html/2607.26181#S1.F2)illustrates the eight\-stage pipeline, which runs in four phases\.*Specification Analysis*covers S1–S2 and parses the specification into a design profile and derives a testpoint inventory across seven behavioral dimensions\.*Environment Construction*covers S3–S5 and builds the coverage model, reference model, and testbench from those testpoints\.*Stochastic Simulation*covers S6–S7 and runs test stimuli under an adaptive seed policy and accumulates code and functional coverage\.*Coverage Closure*covers S8 and diagnoses residual gaps by root cause and drives targeted remediation until closure criteria are met\. Each stage runs as an agentic node\. Deterministic quality checks at each stage boundary block broken outputs before they reach the next stage\.
Reliable closure across this pipeline requires three subsystems, each realizing one design insight from Section[1](https://arxiv.org/html/2607.26181#S1)\.
1. \(1\)An*agentic execution control layer*\(Section[3\.2](https://arxiv.org/html/2607.26181#S3.SS2)\) addresses three independent failure modes through three dedicated layers\. Each layer targets one failure mode: tool\-call boundaries, stage boundaries, and repeated repair attempts\.
2. \(2\)An*evolvable knowledge system*\(Section[3\.3](https://arxiv.org/html/2607.26181#S3.SS3)\) injects the right knowledge at the right stage\. A Skill Library dispatches methodology modules by stage and design category\. A Reference Library selects past environments by structural similarity and suppresses dispatch when no close match exists\.
3. \(3\)*Specification\-grounded coverage closure*\(Section[3\.4](https://arxiv.org/html/2607.26181#S3.SS4)\) anchors every coverage bin to a named specification behavior before simulation begins\. When gaps remain, each uncovered bin has an unambiguous root cause and a targeted remedy drawn from a seven\-category taxonomy\.
The execution control layer wraps all eight stages\. The knowledge system injects domain knowledge at the entry point of each stage\. Coverage closure runs across stages S1–S3 and S6–S8\.
### 3\.2\.Agentic Execution Control
Figure 3\.Three\-layer agentic execution control\. Layer 1 intercepts every tool call with guard, structuring, and validation steps\. Layer 2 enforces quality gates at stage boundaries with deterministic auto\-repair\. Layer 3 escalates repair scope through four levels when earlier fixes fail\. Boundary enforcement applies as a cross\-cutting constraint across all layers\.Reliable loop execution requires that every LLM action passes through a deterministic enforcement boundary\.GoGoTBintroduces three dedicated layers for this \(Fig\.[3](https://arxiv.org/html/2607.26181#S3.F3)\)\. Layer 1 controls every tool call, output, and session boundary so the LLM always reasons from clean, structured information\. Layer 2 checks every stage boundary and blocks broken outputs before they reach the next stage\. Layer 3 classifies each failure pattern and escalates repair scope rather than repeating the same action\.
Layer 1: Tool\-Call Interception\.Without an intermediary between the LLM and its tools, the LLM invokes the wrong tools, reads raw output directly, and decides when a session is done\. All three contact points are sources of unreliable behavior\. Every tool call therefore passes through three steps in sequence \(Fig\.[3](https://arxiv.org/html/2607.26181#S3.F3)\)\. TheGuardstep checks that the LLM is invoking the right tool for the current stage and blocks incorrect calls before execution\. TheStructuringstep converts raw compiler or simulator output into a clean structured diagnostic so the LLM reasons from root causes rather than surface noise\. TheValidationstep verifies at the session boundary that the session produced complete and compilable artifacts before it closes\. If the same command fails three times in a row with no change, a redirect prompt is injected to break the stuck loop\. The structured result is returned to the LLM through the observation loop\. All three steps run independently of LLM decisions\.
Layer 2: Quality Gates\.Poor artifacts from one stage propagate silently to the next, causing failures that are attributed to the wrong stage and making diagnosis impossible\. Layer 2 first runs anAuto\-Repairpass that fixes common structural errors without LLM involvement\. The output then passes through three gates in order of strictness\. TheHard Gatechecks for compilation success and mandatory artifacts and always blocks on failure\. TheSoft Gatechecks for testpoint completeness and signal\-name consistency\. It blocks on failure in early attempts and softens to a warning as attempt count rises\. TheAdvisory Gateperforms a semantic review without blocking progress\. On failure, the output is returned to the LLM through the retry loop and re\-enters the gates on the next attempt\.
Layer 3: Failure Recovery\.When a repair attempt fails, repeating the same action without changing scope makes no progress and exhausts the retry budget\. Layer 3 first passes the failure through a classifier that identifies the error category and injects targeted feedback into the next attempt\. Each time the attempt count crosses a threshold, the repair automatically escalates to the next level along the L1–L4 chain\.L1applies a minimal patch to a single file\.L2reads full cross\-file context before editing\.L3refactors the problematic component entirely\.L4falls back to a minimal\-correct implementation\. Each escalation step expands the repair scope so the loop converges rather than stalls\.
Boundary Enforcement\.Boundary Enforcement applies across all three layers as a cross\-cutting constraint \(Fig\.[3](https://arxiv.org/html/2607.26181#S3.F3)\)\. Each stage can only access the artifacts it needs\. The RTL source and the reference model are immutable to all later stages\. The reference model is derived from the specification alone and has no visibility into RTL implementation details\. This isolation guarantee ensures that any simulation mismatch cannot originate from a corrupted oracle and therefore points to a real design defect\.
### 3\.3\.Evolvable Knowledge System
Figure 4\.Two\-tier evolvable knowledge system\. The Skill Library dispatches methodology modules by a deterministic stage\-level map\. The Reference Library selects historical environments by similarity gating and suppresses dispatch when no close match exists\.Knowledge injection is an on\-demand dispatch problem\. Stage\-level methodology knowledge is predictable from the stage definition and can be dispatched deterministically\. Design\-specific knowledge depends on structural similarity to the current design and must be controlled by similarity gating\. Both dispatch mechanisms require no LLM involvement and cannot silently apply the wrong knowledge\. The two\-tier architecture in Fig\.[4](https://arxiv.org/html/2607.26181#S3.F4)realizes these two mechanisms separately\.
Skill Library\.The Skill Library delivers stage\-level methodology knowledge through deterministic dispatch\. It holds ten skill modules covering testbench construction, test planning, coverage closure, RTL analysis, simulator backend, cryptographic design patterns, serial protocol patterns, and more\. Before each pipeline stage runs, a stage\-level map selects which modules apply \(Fig\.[4](https://arxiv.org/html/2607.26181#S3.F4)\)\. Selection requires no LLM involvement\. To keep context small, only a table of contents arrives at stage start and full module content loads on demand\. Adding a new design category requires only adding a new file to the knowledge directory\.
Reference Library\.The Reference Library delivers design\-specific historical knowledge through similarity gating\. A design\-category trigger activates the library based on the design category detected at specification\-parse time\. Each entry holds the RTL, specification, testbench, and testpoints from a real verified design\. When a new design arrives, the system scores it against the library across three structural levels\. A close match dispatches the reference as a direct template\. A partial match dispatches it as a style guide\. A poor match suppresses dispatch entirely so the LLM receives nothing rather than a misleading reference\. This suppression is the key design choice\. The LLM only receives historical knowledge when that knowledge genuinely applies\.
### 3\.4\.Specification\-Grounded Coverage Closure
Coverage quality is determined at model authoring time, not at measurement time\. Grounding every bin in a specification behavior before simulation makes gaps diagnosable and convergence achievable, because the reason a bin is uncovered must belong to a finite set of root causes\. As listed in Table[1](https://arxiv.org/html/2607.26181#S3.T1), these causes include insufficient stimulus, missing multi\-step sequences, cross\-coverage constraint mismatches, unreached FSM states, timing\-dependent conditions, edge transitions, and inadequate seed sampling\. Each cause demands a different targeted remedy\. The loop terminates when no addressable cause remains\. Without specification grounding, none of these categories is diagnosable and the supplement loop runs without principled convergence\.
Testpoint decomposition \(S1–S2\)\.To anchor every bin to a named specification behavior before simulation begins, S1 parses the RTL and specification into a design profile\. S2 derives testpoints across seven behavioral dimensions\.D1covers data boundaries,D2control flow,D3timing constraints,D4FSM state transitions,D5protocol compliance,D6error injection, andD7microarchitectural interaction\. These seven dimensions cover all verifiable behavior types in a design specification\. The resulting testpoint set establishes the ceiling on verification completeness\.
Coverage model and simulation \(S3, S6–S7\)\.With testpoints in place, S3 converts each one into a coverage model element where every bin carries a named behavioral claim traceable to the specification\. This transforms the coverage model from a signal\-value partition into a set of behavioral assertions\. S6 then generates test cases per testpoint group\. S7 runs them under an adaptive seed policy and stops when consecutive batches yield no gain, marking the boundary of what stochastic simulation can reach on its own\.
Gap diagnosis and closure \(S8\)\.Because every bin is specification\-grounded, each residual gap maps directly to a cause in Table[1](https://arxiv.org/html/2607.26181#S3.T1)\. S8 classifies each uncovered bin and directs it to the targeted remedy\. Structurally unreachable bins are removed from the denominator\. A stall detector terminates the loop when two consecutive rounds show no improvement\. The sign\-off report records evidence for every covered bin and a root\-cause disposition for every residual gap\. The verification outcome is fully auditable with no gap silently dropped\.
Table 1\.Gap root\-cause taxonomy and remediation routing in S8\.
## 4\.Experimental Evaluation
RQ1\(Sec\.[4\.2](https://arxiv.org/html/2607.26181#S4.SS2)\): Does the three\-layer execution control architecture achieve reliable environment generation across diverse designs and LLM backends?
RQ2\(Sec\.[4\.3](https://arxiv.org/html/2607.26181#S4.SS3)\): Does specification\-grounded testpoint decomposition produce a richer coverage model and higher functional coverage than signal\-space generation?
CS1\(Sec\.[4\.4](https://arxiv.org/html/2607.26181#S4.SS4)\): Does domain knowledge injection resolve total functional coverage failure caused by a knowledge gap rather than a model capability gap?
CS2\(Sec\.[4\.5](https://arxiv.org/html/2607.26181#S4.SS5)\): Does specification\-grounded closure converge through root\-cause diagnosis across all seven gap categories?
### 4\.1\.Experimental Setup
Platform\.Intel Xeon Gold 6248R with 48 cores, 512 GB DDR4, and Ubuntu 22\.04 LTS\. Verilator 5\.024\(The Verilator Project,[2026](https://arxiv.org/html/2607.26181#bib.bib21)\)with cocotb 2\.0\.1\(cocotb contributors,[2025](https://arxiv.org/html/2607.26181#bib.bib20)\)\.GoGoTBis implemented on the CodeBuddy Agent SDK\(Tencent Cloud,[2025](https://arxiv.org/html/2607.26181#bib.bib24)\)\. RQ1 evaluates seven backbone LLMs: Kimi\-K2\.5, GLM\-5\.1, Minimax\-M2\.7, Claude Opus 4\.6, Claude Sonnet 4\.6, GPT\-5\.4, and DeepSeek\-V3\.2\. RQ2 and all case studies use Kimi\-K2\.5 as the primary model\.
Benchmark suite\.Table[2](https://arxiv.org/html/2607.26181#S4.T2)lists eight benchmark designs ranging from 291 to 1,424 lines and 1 to 9 modules, spanning combinational datapaths, pipelined cryptographic cores, serial protocols, and stateful memory controllers\.
Table 2\.Benchmark designs used for evaluation \(†= minor Verilator patch\)\.DesignDescriptionSourceModuleLineCountsCountsAES†Encrypts a 128\-bit plaintext with a 128\-bit key using the AES algorithm, outputting 128\-bit ciphertext\.UVM2\(Ye et al\.,[2025](https://arxiv.org/html/2607.26181#bib.bib8)\)8684ALU†A 32\-bit unit performing IEEE 754 FP arithmetic, logical operations \(OR, AND, XOR\), shifts, and FP\-to\-integer conversion\.UVM2\(Ye et al\.,[2025](https://arxiv.org/html/2607.26181#bib.bib8)\)7409UARTFacilitates serial communication between a host and peripherals, supporting configurable baud rates and stop bits\.\(Rudy,[2020](https://arxiv.org/html/2607.26181#bib.bib25)\)4387SPIA lightweight SPI controller supporting Master mode with FIFO\-based data transfer and interrupt\-driven operation\.This work3291I2CAn I2C slave controller handling address detection, data transfer, clock stretching, and ACK/NACK signaling\.\(Forencich,[2017](https://arxiv.org/html/2607.26181#bib.bib26)\)3470SHA256A cryptographic unit that computes a 256\-bit hash value from an input message using the SHA\-256 algorithm\.\(Strömbergson,[2013](https://arxiv.org/html/2607.26181#bib.bib27)\)41,142SM4A synchronous SM4 encryption/decryption core supporting both modes via a 128\-bit key and data interface\.\(gongxunwu,[2020](https://arxiv.org/html/2607.26181#bib.bib28)\)91,424SDRAMA synchronous DRAM controller managing ACTIVATE, READ, WRITE, and PRECHARGE commands with timing enforcement\.\(Terasic Inc\.,[2012](https://arxiv.org/html/2607.26181#bib.bib29)\)1433AES and ALU are taken from the UVM2benchmark suite, the only two designs that UVM2makes publicly available\. The remaining six designs were selected from high\-quality open\-source RTL repositories to cover a broad range of design categories not represented in UVM2\. The complete benchmark package, including specifications, verification environments, and testpoints for all eight designs, will be open\-sourced at a later date\. All experiments run fully automated without human intervention and provide a reference model as input\.
Metrics\.*SRG*measures whole\-environment success requiring zero\-error compilation, deadlock\-free simulation, and non\-zero coverage\. For designdid\_\{i\}under modelℳ\\mathcal\{M\}:
\(2\)SRG\(ℳ\)=1N∑i=1N𝟏\[compile∧simulate∧cov\>0\]\.\\mathrm\{SRG\}\(\\mathcal\{M\}\)=\\frac\{1\}\{N\}\\sum\_\{i=1\}^\{N\}\\mathbf\{1\}\[\\,\\text\{compile\}\\wedge\\text\{simulate\}\\wedge\\text\{cov\}\>0\\,\]\.*Code coverage*reports per\-design line, branch, and toggle ratios under Kimi\-K2\.5\.*Functional coverage*reports the fraction of specification\-derived bins hit after stochastic simulation\.*Coverage model richness*reports coverpoint and cross\-coverage counts per design, compared against UVM2\.
### 4\.2\.RQ1: Environment Generation Reliability
Figure 5\.Environment generation success rate and build time across designs and backbone LLMs\.Fig\.[5](https://arxiv.org/html/2607.26181#S4.F5)reports SRG and wall\-clock build time for both systems\.GoGoTBachieves100% SRGon 8 of 8 designs across all seven backbone LLMs\. Build times range from 5m30s to 48m10s\. The result holds across model families, parameter scales, and training corpora\. The enforcement layers drive convergence, not the backbone LLM\.
SRG measures whole\-environment success\. A high per\-component generation rate does not imply that components assemble into a runnable environment, because components generated in isolation diverge in interface assumptions and the assembled environment fails even when each component passes its own syntax check\.
UVM2achieves0% SRGon all eight designs\. We cloned the UVM2\(Ye et al\.,[2025](https://arxiv.org/html/2607.26181#bib.bib8)\)repository and ran each design four to six times\. UVM2generates each component through a single LLM call with no shared context and no generate\-debug\-iterate loop\. Compilation failures and simulation deadlocks accumulate without any recovery mechanism\. These failures are architectural and follow directly from the absence of enforcement and iteration, not from the capability of any specific LLM\.
The 100% vs\. 0% contrast across seven LLMs confirms that the three\-layer enforcement architecture drives loop convergence regardless of model choice\.
### 4\.3\.RQ2: Coverage Achievement and Model Quality
Table[3](https://arxiv.org/html/2607.26181#S4.T3)presents results under Kimi\-K2\.5\. UVM2is the only prior system that attempts functional coverage, making it the only valid baseline\.GoGoTBderives every coverpoint from a named specification behavior\. Each bin tests whether a specific design intent was exercised\. UVM2partitions transaction signal ranges into bins with no connection to specification meaning\.
Table 3\.RQ2: cross\-system coverage results under Kimi\-K2\.5 \(×\\times= compile/sim failed; — = not available;†UVM2SRG 0/8; UVM2subscriber open\-sourced only for AES and ALU\)\.Cat\.: CC = Comb\. crypto; CS = Comb\./seq\.; SP = Serial proto\.; SC = Sequential crypto; MC = Memory ctrl\. — = UVM2subscriber not open\-sourced for this design\.
GoGoTBaverages98\.4%line,97\.2%branch,97\.0%toggle, and83\.2%functional coverage\. Five designs reach≥99%\{\\geq\}99\\%line coverage and all exceed 91% branch coverage\.
UVM2achieves 0% SRG, so no functional coverage numbers are available\. We compare coverage model quality on AES and ALU, the only two designs UVM2open\-sources\. UVM2produces 3 coverpoints and 1 cross on AES\.GoGoTBproduces 6 coverpoints and 2 crosses\. EachGoGoTBcoverpoint maps to a named specification behavior such as key shape, NIST anchor hit, and plaintext change\. Each UVM2bin is a numeric interval on a raw signal with no behavioral claim\. For the remaining six designs UVM2provides no coverage model at all\.
Code coverage exceeds 97% while functional coverage averages 83\.2%\. This gap reflects the inherent difficulty of specification\-level verification\. Reaching a code line is easier than hitting the precise multi\-signal condition that a specification\-derived bin requires\. Closing this gap requires advances in LLM\-driven directed test generation\.
### 4\.4\.CS1: Knowledge\-Driven Environment Construction on UART
Figure 6\.CS1: Three silent failures \(left\) caused by missing domain knowledge and the corresponding fixes \(right\) fromGoGoTBdomain contracts\. Each row addresses one failure: protocol framing \(top\), observer synchronization \(middle\), and reference model separation \(bottom\)\.CS1 examines whether domain knowledge injection resolves coverage failures caused by a knowledge gap rather than a model capability gap\. We select UART because its protocol framing, baud\-rate timing, and direction\-specific TX/RX behavior are domain conventions that no general\-purpose LLM reliably encodes\.
We first runGoGoTBon UART with the knowledge system disabled\. All other subsystems remain active\. Under this configuration, the LLM agent generates a testbench with three silent failures shown in Fig\.[6](https://arxiv.org/html/2607.26181#S4.F6)\. The stimulus drives raw signal transitions without protocol framing, so the DUT never sees a valid UART frame\. The observer samples asynchronously instead of synchronizing to the START\_BIT boundary, so it produces bit\-shifted data that never matches the expected output\. A single reference model serves both TX and RX paths and generates wrong expectations even when the DUT responds correctly\. These three failures form a closed chain from stimulus to oracle\. The environment compiles and simulates without error, yet the coverage report shows all functional bins at0%\.
GoGoTBinjects three domain contracts before the LLM session to eliminate these failures\. The protocol\-framing contract enforces correct UART frame construction with start bit, parity, and stop bits\. The synchronization contract aligns the observer to the mid\-bit sampling point at0\.5×BAUD\_PERIOD0\.5\\times\\text\{BAUD\\\_PERIOD\}\. The independent\-models contract instructs the LLM agent to model TX and RX behavior separately so that it captures direction\-specific expectations\. Functional coverage rises from 0% to75\.3%\. The only variable changed is the knowledge system, confirming that the failure was a knowledge gap, not a model capability gap\. Functional coverage does not reach 100% because the LLM agent still struggles to generate directed tests for protocol\-specific corner cases such as multi\-frame error recovery and baud\-rate edge conditions\.
### 4\.5\.CS2: Coverage\-Driven Closure on I2C
CS2 examines whether specification\-grounded closure remains principled and convergent across all seven root\-cause categories\. We select I2C because its stateful protocol behavior exercises all seven categories simultaneously\.
After the initial test suite completes regression, functional coverage reaches56\.0%across 132 specification\-derived bins\. The system then classifies every uncovered bin by root cause, routes each to its targeted remedy, and generates new tests\. This first supplement round advances coverage to71\.6%\. The loop repeats the same root\-cause analysis and targeted generation over multiple iterations until coverage converges at84\.1%, a total gain of\+28 pp\. Each decision follows from the root\-cause label on the uncovered bin, not from a generic retry strategy\.
After the loop terminates, the system analyzes the 24 residual bins\. Stimulus Gap and Sequence Gap account for 16, Cross Hole and State Unreached account for 5, and the remaining 3 are Timing Dependent, Transition Gap, and Seed Insufficient\. Stimulus and sequence gaps mean that the system identified the correct bin and knew what stimulus was needed, but the LLM agent could not produce a test that hits the required condition\. The coverage model itself is not the bottleneck\. The bottleneck is the ability of the LLM agent to generate precise directed tests for complex multi\-step protocol scenarios\.
Figure 7\.CS2: I2C coverage progression and root\-cause distribution of 24 residual gaps\. Stimulus and sequence gaps dominate, indicating that the system correctly identifies what to test but the LLM agent cannot yet produce the required directed tests\.
## 5\.Conclusion
GoGoTBdemonstrates that fully automated IC functional verification with coverage tied to specification behaviors is achievable when a single system co\-designs an agentic execution control layer, an evolvable knowledge system, and specification\-grounded coverage closure\.
The three\-layer execution control makes environment generation succeed across all 8 designs and 7 backbone LLMs\. The evolvable knowledge system eliminates domain knowledge failures that would otherwise cause silent coverage collapse, as demonstrated by the 0% to 75\.3% functional coverage recovery on UART\. Specification\-grounded coverage closure drives functional coverage to an average of 83\.2% across all designs, with a \+28 pp gain through iterative root\-cause diagnosis on I2C alone\.
The primary limitation is directed test generation for complex multi\-step protocol scenarios\. Stimulus Gap and Sequence Gap account for the majority of residual bins across designs\. Improving LLM\-driven directed test generation and integrating formal methods to cover structurally hard\-to\-reach states are the most impactful directions for future work\.
On the current benchmark,GoGoTBaverages 98\.4% line, 97\.2% branch, 97\.0% toggle, and 83\.2% functional coverage with fully auditable sign\-off reports\.
## References
- \(1\)
- IEE \(2017\)2017\.IEEE Standard for Universal Verification Methodology Language Reference Manual\.*IEEE Std 1800\.2\-2017*\(2017\), 1–472\.[doi:10\.1109/IEEESTD\.2017\.7932212](https://doi.org/10.1109/IEEESTD.2017.7932212)
- Foster \(2015\)Harry D\. Foster\. 2015\.Trends in functional verification: a 2014 industry study\. In*Proceedings of the 52nd Annual Design Automation Conference*\(San Francisco, California\)*\(DAC ’15\)*\. Association for Computing Machinery, New York, NY, USA, Article 48, 6 pages\.[doi:10\.1145/2744769\.2744921](https://doi.org/10.1145/2744769.2744921)
- Bergeron \(2003\)Janick Bergeron\. 2003\.*Writing Testbenches: Functional Verification of HDL Models*\.Springer, Boston, MA, USA\.[doi:10\.1007/978\-1\-4615\-0302\-6](https://doi.org/10.1007/978-1-4615-0302-6)
- Spear \(2008\)Chris Spear\. 2008\.*SystemVerilog for Verification: A Guide to Learning the Testbench Language Features*\.Springer\.
- Qiu et al\.\(2024\)Ruidi Qiu, Grace Li Zhang, Rolf Drechsler, Ulf Schlichtmann, and Bing Li\. 2024\.AutoBench: Automatic Testbench Generation and Evaluation Using LLMs for HDL Design\. In*Proceedings of the 2024 ACM/IEEE International Symposium on Machine Learning for CAD*\(Salt Lake City, UT, USA\)*\(MLCAD ’24\)*\. Association for Computing Machinery, New York, NY, USA, Article 18, 10 pages\.[doi:10\.1145/3670474\.3685956](https://doi.org/10.1145/3670474.3685956)
- Qiu et al\.\(2025\)Ruidi Qiu, Grace Li Zhang, Rolf Drechsler, Ulf Schlichtmann, and Bing Li\. 2025\.CorrectBench: Automatic Testbench Generation with Functional Self\-Correction using LLMs for HDL Design\. In*2025 Design, Automation & Test in Europe Conference \(DATE\)*\. 1–7\.[doi:10\.23919/DATE64628\.2025\.10992873](https://doi.org/10.23919/DATE64628.2025.10992873)
- Ye et al\.\(2025\)Junhao Ye, Yuchen Hu, Ke Xu, Dingrong Pan, Qichun Chen, Jie Zhou, Shuai Zhao, Xinwei Fang, Xi Wang, Nan Guan, and Zhe Jiang\. 2025\.From Concept to Practice: an Automated LLM\-aided UVM Machine for RTL Verification\. In*2025 IEEE/ACM International Conference On Computer Aided Design \(ICCAD\)*\. 1–9\.[doi:10\.1109/ICCAD66269\.2025\.11240679](https://doi.org/10.1109/ICCAD66269.2025.11240679)arXiv:2504\.19959\.
- Yan et al\.\(2025\)Zhiyuan Yan, Wenji Fang, Mengming Li, Min Li, Shang Liu, Zhiyao Xie, and Hongce Zhang\. 2025\.AssertLLM: Generating Hardware Verification Assertions from Design Specifications via Multi\-LLMs\. In*Proceedings of the 30th Asia and South Pacific Design Automation Conference*\(Tokyo, Japan\)*\(ASPDAC ’25\)*\. Association for Computing Machinery, New York, NY, USA, 614–621\.[doi:10\.1145/3658617\.3697756](https://doi.org/10.1145/3658617.3697756)arXiv:2402\.00386\.
- Wu et al\.\(2025\)Fenghua Wu, Evan Pan, Rahul Kande, Michael Quinn, Aakash Tyagi, David Kebo Houngninou, Jeyavijayan Rajendran, and Jiang Hu\. 2025\.Spec2Assertion: Automatic Pre\-RTL Assertion Generation using Large Language Models with Progressive Regularization\.arXiv:2505\.07995 \[cs\.AR\][https://arxiv\.org/abs/2505\.07995](https://arxiv.org/abs/2505.07995)
- Orenes\-Vera et al\.\(2023\)Marcelo Orenes\-Vera, Margaret Martonosi, and David Wentzlaff\. 2023\.Using LLMs to Facilitate Formal Verification of RTL\.arXiv:2309\.09437 \[cs\.AR\][https://arxiv\.org/abs/2309\.09437](https://arxiv.org/abs/2309.09437)
- Liu et al\.\(2026\)Hongduo Liu, Yuntao Lu, Mingjun Wang, Xufeng Yao, and Bei Yu\. 2026\.LLM\-Assisted Circuit Verification: A Comprehensive Survey\. In*2026 31st Asia and South Pacific Design Automation Conference \(ASP\-DAC\)*\. 439–446\.[doi:10\.1109/ASP\-DAC66049\.2026\.11420692](https://doi.org/10.1109/ASP-DAC66049.2026.11420692)
- Yao et al\.\(2023\)Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik R Narasimhan, and Yuan Cao\. 2023\.ReAct: Synergizing Reasoning and Acting in Language Models\. In*The Eleventh International Conference on Learning Representations*\.[https://openreview\.net/forum?id=WE\_vluYUL\-X](https://openreview.net/forum?id=WE_vluYUL-X)
- Wang et al\.\(2024\)Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Jirong Wen\. 2024\.A survey on large language model based autonomous agents\.*Frontiers of Computer Science*18, 6 \(2024\), 186345\.[doi:10\.1007/s11704\-024\-40231\-1](https://doi.org/10.1007/s11704-024-40231-1)
- Shinn et al\.\(2023\)Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao\. 2023\.Reflexion: Language agents with verbal reinforcement learning\.*Advances in Neural Information Processing Systems*36 \(2023\), 8634–8652\.
- Wile et al\.\(2005\)Bruce Wile, John Goss, and Wolfgang Roesner\. 2005\.*Comprehensive functional verification: The complete industry cycle*\.Morgan Kaufmann\.
- Piziali \(2008\)Andrew Piziali\. 2008\.*Functional verification coverage measurement and analysis*\.Springer\.
- Synopsys, Inc\. \(2023\)Synopsys, Inc\. 2023\.*VCS Functional Verification Solution*\.Synopsys, Inc\.[https://www\.synopsys\.com/content/dam/synopsys/gated\-assets/verification/vcs\-ds\.pdf](https://www.synopsys.com/content/dam/synopsys/gated-assets/verification/vcs-ds.pdf)Official datasheet\. Accessed: 2026\-04\-01\.
- Cadence Design Systems, Inc\. \(2023\)Cadence Design Systems, Inc\. 2023\.*Xcelium Logic Simulator*\.Cadence Design Systems, Inc\.[https://www\.cadence\.com/content/dam/cadence\-www/global/zh\_CN/documents/tools/system\-design\-verification/secured/xcelium\-logic\-simulator\-ds\.pdf](https://www.cadence.com/content/dam/cadence-www/global/zh_CN/documents/tools/system-design-verification/secured/xcelium-logic-simulator-ds.pdf)Official datasheet\. Accessed: 2026\-04\-01\.
- cocotb contributors \(2025\)cocotb contributors\. 2025\.*Welcome to cocotb’s Documentation\!*cocotb\.[https://docs\.cocotb\.org/](https://docs.cocotb.org/)Version 2\.0\.1\. Accessed: 2026\-03\-21\.
- The Verilator Project \(2026\)The Verilator Project 2026\.*Verilator Documentation*\.The Verilator Project\.[https://veripool\.org/verilator/documentation/](https://veripool.org/verilator/documentation/)Version 5\.024\. Accessed: 2026\-04\-01\.
- Hu et al\.\(2025\)Yuchen Hu, Junhao Ye, Ke Xu, Jialin Sun, Shiyue Zhang, Xinyao Jiao, Dingrong Pan, Jie Zhou, Ning Wang, Weiwei Shan, Xinwei Fang, Xi Wang, Nan Guan, and Zhe Jiang\. 2025\.UVLLM: An Automated Universal RTL Verification Framework using LLMs\. In*2025 62nd ACM/IEEE Design Automation Conference \(DAC\)*\. 1–7\.[doi:10\.1109/DAC63849\.2025\.11435108](https://doi.org/10.1109/DAC63849.2025.11435108)arXiv:2411\.16238\.
- Zhang et al\.\(2025\)Zixi Zhang, Balint Szekely, Pedro Gimenes, Greg Chadwick, Hugo McNally, Jianyi Cheng, Robert Mullins, and Yiren Zhao\. 2025\.LLM4DV: Using Large Language Models for Hardware Test Stimuli Generation\. In*2025 IEEE 33rd Annual International Symposium on Field\-Programmable Custom Computing Machines \(FCCM\)*\. 133–137\.[doi:10\.1109/FCCM62733\.2025\.00048](https://doi.org/10.1109/FCCM62733.2025.00048)arXiv:2310\.04535\.
- Tencent Cloud \(2025\)Tencent Cloud\. 2025\.*CodeBuddy Documentation: IDE User Guide Overview*\.Tencent Cloud\.[https://www\.codebuddy\.ai/docs/zh/ide/User\-guide/Overview](https://www.codebuddy.ai/docs/zh/ide/User-guide/Overview)Official documentation\. Accessed: 2026\-04\-01\.
- Rudy \(2020\)Tim Rudy\. 2020\.uart\-verilog\.[https://github\.com/TimRudy/uart\-verilog](https://github.com/TimRudy/uart-verilog)GitHub repository\. Accessed: 2026\-04\-01\.
- Forencich \(2017\)Alex Forencich\. 2017\.verilog\-i2c\.[https://github\.com/alexforencich/verilog\-i2c](https://github.com/alexforencich/verilog-i2c)GitHub repository\. MIT License\. Accessed: 2026\-04\-01\.
- Strömbergson \(2013\)Joachim Strömbergson\. 2013\.sha256\.[https://github\.com/secworks/sha256](https://github.com/secworks/sha256)GitHub repository\. BSD\-2\-Clause License\. Accessed: 2026\-04\-01\.
- gongxunwu \(2020\)gongxunwu\. 2020\.sm4\-verilog\.[https://github\.com/gongxunwu/sm4\-verilog](https://github.com/gongxunwu/sm4-verilog)GitHub repository\. Accessed: 2026\-04\-01\.
- Terasic Inc\. \(2012\)Terasic Inc\. 2012\.*DE0\-Nano SDRAM Controller*\.Terasic Inc\.[https://www\.terasic\.com\.tw/cgi\-bin/page/archive\.pl?Language=English&No=593](https://www.terasic.com.tw/cgi-bin/page/archive.pl?Language=English&No=593)Official reference design archive\. Accessed: 2026\-04\-01\.Similar Articles
Structured Testbench Generation for LLM-Driven HDL Design and Verification-Oriented Data Curation
This paper presents STG, a structured testbench generation framework for LLM-driven hardware design workflows that reduces token cost and improves verification reliability compared to existing prompt-based approaches.
VeriTrace: Human-Like Temporal Exploration Completes Agentic Action Space
VeriTrace, a multi-agent system for automated Verilog RTL generation, introduces Agentic Temporal Exploration that gives debugging agents full control over signal selection, time windows, and iteration depth, achieving 100% Pass@1 on VerilogEval-V2 and outperforming baselines by +5.1%.
Alpha-RTL: Test-Time Training for RTL Hardware Optimization
Alpha-RTL (TTT-RTL) introduces a test-time training framework for RTL hardware optimization, using reinforcement learning with EDA feedback to refine LLM-generated designs. It achieves significant PPA reductions on benchmarks.
Text-Graph Synergy: A Bidirectional Verification and Completion Framework for RAG
This paper introduces TGS-RAG, a bidirectional verification and completion framework that synergizes text-based and graph-based Retrieval-Augmented Generation to improve multi-hop reasoning accuracy.
RTL-BenchMT: Dynamic Maintenance of RTL Generation Benchmark Through Agent-Assisted Analysis and Revision
RTL-BenchMT is an agentic framework that automatically identifies and revises flawed cases and detects overfitting in RTL generation benchmarks, reducing human maintenance effort in EDA research.