Public services are increasingly strained by LLM-written appeals for benefits

Hacker News Top Papers

Summary

This paper characterizes 'agentic flooding' of government services by AI agents, particularly LLMs, analyzes service exposure to such surges, and recommends mitigation strategies to balance accessibility and strain.

No content available
Original Article
View Cached Full Text

Cached at: 08/24/26, 04:50 PM

# Characterizing Agentic Flooding of Government Services
Source: [https://arxiv.org/html/2608.16603](https://arxiv.org/html/2608.16603)
###### Abstract

AI agents are making it easier for the public to interact with government, such as by helping them apply for benefits, understand complex policies, and make their opinions heard\. Although improving service accessibility is beneficial, any resulting surges in demand could strain unprepared government services\. We term such surgesagentic floodingof government services \(“flooding”\) and provide three contributions\. First, based on a collected dataset of 84 potential cases of flooding across 11 jurisdictions, we posit that flooding is likely occurring widely today, mostly through large language models \(LLMs\) generating text cheaply\. Second, we evaluate what services are most exposed to flooding\. We develop a risk matrix to analyze a service’s exposure, and suggest that near\-term risk is highest for financially attractive, but complex services\. Finally, we map possible government responses to flooding\. Precedent suggests these responses will likely be sufficient to stop most cases of flooding, but the fastest to deploy – friction\-inducing measures like fees – often trade off equitable access to public services\. Accordingly, we close by recommending near\-term actions that may allow governments to mitigate flooding without invoking this trade\-off\.

Dataset—https://github\.com/CLSchmitz/flooding\-dataset

Preprint\.This work will appear in the proceedings of the 9th AAAI Conference on AI, Ethics, and Society \(AIES\), October 12–14, 2026\.

## 1Introduction

Figure 1:Annualized submission volumes \(2018–2025\) for 12 government services\. For each, either government officials \(red\) or reputable secondary sources \(blue\) have asserted AI involvement in surges\. Only some surges begin around the 2022 release of ChatGPT \(dashed lines\)\.Table 1:Types of “administrative burden” citizens face when interacting with government\([44](https://arxiv.org/html/2608.16603#bib.bib2)\), the AI agent capabilities that may reduce them, and technologies that enable these capabilities\.AI agents increasingly help the public interact with government\. For example, they can already generate legally coherent text for an official complaint, or parse complex policy documents into plain language\.[30](https://arxiv.org/html/2608.16603#bib.bib3)classify about 1% of requests to Google Gemini as helping with government interaction\. These capabilities reduce submission costs for citizens and make government more accessible, a desirable public value\([43](https://arxiv.org/html/2608.16603#bib.bib4)\)\.

However, such reductions in cost could also cause surges in the complexity or quantity of requests\. Evidence of such surges is emerging: the Australian government has considered reintroducing fees for Freedom\-of\-Information \(FOI\) requests citing a wave of AI\-generated submissions\([9](https://arxiv.org/html/2608.16603#bib.bib5)\); some German social courts largely attribute a 55% year\-on\-year rise in caseload in 2025 to AI\-generated claims\([38](https://arxiv.org/html/2608.16603#bib.bib6)\);[64](https://arxiv.org/html/2608.16603#bib.bib7)find a “dramatic increase” in both the share and average length of self\-represented cases before US federal courts since 2022, presumably driven by AI\-aided citizen self\-representation\.

Key questions about such surges remain open\. In particular, we need to know how widespread they are, how serious a risk they pose, and what governments can do about them\. Given the rapid pace of AI capabilities advancements, there is a clear and pressing need for greater understanding\.

This work provides a structured treatment of the phenomenon we termagentic floodingof government services \(or “flooding” hereafter\): surges in the volume or complexity of requests they receive, which \(1\) strain their capacity and \(2\) are enabled by agents reducing the cost of interacting with such services \([Section2](https://arxiv.org/html/2608.16603#S2)\)\. We address three questions:

##### Is flooding happening – and if so, in what forms? \([Section4](https://arxiv.org/html/2608.16603#S4)\)

We compile a dataset \([Section3](https://arxiv.org/html/2608.16603#S3)\) of 84 recent surges in service demand, each of which officials or reputable secondary sources attribute to AI use \([Figure1](https://arxiv.org/html/2608.16603#S1.F1)\)\. The data is consistent with agentic flooding occurring widely today, though our methodology does not permit causal or quantitative claims about AI’s impact\. Almost all the cases we identify share a mechanism: LLMs generating large amounts of text, paired with humans navigating the rest of the process manually\. More sophisticated use of agents, such as autonomous website navigation, is not yet evident\.

##### How big is the risk? \([Section5](https://arxiv.org/html/2608.16603#S5)\)

The risk of flooding depends on the progress and diffusion of AI capabilities, and varies significantly across countries and services\. Analyzing our data, we posit that the near\-term risk of flooding is highest for financially attractive services where complex submission requirements have historically suppressed demand, such as court claims and tax returns\. These properties can be assessed in services now, and such assessments can be updated as new information about AI capabilities arrives\. We construct a risk matrix to guide more detailed assessment of individual services\.

##### How could governments respond? \([Section6](https://arxiv.org/html/2608.16603#S6)\)

We map government response options, grouped under two strategies: suppressing demand \(e\.g\. fees, rate limits, in\-person requirements\) and increasing capacity \(e\.g\. staffing more people, deploying AI, or redesigning the service interface\)\. Precedent suggests both approaches are sufficient to address flooding, but demand suppression is far faster to deploy\. However, it decreases the quality and accessibility of government services, and can introduce procedural inequality\.

Preparing proactively could help governments avoid this trade\-off\. To that end, we recommend three near\-term actions \([Section7\.1](https://arxiv.org/html/2608.16603#S7.SS1)\): auditing service vulnerability, integrating digital identity into the most exposed services, and clarifying the legality of resilience\-building measures\.

## 2Agentic Flooding of Government Services

This section definesagentic floodingof government services \([Section2\.1](https://arxiv.org/html/2608.16603#S2.SS1)\) and analyzes its potential causes \([Section2\.2](https://arxiv.org/html/2608.16603#S2.SS2)\)\.

### 2\.1Definitions

Throughout this work, we useAI agents\(or “agents”\) to mean LLM\-based systems with some degree of autonomy and tool access\([33](https://arxiv.org/html/2608.16603#bib.bib8)\)\. Our definition includes current\-generation consumer\-facing LLMs, as they typically integrate tools such as web search\([62](https://arxiv.org/html/2608.16603#bib.bib9)\)\.

We usegovernment servicesas a deliberately broad term, to mean both public services – such as welfare or infrastructure – and non\-service channels of government interaction, e\.g\. public participation, judicial procedures, or freedom of information requests\([51](https://arxiv.org/html/2608.16603#bib.bib10)\)\.

We usethe publicandusersinterchangeably to refer to all individuals interacting with government services\.

With these terms, we can define agentic flooding\.

Agentic floodingof a government service is a surge in the volume or complexity of requests it receives, which\(i\)is caused by AI agents interacting with it; and\(ii\)substantially strains the service\.

We distinguish between two types of flooding: increases in the number of requests \(quantitativeflooding\), and increases in the complexity of each request \(qualitativeflooding\)\. For example, agents that help users discover eligible services may increase request volume, whereas agents that draft text may lengthen the average complaint letter\. An instance of flooding could be both qualitative and quantitative\. For example, LLM hallucinations could result in high volumes of complex, but incorrect, applications\. Both of these effects, which[4](https://arxiv.org/html/2608.16603#bib.bib11)term the “extensive” and “intensive” margins of work, increase processing burden\. However, some potential responses only address one type \([Section6](https://arxiv.org/html/2608.16603#S6)\)\. For example, introducing rate limits is unlikely to ease qualitative flooding\.

### 2\.2Potential Causes of Flooding

AI agents could cause a surge in the volume or complexity of requests by \(1\) reducing the costs of interacting with the government service or \(2\) increasing the benefits of doing so\. As such, both users and middle\-man organizations, e\.g\. law firms, have incentives to deploy agents to interact with government services\. Any resulting increases in demand could strain the service because many services are only designed to handle a limited amount of demand\.

##### Agentic Capabilities Could Reduce Costs\.

AI agents could reduceadministrative burdens: the costs of participating in bureaucratic procedure which users face when interacting with government\([23](https://arxiv.org/html/2608.16603#bib.bib12)\)\.[44](https://arxiv.org/html/2608.16603#bib.bib2)group these costs into three types: learning, compliance, and psychological costs\.

In[Table1](https://arxiv.org/html/2608.16603#S1.T1), we map how agent capabilities could reduce all three costs\.Learning costs, the cognitive effort of understanding government, are addressed most directly by agents’ ability to parse long contexts and personalize communication to users\([34](https://arxiv.org/html/2608.16603#bib.bib13)\)\.Compliance costs, the practical effort of completing required steps, fall as agents grow more capable of high\-quality digital work\([13](https://arxiv.org/html/2608.16603#bib.bib14)\)\.Psychological costs, the emotional toll of repeated bureaucratic contact, drop as agents handle more of that contact\.

Reductions in administrative burdens often raise demand for services\. For example, the introduction of an app to request municipal services in Boston led to a 33% increase in reports\([52](https://arxiv.org/html/2608.16603#bib.bib15)\), and online voter registration options consistently increase enrolment numbers\([17](https://arxiv.org/html/2608.16603#bib.bib16)\)\.

At a conceptual level, rational\-actor models propose that users submit requests when the benefit they expect exceeds the cost of submission\.\([48](https://arxiv.org/html/2608.16603#bib.bib17);[76](https://arxiv.org/html/2608.16603#bib.bib18)\)\. Individual users rarely act financially optimally – their behavior is shaped by many other factors, such as psychological frictions or stigma\([11](https://arxiv.org/html/2608.16603#bib.bib19);[3](https://arxiv.org/html/2608.16603#bib.bib20)\)\. But the rational\-actor model often accurately describes aggregate behavior across a population\([71](https://arxiv.org/html/2608.16603#bib.bib21)\)\. We would therefore expect requests to increase when AI agents decrease submission costs\.

##### Agentic Capabilities Could Increase Benefits\.

Agentic capabilities could also increase the benefits of interacting with a service\. For example, agents could rewrite Request for Information \(RFI\) responses to be more convincing, increasing the likelihood that the responses’ views are reflected in policy\. AI agents could also cheaply help to ensure that all information relevant to an application is gathered, ensuring that applicants receive the maximal benefit\. Analogously, tax accountants help to ensure that clients receive their maximum possible tax refund\.

##### Incentives to Deploy Agentic Capabilities\.

Beyond the user\-level incentives, there are strong commercial incentives to deploy agentic capabilities in ways that increase demand for government services\. The business model is well\-established: middle\-man organizations offer to handle complex government interactions on behalf of users in exchange for a cut of the earned entitlements, growing the number of claims to government\. The most prominent example is claims management companies, which have long mass\-submitted benefit and compensation claims through templated submissions, such as for flight delay compensation and UK personal injury claims\. Companies advertising their use of AI to similar effect already exist\([12](https://arxiv.org/html/2608.16603#bib.bib22)\)\.

Further, adversarial actors may explicitly seek to flood services\. For example, they may aim to destabilize critical infrastructure\([57](https://arxiv.org/html/2608.16603#bib.bib23)\), or to incapacitate agencies\. Even good\-faith users may be incentivized to abuse the system\. Analoguously, “vexatious litigants” file large numbers of meritless court claims, each of which they are entitled to file, but which collectively overwhelm the court\([60](https://arxiv.org/html/2608.16603#bib.bib24)\)\. Thus far, transaction costs have kept such cases rare enough for ad\-hoc management; agent capabilities could make these cases much more frequent\.

##### Increased Demand Can Cause Strain\.

Increased demand for government services is not in itself harmful, but agencies often cannot meet it\. Many services run close to operational capacity because of budget constraints, operational inefficiency, or legal obligations to minimize overhead\([18](https://arxiv.org/html/2608.16603#bib.bib25)\)\. Some governments know of and tolerate, or even deliberately maintain, administrative burdens which suppress take\-up\([56](https://arxiv.org/html/2608.16603#bib.bib26)\)\.

Capacity strain can have both direct operational consequences – like growing backlogs or missed response deadlines – and broader fiscal and policy effects, such as requiring budget adjustments \([Section5](https://arxiv.org/html/2608.16603#S5)\)\. Collectively, these effects may erode user experience of public services, which strongly predicts trust in government\([70](https://arxiv.org/html/2608.16603#bib.bib27);[10](https://arxiv.org/html/2608.16603#bib.bib28)\)\.

## 3Evidence Collection

To compile a dataset of real\-life cases where flooding may be occurring, we scan government services in 12 countries \([Section3\.1](https://arxiv.org/html/2608.16603#S3.SS1)\)\. Because AI involvement in demand surges is difficult to prove, we only include cases where a government official or credible third party has asserted such involvement \([Section3\.2](https://arxiv.org/html/2608.16603#S3.SS2)\)\. This enables a broad scan but trades off statistical representativeness\. To collect cases, we employ an LLM\-aided qualitative coding workflow \([Section3\.3](https://arxiv.org/html/2608.16603#S3.SS3)\)\.

### 3\.1Scope

We scan for publicly accessible information about government services in 12 countries: Australia, Brazil, Denmark, Estonia, France, Germany, Japan, the Netherlands, Singapore, South Korea, the United Kingdom, and the United States\. We choose countries where it is common to report publicly on administrative matters, aiming for variation in geography and in two core dimensions of public administration scholarship: administrative tradition\([54](https://arxiv.org/html/2608.16603#bib.bib29)\)and digital government maturity\([50](https://arxiv.org/html/2608.16603#bib.bib30)\)\.

In each of these countries, we scan the same fixed list of 13 government domains\. We set this list by combining public\-facing lists of government services from three studied countries \(UK, US, and France\)\. Ten domains cover public services, e\.g\. benefits and social protection, citizenship, or jobs and pensions\. Three cover government processes that are not public services, but in which flooding may also occur: participatory processes; regulatory complaints and reporting; and transparency and access to information\. Appendix[A](https://arxiv.org/html/2608.16603#A1)describes the detailed methodology with which we set both scopes\.

Domain promptDiscoveryTriageRefinementReviewDecisionRejectAcceptFigure 2:Overview of LLM\-aided pipeline for collection of cases\. Yellow: human step\. Blue: LLM step\.Table 2:Ten illustrative cases ofagentic floodingin our dataset\.
### 3\.2Case Selection and Attribution

Demand for public services is shaped by many interrelated factors including policy changes, spurious events, and broader social and economic developments; processing indicators like backlog sizes and wait times are similarly confounded\([67](https://arxiv.org/html/2608.16603#bib.bib31)\)\. Deeper analysis may mitigate these problems in individual cases, especially where submission contents are public\([64](https://arxiv.org/html/2608.16603#bib.bib7)\), but it is not \(yet\) feasible at scale\.

As a way to partially address this issue, we select cases conservatively\. In addition to evidence of the two criteria in our definition, we require that the affected government body itself – or a reputable, relevant source – attribute a change in demand patterns to AI use by the public\. We pose three separate inclusion criteria per case:

1. 1\.A plausible, specific mechanismby which AI could reduce transaction costs for the service – e\.g\. that LLMs help write a complaint which can then be submitted online\. We judge plausibility ourselves based on the design of the service\.
2. 2\.Evidence of changein demand patterns for the service that is consistent with what the mechanism in \(1\) could cause, found in official statements or actions \(or another high\-quality source\), and reflected in primary \(self\-collected\) volume data if available\. Trends in the volume data alone do not fulfill the criterion\.
3. 3\.External attributionof that change to AI use, by either \(a\) a government source – e\.g\. a press statement, official communication, or policy update – or \(b\) a reputable secondary source – e\.g\. an established news outlet, professional legal or administrative publication, or peer\-reviewed work\.

Our methodology produces a broad dataset but does not afford causal, quantitative, or statistical claims\. We further discuss limitations in[Section8](https://arxiv.org/html/2608.16603#S8)\.

### 3\.3Methodology

##### Data Format\.

We code all cases into a standardized template, provided in the dataset repository and outlined in Appendix[B](https://arxiv.org/html/2608.16603#A2)\. All figures for which we provide statistics are captured in a structured format in this schema\. We do not perform additional analysis on free\-text fields, except to highlight illustrative examples\. For each case, we attempt to collect:

- •Free\-text descriptions of the service and its submission interface, the mechanism by which flooding may occur, and any impact and government response observed\.
- •Structured fields classifying the case type and mechanism along the typologies in[Sections2\.2](https://arxiv.org/html/2608.16603#S2.SS2)and[2\.1](https://arxiv.org/html/2608.16603#S2.SS1), whether it meets each of the three inclusion criteria, and who asserts AI involvement in the change in requests\.
- •Annualized volume statistics from 2018 to 2025 in a case\-specific unit \(e\.g\. requests filed, backlog size\)\.
- •Metadata, e\.g\. country, government body, and domain; and links to all sources used in compiling the case\.

##### Case Collection Pipeline\.

We employ an LLM\-aided research and qualitative coding workflow\([74](https://arxiv.org/html/2608.16603#bib.bib32)\)\. As outlined in[Figure2](https://arxiv.org/html/2608.16603#S3.F2), the pipeline has three stages\. Duringdiscovery, a country\-domain combination \(e\.g\. “France, Health Services”\) is scanned for candidate services\. Initeration, each case cycles between LLM calls to improve case data and review case quality\. Infinalization, a human reviewer validates the case, makes required corrections, and chooses whether to include it in the dataset\. Detailed prompts and scaffolding are listed in Appendix[B](https://arxiv.org/html/2608.16603#A2)\.

##### Preventing Hallucinations\.

Our approach contains three safeguards against the risk of LLM hallucinations\([27](https://arxiv.org/html/2608.16603#bib.bib33)\)\. First, we perform human verification of all cases, before both iteration and finalization\. Second, we force structured responses from all LLM calls, and maintain each case in a standardized, JSON\-based data format\. This format enables us to verify key fields, such as the names of government agencies and numeric case counts, deterministically; it also allows metadata, such as iteration count and review comments, to be tracked and passed to each LLM call\. Third, we extract accessed web URLs deterministically from the metadata of LLM API responses, rather than requiring the models themselves to generate them as part of the response\. This guarantees all cited sources are real, human\-verifiable websites, which we visit to confirm key figures\.

## 4Is Agentic Flooding Happening?

The collected data is consistent with flooding occurring today, across a wide range of government domains and jurisdictions\. Almost all cases occur through one mechanism: LLM\-generated text submitted by users who navigate the rest of the process manually\.

##### Dataset Overview\.

We document 84 cases across 11 jurisdictions and 13 service domains\. Full case data are available in the linked dataset repository\.[Table2](https://arxiv.org/html/2608.16603#S3.T2)contains ten illustrative cases to which the subsequent analysis refers\.

These 84 cases are the product of a broad scan followed by strict inclusion criteria: of the 2288 candidate services named during discovery – roughly 190 per country – fewer than one in twenty met all three criteria in[Section3\.2](https://arxiv.org/html/2608.16603#S3.SS2)\. The biggest constraint on inclusion was requiring official or third\-party attribution of the surge to AI\.

Of the 84 cases, government officials assert AI involvement \(with or without third\-party corroboration\) in 58 \(69%\); solely third\-party sources assert it in the remaining 26 \(31%\)\. Public and judicial services are more represented than non\-service interaction channels, such as participation procedures and transparency requests \(73% and 27% of cases, respectively\)\. The most frequently represented domains are Justice & Legal Services \(23%, N=19\), followed by Regulatory Complaints \(12%, N=10\) and Benefits & Social Protection \(11%, N=9\)\.

##### Strength of Evidence\.

The dataset provides evidence that AI use is increasing the volume and complexity of requests for at least some government services\. The cases span jurisdictions, service types, and reporting outlets; in most cases, the affected government body itself asserts AI involvement\. Any alternative explanation would have to account for independent misattribution across all of these criteria\.

However, our data\-gathering remains exploratory and diagnostic, and affords no causal or quantitative claims about either \(a\) the prevalence of flooding or \(b\) the role of agents\. Demand is driven by many factors, and some volume series in[Figure1](https://arxiv.org/html/2608.16603#S1.F1)begin rising before the release of ChatGPT\. Moreover, because we require explicit public attribution, the dataset likely undercounts the cases that meet our definition of flooding \([Section2\.1](https://arxiv.org/html/2608.16603#S2.SS1)\)\.[Section8](https://arxiv.org/html/2608.16603#S8)discusses these limitations\.

##### Capabilities Enabling Flooding\.

In most of our cases \(87%\), flooding occurs when LLMs help to cheaply generate legally sophisticated text\. Most of the services in the dataset accept text, often of arbitrary length, e\.g\. FOI requests \[Case A\], planning consultations \[Case F\], or public comments \[Case C\]\.

There are two potential explanations for this pattern\. First, text generation is more mature and accessible than agentic capabilities such as browser navigation\([78](https://arxiv.org/html/2608.16603#bib.bib34)\)\. Second, sampling bias: cases largely enter our dataset because officials or domain experts noticed an anomalous change in the pattern of submissions \(69%, N=58\)\. LLM\-generated text can be identified by the content of the submission itself, whereas many other types of agent\-aided submissions may look identical to human ones if the submission format is constrained – e\.g\. by short fields or structured online forms\. Any effect such submissions have on request patterns could only be observed via the quantity of requests, which is also affected by various other factors\([67](https://arxiv.org/html/2608.16603#bib.bib31)\)\. It is plausible that officials are thus far attributing such quantitative changes to AI less frequently\.

##### Types of Flooding Observed\.

We code 60% \(N=50\) of cases as quantitative flooding and 90% \(N=76\) as qualitative, with 50% \(N=42\) exhibiting both types\. These numbers are largely consistent with the above pattern of LLM text generation, which both allowsmoreusers to submit requests to free\-text service interfaces – e\.g\. the less legally literate submitting litigation \[Case H\] – and makes these requestslonger, as in \[Case B\], where some letters spanned over 4000 pages\.

We observe cases of both adversarial and non\-adversarial flooding in the dataset\. While most cases show ordinary users acting in good faith, some involve individuals acting in bad faith for personal gain, e\.g\. attempting to get disability benefits with AI\-generated medical certificates \[Case D\]\. A few involve organized adversarial campaigns, notably mass\-submitting FOI requests or voter roll challenges \[Case J\]\.

## 5How Big Is the Risk?

ABCSeverityLikelihoodSample servicesAPublic comment proceduresBBenefits appealsCEmergency services

Figure 3:Risk matrix of factors which may predict the likelihood and severity ofagentic floodingfor a given government service, with three illustrative services mapped\. Arrows indicate whether a higher value raises \(↑\\uparrow\) or lowers \(↓\\downarrow\) risk\.The risks posed by flooding depend on how quickly AI capabilities progress and diffuse, and vary with the service and jurisdiction affected\. A key question is therefore which services are most at risk\. Analyzing our dataset, we posit that near\-term risk is highest for financially lucrative services with complex applications and legal processing obligations, such as tax returns \([Section5\.1](https://arxiv.org/html/2608.16603#S5.SS1)\)\. We propose a risk matrix \([Section5\.2](https://arxiv.org/html/2608.16603#S5.SS2)\) to more rigorously evaluate the likelihood and severity of a specific service being flooded\.

### 5\.1Which Services Are Most Vulnerable?

The services in our dataset share several commonalities\. Most obviously, they overwhelmingly accept free\-form text through open digital channels, and the agent capability required to flood them – LLM text generation – is freely accessible \([Section4](https://arxiv.org/html/2608.16603#S4)\)\.

The observed impacts of flooding so far are moderate\. While we find evidence of overworked public servants, calls for additional budget, and calls to reform legal response obligations \[Case E\], the strain is modest compared to precedent, e\.g\. from mass comment campaigns\([2](https://arxiv.org/html/2608.16603#bib.bib35)\)\. Where governments respond, e\.g\. by re\-introducing FOI fees \[Case A\] and batch\-dismissing consultation responses \[Cases C, G\], these responses also appear effective\.

There are some higher\-severity cases in our dataset, e\.g\. objections to tax valuations \[Case I\], social court lawsuits \[Case B\], and civil claims \[Cases E, H\]\. These share two properties\. First, each successful submission is individually consequential: it is financially valuable to the claimant if successful, or triggers legally mandated processing effort\. Second, each service’s resilience has historically been owed to friction rather than design\. With few structural limits on submission length or volume, the complexity of submitting – and the legal or professional knowledge it demanded – historically gated who submitted, and how well\.

Given our inclusion criteria \([Section3](https://arxiv.org/html/2608.16603#S3)\), these patterns may also reflect which cases officials notice and report\. Nonetheless the risk of flooding appears highest for services with these two properties: services that are financially attractive, and whose demand has so far been suppressed by friction or tacit knowledge\. These could include tax administration and court systems\. Notably, an initial assessment of services against these two properties does not require forecasting AI progress\. They can be analyzed today, as we discuss in[Section7\.1](https://arxiv.org/html/2608.16603#S7.SS1)\.

### 5\.2Risk Matrix for Flooding

[Figure3](https://arxiv.org/html/2608.16603#S5.F3)describes a risk matrix\([32](https://arxiv.org/html/2608.16603#bib.bib36)\)for the flooding of a given service, containing factors that may predict thelikelihoodthat a service is flooded and theseverityif it is\. We construct this matrix by combining our data with theory and precedent from past demand surges\([2](https://arxiv.org/html/2608.16603#bib.bib35);[72](https://arxiv.org/html/2608.16603#bib.bib37);[73](https://arxiv.org/html/2608.16603#bib.bib38)\)\. Such a matrix could be used to map a government’s exposure and prioritize interventions \([Section7\.1](https://arxiv.org/html/2608.16603#S7.SS1)\)\.

#### Likelihood\.

Two sets of factors may predict the likelihood that a service is flooded: agent capabilities and access, and existing features of the service\.

Agent Capabilities and Access\.Flooding overall becomes more likely as agent capabilities improve \(Factor 1\)\. Some capabilities, e\.g\. browser use, are likely particularly useful for government interaction \([Section2\.2](https://arxiv.org/html/2608.16603#S2.SS2)\)\.

A service is more susceptible to flooding if the relevant agent capabilities are broadly accessible \(Factor 2\)\. Our cases \([Section4](https://arxiv.org/html/2608.16603#S4)\) are driven almost entirely by LLM text generation, a capability now freely available from multiple providers\. However, the most capable agents are paid services\. Although inference is becoming cheaper overall, access costs or restrictions could prevent flooding, or restrict it to cases caused by wealthier users\([28](https://arxiv.org/html/2608.16603#bib.bib39);[66](https://arxiv.org/html/2608.16603#bib.bib40)\)\.

Service Features\.The design of service interfaces can affect how readily certain capabilities translate into flooding \(Factor 3\)\([31](https://arxiv.org/html/2608.16603#bib.bib1)\)\. For example, even if an agent can collect all the relevant information and generate an application, it can have trouble interacting with portals that have complex login flows or JavaScript\-heavy front\-ends\([13](https://arxiv.org/html/2608.16603#bib.bib14);[22](https://arxiv.org/html/2608.16603#bib.bib41)\)\. Similarly, digital agents cannot attend in\-person appointments\.

Services requiring more effort to submit could be more exposed to flooding \(Factor 4\)\. We would expect higher submission costs to suppress a higher share of true demand \([Section2\.2](https://arxiv.org/html/2608.16603#S2.SS2)\)\. For example, voluntary tax returns can take significant effort, disincentivizing those who only expect a small return\. As agents reduce these costs, it is therefore more likely the resulting demand increase causes strain\. Conversely, hard limits on submission volume – e\.g\. one annual tax return per citizen – could limit how much demand agents can add\.

Finally, the expected benefit of submission likely predicts flooding itself, regardless of whether agents affect it \(Factor 5\)\. Falling submission costs alone can unlock latent demand\([11](https://arxiv.org/html/2608.16603#bib.bib19)\)\. For example, in \[Case B\] claimants more frequently respond to financially consequential claim denials because it is now easy, though there is no evidence their chance of success increases\. We would expect this incentive to be stronger for services that offer higher financial reward\.

#### Severity\.

Where flooding occurs, its impact could take two forms: first\-order operational impacts from the demand surge itself, and second\-order impacts and externalities from the responses governments adopt to address it\.

Operational Impact\.Operational impacts, such as backlogs, affect both the agency handling a service and its users\. Their size depends on the efficiency and flexibility of the existing service \(Factors 6–9\), or its “surge capacity”\([25](https://arxiv.org/html/2608.16603#bib.bib42);[5](https://arxiv.org/html/2608.16603#bib.bib43)\)\. For public services, technical indicators of this capacity may include the use of structured and standardized data formats, which can enable “straight\-through” processing of some cases without human intervention\([35](https://arxiv.org/html/2608.16603#bib.bib44)\), or digital identity systems, which can avoid time\-consuming personhood confirmation\([45](https://arxiv.org/html/2608.16603#bib.bib45)\)\.

Response Impact\.Beyond the direct impacts of flooding, the responses governments may choose often have trade\-offs or negative externalities\. We map these in[Section6](https://arxiv.org/html/2608.16603#S6)\. For example, closing a digital submission channel could restrict access to the service\. These trade\-offs depend both on what response is chosen, and how intensely it is implemented\.

Legal constraints can affect which response governments choose\. Rigid budget and entitlement policies \(Factor 10\) and legally mandated processing effort, such as response requirements\([36](https://arxiv.org/html/2608.16603#bib.bib46)\)or right\-to\-explanation legislation\([8](https://arxiv.org/html/2608.16603#bib.bib47)\)\(Factor 11\), narrow the set of feasible responses\.

More intense versions of responses have bigger externalities\. For example, a more aggressive eligibility reduction for a welfare service affects more people\. How strongly a government must intervene depends on the “slack”\([53](https://arxiv.org/html/2608.16603#bib.bib48)\)in the budgets and rules around a service: how much demand could change before they would require updating\([24](https://arxiv.org/html/2608.16603#bib.bib49)\)\. Slack is low where budgets assume take\-up well below full entitlement \(Factor 12\), and where a successful submission has a high downstream cost, such as a benefit payout or a court proceeding \(Factor 13\)\.

Table 3:Overview of possible government responses to agentic flooding, grouped under two strategies\.††nicematrix\-placeholder:NiceTabularX \(nicematrix\)

## 6How Could Governments Respond?

Governments explicitly respond to flooding in over half \(56%\) of cases in our dataset\. However, these actions are usually tightly scoped or non\-binding, like releasing AI use guidance\. Precedent suggests a much broader range of possible future responses, which we map to two categories \([Table3](https://arxiv.org/html/2608.16603#S5.T3)\): suppress demand for the affected services, or increase processing capacity to meet that demand\. These responses would likely be effective, but each induces trade\-offs\. In particular, demand suppression restricts access to public services and can create procedural inequality\.

### 6\.1Suppress Demand

Governments could suppress demand for services by making applications either more difficult or less attractive to complete, and have frequently done so in the past\.

##### Add Friction\.

Governments could \(re\-\)raise the cost of submission until enough users are deterred from filing, thus suppressing quantitative flooding\. Measures taken during previous demand surges include: financial barriers like fees\([36](https://arxiv.org/html/2608.16603#bib.bib46)\), procedural barriers like digital identity verification\([46](https://arxiv.org/html/2608.16603#bib.bib50)\), and direct access limits like per\-claimant rate caps\([60](https://arxiv.org/html/2608.16603#bib.bib24)\)\. An obvious candidate measure for agentic flooding is blocking bots from government websites, or restricting their permissions\([41](https://arxiv.org/html/2608.16603#bib.bib51)\)\.

Most measures to introduce friction require little infrastructural change, have precedent, and can be deployed quickly, including when a surge is acute\. In 14 cases \(17%\) in our dataset, governments have already responded with friction\. For example, fees can be raised wherever a payment system is in place: Australia has considered reintroducing fees for FOI requests in response to a wave of AI\-generated submissions \[Case A\]\. Access can often be restricted with equally little lead time, as when Japanese authorities blocked submissions to a comment procedure by IP address \[Case G\]\. Even where legal change is needed, it can be narrow: in \[Case I\], targeted reform stopped middle\-man organizations from collecting fees on certain types of requests\.

However, introducing friction decreases the accessibility of government services\. Friction disproportionately deters poorer, less digitally literate, and otherwise vulnerable users, effectively shaping who accesses public services\([56](https://arxiv.org/html/2608.16603#bib.bib26)\)\. Fees and in\-person requirements can thereby create procedural inequality and, where legal services are concerned, undermine access to justice\. Friction also worsens user experience of public services and, by extension, trust in government\([70](https://arxiv.org/html/2608.16603#bib.bib27);[10](https://arxiv.org/html/2608.16603#bib.bib28)\)\. Some measures, such as closing digital submission channels, could additionally prevent the accessibility improvements that AI agents promise\([75](https://arxiv.org/html/2608.16603#bib.bib52)\)\.

Friction\-inducing measures may also face practical challenges\. For one, some forms of friction could become less effective as agent capabilities improve, such as how CAPTCHAs no longer reliably identify human website visitors\([58](https://arxiv.org/html/2608.16603#bib.bib53)\)\. They may also be legally prohibited\. For example, in many jurisdictions it is illegal to charge fees for applications to welfare services\([68](https://arxiv.org/html/2608.16603#bib.bib54)\)\. Lastly, friction may not deter adversarial actors or directly address qualitative flooding \(e\.g\. longer text submissions\)\.

##### Reduce the Expected Benefit\.

Governments could also lower what a user stands to gain from submitting to a service\. They could reduce entitlement amounts, narrow eligibility, or increase the cost of unsuccessful applications\. Such revisions often reduce service demand, e\.g\. reforms to poverty assistance in the US in the 90s\([21](https://arxiv.org/html/2608.16603#bib.bib55)\)\.

Entitlement revisions may be required regardless of whether they resolve flooding\. Many services are budgeted assuming only limited take\-up\([48](https://arxiv.org/html/2608.16603#bib.bib17)\)\. For example, the UK Department for Work and Pensions explicitly anchors benefit budgets to historical take\-up rates\([19](https://arxiv.org/html/2608.16603#bib.bib56)\)\. Where budgets cannot be raised to meet full take\-up, reducing benefits is a likely response\.

A key challenge is that reducing benefits is not a purely executive decision, but a legislative decision\. It may be harder to implement because of political priorities, and run counter to other aims, such as reducing poverty or helping the unemployed\.

### 6\.2Increase Capacity

Governments could also respond to flooding by increasing their processing capacity\. Most obviously, they may choose to scale existing resources, e\.g\. by hiring more staff\. However, tight budget constraints are common and existing processes are often inefficient\([15](https://arxiv.org/html/2608.16603#bib.bib57)\)\. We list two alternative options\.

##### Redesign the Service\.

Service redesign can include restructuring service interfaces, integrating identity verification, standardizing data formats, and providing some services proactively without applications\([16](https://arxiv.org/html/2608.16603#bib.bib58)\)\. Such redesign could make government more efficient and free up resources for addressing flooding\. Moreover, as discussed in[Section5\.2](https://arxiv.org/html/2608.16603#S5.SS2), structured forms and identity verification could decrease both the likelihood and the severity of flooding\.

The main disadvantage is the long lead times and proactive financial investment required for implementation\. Structural redesign requires cross\-agency coordination and, in many cases, legislative change\([59](https://arxiv.org/html/2608.16603#bib.bib59);[26](https://arxiv.org/html/2608.16603#bib.bib60)\)\. Bigger projects can take years or even decades, such as introducing central registers or data standards; many Western governments have famously taken decades to get digital infrastructure in place despite well\-documented benefits\. As such, redesign is likely unavailable as a short\-term measure\.

In 13 cases \(15%\), we code the measures governments take in response to flooding as redesign – though they are exclusively small changes, such as introducing a service to verify the legitimacy of case numbers \[Case H\], rather than broad structural reform\.

##### Deploy Agents in Processing\.

Governments could use AI agents themselves – both to accelerate back\-end processing, and to detect and prevent fraud\. Governments are already taking first steps in this direction: in 21 cases \(25%\) in our dataset, they introduce AI tools explicitly in response to flooding, e\.g\. to analyze sentiment across large numbers of public comments \[Case F\] and to improve detection of AI\-generated fake medical certificates \[Case D\]\. In the vast majority of cases these are bounded tools handling one step in the process, consistent with established patterns of slow, piecemeal AI diffusion in public\-sector organizations\([47](https://arxiv.org/html/2608.16603#bib.bib61)\)\.

More comprehensive agent deployment lacks precedent, but evidence suggests it could scale processing throughput drastically\. AI tools are already used in some countries for intake screening, triage, drafting of standard correspondence, and even autonomous handling of routine cases\([69](https://arxiv.org/html/2608.16603#bib.bib62);[49](https://arxiv.org/html/2608.16603#bib.bib63)\)\. Government work is likely suitable for agent deployment because bureaucratic processes and rules are generally well\-documented and exhaustive\([63](https://arxiv.org/html/2608.16603#bib.bib64)\)\. Benchmarks also suggest that agents are becoming increasingly competent in tasks similar to government work\([55](https://arxiv.org/html/2608.16603#bib.bib65)\)\.

However, there remain significant legal and practical risks regarding government use of agents\. Laws often prevent automated government decision\-making\([20](https://arxiv.org/html/2608.16603#bib.bib66)\)and evaluating agents on government tasks remains difficult\([61](https://arxiv.org/html/2608.16603#bib.bib67)\)\. Civil society groups regularly voice concerns about decision transparency, bias, and accountability diffusion\([1](https://arxiv.org/html/2608.16603#bib.bib68)\)\. Careless deployment could further expose government bodies to provider lock\-in or sovereign\-data risk\([7](https://arxiv.org/html/2608.16603#bib.bib69)\)\.

## 7Discussion and Recommendations

We find broad evidence of agentic flooding, but our work does not suggest it poses a severe*operational*risk to governments\. For one, responses to flooding so far are diverse and overall small in scale: piloting AI tools to help with individual steps, issuing guidance on AI use, or precise access limitations\. Looking forward, while uncertainty remains about when particular agent capabilities arrive or diffuse widely, friction\-inducing measures are likely to mitigate the operational impacts of flooding, as they have for previous demand surges following technological shifts\([72](https://arxiv.org/html/2608.16603#bib.bib37);[2](https://arxiv.org/html/2608.16603#bib.bib35)\)\.

However, adding friction worsens user experience of public services, potentially introduces procedural inequality, and prevents some of the accessibility gains that AI agents enable\([29](https://arxiv.org/html/2608.16603#bib.bib70)\)\. In contrast, capacity\-building responses could address flooding without these trade\-offs, but they require lead time, technical investment, and, in many cases, cross\-departmental coordination; they are unlikely to be feasible once a surge is already underway\.

One plausible trajectory is therefore that demand suppression again becomes a default response\. Governments might reach for it because, when a surge occurs, it is the only intervention available on a short enough timeline\. This pattern matches what[37](https://arxiv.org/html/2608.16603#bib.bib71)terms “muddling through” – iterative, short\-term, and mostly reactive adaptation, which keeps government organizations operational, but does not usually optimize outcomes for the public\.

### 7\.1Three Near\-Term Recommendations

Preparing for flooding proactively may avoid governments having to \(re\)introduce friction to address it\. We propose three near\-term actions to this effect\.

##### Auditing Exposure\.

Governments should systematically audit their services for vulnerability to flooding\. Key factors predicting the susceptibility of a service can be assessed today, and updated with future information about AI capabilities \([Section5\.2](https://arxiv.org/html/2608.16603#S5.SS2)\)\. A portfolio of services could be mapped and ranked against these factors, e\.g\. using our risk matrix, yielding a prioritized list of services and the response classes \([Section6](https://arxiv.org/html/2608.16603#S6)\) most applicable to each\.

##### Developing a Digital Identity Strategy\.

Strong identity verification could effectively mitigate quantitative flooding because it allows automated enforcement of per\-claimant rate limits and makes adversarial flooding more difficult\. Digital identity also supports pre\-population of known data, thus lowering the effort required for entitled users to apply, and for governments to process their case\. Over 100 jurisdictions already have digital identity infrastructure, but its maturity varies drastically\([42](https://arxiv.org/html/2608.16603#bib.bib72)\)\. Where feasible, integrating it into the most exposed services is likely to be a high\-leverage action to preempt flooding\.

##### Establishing Legal Certainty\.

Several of the response options in[Section6](https://arxiv.org/html/2608.16603#S6)are either illegal or of unclear legality in some jurisdictions\. For example, many uses of AI in government processing are constrained by AI and administrative law\([8](https://arxiv.org/html/2608.16603#bib.bib47)\), and it can be unclear where restricting free\-form or digital submission channels, or raising fees, is lawful\. This ambiguity may itself be a barrier to preventative action\. Governments should commission internal legal reviews to clarify, for each response class, which forms of it are permissible under current law\. This understanding would be valuable regardless of the order and intensity in which services are flooded\.

## 8Limitations and Future Work

The methodology we employ aims to collect evidence of an emerging phenomenon across jurisdictions and domains\. However, it allows neither statistical claims about prevalence, nor causal claims about AI’s role in flooding\. We suggest future research directions building on this work\. To support such research, our dataset is published with this paper\.

##### Collection Methodology\.

Our semi\-automated pipeline only collects positive instances of potential flooding and does not measure what share of public services \(per jurisdiction\) are assessed\. As such, the statistics presented in[Section4](https://arxiv.org/html/2608.16603#S4)\-[Section6](https://arxiv.org/html/2608.16603#S6)cannot be taken to generalize within or beyond the studied countries, and we do not make comparative claims, e\.g\. of prevalence per country\.

Further sources of uncertainty or bias include the selection of countries and domains studied, the selection of “candidate services” presented during discovery, differing traditions of publicly discussing administrative matters, and the differing performance of LLMs in different languages \(all prompts were given in English\)\.

Further, while we increase our confidence in included cases by requiring explicit third\-party attribution of the surge to AI, that criterion likely excludes many cases that meet our definition of flooding, and could wrongly include some others\. We also do not require time\-series volume data to include a case, and only 25 possess it\. Accordingly, we cannot make even correlative quantitative claims\.

Finally, despite validating all cases with multiple instances of human review, we cannot exclude the risk of LLM hallucination\. We only provide human\-LLM agreement numbers, perform no inter\-rater coding, and do not benchmark the performance of the LLM against humans\.

Based on these limitations, a pressing direction for future work is more rigorous analysis of the relative prevalence of flooding for specific government services\. Candidate methodologies may include quantitative analysis of volume time series, surveys of citizens about their use of AI agents, and analysis of agent usage data, where available\. A further promising direction is more detailed analysis of the relevance of digital\-government maturity for the risk and severity of flooding, e\.g\. analyzing how effective digital identity is at mitigation\.

##### Identification of Risk Factors and Mitigations\.

Our contributions in[Section5](https://arxiv.org/html/2608.16603#S5)and[Section6](https://arxiv.org/html/2608.16603#S6)are largely based on precedent and theory, rather than our data\. We do not place our data in the risk matrix \([Section5\.1](https://arxiv.org/html/2608.16603#S5.SS1)\)\. We also do not attempt to measure the success of any initial government responses we record, and the taxonomy of possible responses abstracts over jurisdiction\-specific legal constraints\.

This suggests a number of natural next steps for future work\. The predictive value of the listed risk factors should be empirically validated within a single jurisdiction\. Government responses should be measured empirically for both their effectiveness and the externalities they produce\. Both the proposed risk matrix and response taxonomies should be continually updated as a result\.

## 9Related Work

##### Concurrent Work\.

[57](https://arxiv.org/html/2608.16603#bib.bib23)theorize “congested bureaucracy” as a result of AI\-driven friction reduction; we validate and extend their framing here\.[64](https://arxiv.org/html/2608.16603#bib.bib7)find evidence that suggests AI use is driving a “dramatic increase” in cases before US federal courts\.

##### Agents and Friction Reduction\.

[65](https://arxiv.org/html/2608.16603#bib.bib73)identify that agents may lower search, communication, and contracting costs in markets\.[6](https://arxiv.org/html/2608.16603#bib.bib74)theorize that agents may reduce the cost of knowledge formalization\.

##### AI and Government Interaction\.

[31](https://arxiv.org/html/2608.16603#bib.bib1)measure the accessibility of government service interfaces to AI agents\.[75](https://arxiv.org/html/2608.16603#bib.bib52)propose using LLMs to improve policy communication\.[40](https://arxiv.org/html/2608.16603#bib.bib75)benchmark LLMs on answering citizen queries\.[77](https://arxiv.org/html/2608.16603#bib.bib76)find that AI tools improve some citizen\-government interactions\.

##### Administrative Burden\.

[39](https://arxiv.org/html/2608.16603#bib.bib77)review how the transaction costs faced by citizens are studied under overlapping terms, including “sludge”, “red tape”, and “ordeals”\.[76](https://arxiv.org/html/2608.16603#bib.bib18)examines a case where burdens are maintained intentionally\.

## 10Conclusion

Agentic floodingis plausibly occurring across a broad range of government services, in the form of LLM\-generated text submitted via permissive interfaces\. Though it may intensify as agent capabilities improve, acute operational collapse appears unlikely\. However, itislikely that – at least in some cases – budgets will need adjusting or user experience will suffer, particularly where governments introduce friction to reduce service demand\.

On a more positive note, mitigating agentic flooding is an opportunity to transform government services more broadly: many governments recognize the value of AI\-enabled government interaction, and the digital\-government literature has long advocated similar structural reforms for the benefit of the public\([15](https://arxiv.org/html/2608.16603#bib.bib57)\)\. While friction\-based responses lock this potential further out of reach, proactive interventions to build capacity and redesign services could instead help to realize the user\-experience benefits that have long been possible\.

Governments can act now\. Our analysis suggests three near\-term, proactive measures whose value does not depend on tracking the progress of AI capabilities: auditing service vulnerability, developing digital\-identity integration strategies for the most exposed services, and clarifying the legal basis for resilience\-building measures\.

## Acknowledgements

CS acknowledges funding by the Dieter Schwarz Foundation via the Hertie School\. This work was partially completed during a seasonal fellowship at GovAI\.

## Appendix AResearch Scope

Table 4:Top\-level citizen service categories on the national service portals of the UK, US, and France, and the domains we derive from them\.##### Country selection\.

We select twelve countries to capture variation geographically and along two institutional dimensions\. First, digital government maturity, as measured by the OECD Digital Government Index\([50](https://arxiv.org/html/2608.16603#bib.bib30)\)\. We include three of the five highest\-rated countries: South Korea, Australia, and the United Kingdom\. Three further included countries score above0\.80\.8, indicating middle\-to\-high maturity: Estonia, Denmark, and France\. We include three that score under0\.80\.8, indicating lower maturity, e\.g\. higher reliance on non\-digital processes: Brazil, Japan, and the Netherlands\. Finally, we include three countries for which OECD data is not available: the United States, Germany, and Singapore\. Beside aggregate maturity, these states vary in specific technical implementation of digital government\. For example, four studied countries \(Estonia, Denmark, Singapore, South Korea\) have mandatory digital identity systems\.

Second, administrative tradition\([54](https://arxiv.org/html/2608.16603#bib.bib29)\): the list includes Anglo\-American \(UK, US, Australia\), Napoleonic \(France, Netherlands\), Germanic \(Germany\), Scandinavian \(Denmark\), East Asian \(Japan, South Korea, Singapore\), Latin American \(Brazil\), and post\-Soviet \(Estonia\) systems\. Legally, these countries are split almost evenly between common\-law and continental systems\.

We select this list of countries to lower the chance of obvious confounders from these two dimensions, especially for our analyses of risk and prevalence of flooding\. However, the list is evidently not representative, such that we do not make comparative or analytical claims about the distribution of flooding cases we find\. Future work analyzing the relevance of these factors for flooding would likely be valuable, as we discuss in[Section8](https://arxiv.org/html/2608.16603#S8)\.

##### Domain selection\.

Within each country we scan a fixed list of government interaction channels\. Existing taxonomies of the functions of government, e\.g\. COFOG by the[50](https://arxiv.org/html/2608.16603#bib.bib30), are not suitable for this because they map government responsibilities, not citizen interfaces\.

Instead, we compile a list of domains to scan by comparing the top\-level citizen service categories used by the national service portals of three of our sample countries: the UK, the US, and France \([Table4](https://arxiv.org/html/2608.16603#A1.T4)\)\. We map ten public service categories that return across all three of these portals\. However, as clarified in our definition \([Section2\.1](https://arxiv.org/html/2608.16603#S2.SS1)\), some points of government interaction are not public services\. We therefore append three domains for such channels: participatory processes, regulatory complaints and reporting, and transparency and access to information\. Of these, only complaints have a top\-level equivalent on any of the three portals\.

This methodology produces a plausible list of 13 government domains, validated in three of the 12 studied countries and enabling the broad and indicative scans for flooding we seek to make\. However, it is unlikely to be optimal\. More detailed refinement of this domain list could improve it, e\.g\. by comparing against existing taxonomies or scanning more countries\. Accordingly, we name this list of domains as a possible source of bias in[Section8](https://arxiv.org/html/2608.16603#S8)\.

## Appendix BCase Collection Pipeline

This appendix documents the pipeline used to collect the dataset presented in the paper\. We describe the three pipeline stages and the prompts used in each \([SectionB\.1](https://arxiv.org/html/2608.16603#A2.SS1)\), the schema each case is coded against \([SectionB\.2](https://arxiv.org/html/2608.16603#A2.SS2)\), and the scaffolding that orchestrates the pipeline \([SectionB\.3](https://arxiv.org/html/2608.16603#A2.SS3)\)\. The full case records, prompts, data schema, and harness code are available in the linked dataset repository\.

### B\.1Pipeline Stages

As depicted in[Figure2](https://arxiv.org/html/2608.16603#S3.F2), the pipeline has three stages: discovery, iteration, and finalization\. An LLM is called during discovery and iteration, and all LLM calls return structured responses\. We here describe each stage, provide summary statistics of cases processed in each, and give an overview of prompts and response schemas passed to the LLM\.

#### Discovery

From a domain prompt combining a country and a domain, e\.g\. “France, Health Services”, an LLM with web search access returns 20–30 candidate services to investigate\. We survey 12 countries and the same 13 domains in each \([AppendixA](https://arxiv.org/html/2608.16603#A1)\)\. Each candidate is returned as a stub, which is expanded into a case template deterministically\. We then manually review the candidates and decide whether to investigate or remove each: of22882288candidates, we pass22102210\(96%\) on to iteration\.

System prompt \(opening\)

You are a research assistant identifying candidate government services which may be experiencing "agentic flooding": surges of demand driven by AI use\. You will receive a user input consisting of a country, and a domain of government services to investigate in that country\. Your task is to return an initial list of 5\-30 specific government services or interfaces which could be flooded in that domain\. This is the first step in the pipeline, and these cases will be explored in detail afterward\.\[\.\.\.\]

User prompt

Country to analyze: \{country\}Government domain to analyze: \{domain\}\.

Response schema

Response: DiscoveryOutputDiscoveryCandidate:service\_name: strcountry: strgovernment\_body: strrationale: strDiscoveryOutput:candidates: list\[DiscoveryCandidate\]

#### Iteration

Each selected case then cycles between a refinement call and a review call, with no human involvement, until the review call recommends finalization or removal, or a maximum of five cycles is reached\.

##### Refinement\.

The refinement call receives the current state of the case, a prompt to improve it by coding it against the case template\([14](https://arxiv.org/html/2608.16603#bib.bib78)\), and any instructions from the previous iteration, review, or human reviewer\. It returns a set of JSON Patch operations to apply to the case file, a recommendation to either continue iterating or pass the case to review, and a description of next steps for further iteration, if any\. The harness applies the patches deterministically to the stored case and proceeds according to the recommendation\.

The LLM does not produce source URLs as part of its response\. Sources are extracted deterministically from the web search grounding metadata returned with each API response, and added to the case data\. For some elements of the case data, the model returns a text description of the source used to aid the human reviewer\. We append any URL returned by API response metadata, regardless of whether the model substantively used it\. We manually verified all cases listed in the final dataset, but the broader source list in each case file often contains irrelevant links\. We choose this approach as generating URLs as part of the response was very susceptible to hallucination\([27](https://arxiv.org/html/2608.16603#bib.bib33)\), while line\-specific grounding is not available when using structured outputs\.

System prompt \(opening\)

You are a research assistant helping build a dataset of "agentic flooding" cases for an academic paper\. Agentic flooding is when AI use by citizens drives a surge in demand on a government service, either in volume, in complexity, or both\. The paper asks whether this is happening, where, and what governments could do about it\.Your job is to improve a single case record by performing web search to fill in or correct fields, and to judge whether the case meets each of three inclusion criteria\.\[\.\.\.\]

User prompt

\{Case JSON\}

Response schema

Response: IterationOutputPatch:path: strvalue: Anyrationale: strIterationOutput:patches: list\[Patch\]summary\_of\_changes: strnext\_steps: strrecommendation: str

FieldTypeIdentificationcase\_idstringservice\_namestringcountrystringjurisdiction\_detailstringgovernment\_bodystringincludebooltldrstringanalysisstringTypes of Flooding\([Section2\.1](https://arxiv.org/html/2608.16603#S2.SS1)\)flooding\.qualitativeboolflooding\.quantitativeboolMechanism\(inclusion criterion 1\)mechanism\.meets\_inclusion\_criteriaboolmechanism\.inclusion\_explanationstringmechanism\.ai\_descriptionstringmechanism\.primary\_is\_text\_generationboolmechanism\.interface\_descriptionstringEvidence of change\(inclusion criterion 2\)evidence\.meets\_inclusion\_criteriaboolevidence\.explanationstringAI attribution\(inclusion criterion 3\)attribution\.meets\_inclusion\_criteriaboolattribution\.explanationstringattribution\.sourcestringattribution\.descriptionstringVolume time seriesvolume\.unitstringvolume\.annual\.\{2018\.\.2025\}numbervolume\.source\_notes\.\{2018\.\.2025\}stringGovernment response\([Section6](https://arxiv.org/html/2608.16603#S6)\)response\.action\_explicitly\_for\_aiboolresponse\.action\_for\_general\_surgeboolresponse\.descriptionstringresponse\.categoriesstring\[\]Sourcessources\[\]\.urlstringsources\[\]\.originstringPipeline metadatapipeline\.statusstringpipeline\.discovery\_countrystringpipeline\.discovery\_domainstringpipeline\.discovery\_rationalestringpipeline\.iteration\_countintegerpipeline\.last\_summary\_of\_changesstringpipeline\.last\_next\_stepsstringpipeline\.last\_review\_concernsstringpipeline\.human\_review\_commentstringpipeline\.last\_errorstringTable 5:Fields collected for each case, grouped by section\. Field descriptions are released with the dataset repository\.
##### Review\.

The review call skeptically evaluates the current case and returns a recommendation offinalize,remove, oriterate, with optional field\-level concerns\. It does not modify the case\. Its output becomes input either to a human decision, when it recommends finalization or removal, or to the next refinement call, when it recommends further iteration\.

System prompt \(opening\)

You are a skeptical reviewer evaluating a candidate case for an academic dataset of "agentic flooding" — situations where AI use by citizens drives surges in demand on government services\. A research assistant has iterated on this case across multiple passes\. Your job is to decide what happens to it next: finalize it into the dataset, send it back for more work, or remove it as unsuitable\.\[\.\.\.\]

User prompt

\{Case JSON\}

Response schema

Response: ReviewOutputReviewConcern:field\_path: strconcern: strReviewOutput:recommendation: strrationale: strspecific\_concerns: list\[ReviewConcern\]

#### Finalization

We manually review every case the review call recommends for finalization or removal\. Each is either added to the dataset, removed, or returned to iteration with concrete improvement instructions\. Of22102210cases, we take the recommended action in21822182\(98%\); the most recommended action is removal \(95% of cases\)\.

### B\.2Case Schema

Each case is stored as a JSON document\.[Table5](https://arxiv.org/html/2608.16603#A2.T5)lists the fields collected and their types, grouped by section\. The full schema, including a description of every field, is available in the dataset repository\. The schema mirrors the structure of our definition \([Section2](https://arxiv.org/html/2608.16603#S2)\)\. The three sections corresponding to our inclusion criteria \([Section3\.2](https://arxiv.org/html/2608.16603#S3.SS2)\) –mechanism,evidence,attribution– each contain ameets\_inclusion\_criteriaboolean, and a case is included in the dataset only if all three are true\.

### B\.3Scaffolding

##### Storage\.

Each case is stored as a JSON file on disk, named by itscase\_id, and versioned in the repository\. An append\-only log records the output of every iteration, including the patches applied, the LLM’s summary of changes, and the recommendation returned\.

##### Human\-in\-the\-loop\.

A minimal web interface allows a human reviewer \(the first author\) to make two decisions: \(i\) which candidates from discovery to advance to iteration, and \(ii\) for each case returned by the review stage, whether to finalize, exclude, or return to iteration with optional written feedback\.

##### Models, cost, and reproducibility\.

Alldiscoveryandreviewcalls usedGemini\-3\.1\-Pro\-Previewvia the Gemini API; alliterationcalls usedGemini\-3\.1\-Flash\-Lite\. Web search grounding was enabled fordiscoveryanditeration, but not forreviewcalls\. Temperature was set to 0 for all calls\. Total token cost for producing the final dataset was approximately$200\. Full prompts and orchestration code are included in the linked dataset repository to support replication\. Both the models used and the web search grounding component cannot be guaranteed to be deterministic; re\-running the pipeline may not produce an identical dataset\.

## References

- Ada Lovelace Institute \(2021\)Ada Lovelace InstituteAlgorithmic accountability for the public sector\.Technical reportAda Lovelace Institute\.Cited by:[§6\.2](https://arxiv.org/html/2608.16603#S6.SS2.SSS0.Px2.p3.1)\.
- Ballaet al\.\(2022\)S\. J\. Balla, R\. Bull, B\. C\.E\. Dooling, E\. Hammond, M\. Herz, M\. Livermore, and B\. S\. NoveckResponding to Mass, Computer\-Generated, and Malattributed Comments\.Administrative Law Review74\(1\),pp\. 95–160\.External Links:27177981,ISSN 00018368, 23269154Cited by:[§5\.1](https://arxiv.org/html/2608.16603#S5.SS1.p2.1),[§5\.2](https://arxiv.org/html/2608.16603#S5.SS2.p1.1),[§7](https://arxiv.org/html/2608.16603#S7.p1.1)\.
- Bhargava and Manoli \(2015\)S\. Bhargava and D\. ManoliPsychological Frictions and the Incomplete Take\-Up of Social Benefits: Evidence from an IRS Field Experiment\.American Economic Review105\(11\),pp\. 3489–3529\.External Links:ISSN 0002\-8282,[Document](https://dx.doi.org/10.1257/aer.20121493)Cited by:[§2\.2](https://arxiv.org/html/2608.16603#S2.SS2.SSS0.Px1.p4.1)\.
- Blundellet al\.\(2013\)R\. Blundell, A\. Bozio, and G\. LaroqueExtensive and Intensive Margins of Labour Supply: Work and Working Hours in the US, the UK and France\*\.Fiscal Studies34\(1\),pp\. 1–29\.External Links:ISSN 0143\-5671, 1475\-5890,[Document](https://dx.doi.org/10.1111/j.1475-5890.2013.00175.x)Cited by:[§2\.1](https://arxiv.org/html/2608.16603#S2.SS1.p6.1)\.
- Bonnettet al\.\(2007\)C\. J\. Bonnett, B\. N\. Peery, S\. V\. Cantrill, P\. T\. Pons, J\. S\. Haukoos, K\. E\. McVaney, and C\. B\. ColwellSurge capacity: a proposed conceptual framework\.The American Journal of Emergency Medicine25\(3\),pp\. 297–306\.External Links:ISSN 0735\-6757,[Document](https://dx.doi.org/10.1016/j.ajem.2006.08.011)Cited by:[§5\.2](https://arxiv.org/html/2608.16603#S5.SS2.SSSx2.p2.1)\.
- Brynjolfsson and Hitzig \(2025\)E\. Brynjolfsson and Z\. HitzigAI’s use of knowledge in society\.InThe Economics of Transformative AI,Cited by:[§9](https://arxiv.org/html/2608.16603#S9.SS0.SSS0.Px2.p1.1)\.
- Burwell and Propp \(2022\)F\. Burwell and K\. ProppDigital sovereignty in practice: The EU’s push to shape the new global economy\.Technical reportAtlantic Council\.Cited by:[§6\.2](https://arxiv.org/html/2608.16603#S6.SS2.SSS0.Px2.p3.1)\.
- Buttaboni and Floridi \(2026\)C\. Buttaboni and L\. FloridiA regulatory taxonomy of AI opacity in the EU: rethinking transparency, traceability, interpretability, and explainability\.AI and Ethics6\(1\),pp\. 100\.External Links:ISSN 2730\-5953, 2730\-5961,[Document](https://dx.doi.org/10.1007/s43681-025-00940-0)Cited by:[§5\.2](https://arxiv.org/html/2608.16603#S5.SS2.SSSx2.p4.1),[§7\.1](https://arxiv.org/html/2608.16603#S7.SS1.SSS0.Px3.p1.1)\.
- Canales \(2025\)S\. B\. Canales‘Addiction to secrecy’: opposition and crossbench slam Labor’s ‘undemocratic’ changes to FoI – including charging fees\.The Guardian\.External Links:ISSN 0261\-3077Cited by:[§1](https://arxiv.org/html/2608.16603#S1.p2.1)\.
- Christensen and Laegreid \(2005\)T\. Christensen and P\. LaegreidTrust in Government: The Relative Importance of Service Satisfaction, Political Factors, and Demography\.Public Performance & Management Review28\(4\),pp\. 487–511\.External Links:ISSN 1530\-9576,[Document](https://dx.doi.org/10.1080/15309576.2005.11051848)Cited by:[§2\.2](https://arxiv.org/html/2608.16603#S2.SS2.SSS0.Px4.p2.1),[§6\.1](https://arxiv.org/html/2608.16603#S6.SS1.SSS0.Px1.p3.1)\.
- Currie \(2004\)J\. CurrieThe Take Up of Social Benefits\.Technical reportTechnical Reportw10488,National Bureau of Economic Research,Cambridge, MA\.External Links:[Document](https://dx.doi.org/10.3386/w10488)Cited by:[§2\.2](https://arxiv.org/html/2608.16603#S2.SS2.SSS0.Px1.p4.1),[§5\.2](https://arxiv.org/html/2608.16603#S5.SS2.SSSx1.p6.1)\.
- DoNotPay \(2026\)DoNotPaySave Time and Money with DoNotPay\!\.Note:https://donotpay\.com/Cited by:[§2\.2](https://arxiv.org/html/2608.16603#S2.SS2.SSS0.Px3.p1.1)\.
- Drouinet al\.\(2024\)A\. Drouin, M\. Gasse, M\. Caccia, I\. H\. Laradji, M\. D\. Verme, T\. Marty, D\. Vazquez, N\. Chapados, and A\. LacosteWorkArena: how capable are web agents at solving common knowledge work tasks?\.InProceedings of the 41st International Conference on Machine Learning,pp\. 11642–11662\.External Links:ISSN 2640\-3498Cited by:[§2\.2](https://arxiv.org/html/2608.16603#S2.SS2.SSS0.Px1.p2.1),[§5\.2](https://arxiv.org/html/2608.16603#S5.SS2.SSSx1.p4.1)\.
- Dunivin \(2025\)Z\. O\. DunivinScaling hermeneutics: a guide to qualitative coding with LLMs for reflexive content analysis\.EPJ Data Science14\(1\)\.External Links:ISSN 2193\-1127,[Document](https://dx.doi.org/10.1140/epjds/s13688-025-00548-8)Cited by:[§B\.1](https://arxiv.org/html/2608.16603#A2.SS1.SSSx2.Px1.p1.1)\.
- Dunleavyet al\.\(2006\)P\. Dunleavy, H\. Margetts, S\. Bastow, and J\. TinklerDigital era governance: IT corporations, the state, and e\-Government\.Oxford University Press\.External Links:[Document](https://dx.doi.org/10.1093/acprof%3Aoso/9780199296194.001.0001)Cited by:[§10](https://arxiv.org/html/2608.16603#S10.p2.1),[§6\.2](https://arxiv.org/html/2608.16603#S6.SS2.p1.1)\.
- Dunleavy and Margetts \(2015\)P\. Dunleavy and H\. MargettsDesign Principles for Essentially Digital Governance\.InAmerican Political Science Association Annual Meeting,USA\.External Links:ISSN 2015\-0903Cited by:[§6\.2](https://arxiv.org/html/2608.16603#S6.SS2.SSS0.Px1.p1.1)\.
- Garnett \(2022\)H\. A\. GarnettRegistration Innovation: The Impact of Online Registration and Automatic Voter Registration in the United States\.Election Law Journal: Rules, Politics, and Policy21\(1\),pp\. 34–45\.External Links:ISSN 1533\-1296, 1557\-8062,[Document](https://dx.doi.org/10.1089/elj.2020.0634)Cited by:[§2\.2](https://arxiv.org/html/2608.16603#S2.SS2.SSS0.Px1.p3.1)\.
- GII \(2026\)GII§ 7 BHO \- Einzelnorm\.Note:https://www\.gesetze\-im\-internet\.de/bho/\_\_7\.htmlCited by:[§2\.2](https://arxiv.org/html/2608.16603#S2.SS2.SSS0.Px4.p1.1)\.
- GOV\.UK \(2026\)GOV\.UKBenefit expenditure and caseload tables: guidance and methodology\.Note:https://www\.gov\.uk/government/publications/benefit\-expenditure\-and\-caseload\-tables\-guidance\-and\-methodologyCited by:[§6\.1](https://arxiv.org/html/2608.16603#S6.SS1.SSS0.Px2.p2.1)\.
- Grimmelikhuijsen and Meijer \(2022\)S\. Grimmelikhuijsen and A\. MeijerLegitimacy of algorithmic decision\-making: six threats and the need for a calibrated institutional response\.Perspectives on Public Management and Governance5\(3\),pp\. 232–242\.External Links:ISSN 2398\-4910,[Document](https://dx.doi.org/10.1093/ppmgov/gvac008)Cited by:[§6\.2](https://arxiv.org/html/2608.16603#S6.SS2.SSS0.Px2.p3.1)\.
- Grogger and Karoly \(2005\)J\. T\. Grogger and L\. A\. KarolyWelfare reform: Effects of a decade of change\.Harvard University Press\.Cited by:[§6\.1](https://arxiv.org/html/2608.16603#S6.SS1.SSS0.Px2.p1.1)\.
- Guret al\.\(2023\)I\. Gur, H\. Furuta, A\. V\. Huang, M\. Safdari, Y\. Matsuo, D\. Eck, and A\. FaustA real\-world WebAgent with planning, long context understanding, and program synthesis\.InThe Twelfth International Conference on Learning Representations,Cited by:[§5\.2](https://arxiv.org/html/2608.16603#S5.SS2.SSSx1.p4.1)\.
- Halling and Baekgaard \(2024\)A\. Halling and M\. BaekgaardAdministrative Burden in Citizen–State Interactions: A Systematic Literature Review\.Journal of Public Administration Research and Theory34\(2\),pp\. 180–195\.External Links:ISSN 1053\-1858,[Document](https://dx.doi.org/10.1093/jopart/muad023)Cited by:[§2\.2](https://arxiv.org/html/2608.16603#S2.SS2.SSS0.Px1.p1.1)\.
- Hendrick \(2006\)R\. HendrickThe Role of Slack in Local Government Finances\.Public Budgeting & Finance26\(1\),pp\. 14–46\.External Links:ISSN 0275\-1100, 1540\-5850,[Document](https://dx.doi.org/10.1111/j.1540-5850.2006.00837.x)Cited by:[§5\.2](https://arxiv.org/html/2608.16603#S5.SS2.SSSx2.p5.1)\.
- Hicket al\.\(2009\)J\. L\. Hick, J\. A\. Barbera, and G\. D\. KelenRefining Surge Capacity: Conventional, Contingency, and Crisis Capacity\.Disaster Medicine and Public Health Preparedness3\(S1\),pp\. S59–S67\.External Links:ISSN 1935\-7893, 1938\-744X,[Document](https://dx.doi.org/10.1097/DMP.0b013e31819f1ae2)Cited by:[§5\.2](https://arxiv.org/html/2608.16603#S5.SS2.SSSx2.p2.1)\.
- Hood \(1991\)C\. HoodA Public Management for All Seasons?\.Public Administration69\(1\),pp\. 3–19\.External Links:ISSN 1467\-9299,[Document](https://dx.doi.org/10.1111/j.1467-9299.1991.tb00779.x)Cited by:[§6\.2](https://arxiv.org/html/2608.16603#S6.SS2.SSS0.Px1.p2.1)\.
- Huanget al\.\(2025\)L\. Huang, W\. Yu, W\. Ma, W\. Zhong, Z\. Feng, H\. Wang, Q\. Chen, W\. Peng, X\. Feng, B\. Qin, and T\. LiuA Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions\.ACM Transactions on Information Systems43\(2\),pp\. 1–55\.External Links:ISSN 1046\-8188,[Document](https://dx.doi.org/10.1145/3703155)Cited by:[§B\.1](https://arxiv.org/html/2608.16603#A2.SS1.SSSx2.Px1.p2.1),[§3\.3](https://arxiv.org/html/2608.16603#S3.SS3.SSS0.Px3.p1.1)\.
- Humlum and Vestergaard \(2025\)A\. Humlum and E\. VestergaardThe unequal adoption of ChatGPT exacerbates existing inequalities among workers\.Proceedings of the National Academy of Sciences122\(1\),pp\. e2414972121\.External Links:ISSN 0027\-8424, 1091\-6490,[Document](https://dx.doi.org/10.1073/pnas.2414972121)Cited by:[§5\.2](https://arxiv.org/html/2608.16603#S5.SS2.SSSx1.p3.1)\.
- Ilveset al\.\(2025\)L\. Ilves, M\. Kilian, T\. Peixoto, and O\. VelsbergThe Agentic State: How Agentic AI will Revamp 10 Functional Layers of Government and Public administration\.Technical reportThe Agentic State\.Note:Whitepaper, Global GovTech CentreCited by:[§7](https://arxiv.org/html/2608.16603#S7.p2.1)\.
- Iscenkoet al\.\(2026\)Z\. Iscenko, S\. Strand, Y\. Chen, G\. Aimard, M\. Codreanu, V\. Sampathkumar, A\. Imas, J\. Jacobs, E\. Muiruri, J\. Mateos\-Garcia, J\. J\. Ng, S\. Javed, J\. Martin, O\. Ajmeri, D\. Calin, A\. Kim, F\. C\. Millet, and J\. ManyikaGoogle’s AI & Economy ATLAS v1\.0: Mapping Gemini Usage in the Economy\.arXiv\.External Links:2608\.00038,[Document](https://dx.doi.org/10.48550/arXiv.2608.00038)Cited by:[§1](https://arxiv.org/html/2608.16603#S1.p1.1)\.
- Jordanet al\.\(2026\)L\. Jordan, T\. C\. Peixoto, and M\. Ramos\-MaquedaRADAR: readiness for AI discovery and agentic reach — measuring the accessibility of digital government services to AI systems\.PreprintThe Agentic State\.External Links:[Link](https://agenticstate.org/assets/RADAR.pdf)Cited by:[§5\.2](https://arxiv.org/html/2608.16603#S5.SS2.SSSx1.p4.1),[§9](https://arxiv.org/html/2608.16603#S9.SS0.SSS0.Px3.p1.1)\.
- Kaplan and Garrick \(1981\)S\. Kaplan and B\. J\. GarrickOn The Quantitative Definition of Risk\.Risk Analysis1\(1\),pp\. 11–27\.External Links:ISSN 0272\-4332, 1539\-6924,[Document](https://dx.doi.org/10.1111/j.1539-6924.1981.tb01350.x)Cited by:[§5\.2](https://arxiv.org/html/2608.16603#S5.SS2.p1.1)\.
- Kasirzadeh and Gabriel \(2025\)A\. Kasirzadeh and I\. GabrielCharacterizing AI Agents for Alignment and Governance\.arXiv\.External Links:2504\.21848,[Document](https://dx.doi.org/10.48550/arXiv.2504.21848)Cited by:[§2\.1](https://arxiv.org/html/2608.16603#S2.SS1.p1.1)\.
- Kasneciet al\.\(2023\)E\. Kasneci, K\. Sessler, S\. Küchemann, M\. Bannert, D\. Dementieva, F\. Fischer, U\. Gasser, G\. Groh, S\. Günnemann, E\. Hüllermeier, S\. Krusche, G\. Kutyniok, T\. Michaeli, C\. Nerdel, J\. Pfeffer, O\. Poquet, M\. Sailer, A\. Schmidt, T\. Seidel, M\. Stadler, J\. Weller, J\. Kuhn, and G\. KasneciChatGPT for good? On opportunities and challenges of large language models for education\.Learning and Individual Differences103,pp\. 102274\.External Links:ISSN 10416080,[Document](https://dx.doi.org/10.1016/j.lindif.2023.102274)Cited by:[§2\.2](https://arxiv.org/html/2608.16603#S2.SS2.SSS0.Px1.p2.1)\.
- A\. Khanna \(Ed\.\) \(2008\)A\. Khanna \(Ed\.\)Straight through processing for financial services: the complete guide\.Complete Technology Guides for Financial Services Series,Academic Press,Amsterdam Boston\.External Links:ISBN 978\-0\-08\-055484\-6 978\-0\-12\-466470\-8Cited by:[§5\.2](https://arxiv.org/html/2608.16603#S5.SS2.SSSx2.p2.1)\.
- Kwoka \(2021\)M\. B\. KwokaSaving the Freedom of Information Act\.Cambridge University Press,Cambridge\.External Links:ISBN 978\-1\-108\-71089\-3 978\-1\-108\-69763\-7Cited by:[§5\.2](https://arxiv.org/html/2608.16603#S5.SS2.SSSx2.p4.1),[§6\.1](https://arxiv.org/html/2608.16603#S6.SS1.SSS0.Px1.p1.1)\.
- Lindblom \(1959\)C\. E\. LindblomThe Science of ”Muddling Through”\.Public Administration Review19\(2\),pp\. 79–88\.External Links:973677,ISSN 0033\-3352,[Document](https://dx.doi.org/10.2307/973677)Cited by:[§7](https://arxiv.org/html/2608.16603#S7.p3.1)\.
- LTO \(2026\)LTOSozialgerichte: Überlastung durch KI\-Schriftsätze\.Legal Tribune Online\.Cited by:[§1](https://arxiv.org/html/2608.16603#S1.p2.1)\.
- Madsenet al\.\(2022\)J\. K\. Madsen, K\. S\. Mikkelsen, and D\. P\. MoynihanBurdens, Sludge, Ordeals, Red tape, Oh My\!: A User’s Guide to the Study of Frictions\.Public Administration100\(2\),pp\. 375–393\.External Links:ISSN 0033\-3298, 1467\-9299,[Document](https://dx.doi.org/10.1111/padm.12717)Cited by:[§9](https://arxiv.org/html/2608.16603#S9.SS0.SSS0.Px4.p1.1)\.
- Majithiaet al\.\(2026\)N\. Majithia, R\. Shinde, Z\. Chapman, P\. Trital, J\. Decker, M\. Maskey, E\. Simperl, and N\. ShadboltThe CitizenQuery Benchmark: A Novel Dataset and Evaluation Pipeline for Measuring LLM Performance in Citizen Query Tasks\.Note:https://arxiv\.org/abs/2602\.04064v1Cited by:[§9](https://arxiv.org/html/2608.16603#S9.SS0.SSS0.Px3.p1.1)\.
- Marroet al\.\(2026\)S\. Marro, A\. Chan, X\. Ren, L\. Hammond, J\. Wright, G\. Wanga, T\. Piccardi, N\. Campos, T\. South, J\. Yu, S\. Sengupta, E\. Sommerlade, A\. Pentland, P\. Torr, and J\. PeiPermission Manifests for Web Agents\.arXiv\.External Links:2601\.02371,[Document](https://dx.doi.org/10.48550/arXiv.2601.02371)Cited by:[§6\.1](https://arxiv.org/html/2608.16603#S6.SS1.SSS0.Px1.p1.1)\.
- Metzet al\.\(2024\)A\. Metz, C\. Casher, and J\. ClarkID4D Global Dataset Volume 2: Digital Identification Progress and Gaps\.Washington, DC: World Bank\.External Links:10986/41076,[Document](https://dx.doi.org/10.1596/41076)Cited by:[§7\.1](https://arxiv.org/html/2608.16603#S7.SS1.SSS0.Px2.p1.1)\.
- Moore \(1995\)M\. H\. MooreCreating public value: strategic management in government\.Harvard University Press,Cambridge, Mass\.External Links:ISBN 978\-0\-674\-17557\-0,LCCN JF1525\.E8 M66 1995Cited by:[§1](https://arxiv.org/html/2608.16603#S1.p1.1)\.
- Moynihanet al\.\(2015\)D\. Moynihan, P\. Herd, and H\. HarveyAdministrative Burden: Learning, Psychological, and Compliance Costs in Citizen\-State Interactions\.Journal of Public Administration Research and Theory25\(1\),pp\. 43–69\.External Links:ISSN 1053\-1858, 1477\-9803,[Document](https://dx.doi.org/10.1093/jopart/muu009)Cited by:[Table 1](https://arxiv.org/html/2608.16603#S1.T1),[§2\.2](https://arxiv.org/html/2608.16603#S2.SS2.SSS0.Px1.p1.1)\.
- Naghmouchiet al\.\(2025\)M\. Naghmouchi, M\. Laurent, C\. Levallois\-Barth, and N\. KaanichePerspectives on National Digital Identity Systems\.Blockchain: Research and Applications,pp\. 100429\.External Links:ISSN 20967209,[Document](https://dx.doi.org/10.1016/j.bcra.2025.100429)Cited by:[§5\.2](https://arxiv.org/html/2608.16603#S5.SS2.SSSx2.p2.1)\.
- National Employment Law Project \(2023\)National Employment Law ProjectID Verification\.Technical reportNational Employment Law Project\.Cited by:[§6\.1](https://arxiv.org/html/2608.16603#S6.SS1.SSS0.Px1.p1.1)\.
- Neumannet al\.\(2024\)O\. Neumann, K\. Guirguis, and R\. SteinerExploring artificial intelligence adoption in public organizations: a comparative case study\.Public Management Review26\(1\),pp\. 114–141\.External Links:ISSN 1471\-9037, 1471\-9045,[Document](https://dx.doi.org/10.1080/14719037.2022.2048685)Cited by:[§6\.2](https://arxiv.org/html/2608.16603#S6.SS2.SSS0.Px2.p1.1)\.
- Nichols and Zeckhauser \(1982\)A\. L\. Nichols and R\. J\. ZeckhauserTargeting Transfers through Restrictions on Recipients\.The American Economic Review72\(2\),pp\. 372–377\.External Links:1802361,ISSN 00028282Cited by:[§2\.2](https://arxiv.org/html/2608.16603#S2.SS2.SSS0.Px1.p4.1),[§6\.1](https://arxiv.org/html/2608.16603#S6.SS1.SSS0.Px2.p2.1)\.
- OECD \(2025a\)OECDGoverning with artificial intelligence: the state of play and way forward in core government functions\.OECD Publishing\.External Links:[Document](https://dx.doi.org/10.1787/795de142-en),ISBN 978\-92\-64\-43767\-8 978\-92\-64\-68405\-8 978\-92\-64\-81828\-6Cited by:[§6\.2](https://arxiv.org/html/2608.16603#S6.SS2.SSS0.Px2.p2.1)\.
- OECD \(2025b\)OECDGovernment at a Glance 2025\.Government at a Glance,OECD Publishing\.External Links:[Document](https://dx.doi.org/10.1787/0efd0bcd-en),ISBN 978\-92\-64\-31390\-3 978\-92\-64\-36321\-2 978\-92\-64\-59508\-8Cited by:[Appendix A](https://arxiv.org/html/2608.16603#A1.SS0.SSS0.Px1.p1.1),[Appendix A](https://arxiv.org/html/2608.16603#A1.SS0.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2608.16603#S3.SS1.p1.1)\.
- Osborne \(2018\)S\. P\. OsborneFrom public service\-dominant logic to public service logic: are public service organizations capable of co\-production and value co\-creation?\.Public Management Review20\(2\),pp\. 225–231\.External Links:ISSN 1471\-9037, 1471\-9045,[Document](https://dx.doi.org/10.1080/14719037.2017.1350461)Cited by:[§2\.1](https://arxiv.org/html/2608.16603#S2.SS1.p2.1)\.
- O’Brienet al\.\(2016\)D\. T\. O’Brien, D\. Offenhuber, J\. Baldwin\-Philippi, M\. Sands, and E\. GordonUncharted Territoriality in Coproduction: The Motivations for 311 Reporting\.Journal of Public Administration Research and Theory\.External Links:ISSN 1053\-1858, 1477\-9803,[Document](https://dx.doi.org/10.1093/jopart/muw046)Cited by:[§2\.2](https://arxiv.org/html/2608.16603#S2.SS2.SSS0.Px1.p3.1)\.
- O’Toole and Meier \(2010\)L\. J\. O’Toole and K\. J\. MeierIn Defense of Bureaucracy: Public managerial capacity, slack and the dampening of environmental shocks\.Public Management Review12\(3\),pp\. 341–361\.External Links:ISSN 1471\-9037, 1471\-9045,[Document](https://dx.doi.org/10.1080/14719030903286599)Cited by:[§5\.2](https://arxiv.org/html/2608.16603#S5.SS2.SSSx2.p5.1)\.
- M\. Painter and B\. G\. Peters \(Eds\.\) \(2010\)M\. Painter and B\. G\. Peters \(Eds\.\)Tradition and Public Administration\.Palgrave Macmillan UK\.External Links:[Document](https://dx.doi.org/10.1057/9780230289635)Cited by:[Appendix A](https://arxiv.org/html/2608.16603#A1.SS0.SSS0.Px1.p2.1),[§3\.1](https://arxiv.org/html/2608.16603#S3.SS1.p1.1)\.
- Patwardhanet al\.\(2025\)T\. Patwardhan, R\. Dias, E\. Proehl, G\. Kim, M\. Wang, O\. Watkins, S\. P\. Fishman, M\. Aljubeh, P\. Thacker, L\. Fauconnet, N\. S\. Kim, P\. Chao, S\. Miserendino, G\. Chabot, D\. Li, M\. Sharman, A\. Barr, A\. Glaese, and J\. TworekGDPval: Evaluating AI Model Performance on Real\-World Economically Valuable Tasks\.Note:https://arxiv\.org/abs/2510\.04374v1Cited by:[§6\.2](https://arxiv.org/html/2608.16603#S6.SS2.SSS0.Px2.p2.1)\.
- Peeters and Widlak \(2023\)R\. Peeters and A\. C\. WidlakAdministrative exclusion in the infrastructure\-level bureaucracy: The case of the Dutch daycare benefit scandal\.Public Administration Review83\(4\),pp\. 863–877\.External Links:ISSN 1540\-6210,[Document](https://dx.doi.org/10.1111/puar.13615)Cited by:[§2\.2](https://arxiv.org/html/2608.16603#S2.SS2.SSS0.Px4.p1.1),[§6\.1](https://arxiv.org/html/2608.16603#S6.SS1.SSS0.Px1.p3.1)\.
- Piedrahitaet al\.\(2026\)D\. G\. Piedrahita, D\. Banerjee, K\. Blin, P\. Cobben, G\. Corsi, X\. A\. Huang, C\. Li, S\. Majumder, P\. S\. Pandey, S\. Simko, I\. Strauss, T\. J\. Zhang, A\. Anderson, Y\. Bengio, M\. Bethge, R\. Grosse, K\. Helbig, D\. Lie, R\. Mallah, R\. Mihalcea, S\. Nesbitt, S\. Perry, P\. Resnick, S\. Russell, M\. Sachan, B\. Schölkopf, A\. Tang, and Z\. JinAI Poses Risks to Democratic and Social Systems\.Cited by:[§2\.2](https://arxiv.org/html/2608.16603#S2.SS2.SSS0.Px3.p2.1),[§9](https://arxiv.org/html/2608.16603#S9.SS0.SSS0.Px1.p1.1)\.
- Plesneret al\.\(2024\)A\. Plesner, T\. Vontobel, and R\. WattenhoferBreaking reCAPTCHAv2\.External Links:2409\.08831Cited by:[§6\.1](https://arxiv.org/html/2608.16603#S6.SS1.SSS0.Px1.p4.1)\.
- Pollitt \(2017\)C\. PollittPublic management reform: a comparative analysis \- into the age of austerity\.Fourth edition edition,Oxford University Press,New York, NY\.External Links:ISBN 978\-0\-19\-879517\-9 978\-0\-19\-879518\-6,LCCN JF1351 \.P665 2017Cited by:[§6\.2](https://arxiv.org/html/2608.16603#S6.SS2.SSS0.Px1.p2.1)\.
- Rust \(2024\)S\. RustThe Vexatious Litigant Problem\.Note:https://houstonlawreview\.org/article/127507\-the\-vexatious\-litigant\-problemCited by:[§2\.2](https://arxiv.org/html/2608.16603#S2.SS2.SSS0.Px3.p2.1),[§6\.1](https://arxiv.org/html/2608.16603#S6.SS1.SSS0.Px1.p1.1)\.
- Rystrømet al\.\(2026\)J\. Rystrøm, C\. Schmitz, K\. Korgul, J\. Batzner, and C\. RussellAgent Benchmarks Fail Public Sector Requirements\.InIASEAI 2026,External Links:2601\.20617,[Document](https://dx.doi.org/10.48550/arXiv.2601.20617)Cited by:[§6\.2](https://arxiv.org/html/2608.16603#S6.SS2.SSS0.Px2.p3.1)\.
- Sadeddineet al\.\(2025\)Z\. Sadeddine, W\. Maxwell, G\. Varoquaux, and F\. M\. SuchanekLarge Language Models as Search Engines: Societal Challenges\.arXiv\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2512.08946)Cited by:[§2\.1](https://arxiv.org/html/2608.16603#S2.SS1.p1.1)\.
- Schmitzet al\.\(2025\)C\. Schmitz, J\. Rystrøm, and J\. BatznerOversight structures for agentic AI in public\-sector organizations\.InProceedings of the 1st Workshop for Research on Agent Language Models \(REALM 2025\),E\. Kamalloo, N\. Gontier, X\. H\. Lu, N\. Dziri, S\. Murty, and A\. Lacoste \(Eds\.\),Vienna, Austria,pp\. 298–308\.External Links:ISBN 979\-8\-89176\-264\-0Cited by:[§6\.2](https://arxiv.org/html/2608.16603#S6.SS2.SSS0.Px2.p2.1)\.
- Shah and Levy \(2026\)A\. V\. Shah and J\. Y\. LevyAccess to Justice in the Age of AI: Evidence from U\.S\. Federal Courts\.Cited by:[§1](https://arxiv.org/html/2608.16603#S1.p2.1),[§3\.2](https://arxiv.org/html/2608.16603#S3.SS2.p1.1),[§9](https://arxiv.org/html/2608.16603#S9.SS0.SSS0.Px1.p1.1)\.
- Shahidiet al\.\(2025\)P\. Shahidi, G\. Rusak, B\. S\. Manning, A\. Fradkin, and J\. J\. HortonThe coasean singularity? Demand, supply, and market design with AI agents\.InThe Economics of Transformative AI,Cited by:[§9](https://arxiv.org/html/2608.16603#S9.SS0.SSS0.Px2.p1.1)\.
- Sharpet al\.\(2025\)M\. Sharp, O\. Bilgin, I\. Gabriel, and L\. HammondAgentic Inequality\.arXiv\.External Links:2510\.16853,[Document](https://dx.doi.org/10.48550/arXiv.2510.16853)Cited by:[§5\.2](https://arxiv.org/html/2608.16603#S5.SS2.SSSx1.p3.1)\.
- L\. Siciliani, M\. Borowitz, and V\. Moran \(Eds\.\) \(2013\)L\. Siciliani, M\. Borowitz, and V\. Moran \(Eds\.\)Waiting Time Policies in the Health Sector: What Works?\.OECD Health Policy Studies,OECD\.External Links:[Document](https://dx.doi.org/10.1787/9789264179080-en),ISBN 978\-92\-64\-17906\-6 978\-92\-64\-17908\-0Cited by:[§3\.2](https://arxiv.org/html/2608.16603#S3.SS2.p1.1),[§4](https://arxiv.org/html/2608.16603#S4.SS0.SSS0.Px3.p2.1)\.
- SSA \(2026\)O\. SSAState plans for medical assistance\.Note:https://www\.ssa\.gov/OP\_Home/ssact/title19/1902\.htmCited by:[§6\.1](https://arxiv.org/html/2608.16603#S6.SS1.SSS0.Px1.p4.1)\.
- Straubet al\.\(2024\)V\. J\. Straub, Y\. Hashem, J\. Bright, S\. Bhagwanani, D\. Morgan, J\. Francis, S\. Esnaashari, and H\. MargettsAI for bureaucratic productivity: Measuring the potential of AI to help automate 143 million UK government transactions\.arXiv\.External Links:2403\.14712,[Document](https://dx.doi.org/10.48550/arXiv.2403.14712)Cited by:[§6\.2](https://arxiv.org/html/2608.16603#S6.SS2.SSS0.Px2.p2.1)\.
- Van De Walle and Bouckaert \(2003\)S\. Van De Walle and G\. BouckaertPublic Service Performance and Trust in Government: The Problem of Causality\.International Journal of Public Administration26\(8\-9\),pp\. 891–913\.External Links:ISSN 0190\-0692, 1532\-4265,[Document](https://dx.doi.org/10.1081/PAD-120019352)Cited by:[§2\.2](https://arxiv.org/html/2608.16603#S2.SS2.SSS0.Px4.p2.1),[§6\.1](https://arxiv.org/html/2608.16603#S6.SS1.SSS0.Px1.p3.1)\.
- van Oorschot \(1991\)W\. van OorschotNon\-Take\-Up of Social Security Benefits in Europe\.Journal of European Social Policy1\(1\),pp\. 15–30\.External Links:ISSN 0958\-9287,[Document](https://dx.doi.org/10.1177/095892879100100103)Cited by:[§2\.2](https://arxiv.org/html/2608.16603#S2.SS2.SSS0.Px1.p4.1)\.
- Verhoeven \(2024\)T\. Verhoeven‘My how I have walked and worked to get those names’: Petitioning and the Women’s Suffrage Movement in the United States, 1908–1920\.\.Women’s History Review33\(5\),pp\. 669–691\.External Links:ISSN 0961\-2025,[Document](https://dx.doi.org/10.1080/09612025.2023.2270361)Cited by:[§5\.2](https://arxiv.org/html/2608.16603#S5.SS2.p1.1),[§7](https://arxiv.org/html/2608.16603#S7.p1.1)\.
- Wood and Lewis \(2017\)A\. K\. Wood and D\. E\. LewisAgency Performance Challenges and Agency Politicization\.Journal of Public Administration Research and Theory27\(4\),pp\. 581–595\.External Links:ISSN 1053\-1858, 1477\-9803,[Document](https://dx.doi.org/10.1093/jopart/mux014)Cited by:[§5\.2](https://arxiv.org/html/2608.16603#S5.SS2.p1.1)\.
- Yeet al\.\(2024\)A\. Ye, A\. Maiti, M\. Schmidt, and S\. J\. PedersenA Hybrid Semi\-Automated Workflow for Systematic and Literature Review Processes with Large Language Model Analysis\.Future Internet16\(5\),pp\. 167\.External Links:ISSN 1999\-5903,[Document](https://dx.doi.org/10.3390/fi16050167)Cited by:[§3\.3](https://arxiv.org/html/2608.16603#S3.SS3.SSS0.Px2.p1.1)\.
- Yunet al\.\(2024\)L\. Yun, S\. Yun, and H\. XueImproving citizen\-government interactions with generative artificial intelligence: novel human\-computer interaction strategies for policy understanding through large language models\.PLOS One19\(12\),pp\. e0311410\.External Links:ISSN 1932\-6203,[Document](https://dx.doi.org/10.1371/journal.pone.0311410)Cited by:[§6\.1](https://arxiv.org/html/2608.16603#S6.SS1.SSS0.Px1.p3.1),[§9](https://arxiv.org/html/2608.16603#S9.SS0.SSS0.Px3.p1.1)\.
- Zeckhauser \(2019\)R\. ZeckhauserStrategic Sorting: The Role of Ordeals in Health Care\.Technical reportTechnical Reportw26041,National Bureau of Economic Research,Cambridge, MA\.External Links:[Document](https://dx.doi.org/10.3386/w26041)Cited by:[§2\.2](https://arxiv.org/html/2608.16603#S2.SS2.SSS0.Px1.p4.1),[§9](https://arxiv.org/html/2608.16603#S9.SS0.SSS0.Px4.p1.1)\.
- Zhang and Nie \(2025\)R\. Zhang and L\. NieEnhancing Citizen\-Government Communication with AI: Evaluating the Impact of AI\-Assisted Interactions on Communication Quality and Satisfaction\.Note:https://arxiv\.org/abs/2501\.10715v3External Links:[Document](https://dx.doi.org/10.1177/15701255261427986)Cited by:[§9](https://arxiv.org/html/2608.16603#S9.SS0.SSS0.Px3.p1.1)\.
- Zhouet al\.\(2024\)S\. Zhou, F\. F\. Xu, H\. Zhu, X\. Zhou, R\. Lo, A\. Sridhar, X\. Cheng, T\. Ou, Y\. Bisk, D\. Fried, U\. Alon, and G\. NeubigWebArena: A Realistic Web Environment for Building Autonomous Agents\.arXiv\.External Links:2307\.13854,[Document](https://dx.doi.org/10.48550/arXiv.2307.13854)Cited by:[§4](https://arxiv.org/html/2608.16603#S4.SS0.SSS0.Px3.p2.1)\.

Similar Articles

LLMs in the Real World: Evaluating "AI" in Emergency Contexts

arXiv cs.AI

This paper examines the deployment of an LLM-based machine translation system for text-to-911 emergency services, highlighting common misconceptions and providing recommendations for stakeholders to ensure safe and effective use of AI in critical contexts.

Your LLM shouldn’t be your coding-agent workflow

Reddit r/openclaw

Argues that LLMs should be used for reasoning within coding-agent workflows, while deterministic infrastructure handles queues, state, retries, and recovery, so the process doesn't break when usage limits hit.