CDEP Agent: Connecting Meteorologically Detected Temporal Compound Events to Real-World Documentary Evidence

arXiv cs.AI Papers

Summary

CDEP Agent is an auditable LLM-agent framework that connects meteorologically detected compound drought-to-extreme-precipitation events to real-world documentary evidence, revealing that most such events go undocumented in current reporting systems.

arXiv:2608.28628v1 Announce Type: new Abstract: Compound drought-to-extreme-precipitation (CDEP) events are recognized in climate science as a growing driver of extreme impact, but whether this recognition carries over into real-world early warning and post-event documentation is unknown, so a meteorologically real CDEP event may pass with neither advance warning nor any later record. Here we present CDEP Agent, an auditable LLM-agent framework that tests this mismatch directly by linking CDEP candidates detected from meteorological reanalysis to real-world hazard and impact evidence across sources with different spatial scales, temporal resolutions, and reporting conventions. Using California as a case study, we identify 408 candidate CDEP events from ERA5 observations during 2021-2025 and evaluate each against the U.S. Drought Monitor, NOAA Storm Events, and public webpages along five dimensions: antecedent drought, extreme rainfall, local impact, hazard-impact attribution, and explicit drought-to-rainfall linkage. Only 34.3% of candidates are corroborated on both hazard components, and just 1.5% are ever explicitly linked to their antecedent drought, indicating that most meteorologically detected CDEP events go undocumented and their compound nature almost never enters the record at all. Our framework gives climate scientists a way to test physical event definitions against what actually gets documented, and gives social scientists, economists, and disaster-response agencies a provenance-linked evidence base for compound events that current warning and reporting systems largely fail to capture.
Original Article
View Cached Full Text

Cached at: 09/01/26, 12:33 PM

# CDEP Agent: Connecting Meteorologically Detected Temporal Compound Events to Real-World Documentary Evidence
Source: [https://arxiv.org/html/2608.28628](https://arxiv.org/html/2608.28628)
###### Abstract

Compound drought\-to\-extreme\-precipitation \(CDEP\) events are recognized in climate science as a growing driver of extreme impact, but whether this recognition carries over into real\-world early warning and post\-event documentation is unknown, so a meteorologically real CDEP event may pass with neither advance warning nor any later record\. Here we present CDEP Agent, an auditable LLM\-agent framework that tests this mismatch directly by linking CDEP candidates detected from meteorological reanalysis to real\-world hazard and impact evidence across sources with different spatial scales, temporal resolutions, and reporting conventions\. Using California as a case study, we identify 408 candidate CDEP events from ERA5 observations during 2021–2025 and evaluate each against the U\.S\. Drought Monitor, NOAA Storm Events, and public webpages along five dimensions: antecedent drought, extreme rainfall, local impact, hazard\-impact attribution, and explicit drought\-to\-rainfall linkage\. Only 34\.3% of candidates are corroborated on both hazard components, and just 1\.5% are ever explicitly linked to their antecedent drought, indicating that most meteorologically detected CDEP events go undocumented and their compound nature almost never enters the record at all\. Our framework gives climate scientists a way to test physical event definitions against what actually gets documented, and gives social scientists, economists, and disaster\-response agencies a provenance\-linked evidence base for compound events that current warning and reporting systems largely fail to capture\.

## Introduction

Climate change is intensifying hydroclimatic extremes worldwide, and a comprehensive understanding of their real\-world consequences is essential for early warning, adaptation planning, and disaster risk management\. A growing body of work points to drought and extreme\-precipitation events as among the clearest signatures of this shift\(Rodell and Li[2023](https://arxiv.org/html/2608.28628#bib.bib1)\), events that are not only becoming more frequent but also more severe\(Guet al\.[2023](https://arxiv.org/html/2608.28628#bib.bib2); Satohet al\.[2022](https://arxiv.org/html/2608.28628#bib.bib3)\)\. What is more striking is how quickly the climate system can swing between these two extremes: rapid transitions from drought to intense rainfall have grown markedly more common worldwide since 1980, a pattern that points to these transitions increasingly behaving as compound events in their own right, rather than as two unrelated hazards that happen to occur close together\(Qinget al\.[2023](https://arxiv.org/html/2608.28628#bib.bib4)\)\. The Intergovernmental Panel on Climate Change \(IPCC\)’s Sixth Assessment Report identifies compound extremes as an important category of climate risk, while recent events illustrate their severe impacts across regions\(IPCC[2021](https://arxiv.org/html/2608.28628#bib.bib5); Zscheischleret al\.[2025](https://arxiv.org/html/2608.28628#bib.bib6)\)\. Yet a compound event being meteorologically real is not the same as it being documented\. A drought\-to\-extreme\-precipitation transition can be clearly present in reanalysis data and still leave little trace in the records that governments, journalists, and disaster databases actually produce, because those records are written around discrete storms, discrete droughts, and discrete administrative jurisdictions, not around the meteorologically defined sequence connecting them\(Joneset al\.[2022](https://arxiv.org/html/2608.28628#bib.bib10); Liet al\.[2026](https://arxiv.org/html/2608.28628#bib.bib24)\)\. This is true even though compound events have by now been carefully formalized as a distinct class of hazard in their own right, with typologies and conceptual frameworks describing how they arise from the joint or sequential occurrence of drivers that individually might not be extreme\(Leonardet al\.[2014](https://arxiv.org/html/2608.28628#bib.bib8); Zscheischleret al\.[2020](https://arxiv.org/html/2608.28628#bib.bib7)\)\. Even recent efforts to make disaster archives more usable, such as GDIS\(Rosvold and Buhaug[2021](https://arxiv.org/html/2608.28628#bib.bib11)\)and Geo\-Disasters\(Teberet al\.[2026](https://arxiv.org/html/2608.28628#bib.bib12)\), add spatial precision to existing single\-hazard entries without asking whether a compound sequence was recognized as such in the first place\. Meanwhile, meteorological detection of compound events has grown increasingly capable\(Yinet al\.[2025](https://arxiv.org/html/2608.28628#bib.bib13)\), which only sharpens the question: when a detector flags hundreds of candidate compound events from meteorological data, how many of them were ever noticed, reported, or acted on by anyone outside the climate science community, and what determines whether an event clears that bar?

To address this, we present CDEP Agent, an auditable LLM\-agent framework that connects meteorologically detected CDEP candidates to the real\-world hazard and impact evidence that would indicate they were recognized\. Language agents show promise for scientific retrieval and synthesis\(Skarlinskiet al\.[2024](https://arxiv.org/html/2608.28628#bib.bib20); Asaiet al\.[2026](https://arxiv.org/html/2608.28628#bib.bib21)\)and domain\-specific decision support\(Xieet al\.[2025](https://arxiv.org/html/2608.28628#bib.bib23)\), but require rigorous task\-level evaluation\(Chenet al\.[2025](https://arxiv.org/html/2608.28628#bib.bib22)\)\. CDEP is one example of the broader class of compound climate events and consists of a prolonged drought period followed rapidly by intense rainfall within a short window\. Given a candidate specified by a county, an antecedent drought period, and a subsequent rainfall window, the agent retrieves and aligns evidence from sources such as the U\.S\. Drought Monitor, NOAA Storm Events, and public webpages, and evaluates each candidate along five explicit dimensions: whether the antecedent drought was recorded, whether the subsequent extreme rainfall was recorded, whether it produced a local impact, whether that impact was attributed to the hazard, and whether any source explicitly links the drought to the later extreme\-precipitation event\. Each decision retains its source provenance, supporting quotations, and date, location, and same\-event matching results, and a conservative labeling protocol ensures that an unresolved or negative outcome is recorded as such rather than treated as evidence that nothing happened\.

CDEP serves as a useful demonstration case for several reasons\. It is well characterized meteorologically\(Qinget al\.[2023](https://arxiv.org/html/2608.28628#bib.bib4); Denget al\.[2024](https://arxiv.org/html/2608.28628#bib.bib29)\), it admits a threshold\-based definition over meteorological data, and California in particular has seen a documented rise in this kind of precipitation volatility \(Swainet al\.[2018](https://arxiv.org/html/2608.28628#bib.bib18);Swainet al\.[2025](https://arxiv.org/html/2608.28628#bib.bib19)\)\. California’s own history illustrates how disruptive this kind of whiplash can be\. Between 2012 and 2017, the state endured a five\-year drought of record severity, only for the drought to end abruptly when an intense storm brought heavy rainfall across the region\. The sudden influx of water damaged the spillway at Oroville Dam so severely that roughly 190,000 residents downstream had to be evacuated\(Vahedifardet al\.[2017](https://arxiv.org/html/2608.28628#bib.bib16)\)\. Determining whether other CDEP candidates were recognized requires matching federal drought monitors, storm logs, and local news to the candidate’s place, time, and hazard\. Manual review does not scale to hundreds of candidates or support efficient reruns as records or detection thresholds change\.

Applying the CDEP Agent to 408 candidate drought\-to\-extreme\-precipitation transitions identified in California produces a structured, provenance\-linked record for each candidate; only 34\.3% are corroborated on both drought and extreme\-precipitation components by official or public records, and just 1\.5% are ever explicitly connected to their antecedent drought by any source\. This gap has direct consequences for the early\-warning, adaptation, and disaster\-response decisions that depend on knowing which hazards were actually recognized on the ground\. Yet where corroborating webpages exist, they typically capture substantive, real\-world consequences, such as flooding, evacuations, power outages, and property damage, showing that the gap lies in whether an event is documented at all, rather than in what can be recovered once it is\. We then use these structured, provenance\-linked records to ask what separates the candidates that entered the documentary record from those that did not: whether meteorological intensity and persistence or local impact predict documentary recognition, as distinct from cases where the gap simply reflects limitations in evidence availability or retrieval\. By making these documentary gaps explicit and auditable, our framework gives climate scientists a systematic way to test and refine meteorological candidate definitions against what actually gets recognized as societally relevant, and gives social scientists and economists a provenance\-linked evidence base for studying how and when compound climate events come to be known\.

## Methods

Our study consisted of two separate stages: reanalysis\-based candidate construction and agent\-based evidence review\. First, we identified historical CDEP candidates in California from ERA5\-Land using a fixed operational definition based on Precipitation–Evapotranspiration Index \(SPEI\-3\), a location\-specific extreme\-precipitation threshold, and a maximum drought\-to\-extreme precipitation interval of three months\. This stage was completed before the agent was applied\. It produced, for each candidate, a county, an antecedent drought period, a subsequent extreme precipitation event window, and the corresponding candidate characteristics\. The agent did not select or modify the meteorological thresholds used to construct these candidates\.

Given these fixed candidate records, the agent first matched each candidate to structured official records and conducted an initial search of public webpages\. It then checked whether qualifying evidence remained missing for the antecedent drought, the subsequent extreme precipitation event, local impacts, or an explicit connection between the drought and the later extreme precipitation event, and directed follow\-up searches to the missing items\. For each retrieved source, the agent extracted the reported dates, locations, drought or extreme precipitation conditions, documented consequences, stated relationships, and supporting quotations\. Deterministic checks then applied the time window and location requirements corresponding to the component under review\. For extreme precipitation\-event impacts and hazard–impact connections, the checks also required the hazard and consequence to refer to the same event\. Figure[1](https://arxiv.org/html/2608.28628#Sx2.F1)summarizes the complete workflow, from ERA5\-Land candidate construction through evidence retrieval, structured source review, candidate\-specific checks, and meteorological–documentary alignment\.

### Constructing Historical CDEP Candidates from Meteorological Data

![Refer to caption](https://arxiv.org/html/2608.28628v1/x1.png)Figure 1:Overview of the CDEP Agent workflow\. Meteorological candidates are first constructed from ERA5\-Land drought and extreme\-precipitation data\. For each fixed candidate, the agent retrieves official records and public webpages, structures the retrieved evidence, and applies temporal, spatial, same\-event, and quotation\-support checks\. The resulting documentary evidence is then aligned with the meteorological candidate to produce provenance\-linked assessments\.We characterized antecedent drought using the three\-month SPEI\-3 derived from ERA5\-Land\(Muñoz Sabater[2019](https://arxiv.org/html/2608.28628#bib.bib25)\)\. SPEI measures the balance between precipitation and atmospheric water demand over a selected accumulation period; the three\-month scale was used here to represent moisture conditions preceding the extreme precipitation transition\(Vicente\-Serranoet al\.[2010](https://arxiv.org/html/2608.28628#bib.bib26),[2020](https://arxiv.org/html/2608.28628#bib.bib27)\)\. For candidate construction, a grid cell was classified as being under drought when SPEI\-3 was below−1\-1, with more negative values indicating more severe drought\(Denget al\.[2024](https://arxiv.org/html/2608.28628#bib.bib29)\)\. This SPEI\-based rule was used only to construct the meteorological candidate set; the later antecedent\-drought assessment used independent documentary evidence from the U\.S\. Drought Monitor and qualifying public webpages\.

Extreme precipitation events were independently detected using daily precipitation data from ERA5\-Land\(Muñoz Sabater[2019](https://arxiv.org/html/2608.28628#bib.bib25)\), based on three\-day accumulated precipitation \(PR3\)\. For each grid cell, the P99 threshold was calculated from the 1985–2014 reference period, and precipitation episodes exceeding this threshold were classified as extreme rainfall events\. This is consistent with event\-based flood classifications that distinguish short\- and long\-rainfall processes\(Steinet al\.[2020](https://arxiv.org/html/2608.28628#bib.bib31)\), and with flood\-risk studies in China that use three\-day rainfall maxima as an indicator\(Liet al\.[2012](https://arxiv.org/html/2608.28628#bib.bib30)\)\. Together, these studies support anchoring the extreme\-precipitation leg of the compound event at a high, flood\-relevant percentile rather than a generic extreme\-rainfall definition\.

The study focuses on detecting extreme precipitation events within three months after drought, following prior CDEP work that evaluates one\-, two\-, and three\-month transition intervals\(Denget al\.[2024](https://arxiv.org/html/2608.28628#bib.bib29)\)\. This enables a more comprehensive identification of CDEP events\.

Together, the SPEI\-3 threshold, the location\-specific 99th percentile of three\-day precipitation, and the maximum three\-month drought\-to\-extreme precipitation interval define the operational CDEP specification used to construct the candidate set in this study\. These literature\-grounded choices provide one consistent way to identify candidate transitions; they are not intended to represent the only valid definition of CDEP\. Because the agent receives candidate locations and event windows as inputs, the same retrieval and evidence\-review procedure can be rerun on candidate sets constructed using alternative thresholds, although the present evaluation uses only the specification reported here\.

### Searching Public Records and Webpages

After the meteorological candidates had been constructed, each fixed candidate record entered the agent\-based retrieval and evidence\-review stage\. The retrieval stage began by aligning each candidate with structured records from the U\.S\. Drought Monitor and NOAA Storm Events using the available county, date, event\-type, and consequence fields\. It also conducted an initial public\-web search covering agency webpages, local and regional news, and retrospective reports\. Web queries combined the candidate location and relevant dates with terms describing the evidence being sought, such as drought, water shortage, drought\-related restrictions, heavy rainfall, flooding, road disruption, property damage, power interruption, evacuation, or agricultural loss\. The purpose was to find records describing the specified place and event period, rather than webpages that discussed drought or flooding only in general terms\.

After the initial search, the agent reviewed the evidence already collected\. Follow\-up searches were selected from four evidence targets\. An antecedent\-drought search was used when no qualifying source documented drought during the candidate drought window\. An extreme\-precipitation\-event search was used when no qualifying source documented the subsequent extreme precipitation event\. Impact\-focused searches were used when consequences of a matched extreme precipitation event were absent or incompletely described\. Once both hazard components were supported, a linkage\-focused search could be used when no source explicitly connected the antecedent drought to the later extreme precipitation event\. The searches therefore differed across candidates according to the evidence already available\.

Follow\-up queries were required to preserve the candidate location, the time period relevant to the selected search target, and the type of evidence being sought\. Drought\-focused queries used the antecedent drought window, whereas extreme precipitation\-event and impact\-focused queries used the subsequent extreme precipitation\-event window\. Linkage\-focused queries could include both periods\. Queries that omitted these elements, repeated an earlier search, remained too broad to identify the candidate event, or introduced unsupported storm names, locations, damage claims, casualties, or other event details were rejected before retrieval\. Previously visited webpages were also excluded from later searches\. These checks kept each search tied to information already present in the candidate record or in an accepted source\.

Search\-result snippets and automatically generated summaries were used only to locate possible sources\. A government record or webpage entered the evidence review only after the underlying record or page text had been retrieved\. The same maximum query and page limits were applied to every candidate so that the search procedure remained comparable across cases\. When a potentially relevant source could not be retrieved or reviewed, that limitation was recorded and carried into the final assessment rather than treated as evidence that no event or impact occurred\.

### Reviewing and Structuring Retrieved Sources

Retrieved sources were reviewed as descriptions of the candidate drought period or extreme precipitation event rather than as collections of matching words\. A page mentioning drought or flooding in the candidate county was not sufficient if it referred to conditions outside the applicable candidate window, described a different event or period, or reported only a forecast or warning\. Each source therefore had to be matched to the applicable candidate window, county, and hazard before it could support a decision\.

Structured government records were read from their fixed fields, including event dates, locations, event type, damage, casualties, and other recorded consequences\. Each webpage was processed separately together with the corresponding candidate information\. The language model determined whether the page contained no relevant record, one clearly identifiable drought period or extreme precipitation event that could be matched to the candidate, or multiple periods or events that could not be separated reliably\. For a clearly identified record, it extracted the reported dates, locations, observed drought or extreme precipitation conditions, local impacts, statements connecting those impacts to the extreme precipitation hazard, any explicit connection between the antecedent drought and the later extreme precipitation event, and the passages supporting those fields\.

Every webpage field used as evidence had to be supported by text that could be located in the retrieved page\. Information from different storms was not combined, and a hazard reported for one event could not be paired with an impact reported for another\. When a page contained several events that could not be separated, it was not used to support a positive decision\.

The language model handled the parts of the review that required interpretation of webpage prose, including identifying the event being described, distinguishing event dates from publication dates, locating the event, identifying observed hazards and consequences, and recognizing relationships explicitly stated by the source\. Deterministic checks then applied the candidate window corresponding to the component under review\. Drought evidence had to overlap the antecedent drought window, whereas extreme precipitation\-event evidence had to overlap the subsequent extreme precipitation\-event window\. In both cases, the source had to identify the candidate county or a clearly located place within it\. For local impact and hazard–impact assessments, the documented consequence and extreme precipitation hazard also had to refer to the same matched event\.

A publication or update date could not substitute for the date on which the event occurred\. Statewide reports, broad regional descriptions, and an agency’s jurisdiction were insufficient unless the source identified the candidate county or a location within it\. Prospective drought risk, general seasonal dryness, forecasts, watches, warnings, preparedness notices, and other statements about possible future conditions did not establish that the corresponding drought or extreme precipitation hazard had occurred during the applicable candidate window\. Similarly, possible losses, exposed asset values, funding announcements, and eligibility for assistance did not establish a local impact\. A single accepted webpage could support several assessments when separate passages documented the antecedent drought, the subsequent extreme precipitation hazard, a local consequence, the connection between that consequence and the hazard, or an explicit connection between the drought and the later extreme precipitation event\.

### Assigning the Five Decisions

The U\.S\. Drought Monitor uses D0 for abnormally dry conditions and D1–D4 for progressively more severe drought, ranging from moderate drought at D1 to exceptional drought at D4\. The antecedent\-drought assessment asked whether qualifying evidence documented drought in the candidate county during the antecedent drought window\. A U\.S\. Drought Monitor county record supported the assessment when the relevant reporting period showed D1\-or\-higher conditions\. A public webpage could also support the assessment when it explicitly documented drought, a drought\-related water shortage, or drought\-related restrictions in the candidate county during the candidate window and provided a verifiable supporting passage\. The subsequent\-extreme precipitation\-event assessment asked whether an observed rain\-, flood\-, or precipitation\-related event occurred in the candidate county during the extreme precipitation\-event window\. The impact assessment asked whether a documented local consequence was reported for the matched extreme precipitation event, such as property or crop damage, road disruption, power interruption, evacuation, debris flow, landslide, casualty, or rescue activity\. The hazard impact assessment asked whether the source connected a documented consequence to rainfall, flooding, or another matched extreme precipitation hazard in the same event\. The drought\-to\-extreme precipitation assessment required a source to explicitly connect the antecedent drought conditions to the subsequent rainfall, flooding, or extreme precipitation event\. The chronological occurrence of drought followed by rainfall was not sufficient by itself\.

Each assessment received a Yes, No, or Unresolved decision\. A Yes decision required at least one eligible source that passed all applicable date, location, event, and content checks\. Because the five questions were assessed separately, the same source could support more than one decision, while evidence relevant to one question did not automatically support another\.

An Unresolved decision was assigned when no source supported Yes, but the available material did not permit a complete judgment\. This included potentially relevant sources whose event, date, or location could not be matched conclusively, sources whose text was insufficient to evaluate, and cases in which the required review could not be completed\. A No decision was assigned only after the planned retrieval and review for that assessment had been completed, no source met the requirements for Yes, and no unresolved source prevented a decision\. A No decision therefore means that the completed review found no qualifying record under the stated evidence requirements; it does not establish that the event, impact, or relationship was absent in the real world\.

For every candidate, the system retained the official\-record identifiers, webpage queries and URLs, relevant source text, supporting quotations, extracted dates and locations, event\-matching results, recorded retrieval limitations, and the rule producing each of the five final decisions\. These retained materials allow each decision to be checked against the exact source information on which it was based\.

### Human Validation of Retrieved Records

To verify that the retrieved records actually documented the events detected from the observed data, we manually audited 40 random cases from the California candidate set\. For each case, the audit examined the antecedent drought and the subsequent extreme precipitation event separately\. A record counted as supporting an event only when its content matched the candidate county, fell within the relevant event window, and described the corresponding drought or extreme precipitation hazard rather than merely mentioning related terms\.

When the pipeline had identified a supporting record, reviewers examined the cited record and its supporting passage or structured entry, checking the location, dates, and hazard type against the case\. When the pipeline had not established support, reviewers searched public records for qualifying sources that retrieval may have missed\. For these searches, the corresponding pipeline materials were withheld until the reviewer submitted the search result\. This procedure tested both whether the records accepted by the pipeline matched the case and whether relevant records had been missed\.

Both reviewers confirmed every drought or extreme precipitation\-event record accepted by the pipeline and independently found the same two additional extreme precipitation\-event records\. Across all 40 audited cases, manual review matched the pipeline for all 40 drought assessments and 38 of 40 extreme precipitation\-event assessments\. This corresponds to agreement on 78 of 80 event assessments \(97\.5%\) and on both events in 38 of 40 cases \(95\.0%\)\. No event supported by the pipeline was rejected during manual review; both differences were extreme precipitation\-event records missed during retrieval\.

## Results

### Record Support for CDEP Candidates

The proportion of California grid cells affected by at least one CDEP event increased during 2000–2025, with a linear trend of1\.03±0\.521\.03\\pm 0\.52percentage points per year \(Figure[2](https://arxiv.org/html/2608.28628#Sx3.F2)\(a\)\)\. We focus the documentary review on 408 candidates identified during 2021–2025\. These candidates were distributed across diverse regions of California \(Figure[2](https://arxiv.org/html/2608.28628#Sx3.F2)\(b\), blue shade\) and had a median drought\-to\-extreme\-precipitation interval of approximately 35 days\.

We applied the same retrieval and labeling procedure to all 408 California drought\-to\-extreme\-precipitation candidates identified from the physical data\. For each candidate, we evaluated the antecedent drought and the subsequent extreme\-precipitation event separately using structured public records and public webpages\. We counted a candidate as corroborated only when both components received a Yes judgment\. At the candidate level, No means that at least one required component received a No judgment\. Unresolved means that neither component received a No judgment, but at least one component remained unresolved\.

Table 1:Record support for the 408 detected drought\-to\-extreme\-precipitation candidates\. A candidate is counted as corroborated only when records support both the antecedent drought and the subsequent extreme\-precipitation event\.AssessmentYesNoUnresolvedAntecedent drought297 \(72\.8%\)111 \(27\.2%\)0 \(0\.0%\)Subsequent extreme\-precipitation event187 \(45\.8%\)160 \(39\.2%\)61 \(15\.0%\)Both components140 \(34\.3%\)217 \(53\.2%\)51 \(12\.5%\)Records supported the antecedent drought in 297 candidates \(72\.8%\) and the subsequent extreme\-precipitation event in 187 candidates \(45\.8%\)\. Both components were supported in 140 candidates \(34\.3%\)\. These 140 cases form the set for which the retrieved records corroborated both parts of the detected drought\-to\-extreme\-precipitation sequence\. Support for only one component was not sufficient for a candidate to enter this set\.

Among the remaining candidates, 217 had no qualifying support for at least one required component\. Another 51 remained unresolved because the retrieved records did not support a firm extreme\-precipitation event judgment for the relevant county and event window\. We evaluated local impacts, hazard–impact connections, and explicit drought\-to\-extreme\-precipitation connections separately\. These results are examined below, with the complete Yes, No, and Unresolved distributions provided in the supplementary material\. Accepted documentary evidence was geographically uneven and occurred at only a subset of detected\-event locations \(Figure[2](https://arxiv.org/html/2608.28628#Sx3.F2)\(b\); orange points, with size proportional to the number of accepted records\)\.

![Refer to caption](https://arxiv.org/html/2608.28628v1/fig5.png)Figure 2:Historical expansion and recent spatial distribution of CDEP events in California\. \(a\) Annual percentage of California grid cells affected by at least one CDEP event during 2000–2025; the dashed line shows the linear trend \(1\.03±0\.521\.03\\pm 0\.52percentage points per year\)\. \(b\) CDEP event frequency during 2021–2025\. Orange points indicate locations with documentary evidence accepted in the final review, with point size proportional to the number of accepted records at each location\.
### Event Characteristics Associated with Public\-Web Corroboration

Because the agent found public webpages for only a subset of the meteorologically detected candidates, we tested whether documentary corroboration varied systematically with measurable characteristics of the detected events\. This analysis characterizes which candidates were more likely to appear in accessible public webpages under a common matching procedure\.

We next compared candidates with and without at least one public webpage that passed the date, location, event, and content checks described above\. We considered four characteristics of the detected sequence: extreme\-precipitation event duration, the interval between the end of the drought and the start of the extreme\-precipitation event, drought extremity, and antecedent drought duration\. Table[2](https://arxiv.org/html/2608.28628#Sx3.T2)reports the group summaries and odds ratios from a logistic regression containing all four characteristics\.

Extreme\-precipitation event duration showed the clearest difference\. Candidates with a qualifying webpage had extreme\-precipitation events lasting 3\.11 days on average, compared with 2\.66 days for the other candidates\. In the joint model, each additional day of extreme\-precipitation event duration was associated with 1\.40 times the odds of finding a qualifying webpage \(95% CI 1\.04–1\.88;p=0\.025p=0\.025\)\. The estimate changed little when standard errors were clustered by date\-defined extreme\-precipitation event group and remained similar after candidates with recorded retrieval or coverage failures were excluded \(OR=1\.37=1\.37\)\.

Candidates with qualifying webpages also had shorter drought\-to\-extreme\-precipitation intervals on average, 41\.9 days compared with 49\.1 days\. The adjusted odds ratio was 0\.72 for each additional 30 days \(95% CI 0\.47–1\.10\)\. Mean drought extremity was 3\.20 in the qualifying\-webpage group and 2\.62 in the other group; the adjusted odds ratio was 1\.21 per unit \(95% CI 0\.94–1\.57\)\. Antecedent drought duration averaged 7\.11 months in the matching\-webpage group and 6\.14 months in the other group, with an adjusted odds ratio of 1\.02 per three months \(95% CI 0\.76–1\.37\)\. Among the four characteristics, extreme\-precipitation event duration showed the clearest relation to whether a qualifying webpage was found\.

This result shows that meteorologically detected candidates were not equally represented in the accessible public\-web record; the observed difference may reflect event salience, reporting practices, source availability, or retrieval coverage\.

Table 2:Candidate characteristics by qualifying\-webpage status\.CharacteristicQualifying webpageNo qualifying webpageAdjusted OR \(95% CI\)Extreme\-precipitation duration \(days\)3\.112\.661\.40 \(1\.04–1\.88\)Transition interval \(days\)41\.949\.10\.72 \(0\.47–1\.10\)Drought extremity3\.20 \(2\.57\)2\.62 \(2\.53\)1\.21 \(0\.94–1\.57\)Antecedent drought duration \(months\)7\.116\.141\.02 \(0\.76–1\.37\)
Note\.Values are means except drought extremity, reported as mean \(median\)\. Transition interval is measured from drought termination to extreme\-precipitation onset\. Adjusted ORs are from the joint four\-predictor model and are scaled per 1 day, 30 days, 1 unit, and 3 months, respectively\.

### Local Impacts of Matched Extreme\-Precipitation Events

The record review extended beyond confirming that an extreme\-precipitation hazard occurred\. It identified local impacts for 78 candidates and hazard–impact connections for 76 candidates\. Qualifying public webpages were found for 35 of the candidates with local impacts and 33 of the candidates with hazard–impact connections\. These records connected the detected extreme\-precipitation events to concrete consequences in the affected county\.

Nearly nine in ten qualifying webpage records \(40 of 45\) documented at least one local impact, and more than four in five \(37 of 45\) explicitly connected that impact to rainfall or flooding\. The recorded consequences included flooding, road and traffic disruption, power outages, evacuations and emergency actions, debris flows and landslides, property damage, and rescue or medical response\. The qualifying webpages therefore did more than note heavy rainfall or flooding: they documented how the event affected residents, roads, homes, businesses, and infrastructure\.

This pattern appeared throughout the study period and across different parts of California\. Qualifying webpages were found in every year from 2021 through 2025, across 26 storm families and 26 counties, including both major urban counties and less\-populated inland and rural counties\. A smaller set of sources explicitly connected the antecedent drought to the later extreme\-precipitation event\. Together, these records described local consequences of the matched extreme\-precipitation events and, in six cases, explicitly connected the extreme\-precipitation event to antecedent drought conditions\.

## Discussion and Conclusion

We developed the CDEP Agent to determine whether drought\-to\-extreme\-precipitation transitions detected from meteorological data also appeared in official records and public webpages\. The agent connected each meteorological candidate to records of the antecedent drought, the subsequent extreme\-precipitation event, and any documented local consequences, while retaining the supporting sources, quotations, dates, locations, and event\-matching results\. Among 408 candidates identified in California, 140 were supported on both the drought and extreme\-precipitation components\. Only six candidates had a source that explicitly connected the antecedent drought conditions to the subsequent rainfall or flooding\. These findings answer the central question posed in this study: a drought\-to\-extreme\-precipitation sequence can be clearly identified from meteorological data even when the two hazards are documented separately, or when the relationship between them is not documented at all\.

The methodological contribution of CDEP Agent is not simply source retrieval, but a candidate\-specific evidence\-review process\. Each candidate fixes the county and event windows before retrieval; follow\-up searches target only missing evidence, while deterministic date, location, and same\-event checks constrain how retrieved text affects the final decisions\. By retaining source passages, matching results, and decision rules, the system produces auditable assessments rather than untraceable model judgments\.

The results also show what different records contribute to the review\. Extreme\-precipitation event duration had the clearest association with the presence of a qualifying public webpage: each additional day of the extreme\-precipitation event was associated with 1\.40 times the odds of finding a matching page, and the estimate remained similar in the sensitivity analyses\. The corresponding relationships for drought extremity, antecedent drought duration, and the drought\-to\-extreme\-precipitation interval were less clear\. When a qualifying webpage was available, however, it usually described more than the occurrence of heavy rainfall or flooding\. Forty of the 45 qualifying webpage records documented at least one local consequence, and 37 explicitly connected that consequence to rainfall or flooding\. These pages recorded effects including road disruption, power outages, evacuations, property damage, debris flows, and rescue activity\. Structured records were therefore useful for confirming drought and extreme\-precipitation hazards, while public webpages often supplied the local consequences needed to understand how a matched event affected residents and infrastructure\.

The broader impact of this work is to make meteorologically detected compound events easier to examine in relation to documented hazards and local consequences\. Climate scientists can apply the same evidence\-review procedure to candidate sets produced by different drought thresholds, rainfall thresholds, or transition windows and compare which definitions identify events that also appear in official or public records\. Because meteorological candidate construction is separate from evidence review, changing an event definition does not require rebuilding the entire review process\. Social scientists, economists, and disaster researchers can use the retained sources, quotations, dates, and locations to study documented consequences of the same set of meteorologically identified events\. The results also show a limitation of existing event records for compound\-event research: droughts and extreme\-precipitation events may both be recorded, while the connection between them remains absent\. A review procedure that preserves both components and their relationship can therefore support future datasets designed specifically for sequential compound climate events\.

The main limitation of the present study is its geographic and temporal scope\. Extending the analysis from California to the entire United States would require substantially more web retrieval, processing, and human review across states with different climates, county structures, government records, and local reporting practices\. Given the computational and review resources available for this study, we evaluated candidates in California during 2021–2025\. The exact support rates and regression estimates reported here should therefore not be assumed to remain unchanged in other regions or over longer periods\. We also evaluated one operational definition based on SPEI\-3, a location\-specific precipitation threshold, and a three\-month transition window\. Alternative definitions may produce different candidate sets and different documentary patterns\. The separation between candidate construction and evidence review nevertheless allows the same procedure to be applied in future work to other thresholds, longer study periods, additional regions, and other types of compound climate events\.

Meteorological detection identifies where a compound event may have occurred; it does not by itself show how that event was recorded or what local consequences were reported\. By connecting each candidate to inspectable hazard and impact evidence, the CDEP Agent provides a repeatable way to examine that missing part of the event record\.

## References

- A\. Asai, J\. He, R\. Shao,et al\.\(2026\)Synthesizing scientific literature with retrieval\-augmented language models\.Nature650,pp\. 857–863\.External Links:[Document](https://dx.doi.org/10.1038/s41586-025-10072-4)Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p2.1)\.
- Z\. Chen, S\. Chen, Y\. Ning,et al\.\(2025\)ScienceAgentBench: toward rigorous assessment of language agents for data\-driven scientific discovery\.InInternational Conference on Learning Representations \(ICLR 2025\),Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p2.1)\.
- S\. Deng, D\. Zhao, Z\. Chen,et al\.\(2024\)Global distribution and projected variations of compound drought\-extreme precipitation events\.Earth’s Future12\(7\),pp\. e2024EF004809\.External Links:[Document](https://dx.doi.org/10.1029/2024EF004809)Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p3.1),[Constructing Historical CDEP Candidates from Meteorological Data](https://arxiv.org/html/2608.28628#Sx2.SSx1.p1.1),[Constructing Historical CDEP Candidates from Meteorological Data](https://arxiv.org/html/2608.28628#Sx2.SSx1.p3.1)\.
- L\. Gu, J\. Yin, P\. Gentine,et al\.\(2023\)Large anomalies in future extreme precipitation sensitivity driven by atmospheric dynamics\.Nature Communications14,pp\. 3197\.External Links:[Document](https://dx.doi.org/10.1038/s41467-023-39039-7)Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p1.1)\.
- IPCC \(2021\)Climate change 2021: the physical science basis\. chapter 11: weather and climate extreme events in a changing climate\.Technical reportWorking Group I contribution to the Sixth Assessment Report\.Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p1.1)\.
- R\. L\. Jones, D\. Guha\-Sapir, and S\. Tubeuf \(2022\)Human and economic impacts of natural disasters: can we trust the global data?\.Scientific Data9,pp\. 572\.External Links:[Document](https://dx.doi.org/10.1038/s41597-022-01667-x)Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p1.1)\.
- M\. Leonard, S\. Westra, A\. Phatak,et al\.\(2014\)A compound event framework for understanding extreme impacts\.WIREs Climate Change5\(1\),pp\. 113–128\.External Links:[Document](https://dx.doi.org/10.1002/wcc.252)Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p1.1)\.
- K\. Li, S\. Wu, E\. Dai, and Z\. Xu \(2012\)Flood loss analysis and quantitative risk assessment in China\.Natural Hazards63\(2\),pp\. 737–760\.External Links:[Document](https://dx.doi.org/10.1007/s11069-012-0180-y)Cited by:[Constructing Historical CDEP Candidates from Meteorological Data](https://arxiv.org/html/2608.28628#Sx2.SSx1.p2.1)\.
- N\. Li, W\. Thiery, S\. Zahra,et al\.\(2026\)Wikimpacts 1\.0: a new global climate impact database based on automated information extraction from Wikipedia\.Natural Hazards and Earth System Sciences26\(6\),pp\. 2609–2636\.External Links:[Document](https://dx.doi.org/10.5194/nhess-26-2609-2026)Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p1.1)\.
- J\. Muñoz Sabater \(2019\)ERA5\-land hourly data from 1950 to present\.Copernicus Climate Change Service Climate Data Store\.External Links:[Document](https://dx.doi.org/10.24381/cds.e2161bac)Cited by:[Constructing Historical CDEP Candidates from Meteorological Data](https://arxiv.org/html/2608.28628#Sx2.SSx1.p1.1),[Constructing Historical CDEP Candidates from Meteorological Data](https://arxiv.org/html/2608.28628#Sx2.SSx1.p2.1)\.
- Y\. Qing, S\. Wang, Z\.\-L\. Yang, and P\. Gentine \(2023\)Soil moisture\-atmosphere feedbacks have triggered the shifts from drought to pluvial conditions since 1980\.Communications Earth & Environment4,pp\. 254\.External Links:[Document](https://dx.doi.org/10.1038/s43247-023-00922-2)Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p1.1),[Introduction](https://arxiv.org/html/2608.28628#Sx1.p3.1)\.
- M\. Rodell and B\. Li \(2023\)Changing intensity of hydroclimatic extreme events revealed by GRACE and GRACE\-FO\.Nature Water1,pp\. 241–248\.External Links:[Document](https://dx.doi.org/10.1038/s44221-023-00040-5)Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p1.1)\.
- E\. L\. Rosvold and H\. Buhaug \(2021\)GDIS, a global dataset of geocoded disaster locations\.Scientific Data8,pp\. 61\.External Links:[Document](https://dx.doi.org/10.1038/s41597-021-00846-6)Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p1.1)\.
- Y\. Satoh, K\. Yoshimura, Y\. Pokhrel,et al\.\(2022\)The timing of unprecedented hydrological drought under climate change\.Nature Communications13,pp\. 3287\.External Links:[Document](https://dx.doi.org/10.1038/s41467-022-30729-2)Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p1.1)\.
- M\. D\. Skarlinski, S\. Cox, J\. M\. Laurent, J\. D\. Braza, M\. M\. Hinks, M\. J\. Hammerling, M\. Ponnapati, S\. G\. Rodriques, and A\. D\. White \(2024\)Language agents achieve superhuman synthesis of scientific knowledge\.Note:arXiv:2409\.13740External Links:[Document](https://dx.doi.org/10.48550/arXiv.2409.13740)Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p2.1)\.
- L\. Stein, F\. Pianosi, and R\. Woods \(2020\)Event\-based classification for global study of river flood generating processes\.Hydrological Processes34\(7\),pp\. 1514–1529\.External Links:[Document](https://dx.doi.org/10.1002/hyp.13678)Cited by:[Constructing Historical CDEP Candidates from Meteorological Data](https://arxiv.org/html/2608.28628#Sx2.SSx1.p2.1)\.
- D\. L\. Swain, B\. Langenbrunner, J\. D\. Neelin, and A\. Hall \(2018\)Increasing precipitation volatility in twenty\-first\-century California\.Nature Climate Change8,pp\. 427–433\.External Links:[Document](https://dx.doi.org/10.1038/s41558-018-0140-y)Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p3.1)\.
- D\. L\. Swain, A\. F\. Prein, J\. T\. Abatzoglou,et al\.\(2025\)Hydroclimate volatility on a warming Earth\.Nature Reviews Earth & Environment6,pp\. 35–50\.External Links:[Document](https://dx.doi.org/10.1038/s43017-024-00624-z)Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p3.1)\.
- K\. Teber, M\. Weynants, F\. Gans, and M\. D\. Mahecha \(2026\)Geo\-disasters: geocoding climate\-related events in the international disaster database EM\-DAT\.Big Earth Data10\(1\),pp\. 303–318\.External Links:[Document](https://dx.doi.org/10.1080/20964471.2025.2576274)Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p1.1)\.
- F\. Vahedifard, A\. AghaKouchak, E\. Ragno, S\. Shahrokhabadi, and I\. Mallakpour \(2017\)Lessons from the Oroville dam\.Science355\(6330\),pp\. 1139–1140\.External Links:[Document](https://dx.doi.org/10.1126/science.aan0171)Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p3.1)\.
- S\. M\. Vicente\-Serrano, S\. Beguéría, and J\. I\. López\-Moreno \(2010\)A multiscalar drought index sensitive to global warming: the standardized precipitation evapotranspiration index\.Journal of Climate23,pp\. 1696–1718\.External Links:[Document](https://dx.doi.org/10.1175/2009JCLI2909.1)Cited by:[Constructing Historical CDEP Candidates from Meteorological Data](https://arxiv.org/html/2608.28628#Sx2.SSx1.p1.1)\.
- S\. M\. Vicente\-Serrano, F\. Domínguez\-Castro, T\. R\. McVicar,et al\.\(2020\)Global characterization of hydrological and meteorological droughts under future climate change: the importance of timescales, vegetation–CO2 feedbacks and changes to distribution functions\.International Journal of Climatology40\(5\),pp\. 2557–2567\.External Links:[Document](https://dx.doi.org/10.1002/joc.6350)Cited by:[Constructing Historical CDEP Candidates from Meteorological Data](https://arxiv.org/html/2608.28628#Sx2.SSx1.p1.1)\.
- Y\. Xie, B\. L\. Jiang, T\. Mallick, J\. Bergerson, J\. Hutchison, D\. Verner, J\. Branham, M\. R\. Alexander, R\. Ross, Y\. Feng, L\.\-A\. Levy, W\. Su, and C\. Taylor \(2025\)MARSHA: multi\-agent RAG system for hazard adaptation\.npj Climate Action4,pp\. 70\.External Links:[Document](https://dx.doi.org/10.1038/s44168-025-00254-1)Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p2.1)\.
- C\. Yin, M\. Ting, K\. Kornhuber, R\. M\. Horton, Y\. Yang, and Y\. Jiang \(2025\)CETD: a global compound events detection and visualisation toolbox and dataset\.Scientific Data12,pp\. 356\.External Links:[Document](https://dx.doi.org/10.1038/s41597-025-04530-x)Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p1.1)\.
- J\. Zscheischler, O\. Martius, S\. Westra,et al\.\(2020\)A typology of compound weather and climate events\.Nature Reviews Earth & Environment1,pp\. 333–347\.External Links:[Document](https://dx.doi.org/10.1038/s43017-020-0060-z)Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p1.1)\.
- J\. Zscheischler, C\. Raymond, Y\. Chen,et al\.\(2025\)Compound weather and climate events in 2024\.Nature Reviews Earth & Environment6\(4\),pp\. 240–242\.External Links:[Document](https://dx.doi.org/10.1038/s43017-025-00657-y)Cited by:[Introduction](https://arxiv.org/html/2608.28628#Sx1.p1.1)\.

## Supplementary Material

Table S1:Evidence required and evidence treated as insufficient for the five candidate assessments\.AssessmentEvidence required forYesEvidence that does not supportYesAntecedent droughtA U\.S\. Drought Monitor county record reported D1\-or\-higher conditions for the relevant reporting period, or a qualifying public webpage documented drought or drought\-related dry conditions, water shortage, or restrictions in the candidate county during the antecedent drought window and provided a verifiable supporting quotation\.General references to dry weather; prospective drought risk; conditions outside the candidate drought window; statewide or regional descriptions without a county\-level location match; or a drought\-focused query without source\-grounded evidence\.Subsequent extreme\-precipitation eventA rain\-, flood\-, or precipitation\-related event occurred during the candidate extreme\-precipitation event window and in the candidate county\. Evidence from a public webpage also required a verifiable quotation describing the event\.Forecasts, watches, warnings, or preparedness notices; publication dates without corresponding event dates; events outside the candidate window; or statewide or regional descriptions without a county\-level location match\.Local impactA documented consequence was reported for the matched extreme\-precipitation event, such as property or crop damage, road disruption, power interruption, evacuation, debris flow, landslide, casualty, rescue activity, or another local impact\.Potential losses; exposed asset values; assistance eligibility; funding announcements; preparedness notices; actions that were only proposed or recommended; or statements that impacts were possible without documenting an observed consequence\.Hazard–impact connectionA consequence was recorded for the same NOAA event, or a public webpage explicitly connected a local impact to rainfall, flooding, or another matched extreme\-precipitation hazard\.For public webpages, a hazard and an impact that appeared in the same document but were not explicitly connected; references to different events; or a connection inferred by combining separate sources\.Explicit drought\-to\-extreme\-precipitation connectionA webpage explicitly connected the antecedent drought or dry conditions to the subsequent rainfall, flooding, or extreme\-precipitation event\.A drought record followed by an extreme\-precipitation event record; general climatic background; or a connection inferred by combining separate sources\.### Retrieval Planning and Query Controls

Each candidate entered the retrieval stage with a county, an antecedent drought period, a subsequent extreme\-precipitation event window, and the corresponding physical\-event information\. The system first matched the candidate against the U\.S\. Drought Monitor and NOAA Storm Events and stored the resulting official records in the candidate record\. It then conducted an initial public\-web search using the candidate location, event period, and terms describing the relevant hazard or consequence\.

After the initial pass, the agent examined the accepted evidence already available for the candidate and selected additional searches according to the information that remained unsupported\. A drought\-focused search was used when no qualifying source documented the antecedent drought\. An extreme\-precipitation hazard search was used when no qualifying source documented the subsequent extreme\-precipitation event\. An impact\-focused search was used when the available sources did not document a local consequence or provided only limited information about the consequences of a matched extreme\-precipitation event\. Once both the antecedent drought and the subsequent extreme\-precipitation event were supported, a linkage\-focused search could be used to locate a source that explicitly connected the dry period to the later rainfall or flooding\.

Table S2:Targeted follow\-up searches selected from each candidate’s current evidence record\. The same selection policy and retrieval limits were applied to all candidates; only the search targets relevant to the remaining evidence gaps were activated\.Evidence state after initial retrievalSearch targetInformation soughtNo qualifying source documented the antecedent droughtDrought hazardObserved drought or drought\-related dry conditions, water shortage, or restrictions in the candidate county during the antecedent drought window\.No qualifying source documented the subsequent extreme\-precipitation eventExtreme\-precipitation hazardObserved heavy rainfall, flooding, or another compatible extreme\-precipitation hazard in the candidate county during the extreme\-precipitation event window\.A matched extreme\-precipitation event was available, but its local consequences were absent or incompletely describedImpactProperty or crop damage, road disruption, power interruption, evacuation, debris flow, landslide, casualty, rescue activity, or another documented local consequence\.The antecedent drought and subsequent extreme\-precipitation event were both supported, but no source explicitly connected themExplicit drought\-to\-extreme\-precipitation connectionA source statement connecting antecedent drought or dry conditions to the subsequent rainfall, flooding, or extreme\-precipitation event\.Table S3:Complete decision distributions for the 408 drought\-to\-extreme\-precipitation candidates\.AssessmentYesNoUnresolvedAntecedent drought297 \(72\.8%\)111 \(27\.2%\)0 \(0\.0%\)Subsequent extreme\-precipitation event187 \(45\.8%\)160 \(39\.2%\)61 \(15\.0%\)Both drought and extreme\-precipitation components140 \(34\.3%\)217 \(53\.2%\)51 \(12\.5%\)Local impact78 \(19\.1%\)208 \(51\.0%\)122 \(29\.9%\)Hazard–impact connection76 \(18\.6%\)214 \(52\.5%\)118 \(28\.9%\)Explicit drought\-to\-extreme\-precipitation connection6 \(1\.5%\)336 \(82\.4%\)66 \(16\.2%\)Table S4:Joint logistic regression for the presence of at least one matching public webpage\. All four predictors were included in the same model\.PredictorCoefficientSEAdjusted OR95% CIpp\-valueExtreme\-precipitation event duration, per day0\.3360\.1501\.401\.04–1\.880\.025Drought\-to\-extreme\-precipitation interval, per 30 days−0\.331\-0\.3310\.2170\.720\.47–1\.100\.127Drought extremity, per unit0\.1910\.1311\.210\.94–1\.570\.144Antecedent drought duration, per 3 months0\.0220\.1501\.020\.76–1\.370\.884The query\-generation model received the candidate location, relevant dates, current evidence target, and the evidence already accepted for the case\. Every generated query was checked before execution\. A query had to retain the candidate county or another location already grounded in the case, include the relevant event period, and contain terms appropriate for the evidence being sought\. Queries were rejected if they were too broad to identify the candidate event, duplicated or closely repeated an earlier query, or introduced storm names, roads, damage amounts, casualties, declarations, or other specific details that were not present in the candidate information or accepted evidence\. URLs were normalized and deduplicated so that a page already considered in the initial pass was not opened again during follow\-up retrieval\.

The retrieval budget was fixed across candidates to keep the procedure comparable\. The initial stage used at most one query and opened at most two webpages\. The follow\-up stage used at most four additional queries and opened at most three additional webpages\. At most one follow\-up query could target the antecedent drought and at most one could target the subsequent extreme\-precipitation event\. Impact\-focused and linkage\-focused searches retained maximums of three and one queries, respectively, but all search types shared the same four\-query total\. When both hazard\-focused searches were activated, at most two follow\-up queries remained for impact and linkage searches combined\. Thus, each candidate used no more than five public\-web queries and five opened webpages\.

Search\-result snippets and provider\-generated summaries were used only to identify possible sources\. They did not enter the evidence record\. A webpage became eligible for review only after the underlying page text had been retrieved and stored\. Pages that could not be retrieved or did not contain enough text for review were recorded in the candidate trace rather than treated as evidence that the event or impact was absent\.

For each candidate, the running record stored the official evidence, generated queries, retrieved URLs, page\-level review results, accepted and unresolved evidence, previously visited pages, and the remaining retrieval budget\. This record allowed the follow\-up search to depend on evidence already obtained for that candidate and preserved the actions leading to the final assessments\.

### Page\-Level Review and Structured Evidence

Each retrieved webpage was reviewed separately together with the corresponding candidate information\. The page\-review model first determined whether the page contained no relevant record, one clearly identifiable drought period or extreme\-precipitation event that could be matched to the candidate, or several periods or events that could not be separated reliably\. Only a clearly matched record could contribute supporting evidence\.

For a potentially matching record, the model returned a structured record containing the assessment being reviewed, the reported dates and locations, the observed drought or extreme\-precipitation conditions relevant to that assessment, local impacts when applicable, stated relationships, and the passages supporting those fields\.

The extracted record then passed explicit candidate\-matching checks\. The reported dates had to overlap the candidate window corresponding to the assessment under review: the antecedent drought window for drought evidence and the subsequent extreme\-precipitation event window for extreme\-precipitation event evidence\. The source had to identify the candidate county or a clearly located place within that county; statewide descriptions, broad regional language, and an agency’s jurisdiction were not sufficient by themselves\. The page also had to describe an observed hazard rather than only a forecast, warning, preparedness notice, or statement of seasonal risk\.

The same\-event check required the hazard, consequence, and stated relationship to refer to one identifiable event\. Information from different storms was not combined\. When several events appeared on the same page but could not be separated reliably, the page was retained for review but did not support a positive decision\. A documented consequence and an extreme\-precipitation hazard appearing in the same page were also insufficient for the hazard–impact assessment unless the source explicitly connected them\.

The language model therefore handled source\-level interpretation, including event identification, date and location extraction, identification of observed hazards and consequences, and recognition of relationships explicitly stated in the source\. Deterministic checks enforced date overlap, county matching, same\-event consistency, quotation support, and the evidence requirements shown in Table[S1](https://arxiv.org/html/2608.28628#Sx5.T1)\. This division allowed webpage prose to be interpreted while preventing an extracted field from supporting a decision unless it satisfied the candidate\-specific checks\.

The query\-generation and webpage\-review stages used GPT\-5\.5\-2026\-04\-23\. Query generation used temperature 0 and a maximum output length of 2,048 tokens\. Page review used high reasoning effort and a maximum output length of 12,000 tokens\. Both stages returned structured outputs\. An unusable model response did not contribute evidence and was recorded as an incomplete review\.

### Assessment Aggregation and Coverage Handling

The accepted official and webpage records were aggregated separately for antecedent drought, subsequent extreme\-precipitation event, local impact, hazard–impact connection, and explicit drought\-to\-extreme\-precipitation connection\. Each source retained its source identifier, relevant text or structured fields, extracted dates and locations, supporting passages, and matching results\. A webpage could support more than one assessment when separate passages independently established the antecedent drought, the subsequent extreme\-precipitation hazard, a local consequence, the connection between the consequence and the hazard, or an explicit drought\-to\-extreme\-precipitation relationship\.

Webpage evidence for antecedent drought and the subsequent extreme\-precipitation event was kept separate\. Evidence accepted or left uncertain for the drought assessment affected only the drought decision, and evidence accepted or left uncertain for the extreme\-precipitation event assessment affected only the extreme\-precipitation event decision\. A webpage could support more than one assessment only when separate source\-grounded passages independently met the requirements for each assessment; the search target alone did not count as evidence\.

A Yes decision was assigned when at least one eligible source satisfied all requirements for the corresponding assessment\. Once qualifying evidence established Yes, uncertainty in another retrieved source did not remove that supporting evidence\. The candidate record nevertheless retained the additional source and its review status\.

When no source supported Yes, the system distinguished a completed negative review from a review that could not be completed\. A No decision required the relevant official and webpage checks to be complete, with no qualifying source, no potentially relevant source awaiting resolution, and no retrieval or processing failure that prevented evaluation\. The scope of this decision was limited to the completed retrieval and review procedure\.

An Unresolved decision was assigned when no source supported Yes but the available evidence did not permit a complete judgment\. This included a potentially relevant source with an uncertain event match, insufficient page text, several inseparable events, an unavailable required record, or a retrieval or processing failure\. An inaccessible or unreadable page could therefore not be used to produce No\.

The final candidate record stored the five decisions together with the supporting and contextual source identifiers, the strongest date and location matches, coverage information, recorded review limitations, and the rule that produced each decision\. These records were used to generate the decision distributions reported in Table[S3](https://arxiv.org/html/2608.28628#Sx5.T3)and permit each result to be traced back to its source material\.

### Representative Retrieval and Review Trace

To illustrate how the stages operated together, consider the Sonoma County candidate with an antecedent drought extending from August 2020 through September 30, 2021, followed by an extreme\-precipitation event window from October 24 to October 26, 2021\. The U\.S\. Drought Monitor record supported the antecedent\-drought assessment, while the initial structured records did not provide accepted evidence for the subsequent extreme\-precipitation event or its local impacts\.

The initial public\-web query searched for flooding and emergency\-management information in Sonoma County during the relevant period\. The initially retrieved materials did not contain a sufficiently matched event\. The agent therefore selected extreme\-precipitation hazard and impact\-focused follow\-up searches while preserving the candidate location and event dates\.

A subsequently retrieved local report described atmospheric\-river rainfall and flooding in Sonoma County and Santa Rosa on October 23–24, 2021\. The page reported water rescues, flooding of buildings and roads, and evacuation of residents\. It explicitly connected the rainfall and overflowing waterways to these consequences\. The same report described the event in the context of a drought\-stricken year and stated that the downpour replenished waterways depleted during the drought\.

The page\-review model extracted the event dates, Sonoma County locations, observed rainfall and flooding, local consequences, hazard–impact statements, drought\-to\-extreme\-precipitation statement, and the supporting passages\. The date, county, and same\-event checks accepted the event as a match to the candidate\. Together with the U\.S\. Drought Monitor record, the accepted evidence supportedYesdecisions for antecedent drought, subsequent extreme\-precipitation event, local impact, hazard–impact connection, and Explicit drought\-to\-extreme\-precipitation connection\. Other retrieved pages that did not pass the event\-matching requirements remained in the candidate trace but did not contribute to these positive decisions\.

### Complete Decision Distributions

Table[S3](https://arxiv.org/html/2608.28628#Sx5.T3)reports the complete Yes, No, and Unresolved distributions for all five assessments\. It also reports whether the retrieved records supported both the antecedent drought and the subsequent extreme\-precipitation event\. Percentages use all 408 candidates as the denominator\.

For the both\-components result, Yes required both the antecedent drought and the subsequent extreme\-precipitation event to receive Yes\. No indicates that at least one of the two components received No\. Unresolved indicates that neither component received No but at least one remained Unresolved\. Local impact, hazard–impact connection, and explicit drought\-to\-extreme\-precipitation connection were assessed separately and did not affect the both\-components result\.

### Regression Specification and Sensitivity Analyses

We modeled whether each candidate had at least one public webpage that passed the date, location, same\-event, and content requirements described above\. The analysis included all 408 candidates, and no candidate was excluded because of missing outcome or predictor values\. We fitted an unpenalized logistic regression with an intercept and included four candidate characteristics in the same model: drought extremity, antecedent drought duration, extreme\-precipitation event duration, and the drought\-to\-extreme\-precipitation interval\.

Extreme\-precipitation event duration was measured as the inclusive number of calendar days in the candidate rainfall window\. The drought\-to\-extreme\-precipitation interval was measured from the final day of the drought\-ending month to the start of the extreme\-precipitation event window\. Antecedent drought duration was the inclusive number of months between the beginning and end of the drought episode\. Drought extremity was defined as the absolute value of the minimum SPEI\-3 value during the antecedent drought episode, so larger values indicated more extreme drought\. Odds ratios are reported for a one\-day increase in extreme\-precipitation event duration, a 30\-day increase in the drought\-to\-extreme\-precipitation interval, a one\-unit increase in drought extremity, and a three\-month increase in antecedent drought duration\.

The confidence intervals andpp\-values in Table[S4](https://arxiv.org/html/2608.28628#Sx5.T4)were calculated using model\-based standard errors and normal\-Wald inference\. We conducted two sensitivity analyses for the association between extreme\-precipitation event duration and the presence of a matching public webpage\.

First, we recalculated the uncertainty using finite\-sample\-corrected cluster\-robust standard errors\. Candidates with identical extreme\-precipitation event start and end dates were assigned to the same group, producing 131 date\-defined storm groups\. The adjusted odds ratio for extreme\-precipitation event duration remained 1\.40, with a clustered 95% confidence interval of 1\.04–1\.88 andp=0\.026p=0\.026\.

Second, we refitted the same four\-variable model after excluding 68 candidates with recorded extreme\-precipitation event retrieval, source, access, or coverage failures\. The remaining analysis included 340 candidates\. The adjusted odds ratio for extreme\-precipitation event duration was 1\.37 \(95% CI 1\.02–1\.85;p=0\.038p=0\.038\)\. The direction and magnitude of the association were similar in both sensitivity analyses\.

Similar Articles

Agent-MD: Selective LLM Intervention with Event-Driven Escalation for Stateful GCMC--MD Campaigns

arXiv cs.AI

Agent-MD is a framework that selectively applies LLM reasoning to long-running molecular simulation campaigns, using a deterministic rule-based agent for routine tasks and event-triggered LLM review for exceptional conditions. Demonstrated in GCMC–MD water-vapor desorption simulations, it shows that auditable, reproducible scientific workflows can avoid placing every operation inside an LLM reasoning loop.

Evidence Core

Product Hunt

Evidence Core is a product that enables building live analytics using coding agents.

Didact: A Cross-Domain Capability Discovery System for Defence

arXiv cs.CL

Didact is a cross-domain capability discovery system for defence that integrates defence reports and policy documents with a knowledge graph from research publications, using a composite RAG pipeline for natural language conversations and an interactive Evidence Rail for source visualization.