EAGER: Enhancing Generative Event Extraction via Reinforcement Learning with Verifiable Rewards

arXiv cs.CL Papers

Summary

EAGER is a reinforcement learning framework that improves generative event extraction through fine-grained verifiable rewards and schema-contrastive advantage estimation, outperforming prompting, fine-tuning, and prior RL methods on seven benchmark datasets.

arXiv:2609.29230v1 Announce Type: new Abstract: End-to-end event extraction remains challenging for large language models as it requires simultaneous identification of event triggers, classification of event types, and extraction of schema-grounded argument spans. We present EAGER, a reinforcement learning framework for generative event extraction that combines fine-grained verifiable rewards with Schema-Contrastive Advantage Estimation to alleviate advantage collapse under sparse binary rewards. Our reward design explicitly targets structural validity, extraction accuracy, groundedness, coverage, over-generation, and span precision. Experiments across seven benchmark datasets show that EAGER consistently outperforms prompting, supervised fine-tuning, and prior reinforcement learning baselines, achieving a substantial improvement over the strongest prior method. Results demonstrate that task-aligned verifiable rewards and contrastive advantage estimation substantially improve structured extraction.
Original Article
View Cached Full Text

Cached at: 09/25/26, 09:18 AM

# EAGER: Enhancing Generative Event Extraction via Reinforcement Learning with Verifiable Rewards
Source: [https://arxiv.org/html/2609.29230](https://arxiv.org/html/2609.29230)
\\DeclareCaptionType

listing\[Figure\]\[List of Figures\]

Omar AdjaliAffiliation:German Research Center for Artificial Intelligence \(DFKI\), GermanyEmail:[omar\.adjali@dfki\.de](mailto:[email protected])Siting LiangAffiliation:German Research Center for Artificial Intelligence \(DFKI\), GermanyAffiliation:Carl von Ossietzky Universität Oldenburg, GermanyEmail:[siting\.liang@dfki\.de](mailto:[email protected])Omair Shahzad BhattiAffiliation:German Research Center for Artificial Intelligence \(DFKI\), GermanyEmail:[omairshahzad\.bhatti@dfki\.de](mailto:[email protected])Daniel SonntagAffiliation:German Research Center for Artificial Intelligence \(DFKI\), GermanyAffiliation:Carl von Ossietzky Universität Oldenburg, GermanyEmail:[daniel\.sonntag@dfki\.de](mailto:[email protected])

###### Abstract

End\-to\-end event extraction remains challenging for large language models as it requires simultaneous identification of event triggers, classification of event types, and extraction of schema\-grounded argument spans\. We present EAGER, a reinforcement learning framework for generative event extraction that combines fine\-grained verifiable rewards with Schema\-Contrastive Advantage Estimation to alleviate advantage collapse under sparse binary rewards\. Our reward design explicitly targets structural validity, extraction accuracy, groundedness, coverage, over\-generation, and span precision\. Experiments across seven benchmark datasets show that EAGER consistently outperforms prompting, supervised fine\-tuning, and prior reinforcement learning baselines, achieving a substantial improvement over the strongest prior method\. Results demonstrate that task\-aligned verifiable rewards and contrastive advantage estimation substantially improve structured extraction\.

## 1Introduction

Event extraction \(EE\) is a fundamental and challenging Information Extraction \(IE\) task that aims to identify event triggers, classify event types, and assign semantic roles to extracted arguments from unstructured text\. With the advent of large language models, a growing body of work[Li et al\. \(2023a\)](https://arxiv.org/html/2609.29230#bib.bib27);[Lu et al\. \(2021\)](https://arxiv.org/html/2609.29230#bib.bib28);[Hsu et al\. \(2022\)](https://arxiv.org/html/2609.29230#bib.bib29);[Ma et al\. \(2022\)](https://arxiv.org/html/2609.29230#bib.bib30);[Wang et al\. \(2023a\)](https://arxiv.org/html/2609.29230#bib.bib2);[Ren et al\. \(2023\)](https://arxiv.org/html/2609.29230#bib.bib31)has reformulated EE as a generative problem, leveraging sequence\-to\-sequence models to directly produce structured event representations under flexible annotation schemes\. More recent approaches have further adopted LLMs through prompt engineering and chain\-of\-thought reasoning[Cai et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib32);[Gao et al\. \(2023\)](https://arxiv.org/html/2609.29230#bib.bib37);[Hong and Liu \(2024\)](https://arxiv.org/html/2609.29230#bib.bib33);[Ma et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib34), aiming to reduce the computational overhead and domain\-specific overfitting associated with fully supervised training\. Nevertheless, as depicted in Figure[1](https://arxiv.org/html/2609.29230#S1.F1), even frontier general\-purpose LLMs still fall short of achieving competitive performance on end\-to\-end EE benchmarks\. A complementary line of work[Sainz et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib3);[Srivastava et al\. \(2025\)](https://arxiv.org/html/2609.29230#bib.bib9)has highlighted the critical role of augmenting training data with structured annotation guidelines, enabling instruction\-tuned LLMs to better adhere to predefined event schemas and improve extraction fidelity\. Despite these advances, significant performance gaps remain across diverse event types and domains\.

Figure 1:Average\-F1 performance of frontier LLMs on end\-to\-end event extraction across 7 datasets and domains\.To address these limitations,[Gao et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib10)proposed EventRL, which enhances LLM\-based event extraction via outcome\-supervised reinforcement learning \(RL\), optimizing the model based on the quality of final extracted event structures rather than relying solely on token\-level supervision\. This is further motivated by broader findings in the literature: comparative studies of supervised fine\-tuning \(SFT\) versus RL\-based fine\-tuning[Huan et al\. \(2025\)](https://arxiv.org/html/2609.29230#bib.bib35);[Chu et al\. \(2025\)](https://arxiv.org/html/2609.29230#bib.bib36)consistently show that RL\-tuned models exhibit stronger cross\-domain generalization and greater adaptability, whereas SFT\-trained models are more susceptible to catastrophic forgetting, often degrading previously acquired general capabilities\.

However, a key limitation in[Gao et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib10)and of existing training strategies more broadly, is that they provide onlycoarse\-grainedsupervision over extraction quality\. Supervised fine\-tuning optimizes next\-token likelihood rather than extraction quality directly, while standard RL objectives typically rely on format validity and task\-level accuracy alone\. In practice, these signals are often too sparse to distinguish among distinct classes of extraction errors\. Consequently, outcome\-level reward signals struggle to explicitly target the issues inherent to generative EE such as hallucinated triggers, unsupported arguments, over\-generation, and incomplete event coverage\. A further structural limitation arises in group\-based policy optimization: when all sampled completions within an optimization group are conditioned on the same prompt, including an identical set of negative schemas outputs tend to be homogeneous, producing near\-zero reward variance and uninformative gradient signal\.

In this work, we present EAGER, a post\-training framework that addresses both limitations through two complementary contributions\. First, we introduce a task\-aligned, decomposed reward framework that explicitly models distinct quality dimensions of event extraction:validity,extraction accuracy,groundedness,over\-generation control,coverage, andspan precision\. By decomposing the reward signal along these axes, our framework provides fine\-grained optimization guidance that better aligns RL with the constraints of generative EE\. Second, we proposeSchema\-Contrastive Advantage Estimation\(SCAE\), which alleviates advantage collapse by independently sampling distinct sets of negative schemas across completions within each optimization group\. This structural diversity induces variability in event type discrimination difficulty, increasing intra\-group reward variance and enabling more informative policy gradients throughout training\.

## 2Related Work

Large language models have increasingly been adapted to structured information extraction through instruction tuning and schema\-guided prompting[Jiao et al\. \(2023\)](https://arxiv.org/html/2609.29230#bib.bib1);[Lu et al\. \(2023\)](https://arxiv.org/html/2609.29230#bib.bib4)\. Prior work has explored instruction\-based IE frameworks such as InstructUIE[Wang et al\. \(2023a\)](https://arxiv.org/html/2609.29230#bib.bib2), annotation\-guided prompting in GoLLIE[Sainz et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib3), and large\-scale instruction resources like IEPile[Gui et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib11)to improve schema adherence and cross\-domain generalization\. Related work has also reformulated IE as code generation, where structured outputs benefit from the syntactic constraints of programming languages, as explored in CodeIE[Li et al\. \(2023b\)](https://arxiv.org/html/2609.29230#bib.bib7)and KnowCoder[Li et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib8)\.

### 2\.1Event Extraction with Large Language Models

Traditional event extraction approaches encompass both pipeline\-based and joint modeling paradigms\.[Huang et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib5)conducted a comprehensive reevaluation of major EE paradigms on a standardized benchmark and found that current LLMs still fall short on several core EE subtasks, underscoring the gap between general instruction\-following ability and task\-specific extraction precision\.[Chen et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib6)explored leveraging LLMs as automated annotators to generate high\-quality event labels, effectively bootstrapping downstream model training while reducing human annotation costs\. More recent work has examined how instruction tuning and annotation guidelines shape EE performance\.[Srivastava et al\. \(2025\)](https://arxiv.org/html/2609.29230#bib.bib9)demonstrated that providing detailed textual descriptions of event types and argument roles during instruction tuning yields improved generalization to low\-frequency and cross\-schema event types\. These studies highlight the persistent challenges of schema adherence, output validity, and domain generalization in event extraction\.

### 2\.2Preference and Reinforcement Learning for Structured Extraction

Beyond supervised instruction tuning, a growing body of work has explored aligning LLMs to structured extraction objectives through preference learning and reinforcement learning\. EventRL[Gao et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib10)introduced an RL\-based framework with outcome\-driven reward functions targeting event identification and argument extraction, achieving gains in structural fidelity and generalization to novel event types\. More recently,[Adjali et al\. \(2026\)](https://arxiv.org/html/2609.29230#bib.bib47)extended this line toward multi\-objective alignment for event extraction, combining task\-level, format, and retrieval rewards within a GRPO framework\. This contrasts with a closely related research direction which focuses on preference optimization tailored to IE tasks\. In particular, ADELIE[Qi et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib12)combines supervised fine\-tuning on curated instruction datasets with Direct Preference Optimization \(DPO\) using IE\-specific comparison pairs, explicitly targeting schema adherence, argument completeness, and format validity\. This approach achieves strong performance across closed, open, and on\-demand IE settings while preserving general reasoning capability, suggesting that decomposing alignment objectives along task\-relevant dimensions yields more reliable structured outputs\. These work show that RL\-based approaches to event extraction remain notably underexplored and suggest that fine\-grained alignment via reward design and preference learning can substantially improve the reliability of LLMs for structured IE, complementing purely supervised approaches\.

## 3Method

![Refer to caption](https://arxiv.org/html/2609.29230v1/latex/prompt.png)Figure 2:Structure of a prompt for end\-to\-end Event Extraction, comprising a task instruction a set of event schemas and the input text\.### 3\.1Task Formulation

Event extraction is a structured prediction task that encompasses four interdependent subtasks\.Trigger Identification \(TI\)detects event\-denoting spans within an input textXX\.Trigger Classification \(TC\)assigns a semantic event type to each identified trigger\.Argument Identification \(AI\)locates textual spans that participate as event arguments\.Argument Classification \(AC\)maps each identified argument span to a predefined semantic role\. Formally, letℰ=\{Ei\}i=1n\\mathcal\{E\}=\\\{E\_\{i\}\\\}\_\{i=1\}^\{n\}denote a predefined event schema, where each schemaEiE\_\{i\}specifies a set of permissible argument roles\. Given an input textXX, the objective is to produce a structured event representation:

Y=\{\(t,r,a\)\},Y=\\\{\(t,\\,r,\\,a\)\\\},\(1\)wheret∈Xt\\in Xdenotes a trigger span,rran argument role defined inℰ\\mathcal\{E\}, anda∈Xa\\in Xan argument span\. EE thus learns a mappingfθ:X→Yf\_\{\\theta\}:X\\rightarrow Ysubject to the schema constraints imposed byℰ\\mathcal\{E\}\.

### 3\.2Code\-Based Input and Output Representations

Following GoLLIE[Sainz et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib3), we cast EE as Python code generation \(Fig\.[2](https://arxiv.org/html/2609.29230#S3.F2)\)\. This leverages LLMs’ structural understanding of code[Wang et al\. \(2023b\)](https://arxiv.org/html/2609.29230#bib.bib13);[Li et al\. \(2023b\)](https://arxiv.org/html/2609.29230#bib.bib7)while providing a unified, less ambiguous representation for structured prediction[Sainz et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib3);[Srivastava et al\. \(2025\)](https://arxiv.org/html/2609.29230#bib.bib9)\. Python syntax also guarantees well\-formed outputs and simplifies parsing\. Each event schemaEi∈ℰE\_\{i\}\\in\\mathcal\{E\}is defined as a@dataclass\(Fig\.[6](https://arxiv.org/html/2609.29230#A1.F6)\), and extracted events are represented as class instances\.

### 3\.3Annotation Guideline Generation

Although code\-based schemas are compact and structured, they lack the semantic detail of annotation manuals\. Prior work shows that LLM\-generated guidelines can match manually written annotations for enriching event representations in instruction tuning[Srivastava et al\. \(2025\)](https://arxiv.org/html/2609.29230#bib.bib9)\. We therefore augment each Python class schema with natural\-language guidelines generated via instruction\-tuned prompting; details are provided in Appendix[B](https://arxiv.org/html/2609.29230#A2)\.

### 3\.4Prompt Structure

Each training instance is constructed as a structured prompt sequence as illustrated in Figure[2](https://arxiv.org/html/2609.29230#S3.F2):

P=I⊕EeG⊕X,P=I\\oplus E\_\{e\}^\{\\text\{G\}\}\\oplus X,\(2\)whereIIdenotes a natural language task instruction,EeGE\_\{e\}^\{\\text\{G\}\}is the gold event schema of typeeeaugmented with its automatically generated annotation guideline, andXXis the input text\. This formulation encourages the model to jointly attend to schema constraints, including the set of permissible argument roles foree, and the contextual semantics conveyed by the annotation guidelines\.

As shown in[Gui et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib11), to further improve event type discrimination, we additionally sample a set of negative schemas, i\.e\., schemas corresponding to event types not present inXX, and include them in the training prompt\. See Figure[7](https://arxiv.org/html/2609.29230#A8.F7)for an example of an input prompt\.

### 3\.5Post\-Training Framework

Reinforcement learning with verifiable rewards \(RLVR\), powered by algorithms such as GRPO[Shao et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib14)and DAPO[Yu et al\. \(2026\)](https://arxiv.org/html/2609.29230#bib.bib39), has demonstrated strong effectiveness across a range of reasoning tasks[Guo et al\. \(2025\)](https://arxiv.org/html/2609.29230#bib.bib41), however the application of RLVR to structured information extraction tasks such as end\-to\-end event extraction, which requires simultaneously identifying event triggers, classifying event types, and extracting schema\-grounded argument spans remains largely unexplored\. To bridge this gap, we investigate DAPO[Yu et al\. \(2026\)](https://arxiv.org/html/2609.29230#bib.bib39)to enhance the reasoning ability of LLMs for EE, leveraging its verifiable, schema\-based reward signal to guide structured output generation\.

Given an input promptXX, the model samples a group of candidate outputs and optimizes the DAPO objective

𝒥DAPO\(θ\)=𝔼\[1∑i\|oi\|∑i=1G∑t=1\|oi\|min\(rtiA^ti,clip\(rti,1−ϵlow,1\+ϵhigh\)A^ti\)\]\\mathcal\{J\}\_\{\\mathrm\{DAPO\}\}\(\\theta\)=\\mathbb\{E\}\\\!\\left\[\\frac\{1\}\{\\sum\_\{i\}\|o\_\{i\}\|\}\\sum\_\{i=1\}^\{G\}\\sum\_\{t=1\}^\{\|o\_\{i\}\|\}\\min\\\!\\Bigl\(r\_\{t\}^\{i\}\\hat\{A\}\_\{t\}^\{i\},\\right\.\\\\ \\left\.\\operatorname\{clip\}\(r\_\{t\}^\{i\},\\,1\{\-\}\\epsilon\_\{\\mathrm\{low\}\},\\,1\{\+\}\\epsilon\_\{\\mathrm\{high\}\}\)\\,\\hat\{A\}\_\{t\}^\{i\}\\Bigr\)\\right\]\(3\)whererti​\(θ\)=πθ​\(oi,t∣X,oi,<t\)/πθold​\(oi,t∣X,oi,<t\)r\_\{t\}^\{i\}\(\\theta\)=\{\\pi\_\{\\theta\}\(o\_\{i,t\}\\mid X,o\_\{i,<t\}\)\}/\{\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\(o\_\{i,t\}\\mid X,o\_\{i,<t\}\)\}is the importance sampling ratio andA^ti\\hat\{A\}\_\{t\}^\{i\}is the group\-normalized advantage\.

### 3\.6Reward\-Decoupled Normalization

Weighted reward sums can be dominated by high\-variance components, suppressing weaker signals and requiring manual tuning\. Following GDPO[Liu et al\. \(2026\)](https://arxiv.org/html/2609.29230#bib.bib46), we use Reward\-Decoupled Normalization \(RDN\), normalizing each reward independently within the sampled group asR~i\(m\)=Ri\(m\)−μ\(m\)σ\(m\)\+ϵ\\tilde\{R\}\_\{i\}^\{\(m\)\}=\\frac\{R\_\{i\}^\{\(m\)\}\-\\mu^\{\(m\)\}\}\{\\sigma^\{\(m\)\}\+\\epsilon\}, whereμ\(m\)\\mu^\{\(m\)\}andσ\(m\)\\sigma^\{\(m\)\}are the group mean and standard deviation for rewardmm\. The final reward isR^i=∑m=1MR~i\(m\)\\hat\{R\}\_\{i\}=\\sum\_\{m=1\}^\{M\}\\tilde\{R\}\_\{i\}^\{\(m\)\}\. This avoids manual weighting, balances reward contributions, and alleviates reward hacking by preventing any single component from dominating the optimization\.

### 3\.7SCAE: Schema\-Contrastive Advantage Estimation

Similar to GRPO, DAPO discards the value network of PPO[Schulman et al\. \(2017\)](https://arxiv.org/html/2609.29230#bib.bib15)by computing advantages directly from group\-level outcome rewards\. For each input textXXand its gold event annotationYY, we samples a group ofGGoutputs\{oi\}i=1G\\\{o\_\{i\}\\\}\_\{i=1\}^\{G\}from the old policyπθold\\pi\_\{\\theta\_\{\\text\{old\}\}\}, assigns binary outcome rewards\{Ri\}i=1G\\\{R\_\{i\}\\\}\_\{i=1\}^\{G\}, and estimates the per\-token advantage as the group\-normalized reward:

A^ti=Ri−mean⁡\(\{Ri\}i=1G\)std⁡\(\{Ri\}i=1G\),\\hat\{A\}\_\{t\}^\{i\}=\\frac\{R\_\{i\}\-\\mathrm\{mean\}\\\!\\left\(\\\{R\_\{i\}\\\}\_\{i=1\}^\{G\}\\right\)\}\{\\mathrm\{std\}\\\!\\left\(\\\{R\_\{i\}\\\}\_\{i=1\}^\{G\}\\right\)\},\(4\)However, a critical issue arises when all sampled completions within a group receive identical rewards leading to near\-zero policy gradients and stalled learning[Zhang et al\. \(2025\)](https://arxiv.org/html/2609.29230#bib.bib45);[Yu et al\. \(2026\)](https://arxiv.org/html/2609.29230#bib.bib39);[He et al\. \(2026\)](https://arxiv.org/html/2609.29230#bib.bib44)\. In standard DAPO sampling setting, allGGcompletions are conditioned on the same promptP=I⊕EeG⊕XP=I\\oplus E\_\{e\}^\{\\text\{G\}\}\\oplus X, including an identical set of negative schemas\. Since the negative schemas strongly constrain the model’s event type discrimination signal, sampled outputs tend to exhibit low diversity, such thatVar⁡\(\{Ri\}i=1G∣P\)≈0\\mathrm\{Var\}\(\\\{R\_\{i\}\\\}\_\{i=1\}^\{G\}\\mid P\)\\approx 0\. Combined with sparse binary rewards, this frequently produces homogeneous reward groups with uninformative gradient signal\.

We propose SCAE, a schema\-contrastive formulation of advantage estimation that structurally injects reward variance by varying the set of negative schemas across completions within the same optimization group, rather than holding the full prompt fixed\. During training, each promptPPincludes not only the gold schemaEeGE\_\{e\}^\{\\text\{G\}\}for the target event typeee, but also a set ofKKnegative schemas sampled fromℰ∖\{e\}\\mathcal\{E\}\\setminus\\\{e\\\}\. Formally, let𝒩=ℰ∖\{e\}\\mathcal\{N\}=\\mathcal\{E\}\\setminus\\\{e\\\}denote the pool of available negative schemas\. Rather than fixing a single negative set across all group members, we construct a*schema\-contrastive group*by independently sampling a distinct subset𝒮i⊂𝒩\\mathcal\{S\}\_\{i\}\\subset\\mathcal\{N\},\|𝒮i\|=K\|\\mathcal\{S\}\_\{i\}\|=K, for each completionii, yielding group\-specific prompts:

Pi=I⊕EeG⊕𝒮i⊕X,𝒮i​∼i\.i\.d\.​\(𝒩K\)P\_\{i\}=I\\oplus E\_\{e\}^\{\\text\{G\}\}\\oplus\\mathcal\{S\}\_\{i\}\\oplus X,\\quad\\mathcal\{S\}\_\{i\}\\overset\{\\text\{i\.i\.d\.\}\}\{\\sim\}\\binom\{\\mathcal\{N\}\}\{K\}\(5\)where𝒮i≠𝒮j\\mathcal\{S\}\_\{i\}\\neq\\mathcal\{S\}\_\{j\}fori≠ji\\neq jwith high probability when\|𝒩\|≫K\|\\mathcal\{N\}\|\\gg K\. The contrastive group is then defined as:

𝒢e=\{\(Pi,o^i\)\}i=1G,o^i∼πθ\(⋅∣Pi\)\\mathcal\{G\}\_\{e\}=\\left\\\{\(P\_\{i\},\\,\\hat\{o\}\_\{i\}\)\\right\\\}\_\{i=1\}^\{G\},\\quad\\hat\{o\}\_\{i\}\\sim\\pi\_\{\\theta\}\(\\cdot\\mid P\_\{i\}\)\(6\)where each sampled completiono^i\\hat\{o\}\_\{i\}receives an independent rewardRiR\_\{i\}computed against the gold annotationYY\. By varying the negative schema context across group members, different completions are exposed to different distractor event types, inducing variability in the difficulty of event type discrimination and thereby increasing intra\-group reward variance:

Var⁡\(\{Ri\}i=1G\)≫Var⁡\(\{Ri\}i=1G∣P\)\\mathrm\{Var\}\\\!\\left\(\\\{R\_\{i\}\\\}\_\{i=1\}^\{G\}\\right\)\\gg\\mathrm\{Var\}\\\!\\left\(\\\{R\_\{i\}\\\}\_\{i=1\}^\{G\}\\mid P\\right\)\(7\)
Advantages are then estimated by normalizing rewards within each schema\-contrastive group and substituted into the DAPO objective \(Eq\.[3](https://arxiv.org/html/2609.29230#S3.E3)\)\. \(See Algorithm[1](https://arxiv.org/html/2609.29230#alg1)for the detailed DAPO with SCAE pseudo\-code\.\)

Figure 3:Cumulative Mean ACR during trainingTo quantify the effectiveness of our schema\-contrastive group construction, we compute the Advantage Collapse Rate \(ACR\)[He et al\. \(2026\)](https://arxiv.org/html/2609.29230#bib.bib44), which measures the proportion of training groups exhibiting near\-zero reward variance:

ACR=1N​∑j=1N𝕀⁡\(σℛj<τ\),\\mathrm\{ACR\}=\\frac\{1\}\{N\}\\sum\_\{j=1\}^\{N\}\\mathbb\{I\}\\\!\\left\(\\sigma\_\{\\mathcal\{R\}\_\{j\}\}<\\tau\\right\),\(8\)whereσℛj\\sigma\_\{\\mathcal\{R\}\_\{j\}\}is the reward standard deviation within groupjjandτ\\tauis a small numerical threshold \(τ≈10−6\\tau\\approx 10^\{\-6\}\)\. An ACR close to 0 indicates that most groups produce informative gradient signals, while ACR≈1\\approx 1signals complete gradient stagnation\. As shown in Figure[3](https://arxiv.org/html/2609.29230#S3.F3), training with our Schema\-Contrastive Advantage Estimation \(w/ SCAE\) consistently achieves a substantially lower cumulative ACR compared to the standard fixed\-prompt baseline \(w/o SCAE\), confirming that varying the negative schema set across group members effectively alleviates reward homogeneity and produces more informative optimization signals throughout training\.

### 3\.8Verifiable Rewards Modeling

Table 1:Summary of the proposed verifiable rewards used in our RL framework for generative event extraction\. Each reward targets a distinct aspect: structural invalidity, extraction inaccuracy, hallucination, over\-generation, under\-coverage, and boundary imprecision\.Rather than relying solely on outcome reward using aggregate F1 scores such in[Gao et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib10), we define specialized rewards described in Table[1](https://arxiv.org/html/2609.29230#S3.T1)that target complementary aspects of model behavior, including output validity, extraction correctness, groundedness, prediction balance, and span precision\. This modular design provides more interpretable and fine\-grained optimization signals\. See Appendix[F](https://arxiv.org/html/2609.29230#A6)for more formal reward definitions\.

Table 2:Mean F1 results for end\-to\-end event extraction on the test split of the seven benchmark datasets\.Avg\.denotes the average F1 performance across datasets and EE subtasks\. Best results are shown in bold\. Green values indicate absolute improvement over the strongest prior baseline\.†denotes literature methods we re\-implemented\.Table 3:Ablation study on the contribution of each proposed individual reward to the global end\-to\-end event extraction performance\. Blue shading denotes improvement relative to the baseline configurationREE\+RFMTR\_\{\\text\{EE\}\}\+R\_\{\\text\{FMT\}\}, while amber shading denotes degradation\. Darker shading indicates larger absolute change\. Bold indicates the best score per dataset\. This color scheme logic applies to the subsequent tables\.Table 4:Impact of augmenting event schemes with annotation guidelines\.

## 4Experimental Setup

##### Baselines

We compare EAGER against representative methods spanning prompting, supervised fine\-tuning, and reinforcement learning paradigms for end\-to\-end event extraction\. We include proprietary large language models evaluated under few\-shot prompting, including o1[Jaech et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib43), GPT\-4o[Hurst et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib38), GPT\-5\.4\-mini, and GPT\-5\.4[Singh et al\. \(2025\)](https://arxiv.org/html/2609.29230#bib.bib42), to assess the effectiveness of general\-purpose reasoning\-oriented LLMs without task\-specific adaptation\. We additionally compare against GoLLIE[Sainz et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib3), a code\-oriented instruction framework for information extraction\. To isolate the contribution of reinforcement learning beyond standard instruction tuning, we fine\-tune GoLLIE\-7B and GoLLIE\-13B using SFT\. We also compare against the annotation\-guideline augmented instruction tuning framework of[Srivastava et al\. \(2025\)](https://arxiv.org/html/2609.29230#bib.bib9), which similarly enriches event schemas with LLM\-generated task descriptions\. Finally, we compare against prior RL\-based event extraction methods, including ADELIEDPO\{\}\_\{\\textnormal\{DPO\}\}[Qi et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib12), which applies direct preference optimization for generative information extraction, and EventRL[Gao et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib10), which optimizes event extraction performance using reinforcement learning with outcome\-based rewards\. These baselines represent the closest prior approaches to post\-training optimization for structured extraction\. For fair comparison, all reproduced baselines marked with†are evaluated using the same evaluation protocols and under a unified experimental setup\.

Table 5:Detailed results of event trigger and argument performance including TI, TC, AI, AC, AI\+, and AC\+\. Bold indicates the best score within each dataset\. Improvement/Degradations are highlighted relative to the baseline configurationREE\+RfmtR\_\{\\text\{EE\}\}\+R\_\{\\text\{fmt\}\}\.
##### Evaluation Datasets

To evaluate the proposed approach, we conducted experiments on 7 standard end\-to\-end EE datasets of different domains: WikiEvents[Li et al\. \(2021\)](https://arxiv.org/html/2609.29230#bib.bib16), PHEE[Sun et al\. \(2022\)](https://arxiv.org/html/2609.29230#bib.bib17), CASIE[Satyapanich et al\. \(2020\)](https://arxiv.org/html/2609.29230#bib.bib18), Genia2011[Kim et al\. \(2011\)](https://arxiv.org/html/2609.29230#bib.bib24), Genia2013[Kim et al\. \(2013\)](https://arxiv.org/html/2609.29230#bib.bib19), MLEE[Pyysalo et al\. \(2012\)](https://arxiv.org/html/2609.29230#bib.bib25), and M2E2[Li et al\. \(2020\)](https://arxiv.org/html/2609.29230#bib.bib26)\. We follow standard splits as in[Huang et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib5)where we use the “split 1” data split\. See Appendix[L](https://arxiv.org/html/2609.29230#A12)and Table[10](https://arxiv.org/html/2609.29230#A12.T10)for more details\.

##### Evaluation Metrics

Following previous work[Huang et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib5);[Sainz et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib3);[Srivastava et al\. \(2025\)](https://arxiv.org/html/2609.29230#bib.bib9), we report F1 scores for both trigger\- and argument\-level identification and classification subtasks including respectivelyTI,TC\),AIandAC\. We also report the attached versionAI\+,AC\+\.

We report Mean\-F1 over the six subtask metrics as the main aggregate measure of end\-to\-end extraction quality\. We additionally report full argument and trigger extraction results in Appendix[M](https://arxiv.org/html/2609.29230#A13)\.

## 5Results and Discussion

Figure 4:Training rewards dynamics\.### 5\.1Performance Analysis

As shown in Table[2](https://arxiv.org/html/2609.29230#S3.T2), despite their strong general reasoning capabilities, frontier models such as GPT\-4o \(9\.20\), o1 \(9\.52\), and GPT\-5\.4 \(9\.47\) evaluated under few\-shot prompting lag far behind task\-adapted smaller models\. Even the best performing GPT\-5\.4\-mini \(11\.70\) fails to approach the performance of the smallest GoLLIE\-7B baseline \(16\.35\), underscoring the difficulty of end\-to\-end event extraction and confirming the finding in[Huang et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib5)that generative event extraction requires task\-specific alignment beyond prompting\. Moreover, SFT provides marginal gains over zero\-shot suggesting that SFT saturates quickly and does not adequately address structured extraction in cross\-domain and \-schema settings\. Comparing SFT\-based approaches against our RL\-trained model \(\+14\.28\) confirms the advantage of reinforcement learning for generative EE\. Finally,EAGERachieves the highest average Mean\-F1 of 30\.79, surpassing the strongest prior RL baseline by \+5\.15\. The gain is consistent across datasets demonstrating that the proposed approach generalizes well across diverse domains and annotation schemes\.

##### Impact of Annotation Guidelines\.

Table[4](https://arxiv.org/html/2609.29230#S3.T4)shows that removing annotation guidelines fromEAGERreduces average F1 from 30\.79 to 22\.61 and GoLLIE\-7B from 16\.35 to 5\.81 which is consistent with the hypothesis that annotation guidelines provide the descriptive grounding needed for schema generalization\([Srivastava et al\., 2025](https://arxiv.org/html/2609.29230#bib.bib9);[Sainz et al\., 2024](https://arxiv.org/html/2609.29230#bib.bib3)\)\.

### 5\.2Ablation Study

We organize our ablation evaluation around the following research questions\.

##### RQ1: How can EE task\-specific reward signals be effectively leveraged for RL ?

While verifiable reward functions provide task\-specific supervision across complementary aspects of event extraction, Table[6](https://arxiv.org/html/2609.29230#S5.T6)shows that reward design alone is insufficient\. This is reflected in the substantial performance drop from 30\.79 to 22\.54 average F1 when SCAE is removed\. The same ablation applied to theREE\+RfmtR\_\{\\text\{EE\}\}\+R\_\{\\text\{fmt\}\}configuration \(27\.36 vs\. 23\.56\) confirms that the benefit of SCAE is not due to the richer reward set\. As shown in Figure[3](https://arxiv.org/html/2609.29230#S3.F3), by introducing structural diversity through varying negative schemas across sampled completions, SCAE increases advantage variance and enables more meaningful gradient updates, effectively activating the fine\-grained reward signals\. This suggests that without SCAE, grouped policy optimization frequently produces homogeneous outputs that prevent the model from effectively leveraging these training signals\. These findings highlight that in structured extraction tasks, informative reward modeling must be coupled with optimization strategies that preserve reward diversity\.

Table 6:Ablation study on the contribution of SCAE\. Improvement/Degradations are highlighted relative to the corresponding configuration with SCAE\.
##### RQ2: How does each reward contribute to extraction quality?

Starting from the baselineREE\+RfmtR\_\{\\text\{EE\}\}\+R\_\{\\text\{fmt\}\}\(27\.36 avg\. F1\), we can see in Table[3](https://arxiv.org/html/2609.29230#S3.T3)that adding any individual specialized reward improves overall performance, with the full combination of all six components reaching the best performance\. In contrast, Table[5](https://arxiv.org/html/2609.29230#S4.T5)shows that trigger metrics are relatively stable across reward configurations, as theREE\+RfmtR\_\{\\text\{EE\}\}\+R\_\{\\text\{fmt\}\}baseline already provides a reasonable foundation for trigger extraction while argument metrics are far more sensitive to reward design\. Table[7](https://arxiv.org/html/2609.29230#S5.T7)further reports the average gain of each additional reward over theREE\+RfmtR\_\{\\text\{EE\}\}\+R\_\{\\text\{fmt\}\}baseline, separately for trigger and argument subtasks, averaged across all seven datasets\.

Table 7:Average gain over theREE\+RfmtR\_\{\\text\{EE\}\}\+R\_\{\\text\{fmt\}\}baseline for trigger\-levelAvg\.Δ\\DeltaTrig\.\(TI \+ TC\) and argument\-levelAvg\.Δ\\DeltaArg\.\(∑\\sumAI,AC,AI\+,AC\+\) subtasks, averaged across seven datasets\.The results suggest that aggregate Mean\-F1 gains reported in Table[2](https://arxiv.org/html/2609.29230#S3.T2)are driven by argument\-level improvements\. Indeed, the full reward against theREE\+RfmtR\_\{\\text\{EE\}\}\+R\_\{\\text\{fmt\}\}configuration raises argument metrics by 7\.4 pts on average versus only 2\.1 pts for trigger metrics, suggesting the importance of EE tailored reward design\. Additionally, no single reward dominates both subtasks simultaneously:RcovR\_\{\\text\{cov\}\}andRspanR\_\{\\text\{span\}\}are the best single\-reward choices for triggers and arguments respectively, yet their combination in the full reward yields the best gains validating the proposed reward framework\. Figure[5](https://arxiv.org/html/2609.29230#S5.F5)reveals complementary reward learning dynamics:RfmtR\_\{\\text\{fmt\}\}converges rapidly whileRgrdR\_\{\\text\{grd\}\},RcovR\_\{\\text\{cov\}\}, andRspanR\_\{\\text\{span\}\}exhibit gradual improvements throughout training\. This asymmetry suggests that structural validity is a prerequisite condition that is quickly satisfied, after which the model shifts its optimization toward semantic precision\.

### 5\.3Error Analysis

![Refer to caption](https://arxiv.org/html/2609.29230v1/error_delta_heatmap.png)Figure 5:Error Change Relative to Baseline \(RE​E\+Rf​m​tR\_\{EE\}\+R\_\{fmt\}\)\.Figure[5](https://arxiv.org/html/2609.29230#S5.F5)reports error changes relative to the baseline \(REE\+RfmtR\_\{\\text\{EE\}\}\+R\_\{\\text\{fmt\}\}\), revealing that the dominant error category of the baseline is over\-generation\. It is worth noting that the list of error categories111Error categories are defined in Table[11](https://arxiv.org/html/2609.29230#A13.T11)\.is not exhaustive and does not cover all possible error types\. Moreover, all reward configurations substantially reduce extra arguments \(−2,086\-2\{,\}086to−3,917\-3\{,\}917\), extra roles \(−1,988\-1\{,\}988to−3,600\-3\{,\}600\), role confusion \(−1,693\-1\{,\}693to−2,534\-2\{,\}534\), and hallucinations \(−152\-152to−196\-196\)\. However, every configuration trades these gains for an increase in missing arguments \(\+430\+430to\+2,422\+2\{,\}422\) and missing roles \(\+222\+222to\+1,080\+1\{,\}080\), exposing a consistent precision\-recall tension across all reward designs\.RovrR\_\{\\text\{ovr\}\}achieves the largest total error reduction \(−8,315\-8\{,\}315\) by most aggressively suppressing false positives, whileRgrdR\_\{\\text\{grd\}\}most severely increases missing arguments \(\+2,422\+2\{,\}422\), reflecting over\-conservative extraction\. The full reward configuration \(All\) balances these competing pressures, yielding the second\-largest total reduction \(−6,639\-6\{,\}639\) while keeping missing argument growth lower \(\+956\+956\) than any individual precision\-focused reward alone\. Span boundary and parsing errors remain marginal and stable across all configurations, confirming that low\-level structural quality is not a primary bottleneck\.

## 6Conclusion

We presented EAGER, a RL framework for end\-to\-end event extraction that combines task\-aligned verifiable rewards with SCAE\. By decomposing reward supervision into complementary extraction objectives and introducing schema\-contrastive grouping to mitigate reward variance collapse, our approach provides more informative optimization signals for structured extraction tasks\. Experimental results across seven EE benchmark datasets demonstrate consistent improvements over strong baselines, particularly on argument\-level extraction quality\. Our findings further highlight the importance of combining fine\-grained reward modeling with optimization strategies that preserve reward diversity in generative information extraction\.

## 7Limitations

Our approach relies on automatically generated annotation guidelines whose quality may vary across LLMs, domains and schemas\. Investigating unified and hierarchical event schema across datasets may reduce annotation inconsistencies and improve cross\-domain transfer, enabling models to better generalize across heterogeneous event definitions\.

Additionally, we primarily assessed error categories that captures surface\-level extraction anomaly and does not fully model complex phenomena such as coreference\. Incorporating coreference modeling may help address cross\-sentence arguments and nested event structures that remain challenging for generative EE systems\.

While the proposed framework improves event extraction performance, balancing precision and recall remains challenging\. Investigating curriculum\-based reinforcenment learning could better balance precision and recall during optimization\.

Finally, our evaluation remains limited to English event extraction with predefined schemas, leaving multilingual and open\-schema settings for future work\.

## Acknowledgment

This work was funded by the Federal Ministry of Research, Technology and Space \(BMFTR\) under grant number 16IW24006 \(NoIDLEChatGPT\) and grant number 25361 \(RV\-NI\-2024–2029\-K\-IML\), the Lower Saxony Ministry of Science and Culture \(MWK\) in the zukunft\.niedersachsen program, and the Endowed Chair of AAI at University of Oldenburg\. We also gratefully acknowledge support from the hessian\.AI Service Center \(funded by the Federal Ministry of Research, Technology and Space, BMFTR, grant no\. 16IS22091\) and the hessian\.AI Innovation Lab \(funded by the Hessian Ministry for Digital Strategy and Innovation, grant no\. S\-DIW04/0013/003\)\.

## References

- O\. Adjali, S\. Liang, O\. S\. Bhatti, and D\. SonntagAligning instruction\-tuned llms for event extraction with multi\-objective reinforcement learning\.InEuropean Conference on Information Retrieval,pp\. 586–595\.Cited by:[§2\.2](https://arxiv.org/html/2609.29230#S2.SS2.p1.1)\.
- Caiet al\.\(2024\)Z\. Cai, P\. Kung, A\. Suvarna, M\. Ma, H\. Bansal, B\. Chang, P\. J\. Brantingham, W\. Wang, and N\. PengImproving event definition following for zero\-shot event detection\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 2842–2863\.Cited by:[§1](https://arxiv.org/html/2609.29230#S1.p1.1)\.
- Chenet al\.\(2024\)R\. Chen, C\. Qin, W\. Jiang, and D\. ChoiIs a large language model a good annotator for event extraction?\.InProceedings of the AAAI conference on artificial intelligence,Vol\.38,pp\. 17772–17780\.Cited by:[§2\.1](https://arxiv.org/html/2609.29230#S2.SS1.p1.1)\.
- Chuet al\.\(2025\)T\. Chu, Y\. Zhai, J\. Yang, S\. Tong, S\. Xie, D\. Schuurmans, Q\. V\. Le, S\. Levine, and Y\. MaSft memorizes, rl generalizes: a comparative study of foundation model post\-training\.arXiv preprint arXiv:2501\.17161\.Cited by:[§1](https://arxiv.org/html/2609.29230#S1.p2.1)\.
- Gaoet al\.\(2024\)J\. Gao, H\. Zhao, W\. Wang, C\. Yu, and R\. XuEventrl: enhancing event extraction with outcome supervision for large language models\.arXiv preprint arXiv:2402\.11430\.Cited by:[Table 12](https://arxiv.org/html/2609.29230#A13.T12.2.1.13.1),[Table 13](https://arxiv.org/html/2609.29230#A13.T13.2.1.13.1),[Table 14](https://arxiv.org/html/2609.29230#A13.T14.2.1.13.1),[§1](https://arxiv.org/html/2609.29230#S1.p2.1),[§1](https://arxiv.org/html/2609.29230#S1.p3.1),[§2\.2](https://arxiv.org/html/2609.29230#S2.SS2.p1.1),[§3\.8](https://arxiv.org/html/2609.29230#S3.SS8.p1.1),[Table 2](https://arxiv.org/html/2609.29230#S3.T2.2.15.1),[§4](https://arxiv.org/html/2609.29230#S4.SS0.SSS0.Px1.p1.1)\.
- Gaoet al\.\(2023\)J\. Gao, H\. Zhao, C\. Yu, and R\. XuExploring the feasibility of chatgpt for event extraction\.arXiv preprint arXiv:2303\.03836\.Cited by:[§1](https://arxiv.org/html/2609.29230#S1.p1.1)\.
- Grattafioriet al\.\(2024\)A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan,et al\.The llama 3 herd of models\.arXiv preprint arXiv:2407\.21783\.Cited by:[Appendix B](https://arxiv.org/html/2609.29230#A2.p1.1)\.
- Guiet al\.\(2024\)H\. Gui, L\. Yuan, H\. Ye, N\. Zhang, M\. Sun, L\. Liang, and H\. ChenIEPile: unearthing large scale schema\-conditioned information extraction corpus\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 2: Short Papers\),pp\. 127–146\.Cited by:[§2](https://arxiv.org/html/2609.29230#S2.p1.1),[§3\.4](https://arxiv.org/html/2609.29230#S3.SS4.p2.1)\.
- Guoet al\.\(2025\)D\. Guo, D\. Yang, H\. Zhang, J\. Song, P\. Wang, Q\. Zhu, R\. Xu, R\. Zhang, S\. Ma, X\. Bi,et al\.Deepseek\-r1: incentivizing reasoning capability in llms via reinforcement learning\.arXiv preprint arXiv:2501\.12948\.Cited by:[§3\.5](https://arxiv.org/html/2609.29230#S3.SS5.p1.1)\.
- Heet al\.\(2026\)X\. He, Q\. Sun, A\. Cheng, X\. Li, X\. Ji, H\. Lu, R\. Huang, and Q\. HuAdvantage collapse in group relative policy optimization: diagnosis and mitigation\.InInternational Conference on Machine Learning,Cited by:[§3\.7](https://arxiv.org/html/2609.29230#S3.SS7.p2.1),[§3\.7](https://arxiv.org/html/2609.29230#S3.SS7.p5.1)\.
- Hong and Liu \(2024\)Z\. Hong and J\. LiuTowards better question generation in qa\-based event extraction\.InFindings of the Association for Computational Linguistics: ACL 2024,pp\. 9025–9038\.Cited by:[§1](https://arxiv.org/html/2609.29230#S1.p1.1)\.
- Hsuet al\.\(2022\)I\. Hsu, K\. Huang, E\. Boschee, S\. Miller, P\. Natarajan, K\. Chang, and N\. PengDEGREE: a data\-efficient generation\-based event extraction model\.InProceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,pp\. 1890–1908\.Cited by:[§1](https://arxiv.org/html/2609.29230#S1.p1.1)\.
- Huet al\.\(2022\)E\. J\. Hu, yelong shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. ChenLoRA: low\-rank adaptation of large language models\.InInternational Conference on Learning Representations,External Links:[Link](https://openreview.net/forum?id=nZeVKeeFYf9)Cited by:[Appendix G](https://arxiv.org/html/2609.29230#A7.p1.1)\.
- Huanet al\.\(2025\)M\. Huan, Y\. Li, T\. Zheng, X\. Xu, S\. Kim, M\. Du, R\. Poovendran, G\. Neubig, and X\. YueDoes math reasoning improve general llm capabilities? understanding transferability of llm reasoning\.arXiv preprint arXiv:2507\.00432\.Cited by:[§1](https://arxiv.org/html/2609.29230#S1.p2.1)\.
- Huanget al\.\(2024\)K\. Huang, I\. Hsu, T\. Parekh, Z\. Xie, Z\. Zhang, P\. Natarajan, K\. Chang, N\. Peng, and H\. JiTextEE: benchmark, reevaluation, reflections, and future challenges in event extraction\.InFindings of the Association for Computational Linguistics ACL 2024,pp\. 12804–12825\.Cited by:[§2\.1](https://arxiv.org/html/2609.29230#S2.SS1.p1.1),[§4](https://arxiv.org/html/2609.29230#S4.SS0.SSS0.Px2.p1.1),[§4](https://arxiv.org/html/2609.29230#S4.SS0.SSS0.Px3.p1.1),[§5\.1](https://arxiv.org/html/2609.29230#S5.SS1.p1.1)\.
- Hurstet al\.\(2024\)A\. Hurst, A\. Lerer, A\. P\. Goucher, A\. Perelman, A\. Ramesh, A\. Clark, A\. Ostrow, A\. Welihinda, A\. Hayes, A\. Radford,et al\.Gpt\-4o system card\.arXiv preprint arXiv:2410\.21276\.Cited by:[Table 12](https://arxiv.org/html/2609.29230#A13.T12.2.1.3.1),[Table 13](https://arxiv.org/html/2609.29230#A13.T13.2.1.3.1),[Table 14](https://arxiv.org/html/2609.29230#A13.T14.2.1.3.1),[Table 2](https://arxiv.org/html/2609.29230#S3.T2.2.4.1),[§4](https://arxiv.org/html/2609.29230#S4.SS0.SSS0.Px1.p1.1)\.
- Jaechet al\.\(2024\)A\. Jaech, A\. Kalai, A\. Lerer, A\. Richardson, A\. El\-Kishky, A\. Low, A\. Helyar, A\. Madry, A\. Beutel, A\. Carney,et al\.Openai o1 system card\.arXiv preprint arXiv:2412\.16720\.Cited by:[Table 2](https://arxiv.org/html/2609.29230#S3.T2.2.3.1),[§4](https://arxiv.org/html/2609.29230#S4.SS0.SSS0.Px1.p1.1)\.
- Jiaoet al\.\(2023\)Y\. Jiao, M\. Zhong, S\. Li, R\. Zhao, S\. Ouyang, H\. Ji, and J\. HanInstruct and extract: instruction tuning for on\-demand information extraction\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 10030–10051\.Cited by:[§2](https://arxiv.org/html/2609.29230#S2.p1.1)\.
- Kimet al\.\(2011\)J\. Kim, Y\. Wang, T\. Takagi, and A\. YonezawaOverview of genia event task in bionlp shared task 2011\.InProceedings of BioNLP shared task 2011 workshop,pp\. 7–15\.Cited by:[Appendix L](https://arxiv.org/html/2609.29230#A12.SS0.SSS0.Px4.p1.1),[§4](https://arxiv.org/html/2609.29230#S4.SS0.SSS0.Px2.p1.1)\.
- Kimet al\.\(2013\)J\. Kim, Y\. Wang, and Y\. YasunoriThe genia event extraction shared task, 2013 edition\-overview\.InProceedings of the BioNLP shared task 2013 workshop,pp\. 8–15\.Cited by:[§4](https://arxiv.org/html/2609.29230#S4.SS0.SSS0.Px2.p1.1)\.
- Liet al\.\(2023a\)H\. Li, Y\. Cao, Y\. Ren, F\. Fang, L\. Zhang, Y\. Li, and S\. WangIntra\-event and inter\-event dependency\-aware graph network for event argument extraction\.InFindings of the Association for Computational Linguistics: EMNLP 2023,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),Singapore,pp\. 6362–6372\.External Links:[Link](https://aclanthology.org/2023.findings-emnlp.421/),[Document](https://dx.doi.org/10.18653/v1/2023.findings-emnlp.421)Cited by:[§1](https://arxiv.org/html/2609.29230#S1.p1.1)\.
- Liet al\.\(2020\)M\. Li, A\. Zareian, Q\. Zeng, S\. Whitehead, D\. Lu, H\. Ji, and S\. ChangCross\-media structured common space for multimedia event extraction\.arXiv preprint arXiv:2005\.02472\.Cited by:[Appendix L](https://arxiv.org/html/2609.29230#A12.SS0.SSS0.Px6.p1.1),[§4](https://arxiv.org/html/2609.29230#S4.SS0.SSS0.Px2.p1.1)\.
- Liet al\.\(2023b\)P\. Li, T\. Sun, Q\. Tang, H\. Yan, Y\. Wu, X\. Huang, and X\. QiuCodeIE: large code generation models are better few\-shot information extractors\.InThe 61st Annual Meeting Of The Association For Computational Linguistics,Cited by:[Appendix A](https://arxiv.org/html/2609.29230#A1.p1.1),[§2](https://arxiv.org/html/2609.29230#S2.p1.1),[§3\.2](https://arxiv.org/html/2609.29230#S3.SS2.p1.1)\.
- Liet al\.\(2021\)S\. Li, H\. Ji, and J\. HanDocument\-level event argument extraction by conditional generation\.InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies,pp\. 894–908\.Cited by:[Appendix L](https://arxiv.org/html/2609.29230#A12.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2609.29230#S4.SS0.SSS0.Px2.p1.1)\.
- Liet al\.\(2024\)Z\. Li, Y\. Zeng, Y\. Zuo, W\. Ren, W\. Liu, M\. Su, Y\. Guo, Y\. Liu, L\. Lixiang, Z\. Hu,et al\.KnowCoder: coding structured knowledge into llms for universal information extraction\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 8758–8779\.Cited by:[§2](https://arxiv.org/html/2609.29230#S2.p1.1)\.
- Liuet al\.\(2026\)S\. Liu, X\. Dong, X\. Lu, S\. Diao, P\. Belcak, M\. Liu, M\. Chen, H\. Yin, Y\. F\. Wang, K\. Cheng,et al\.Gdpo: group reward\-decoupled normalization policy optimization for multi\-reward rl optimization\.arXiv preprint arXiv:2601\.05242\.Cited by:[§3\.6](https://arxiv.org/html/2609.29230#S3.SS6.p1.1)\.
- Luet al\.\(2023\)K\. Lu, X\. Pan, K\. Song, H\. Zhang, D\. Yu, and J\. ChenPivoine: instruction tuning for open\-world information extraction\.arXiv preprint arXiv:2305\.14898\.Cited by:[§2](https://arxiv.org/html/2609.29230#S2.p1.1)\.
- Luet al\.\(2021\)Y\. Lu, H\. Lin, J\. Xu, X\. Han, J\. Tang, A\. Li, L\. Sun, M\. Liao, and S\. ChenText2Event: controllable sequence\-to\-structure generation for end\-to\-end event extraction\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),C\. Zong, F\. Xia, W\. Li, and R\. Navigli \(Eds\.\),Online,pp\. 2795–2806\.External Links:[Link](https://aclanthology.org/2021.acl-long.217/),[Document](https://dx.doi.org/10.18653/v1/2021.acl-long.217)Cited by:[§1](https://arxiv.org/html/2609.29230#S1.p1.1)\.
- Maet al\.\(2024\)M\. D\. Ma, X\. Wang, P\. Kung, P\. J\. Brantingham, N\. Peng, and W\. WangSTAR: boosting low\-resource information extraction by structure\-to\-text data generation with large language models\.InProceedings of the AAAI conference on artificial intelligence,Vol\.38,pp\. 18751–18759\.Cited by:[§1](https://arxiv.org/html/2609.29230#S1.p1.1)\.
- Maet al\.\(2022\)Y\. Ma, Z\. Wang, Y\. Cao, M\. Li, M\. Chen, K\. Wang, and J\. ShaoPrompt for extraction? paie: prompting argument interaction for event argument extraction\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 6759–6774\.Cited by:[§1](https://arxiv.org/html/2609.29230#S1.p1.1)\.
- Ouyanget al\.\(2022\)L\. Ouyang, J\. Wu, X\. Jiang, D\. Almeida, C\. Wainwright, P\. Mishkin, C\. Zhang, S\. Agarwal, K\. Slama, A\. Ray,et al\.Training language models to follow instructions with human feedback\.Advances in neural information processing systems35,pp\. 27730–27744\.Cited by:[Appendix E](https://arxiv.org/html/2609.29230#A5.p1.1)\.
- Pyysaloet al\.\(2012\)S\. Pyysalo, T\. Ohta, M\. Miwa, H\. Cho, J\. Tsujii, and S\. AnaniadouEvent extraction across multiple levels of biological organization\.Bioinformatics28\(18\),pp\. i575–i581\.Cited by:[Appendix L](https://arxiv.org/html/2609.29230#A12.SS0.SSS0.Px4.p1.1),[Appendix L](https://arxiv.org/html/2609.29230#A12.SS0.SSS0.Px5.p1.1),[§4](https://arxiv.org/html/2609.29230#S4.SS0.SSS0.Px2.p1.1)\.
- Qiet al\.\(2024\)Y\. Qi, H\. Peng, X\. Wang, B\. Xu, L\. Hou, and J\. LiADELIE: aligning large language models on information extraction\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,pp\. 7371–7387\.Cited by:[Table 12](https://arxiv.org/html/2609.29230#A13.T12.2.1.12.1),[Table 13](https://arxiv.org/html/2609.29230#A13.T13.2.1.12.1),[Table 14](https://arxiv.org/html/2609.29230#A13.T14.2.1.12.1),[§2\.2](https://arxiv.org/html/2609.29230#S2.SS2.p1.1),[Table 2](https://arxiv.org/html/2609.29230#S3.T2.2.16.1),[§4](https://arxiv.org/html/2609.29230#S4.SS0.SSS0.Px1.p1.1)\.
- Renet al\.\(2023\)Y\. Ren, Y\. Cao, P\. Guo, F\. Fang, W\. Ma, and Z\. LinRetrieve\-and\-sample: document\-level event argument extraction via hybrid retrieval augmentation\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 293–306\.Cited by:[§1](https://arxiv.org/html/2609.29230#S1.p1.1)\.
- Roziereet al\.\(2023\)B\. Roziere, J\. Gehring, F\. Gloeckle, S\. Sootla, I\. Gat, X\. E\. Tan, Y\. Adi, J\. Liu, R\. Sauvestre, T\. Remez,et al\.Code llama: open foundation models for code\.arXiv preprint arXiv:2308\.12950\.Cited by:[Appendix G](https://arxiv.org/html/2609.29230#A7.p1.1)\.
- Sainzet al\.\(2024\)O\. Sainz, I\. García\-Ferrero, R\. Agerri, O\. Lacalle, G\. Rigau, and E\. AgirreGollie: annotation guidelines improve zero\-shot information\-extraction\.InInternational Conference on Learning Representations,Vol\.2024,pp\. 47083–47107\.Cited by:[Appendix A](https://arxiv.org/html/2609.29230#A1.p1.1),[Table 12](https://arxiv.org/html/2609.29230#A13.T12.2.1.6.1),[Table 12](https://arxiv.org/html/2609.29230#A13.T12.2.1.7.1),[Table 12](https://arxiv.org/html/2609.29230#A13.T12.2.1.8.1),[Table 13](https://arxiv.org/html/2609.29230#A13.T13.2.1.6.1),[Table 13](https://arxiv.org/html/2609.29230#A13.T13.2.1.7.1),[Table 13](https://arxiv.org/html/2609.29230#A13.T13.2.1.8.1),[Table 14](https://arxiv.org/html/2609.29230#A13.T14.2.1.6.1),[Table 14](https://arxiv.org/html/2609.29230#A13.T14.2.1.7.1),[Table 14](https://arxiv.org/html/2609.29230#A13.T14.2.1.8.1),[Appendix G](https://arxiv.org/html/2609.29230#A7.p1.1),[§1](https://arxiv.org/html/2609.29230#S1.p1.1),[§2](https://arxiv.org/html/2609.29230#S2.p1.1),[§3\.2](https://arxiv.org/html/2609.29230#S3.SS2.p1.1),[Table 2](https://arxiv.org/html/2609.29230#S3.T2.2.7.1),[Table 2](https://arxiv.org/html/2609.29230#S3.T2.2.8.1),[Table 2](https://arxiv.org/html/2609.29230#S3.T2.2.9.1),[§4](https://arxiv.org/html/2609.29230#S4.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2609.29230#S4.SS0.SSS0.Px3.p1.1),[§5\.1](https://arxiv.org/html/2609.29230#S5.SS1.SSS0.Px1.p1.1)\.
- Satyapanichet al\.\(2020\)T\. Satyapanich, F\. Ferraro, and T\. FininCasie: extracting cybersecurity event information from text\.InProceedings of the AAAI conference on artificial intelligence,Vol\.34,pp\. 8749–8757\.Cited by:[Appendix L](https://arxiv.org/html/2609.29230#A12.SS0.SSS0.Px3.p1.1),[§4](https://arxiv.org/html/2609.29230#S4.SS0.SSS0.Px2.p1.1)\.
- Schulmanet al\.\(2017\)J\. Schulman, F\. Wolski, P\. Dhariwal, A\. Radford, and O\. KlimovProximal policy optimization algorithms\.arXiv preprint arXiv:1707\.06347\.Cited by:[§3\.7](https://arxiv.org/html/2609.29230#S3.SS7.p1.1)\.
- Shaoet al\.\(2024\)Z\. Shao, P\. Wang, Q\. Zhu, R\. Xu, J\. Song, X\. Bi, H\. Zhang, M\. Zhang, Y\. Li, Y\. Wu,et al\.Deepseekmath: pushing the limits of mathematical reasoning in open language models, 2024\.URL https://arxiv\. org/abs/2402\.033002\(3\),pp\. 5\.Cited by:[§3\.5](https://arxiv.org/html/2609.29230#S3.SS5.p1.1)\.
- Singhet al\.\(2025\)A\. Singh, A\. Fry, A\. Perelman, A\. Tart, A\. Ganesh, A\. El\-Kishky, A\. McLaughlin, A\. Low, A\. Ostrow, A\. Ananthram,et al\.Openai gpt\-5 system card\.arXiv preprint arXiv:2601\.03267\.Cited by:[Table 12](https://arxiv.org/html/2609.29230#A13.T12.2.1.4.1),[Table 12](https://arxiv.org/html/2609.29230#A13.T12.2.1.5.1),[Table 13](https://arxiv.org/html/2609.29230#A13.T13.2.1.4.1),[Table 13](https://arxiv.org/html/2609.29230#A13.T13.2.1.5.1),[Table 14](https://arxiv.org/html/2609.29230#A13.T14.2.1.4.1),[Table 14](https://arxiv.org/html/2609.29230#A13.T14.2.1.5.1),[Table 2](https://arxiv.org/html/2609.29230#S3.T2.2.5.1),[Table 2](https://arxiv.org/html/2609.29230#S3.T2.2.6.1),[§4](https://arxiv.org/html/2609.29230#S4.SS0.SSS0.Px1.p1.1)\.
- Srivastavaet al\.\(2025\)S\. Srivastava, S\. Pati, and Z\. YaoInstruction\-tuning LLMs for event extraction with annotation guidelines\.InFindings of the Association for Computational Linguistics: ACL 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 13055–13071\.External Links:[Link](https://aclanthology.org/2025.findings-acl.677/),[Document](https://dx.doi.org/10.18653/v1/2025.findings-acl.677),ISBN 979\-8\-89176\-256\-5Cited by:[Appendix A](https://arxiv.org/html/2609.29230#A1.p1.1),[Appendix B](https://arxiv.org/html/2609.29230#A2.p1.1),[§1](https://arxiv.org/html/2609.29230#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.29230#S2.SS1.p1.1),[§3\.2](https://arxiv.org/html/2609.29230#S3.SS2.p1.1),[§3\.3](https://arxiv.org/html/2609.29230#S3.SS3.p1.1),[Table 2](https://arxiv.org/html/2609.29230#S3.T2.2.13.1),[§4](https://arxiv.org/html/2609.29230#S4.SS0.SSS0.Px1.p1.1),[§4](https://arxiv.org/html/2609.29230#S4.SS0.SSS0.Px3.p1.1),[§5\.1](https://arxiv.org/html/2609.29230#S5.SS1.SSS0.Px1.p1.1)\.
- Sunet al\.\(2022\)Z\. Sun, J\. Li, G\. Pergola, B\. C\. Wallace, B\. John, N\. Greene, J\. Kim, and Y\. HePHEE: a dataset for pharmacovigilance event extraction from text\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,pp\. 5571–5587\.Cited by:[Appendix L](https://arxiv.org/html/2609.29230#A12.SS0.SSS0.Px2.p1.1),[§4](https://arxiv.org/html/2609.29230#S4.SS0.SSS0.Px2.p1.1)\.
- von Werraet al\.\(2020\)L\. von Werra, Y\. Belkada, L\. Tunstall, E\. Beeching, T\. Thrush, N\. Lambert, S\. Huang, K\. Rasul, and Q\. GallouédecTRL: transformer reinforcement learning\.GitHub\.Note:[https://github\.com/huggingface/trl](https://github.com/huggingface/trl)Cited by:[Appendix G](https://arxiv.org/html/2609.29230#A7.p1.1)\.
- Wanget al\.\(2023a\)X\. Wang, W\. Zhou, C\. Zu, H\. Xia, T\. Chen, Y\. Zhang, R\. Zheng, J\. Ye, Q\. Zhang, T\. Gui,et al\.Instructuie: multi\-task instruction tuning for unified information extraction\.arXiv preprint arXiv:2304\.08085\.Cited by:[§1](https://arxiv.org/html/2609.29230#S1.p1.1),[§2](https://arxiv.org/html/2609.29230#S2.p1.1)\.
- Wanget al\.\(2023b\)X\. Wang, S\. Li, and H\. JiCode4Struct: code generation for few\-shot event structure prediction\.InThe 61st Annual Meeting Of The Association For Computational Linguistics,Cited by:[Appendix A](https://arxiv.org/html/2609.29230#A1.p1.1),[§3\.2](https://arxiv.org/html/2609.29230#S3.SS2.p1.1)\.
- Yuet al\.\(2026\)Q\. Yu, Z\. Zhang, R\. Zhu, Y\. Yuan, X\. Zuo, Y\. Yue, W\. Dai, T\. Fan, G\. Liu, L\. Liu,et al\.Dapo: an open\-source llm reinforcement learning system at scale\.Advances in Neural Information Processing Systems38,pp\. 113222–113244\.Cited by:[§J\.2](https://arxiv.org/html/2609.29230#A10.SS2.p1.1),[§J\.3](https://arxiv.org/html/2609.29230#A10.SS3.p1.1),[Appendix G](https://arxiv.org/html/2609.29230#A7.p1.1),[Appendix H](https://arxiv.org/html/2609.29230#A8.p1.1),[§3\.5](https://arxiv.org/html/2609.29230#S3.SS5.p1.1),[§3\.7](https://arxiv.org/html/2609.29230#S3.SS7.p2.1)\.
- Zhanget al\.\(2025\)X\. Zhang, S\. Wu, Y\. Zhu, H\. Tan, S\. Yu, Z\. He, and J\. JiaScaf\-grpo: scaffolded group relative policy optimization for enhancing llm reasoning\.arXiv preprint arXiv:2510\.19807\.Cited by:[§3\.7](https://arxiv.org/html/2609.29230#S3.SS7.p2.1)\.

## Appendix ACode\-based Representation

1@dataclass

2classConflict\_Attack:

3"""Conflict:Attackevent"""

4mention:str

5Attacker:List\[str\]

6Instrument:List\[str\]

7Place:List\[str\]

8Target:List\[str\]

Figure 6:Example of an event schema as python classWe formulate EE as a code generation problem where both the input and output are formatted using Python code similar to Following[Sainz et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib3);[Srivastava et al\. \(2025\)](https://arxiv.org/html/2609.29230#bib.bib9)\. Indeed, IE tasks benefit on the first hand from the strong code understanding capabilities of large language models since code data is widely included in their pre\-training corpora[Wang et al\. \(2023b\)](https://arxiv.org/html/2609.29230#bib.bib13);[Li et al\. \(2023b\)](https://arxiv.org/html/2609.29230#bib.bib7)\. On the other hand, code\-based structure representation provides a unified and human\-readable framework for information extraction tasks, while mitigating ambiguities often encountered in natural language instructions\. The code format ensures that outputs are syntactically well\-formed facilitating output parsing[Sainz et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib3);[Srivastava et al\. \(2025\)](https://arxiv.org/html/2609.29230#bib.bib9)\. In particular, event schemes are expressed as Python classes \(@dataclasstype definitions\) and the extracted output events as instances of these classes\.

## Appendix BGuidelines Annotation Generation

Using LLaMA\-3\.1\-8B\-Instruct model[Grattafiori et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib23), we adopt theGuideline\-PN\(Positive \+ Negative\) generation protocol of[Srivastava et al\. \(2025\)](https://arxiv.org/html/2609.29230#bib.bib9), wherein the LLM is conditioned on a contrastive set of examples to produce guidelines for each event typee∈ℰe\\in\\mathcal\{E\}\. Specifically, the prompt consists of positive examples: 10 annotated instances of event typeeepaired with their source texts and negative examples: 15 texts containing other event types but no instance ofee\. This contrastive design encourages the model to identify definitional boundaries and discriminative features that distinguisheefrom related event types\. The generation prompt instructs the LLM to: \(1\) enumerate all unique argument roles foree; \(2\) provide a precise definition of the event type; and \(3\) characterize each argument role, emphasizing its semantic function, representative mentions, and potential edge cases\. The resulting guidelines are incorporated into the Python class schema via docstrings and inline comments, yielding an augmented schemaEeGE\_\{e\}^\{\\text\{G\}\}that integrates both structural and semantic information\. See Figures[14](https://arxiv.org/html/2609.29230#A14.F14),[15](https://arxiv.org/html/2609.29230#A14.F15),[16](https://arxiv.org/html/2609.29230#A14.F16),[17](https://arxiv.org/html/2609.29230#A14.F17),[18](https://arxiv.org/html/2609.29230#A14.F18),[19](https://arxiv.org/html/2609.29230#A14.F19)and[20](https://arxiv.org/html/2609.29230#A14.F20)for examples of event scheme augmented with annotation guidelines for each dataset\.

### B\.1Additional Annotation Guidelines Analysis

Tables[8](https://arxiv.org/html/2609.29230#A2.T8)reports the the effect of AG on the GoLLIE backbone model\. We can see that augmenting event schemes with annotation guidelines improves event extraction performance\.

Table 8:Performance comparison of GoLLIE models without AG\.

## Appendix CSupervised Fine\-Tuning

To effectively train medium\-scale LLMs while preserving their foundational capabilities, we carry out instruction\-based supervised fine\-tuning to adapt the model to event extraction using code\-style schema representations\. This enables the model to follow natural\-language task instructions while generating syntactically valid and semantically grounded event representations in Python code\. The LLM is trained using supervised fine\-tuning on a set of annotated prompt/response pairs\{\(Pi,Yi\)\}i=1N\\\{\(P\_\{i\},Y\_\{i\}\)\\\}\_\{i=1\}^\{N\}, whereYiY\_\{i\}denotes the target structured event instance\. The fine\-tuning objective follows a standard autoregressive likelihood formulation:

ℒ\(θ\)=−∑i∑jlogpθ\(Yi,j∣Pi,Yi,<j\),\\mathcal\{L\}\(\\theta\)=\-\\sum\_\{i\}\\sum\_\{j\}\\log p\_\{\\theta\}\(Y\_\{i,j\}\\mid P\_\{i\},Y\_\{i,<j\}\),whereYi,<jY\_\{i,<j\}denotes previously generated tokens in the output sequence\.

## Appendix DInference

At inference time, given an input textXX, allnnevent schemas inℰ=\{Ei\}i=1n\\mathcal\{E\}=\\\{E\_\{i\}\\\}\_\{i=1\}^\{n\}are provided jointly in the prompt, enabling the model to perform end\-to\-end event extraction and produce a Python list of instantiated event objects in a single forward pass\.

## Appendix EDAPO Background

DAPO builds upon GRPO and addresses its key limitations i\.e\., entropy collapse, reward noise, and sequence\-level length bias through four modifications: \(1\) Clip\-Higher, which uses asymmetric clipping bounds\[ϵlow,ϵhigh\]\[\\epsilon\_\{\\mathrm\{low\}\},\\epsilon\_\{\\mathrm\{high\}\}\]withϵhigh\>ϵlow\\epsilon\_\{\\mathrm\{high\}\}\>\\epsilon\_\{\\mathrm\{low\}\}to prevent entropy collapse; \(2\) Token\-Level Loss, which normalizes the policy gradient over all active tokens in the batch rather than per sequence, eliminating length bias; \(3\) Dynamic Sampling, which filters out zero\-variance groups \(all\-correct or all\-incorrect\) and resamples until every batch contains informative gradient signal; and \(4\) KL Removal, which setsβ=0\\beta\{=\}0\. Unlike standard RLHF[Ouyang et al\. \(2022\)](https://arxiv.org/html/2609.29230#bib.bib40), where KL regularization prevents the policy from deviating too far from a supervised baseline, EE requires the model to diverge from its initial distribution to acquire precise schema\-grounded output patterns\.

## Appendix FVerifiable Rewards Formulation

We propose a decomposed, task\-aligned reward formulation for generative EE\. Rather than relying solely on aggregate extraction F1, we organize the reward signal around three complementary objectives:validity, requiring outputs to be parseable and schema\-compatible;extraction accuracy, requiring outputs to match gold event annotations; andgeneration behavior, requiring outputs to remain grounded in the source text, avoid spurious predictions, and maintain adequate event coverage\.

### F\.1Validity Reward

As a prerequisite to semantic evaluation, we reward outputs that are syntactically well\-formed and successfully parse into the code\-based representation introduced in Section[3\.2](https://arxiv.org/html/2609.29230#S3.SS2):

Rfmt​\(oi\)=\{1,if​oi​is syntactically parseable,0,otherwise\.R\_\{\\text\{fmt\}\}\(o\_\{i\}\)=\\begin\{cases\}1,&\\text\{if \}o\_\{i\}\\text\{ is syntactically parseable,\}\\\\ 0,&\\text\{otherwise\.\}\\end\{cases\}\(9\)

### F\.2Extraction Accuracy Reward

We define the extraction reward as the sum of F1 scores across the six EE evaluation metrics spanning the four canonical subtasks TI, TC, AI, AC ‚and their trigger\-aware variants AI\+and AC\+:

REE=∑k∈\{TI, TC, AI, AC, AI\+,AC\+\}F1k\.R\_\{\\text\{EE\}\}=\\sum\_\{k\\,\\in\\,\\\{\\text\{TI, TC, AI, AC, AI\}^\{\+\}\\\!,\\,\\text\{AC\}^\{\+\}\\\}\}F\_\{1\}^\{k\}\.\(10\)This reward captures holistic extraction quality across all subtasks but remains coarse with respect to specific generation issues, motivating the supplementary signals below\.

### F\.3Groundedness Reward

To penalize hallucinated triggers and argument spans unsupported by the source text, we introduce a groundedness reward that measures whether predicted mentions appear verbatim in the inputXX\. Let\{ej\}\\\{e\_\{j\}\\\}denote the set of predicted events,ℛ⁡\(ej\)\\mathcal\{R\}\(e\_\{j\}\)the set of argument roles for eventeje\_\{j\}, andTiT\_\{i\}the total number of predicted arguments across all events inoio\_\{i\}\. We compute separate support scores for triggers and arguments:

sm\\displaystyle s\_\{\\text\{m\}\}=1\|\{ej\}\|∑ej𝕀\[contains\(X,ej\.trigger\)\],\\displaystyle=\\frac\{1\}\{\|\\\{e\_\{j\}\\\}\|\}\\sum\_\{e\_\{j\}\}\\mathbb\{I\}\\bigl\[\\mathrm\{contains\}\(X,\\,e\_\{j\}\.\\mathrm\{trigger\}\)\\bigr\],\(11\)sarg\\displaystyle s\_\{\\text\{arg\}\}=1Ti​∑ej∑r∈ℛ⁡\(ej\)∑v∈ej​\[r\]𝕀⁡\[contains⁡\(X,v\)\],\\displaystyle=\\frac\{1\}\{T\_\{i\}\}\\sum\_\{e\_\{j\}\}\\sum\_\{r\\in\\mathcal\{R\}\(e\_\{j\}\)\}\\sum\_\{v\\in e\_\{j\}\[r\]\}\\mathbb\{I\}\\bigl\[\\mathrm\{contains\}\(X,\\,v\)\\bigr\],\(12\)withsm=0s\_\{\\text\{m\}\}=0if no events are predicted andsarg=0s\_\{\\text\{arg\}\}=0ifTi=0T\_\{i\}=0\. The groundedness reward is their average:

Rgrd=12​\(sm\+sarg\)\.R\_\{\\text\{grd\}\}=\\tfrac\{1\}\{2\}\\bigl\(s\_\{\\text\{m\}\}\+s\_\{\\text\{arg\}\}\\bigr\)\.\(13\)This signal penalizes fabricated spans while remaining agnostic to event type correctness\.

### F\.4Over\-generation Reward

Since extraction F1 does not explicitly penalize spurious predictions beyond the gold annotation, we introduce a complementary over\-generation penalty\. LetEi\+E\_\{i\}^\{\+\}andAi\+A\_\{i\}^\{\+\}denote the number of predicted events and arguments inoio\_\{i\}that exceed the gold counts inY∗Y^\{\\ast\}\. We define:

Rovr=max⁡\(0,1−Ei\+\+Ai\+n∗\+a∗\)R\_\{\\mathrm\{ovr\}\}=\\max\\\!\\left\(0,\\;1\-\\frac\{E^\{\+\}\_\{i\}\+A^\{\+\}\_\{i\}\}\{n^\{\*\}\+a^\{\*\}\}\\right\)\(14\)whereEi\+E^\{\+\}\_\{i\}andAi\+A^\{\+\}\_\{i\}are the predicted events and arguments exceeding the gold counts, andn∗n^\{\*\},a∗a^\{\*\}are the gold event and argument counts\. This reward equals 1 when no excess predictions are made and decreases linearly as spurious spans accumulate, saturating at 0\.

### F\.5Coverage Reward

To counterbalanceRovrR\_\{\\text\{ovr\}\}and prevent the model from adopting an overly conservative decoding strategy, we reward proportional event and argument coverage relative to the gold annotation\. Letn∗=\|Y∗\|n^\{\\ast\}=\|Y^\{\\ast\}\|andng=\|\{e∈oi:e\.type∈𝒮\}\|n^\{g\}=\|\\\{e\\in o\_\{i\}:e\.\\text\{type\}\\in\\mathcal\{S\}\\\}\|denote the gold and schema\-valid predicted event counts, respectively, and leta∗a^\{\\ast\}andaga^\{g\}be the corresponding argument counts\. We define:

Rcov=12​\(min⁡\(1,ngn∗\)\+min⁡\(1,aga∗\)\),R\_\{\\text\{cov\}\}=\\frac\{1\}\{2\}\\\!\\left\(\\min\\\!\\left\(1,\\frac\{n^\{g\}\}\{n^\{\\ast\}\}\\right\)\+\\min\\\!\\left\(1,\\frac\{a^\{g\}\}\{a^\{\\ast\}\}\\right\)\\right\),\(15\)where𝒮\\mathcal\{S\}is the event schema registry\. Themin⁡\(⋅,1\)\\min\(\\cdot,1\)clamping ensures that over\-prediction does not inflate the coverage score, maintaining complementarity withRovrR\_\{\\text\{ovr\}\}\.

### F\.6Span Precision Reward

It addresses the case where the model correctly identifies an argument’s semantic content but extracts a superset of the minimal gold span\. Standard exact\-match rewards treat such predictions as fully incorrect, providing no useful gradient signal for boundary refinement\. We therefore introduce a complementary Span Precision RewardRspanR\_\{\\mathrm\{span\}\}that softly penalizes overpredicted spans while rewarding near\-correct predictions\. For each predicted argumenta^\\hat\{a\}, we retrieve its best\-matching gold spana∗=arg​maxa∈𝒜r⁡𝒥​\(a^,a\)a^\{\*\}=\\argmax\_\{a\\in\\mathcal\{A\}\_\{r\}\}\\,\\mathcal\{J\}\(\\hat\{a\},a\)via token Jaccard similarity, and compute a per\-argument score:

s⁡\(a^,a∗\)=\{max⁡\(0,𝒥⁡\(a^,a∗\)CLOSEOPEN−\|a^\|−\|a∗\|max⁡\(1,\|a∗\|\)\)if​a^⊃a∗,𝒥⁡\(a^,a∗\)otherwise\.s\(\\hat\{a\},a^\{\*\}\)=\\begin\{cases\}\\max\\\!\\left\(0,\\;\\mathcal\{J\}\(\\hat\{a\},a^\{\*\}\)\\right\.\\\\ \\left\.\-\\dfrac\{\|\\hat\{a\}\|\-\|a^\{\*\}\|\}\{\\max\(1,\|a^\{\*\}\|\)\}\\right\)\\quad\\text\{if \}\\hat\{a\}\\supset a^\{\*\},\\\\\[4\.0pt\] \\mathcal\{J\}\(\\hat\{a\},a^\{\*\}\)\\quad\\text\{otherwise\.\}\\end\{cases\}\(16\)The reward is the mean score over all aligned predicted arguments, and is gated to zero when no overprediction occurs, decoupling it from the outcome reward in well\-behaved cases\.

## Appendix GImplementation Details

We conducted experiments using the GoLLIE\-7B model[Sainz et al\. \(2024\)](https://arxiv.org/html/2609.29230#bib.bib3), a fine\-tuned version of Code\-LLaMA[Roziere et al\. \(2023\)](https://arxiv.org/html/2609.29230#bib.bib20), which provides a strong backbone for code\-based information extraction\. For efficient optimization, we employed parameter\-efficient fine\-tuning via QLoRA[Hu et al\. \(2022\)](https://arxiv.org/html/2609.29230#bib.bib21), applying low\-rank adapters to the attention projection layers with LoRA rank 8, scaling factorα=16\\alpha=16, learning rate1×10−61\\times 10^\{\-6\}, and batch size 1 for SFT\. For RL optimization, we used the TRL library[von Werra et al\. \(2020\)](https://arxiv.org/html/2609.29230#bib.bib22)with learning rate1×10−61\\times 10^\{\-6\}, batch size 4, and 4 sampled completions per input using nucleus sampling \(p=0\.9p=0\.9,τ=0\.6\\tau=0\.6\)\. DAPO Clipping bounds are set toϵlow=0\.2\\epsilon\_\{\\text\{low\}\}=0\.2andϵhigh=0\.28\\epsilon\_\{\\text\{high\}\}=0\.28following[Yu et al\. \(2026\)](https://arxiv.org/html/2609.29230#bib.bib39), and the KL penalty coefficient is set toβ=0\\beta=0\. Models were trained for up to 10 epochs on 4 NVIDIA H100 GPUs for RL optimization, with early stopping triggered after 3 consecutive non\-improving validation steps\. The best\-performing validation checkpoint was selected for testing\. At inference time, we used greedy decoding\. Our code will be released at:[https://github\.com/OA256864/EE\_RL](https://github.com/OA256864/EE_RL)\.

## Appendix HTraining Protocol

We train the GoLLIE\-7B LLM on a training set constructed by concatenating the training sets of the seven datasets used in our experiments and then we shuffled the result set to have cross\-domain and \-schema training batches\. We conducted DAPO[Yu et al\. \(2026\)](https://arxiv.org/html/2609.29230#bib.bib39)with our proposed SCAE and our reward formulation without resorting to an SFT checkpoint for warm\-starting as SFT provided only marginal improvements on the validation sets\.

##### Impact of number of negative schemas K

Our preliminary experiments indicated that increasing the number of negative schemas K sampled using SCAE consistently improved performance\. Consequently, we set K=10, the largest value permitted by both our GPU memory constraints and the 16k\-token maximum prompt length budget\.

##### Impact of group size G

Similarly, preliminary experiments showed that increasing the group size fromG=2G=2toG=4G=4sampled completions per input improved performance while also accelerating convergence\. These results suggest that scaling toG=8G=8could yield further performance gains given a very large computational budget\.

1’’’Thisisaneventextractiontaskwherethegoalistoextractstructuredeventsfromthetext\.Astructuredeventcontainsaneventtriggerword,aneventtype,theargumentsparticipatingintheevent,andtheirrolesintheevent\.Foreachdifferenteventtype,pleaseoutputtheextractedinformationfromthetextintopython\-styledictionarieswherethefirstkeywillbe’mention’withthevalueoftheeventtrigger\.Next,pleaseoutputtheargumentsandtheirrolesfollowingthesameformat\.Theeventtypedefinitionsandtheirargumentrolesaredefinednext\.’’’

2

3@dataclass

4classCognitive\_Inspection\_SensoryObserve:

5"""Theeventistriggeredbytheactofobservingorinspectingsomethingwithone’ssenses,suchassight,sound,touch,taste,orsmell\.Itinvolvestheuseofaninstrument,suchasacamera,microscope,orotherdevice,togatherinformationaboutanentity,whichcanbeaperson,place,orobject\.Theeventistypicallyperformedbyanobserver,whomaybeanindividualoragroup,andtakesplaceataspecificlocation\.Unlikeotherevents,suchasJustice\_Convict\_UnspecifiedorTransaction\_ExchangeBuySell\_Unspecified,

6Cognitive\_Inspection\_SensoryObservedoesnotinvolvealegalorfinancialtransaction\.Triggerssuchas’searched’,’found’,’looked’,or’reviewing’areindicativeofCognitive\_Inspection\_SensoryObserve,notothereventtypes\."""

7

8mention:str

9

10

11

12

13Instrument:List\[str\]

14

15

16

17

18ObservedEntity:List\[str\]

19

20

21

22

23Observer:List\[str\]

24

25

26

27

28Place:List\[str\]

29

30

31

32@dataclass\.\.\.

33

34

35

36text=’BostonBombSuspectSenttoFederalMedicalDetentionBostonMarathonbombingsuspectDzhokharTsarnaevhasbeenmovedtoaprisonmedicalfacilityasauthoritiescontinuetosearchforanswersabouttheattack\.TheU\.S\.MarshalsServicesaidFridaythatTsarnaevwasmovedtotheFederalMedicalCenterDevens,aBureauofPrisonsfacilityinthenortheasternstateofMassachusetts\.HewastransferredtherefromaBostonhospitalwherehehadbeenreceivingtreatmentforinjuriessustainedduringhiscapturelastweek\.FederalMedicalCenterDevens,Ayer,MassachusettsAspokesmandidnotgivedetailsabouttheconditionofthe19\-year\-old,whoofficialssayisrecoveringfromaneckwound\.\.\.’

37

38

39

40result=

Figure 7:Training Input Prompt Example

## Appendix IPseudo Code of DAPO with SCAE

To provide an overview of the training process, we present the pseudo code of DAPO with SCAE in Algorithm[1](https://arxiv.org/html/2609.29230#alg1)\.

Algorithm 1DAPO with SCAE1:Initial policy

πθ\\pi\_\{\\theta\}; training set

𝒟\\mathcal\{D\}; event schema pool

ℰ\\mathcal\{E\}; group size

GG; negative schema count

KK; clipping bounds

ϵlow,ϵhigh\\epsilon\_\{\\mathrm\{low\}\},\\epsilon\_\{\\mathrm\{high\}\}
2:forstep

=1,…,n=1,\\ldots,ndo

3:Sample a batch

𝒟b\\mathcal\{D\}\_\{b\}from

𝒟\\mathcal\{D\}
4:Update old policy

πθold←πθ\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\\leftarrow\\pi\_\{\\theta\}
5:foreach instance

\(X,e,Y\)∈𝒟b\(X,e,Y\)\\in\\mathcal\{D\}\_\{b\}do

6:Let

𝒩=ℰ∖\{e\}\\mathcal\{N\}=\\mathcal\{E\}\\setminus\\\{e\\\}be the negative schema pool

7:for

i=1,…,Gi=1,\\ldots,Gdo

8:Sample a distinct negative schema subset

𝒮i​∼i\.i\.d\.​\(𝒩K\)\\mathcal\{S\}\_\{i\}\\overset\{\\text\{i\.i\.d\.\}\}\{\\sim\}\\binom\{\\mathcal\{N\}\}\{K\}
9:Construct contrastive prompt

Pi=I⊕EeG⊕𝒮i⊕XP\_\{i\}=I\\oplus E\_\{e\}^\{\\mathrm\{G\}\}\\oplus\\mathcal\{S\}\_\{i\}\\oplus X
10:Sample completion

o^i∼πθold\(⋅∣Pi\)\\hat\{o\}\_\{i\}\\sim\\pi\_\{\\theta\_\{\\mathrm\{old\}\}\}\(\\cdot\\mid P\_\{i\}\)
11:Compute reward

Ri=R⁡\(o^i,Y\)R\_\{i\}=R\(\\hat\{o\}\_\{i\},Y\)
12:endfor

13:Compute group statistics

μR=mean⁡\(\{Ri\}i=1G\)\\mu\_\{R\}=\\mathrm\{mean\}\(\\\{R\_\{i\}\\\}\_\{i=1\}^\{G\}\),

σR=std⁡\(\{Ri\}i=1G\)\\sigma\_\{R\}=\\mathrm\{std\}\(\\\{R\_\{i\}\\\}\_\{i=1\}^\{G\}\)
14:if

σR=0\\sigma\_\{R\}=0then

15:skipgroup \(Dynamic Sampling\)

16:endif

17:Compute per\-token advantages

A^ti=\(Ri−μR\)/σR\\hat\{A\}\_\{t\}^\{i\}=\(\{R\_\{i\}\-\\mu\_\{R\}\}\)/\{\\sigma\_\{R\}\}
18:endfor

19:Update

πθ\\pi\_\{\\theta\}by minimizing

ℒDAPO​\(θ\)\\mathcal\{L\}\_\{\\mathrm\{DAPO\}\}\(\\theta\)\(Eq\.[3](https://arxiv.org/html/2609.29230#S3.E3)\) with advantages

\{A^ti\}\\\{\\hat\{A\}\_\{t\}^\{i\}\\\}
20:endfor

21:return

πθ\\pi\_\{\\theta\}

## Appendix JTraining Dynamics Analysis

Figure 8:Analysis of the main metrics for monitoring our RL training dynamics including the mean completion length, the mean reward score, and generation entropy\.Reinforcement learning over structured prediction tasks such as EE introduces additional complexity beyond standard reasoning benchmarks, as the reward signal must simultaneously supervise trigger identification, event classification, and schema\-grounded argument extraction\. Given this interdependence, monitoring key intermediate metrics throughout training is essential for diagnosing undesired behavior and validating that each reward component contributes as intended\. Figure[8](https://arxiv.org/html/2609.29230#A10.F8)reports the three principal indicators we track across training in addition to the EE validation performance\.

### J\.1Mean Completion Length Analysis

The mean completion length increases steadily over the course of training, reflecting the model’s growing tendency to produce more complete structured outputs as training progresses\. In the context of code\-based EE, this growth is consistent with the model learning to instantiate a larger and more complete set of event objects and argument slots, rather than defaulting to under\-populated outputs\. We observe no prolonged stagnation or decline in length, suggesting that our training signal discourages overly conservative decoding throughout training\.

### J\.2Mean Reward Analysis

The mean reward increases monotonically with a smooth trajectory and no sign of instability or collapse\. This stability indicates that the decomposed reward formulationRfullR\_\{\\text\{full\}\}provides a reliable and consistent training signal, allowing the model to robustly fit the distribution of the training set\. Consistent with findings reported in prior DAPO work[Yu et al\. \(2026\)](https://arxiv.org/html/2609.29230#bib.bib39), we observe that the final reward on the training set correlates imperfectly with held\-out extraction performance, underscoring the importance of complementing reward monitoring with validation\-set evaluation to detect potential overfitting\.

### J\.3Generation Entropy Analysis

The generation entropy exhibits a slow but consistent downward trend throughout training until it stabilizes then it starts increasing slightly\. As shown in[Yu et al\. \(2026\)](https://arxiv.org/html/2609.29230#bib.bib39), too high entropy indicates over\-exploration of the model, while too low entropy leads to a loss of exploration capability suggesting that the model’s entropy needs to be maintained within an appropriate range\. In our setting, an exploratory behavior is particularly important, as the model must explore diverse span\-selection and role\-assignment strategies before converging on schema\-grounded outputs\. The absence of any entropy spike or erratic fluctuation further confirms that the removal of KL regularization \(β=0\\beta\{=\}0\) does not destabilize training in our setting\.

## Appendix KTraining Computational Cost

Table[9](https://arxiv.org/html/2609.29230#A11.T9)reports the computational training cost\.

Table 9:Computational cost of EAGER training on 4 NVIDIA H100 GPUs\. Wall time is reported for the full train dataset, which includes the seven dataset training sets\.
## Appendix LDatasets Details

We evaluate EAGER on seven benchmark datasets spanning diverse domains, annotation schemes, and event ontologies\.

##### WikiEvents

[Li et al\. \(2021\)](https://arxiv.org/html/2609.29230#bib.bib16)is a large\-scale news\-domain benchmark containing richly annotated real\-world event mentions with complex argument structures\.

##### PHEE

[Sun et al\. \(2022\)](https://arxiv.org/html/2609.29230#bib.bib17)focuses on the biomedical domain, specifically pharmacovigilance event extraction from medical case reports, requiring fine\-grained reasoning over domain\-specific terminology\.

##### CASIE

[Satyapanich et al\. \(2020\)](https://arxiv.org/html/2609.29230#bib.bib18)is a cybersecurity event extraction benchmark centered on security incident reports, featuring highly specialized event schemas and technical vocabulary\.

##### GENIA2011 and GENIA2013

[Kim et al\. \(2011\)](https://arxiv.org/html/2609.29230#bib.bib24);[Pyysalo et al\. \(2012\)](https://arxiv.org/html/2609.29230#bib.bib25)are biomedical event extraction datasets derived from PubMed abstracts, involving nested event structures and biologically grounded argument semantics\.

##### MLEE

[Pyysalo et al\. \(2012\)](https://arxiv.org/html/2609.29230#bib.bib25)extends biomedical event extraction to molecular\-level event understanding with more diverse biological interaction types\.

##### M2E2

[Li et al\. \(2020\)](https://arxiv.org/html/2609.29230#bib.bib26)is a multimodal event extraction benchmark originally designed for joint text\-image event understanding; following prior text\-only work, we evaluate exclusively on the textual component\.

These datasets collectively cover news, biomedical, and cybersecurity domains, providing a comprehensive testbed for evaluating robustness across heterogeneous event schemas and extraction difficulty levels\.

Table 10:Statistics of the end\-to\-end event extraction datasets used in our experiments\.

## Appendix MAdditional Results and Analysis

Table[12](https://arxiv.org/html/2609.29230#A13.T12), Table[13](https://arxiv.org/html/2609.29230#A13.T13)and Table[14](https://arxiv.org/html/2609.29230#A13.T14)show the detailed results of EE\.

### M\.1Trigger and Argument Performance Analysis

We decompose event extraction quality into trigger\-level \(TI, TC\) and argument\-level \(AI, AC, AI\+, AC\+\) subtasks using the detailed results and isolating how each reward contributes to the global EE quality\.

#### M\.1\.1Event Trigger Performance \(TI,TC\)

TheREE\+RfmtR\_\{\\text\{EE\}\}\+R\_\{\\text\{fmt\}\}baseline achieves competitive TI and TC scores across most datasets indicating that schema\-grounded instruction tuning already provides a reasonable foundation for trigger localisation and type assignment\.RcovR\_\{\\text\{cov\}\}yields the most consistent TI gains, by penalising missed gold events and directly improving trigger recall\.RspanR\_\{\\text\{span\}\}contributes primarily to TC by refining span boundaries and stabilising type\-assignment confidence\.RovrR\_\{\\text\{ovr\}\}introduces a precision\-recall trade\-off, reducing spurious triggers but degrading TI on datasets where the baseline already under\-generates\. Combining all rewards yields the strongest TI and TC on most datasets, with the largest gains on M2E2 \(TI:\+11\.17\+11\.17, TC:\+14\.29\+14\.29over baseline\) and MLEE \(TI:\+3\.22\+3\.22, TC:\+4\.22\+4\.22\), where multi\-event documents benefit most from joint coverage and precision supervision\.

#### M\.1\.2Argument Performance \(AI, AC, AI\+, AC\+\)

Argument metrics are significantly lower than trigger metrics, and the anchored variants AI\+ and AC\+ collapse further exposing systematic trigger\-argument misalignment that is not apparent using mean\-F1\.RspanR\_\{\\text\{span\}\}is the single most impactful reward for argument tasks, delivering the largest per\-dataset gains across all four subtasks \(PHEE: AC\+14\.15\+14\.15, AC\+\+9\.64\+9\.64; M2E2: AC\+5\.54\+5\.54, AC\+\+4\.63\+4\.63\), as boundary\-level supervision directly addresses the span imprecision that drives argument scoring errors\.RovrR\_\{\\text\{ovr\}\}provides strong corrections in over\-extraction behaviors \(M2E2: AI\+\+8\.23\+8\.23, AC\+\+9\.36\+9\.36\) but degrades recall in domains such as CASIE and MLEE\.RgrdR\_\{\\text\{grd\}\}improves AI and AC where hallucinated spans are prevalent \(WikiEvents, PHEE\) but shows limited benefit for AI\+/AC\+, indicating that verbatim\-span enforcement alone does not resolve trigger\-argument misalignment\. When resorting to the full reward set produces the highest AI\+/AC\+ scores on the majority of datasets, with the most pronounced gains on M2E2 \(AI\+:\+14\.87\+14\.87, AC\+:\+17\.65\+17\.65over baseline\) and PHEE \(AC\+:\+11\.02\+11\.02\)\. These improvements suggest that the combined reward induces trigger\-argument alignment beyond what individual objectives achieve in isolation\.

### M\.2Additional Error Analysis

Figures[9](https://arxiv.org/html/2609.29230#A13.F9),[10](https://arxiv.org/html/2609.29230#A13.F10),[11](https://arxiv.org/html/2609.29230#A13.F11),[12](https://arxiv.org/html/2609.29230#A13.F12)and[13](https://arxiv.org/html/2609.29230#A13.F13)show a detailed comparative evaluation of error categories across datasets for each proposed reward\. The acronyms of error categories are defined as follows:ME: Missing events,EE: Extra events,EE: Extra events,MA: Missing arguments,EA: Extra arguments,RC: Role confusion,HA: Hallucinations,TS: Trigger span boundary errors,AS: Argument span boundary errors,ER: Extra roles,EE: Extra events,MR: Missing roles\. Error categories are defined in Table[11](https://arxiv.org/html/2609.29230#A13.T11)\.

AcronymError CategoryDefinitionEvent\-level ErrorsMEMissing EventsA gold event instance is absent from the model’s output; the trigger and all its associated arguments are undetected\.EEExtra EventsThe model predicts an event instance with no corresponding gold event; a spurious trigger is generated that is not grounded in the annotation\.Argument\-level ErrorsMAMissing ArgumentsA gold argument span is not predicted for an otherwise correctly identified event; the event is detected but one or more of its argument slots are left unfilled\.EAExtra ArgumentsThe model predicts one or more argument spans that have no corresponding gold argument; spurious fillers are generated beyond the gold annotation\.ERExtra RolesThe model populates an argument role that does not exist in the gold annotation for the predicted event type, introducing schema\-inconsistent role assignments\.MRMissing RolesA role defined in the gold event schema is entirely absent from the predicted event instance, resulting in incomplete role coverage\.RCRole ConfusionAn argument span is correctly extracted from the source text but assigned to an incorrect semantic role within the event schema; the span is right but the role label is wrong\.Span\-level ErrorsTSTrigger Span ErrorThe predicted trigger span does not exactly match the gold trigger boundary; includes both under\-specified spans \(partial overlap\) and over\-specified spans \(superset of the gold mention\)\.ASArgument Span ErrorThe predicted argument span does not exactly match the gold argument boundary; includes both partial and over\-extended span predictions relative to the minimal gold span\.Faithfulness ErrorsHAHallucinationThe predicted trigger or argument span does not appear verbatim in the source input text; the model generates mentions unsupported by the source document\.Structural ErrorsPEParsing ErrorThe model output cannot be parsed into the required code\-based structured representation; the generated Python code is syntactically malformed or fails to instantiate valid event objects\.Table 11:Definitions of all error categories used in the error analysis \(Section[5\.3](https://arxiv.org/html/2609.29230#S5.SS3)\)\. Categories are grouped into five types: event\-level errors concerning the detection of event instances, argument\-level errors concerning role assignment and completeness, span\-level errors concerning boundary precision, faithfulness errors concerning grounding in the source text, and structural errors concerning the syntactic validity of the generated output\.Figure 9:Comparative evaluation of error categories across datasets\.REER\_\{\\text\{EE\}\}\+RfmtR\_\{\\text\{fmt\}\}Vs\. \+RspanR\_\{\\text\{span\}\}Figure 10:Comparative evaluation of error categories across datasets\.REER\_\{\\text\{EE\}\}\+RfmtR\_\{\\text\{fmt\}\}Vs\. \+RovrR\_\{\\text\{ovr\}\}Figure 11:Comparative evaluation of error categories across datasets\.REER\_\{\\text\{EE\}\}\+RfmtR\_\{\\text\{fmt\}\}Vs\. \+RgrdR\_\{\\text\{grd\}\}Figure 12:Comparative evaluation of error categories across datasets\.REER\_\{\\text\{EE\}\}\+RfmtR\_\{\\text\{fmt\}\}Vs\. \+RcovR\_\{\\text\{cov\}\}Figure 13:Comparative evaluation of error categories across datasets\.REER\_\{\\text\{EE\}\}\+RfmtR\_\{\\text\{fmt\}\}Vs\. AllTable 12:Full results on WikiEvents, PHEE and CASIE datasets\.Table 13:Full results on Genia2011 and Genia2013 datasets\.Table 14:Full results on MLEE and M2E2 datasets\.

## Appendix NExamples Event Schema with Generated Annotation Guidelines

Figures[14](https://arxiv.org/html/2609.29230#A14.F14),[15](https://arxiv.org/html/2609.29230#A14.F15),[16](https://arxiv.org/html/2609.29230#A14.F16),[17](https://arxiv.org/html/2609.29230#A14.F17),[18](https://arxiv.org/html/2609.29230#A14.F18),[19](https://arxiv.org/html/2609.29230#A14.F19)and[20](https://arxiv.org/html/2609.29230#A14.F20), illustrate examples of one event scheme augmented with annotation guidelines for each dataset\.

1@dataclass

2classArtifactExistence\_DamageDestroyDisableDismantle\_Damage:

3"""TheeventtypeArtifactExistence\_DamageDestroyDisableDismantle\_Damageistriggeredbythementionofdamageordestructionofanartifact,whichcanbeaphysicalobject,structure,orentity\.Itinvolvestheconceptofcausingharmordamagetosomething,resultinginitsdegradationorlossoffunctionality\.UnlikeArtifactExistence\_DamageDestroyDisableDismantle\_Unspecified,thiseventtypespecificallyfocusesondamageordestruction,excludingotherformsofdisablementordismantling\.Triggerssuchas’damage’,’destroy’,’disable’,or’dismantle’areindicativeofthiseventtype\.Examplesinclude:’Thebuildingwasdamagedintheearthquake\.’,’Thecarwasdestroyedintheaccident\.’,’Thebridgewasdisabledbytheprotesters\.’,’Theoldfactorywasdismantledforredevelopment\.’"""

4mention:str

5Artifact:List

6DamagerDestroyer:List

7Instrument:List

8Place:List

9Damager:List

Figure 14:WikiEventsevent schema python class example1@dataclass

2classAdverse\_event:

3"""Theeventistriggeredbythementionofanadverseevent,whichisaharmfulorundesirableeffectthatoccursasaresultofatreatment,medication,orotherintervention\.Adverseeventscanbecausedbyacombinationofdrugs,asingledrug,oratreatmentduration\.Examplesareintravenousazithromycin\-inducedototoxicity,unaccountableseverehypercalcemiainapatienttreatedforhypoparathyroidismwithdihydrotachysterol,andprolongedsevere5\-fluorouracil\-associatedneurotoxicityinapatientwithdihydropyrimidinedehydrogenasedeficiency\.Unlikeotherevents,adverseeventsarenottherapeuticorbeneficial,andtheyareoftencharacterizedbynegativeoutcomessuchashematologicadversereactions,pulmonarytoxicity,orsupravenoushyperpigmentation\.Triggerssuchas’developed’,’induced’,’become’,’on’,and’by’areindicativeofadverseevents,notothereventtypes\."""

4mention:str

5Combination\_Drug:List

6Effect:List

7Subject:List

8Subject\_Age:List

9Subject\_Disorder:List

10Subject\_Gender:List

11Subject\_Population:List

12Subject\_Race:List

13Treatment:List

14Treatment\_Disorder:List

15Treatment\_Dosage:List

16Treatment\_Drug:List

17Treatment\_Duration:List

18Treatment\_Freq:List

19Treatment\_Route:List

20Treatment\_Time\_elapsed:List

Figure 15:PHEEevent schema python class example1@dataclass

2classAttack\_Ransom:

3"""Theeventistriggeredbytheoccurrenceofaransomwareattack,whereanattackerdemandspaymentinexchangeforrestoringaccesstoencrypteddata\.Theeventischaracterizedbytheuseofransomware,encryptionofdata,andthedemandforpayment\.Unlikeothertypesofattacks,Attack\_Ransomeventsinvolvetheuseofransomwareandthedemandforpaymenttorestoreaccesstodata\.Triggerssuchas’demandedaransom’and’ransomwareattacks’areindicativeofAttack\_Ransomevents\.ExamplesofAttack\_Ransomeventsincludeinstanceswheredataisencryptedandaransomisdemanded,suchas’demandedinpayment’and’demandedaransom’\."""

4mention:str

5Attack\_Pattern:List

6Attacker:List

7Damage\_Amount:List

8Payment\_Method:List

9Place:List

10Price:List

11Time:List

12Tool:List

13Victim:List

Figure 16:CASIEevent schema python class example1@dataclass

2classProtein\_catabolism:

3"""Theeventistriggeredbytheprocessofbreakingdownordegradingproteins\.Thisprocesscanbeinducedbyvariousmechanisms,includingproteasomaldegradation,ubiquitination,andphosphorylation\.Thekeycharacteristicsofthiseventincludethedegradationofproteins,whichcanbearesultofvariouscellularprocesses\.UnlikeothereventssuchasGene\_expression,thiseventisfocusedonthebreakdownofproteinsratherthantheirsynthesis\.Triggerssuchas’degradation’,’proteolyticallydegraded’,and’ubiquitination’areindicativeofthiseventtype\.ExamplesofproteindegradationincludethebreakdownofIkappaBalpha,A3G,andp27kip1\.Thescopeofthiseventislimitedtothedegradationofproteins,anditdoesnotincludeothercellularprocessessuchastranscriptionorphosphorylation\."""

4mention:str

5Theme:List

Figure 17:Genia2011event schema python class example1@dataclass

2classPhosphorylation:

3"""Theeventistriggeredbythephosphorylationofamolecule,typicallyaprotein\.Phosphorylationisachemicalreactionthataddsaphosphategrouptoamolecule,oftenalteringitsfunctionoractivity\.Thiseventischaracterizedbythetransferofaphosphategroupfromaphosphatedonortoaprotein,resultinginachangeintheprotein’sconformationoractivity\.Examplesoftriggersinclude’phosphorylated’,’phosphorylate’,and’phosphorylation’\.UnlikeProtein\_catabolism,thiseventdoesnotinvolvethedegradationofaprotein,andunlikeProtein\_modification,itdoesnotinvolvetheadditionorremovalofanon\-phosphategroup\.Triggerssuchas’phosphorylation’areindicativeofPhosphorylation,notProtein\_modification\."""

4mention:str

5Cause:List

6Site:List

7Theme:List

Figure 18:Genia2013event schema python class example1@dataclass

2classDephosphorylation:

3"""Theeventistriggeredbytheremovalofaphosphategroupfromaprotein,typicallyasaresultoftheactionofaphosphataseenzyme\.Theeventischaracterizedbythedephosphorylationofaspecificproteinsite\.UnlikePhosphorylation,thiseventdoesnotinvolvetheadditionofaphosphategroup\.Triggerssuchas’dephosphorylation’,’phosphatase’,or’dephosphorylated’areindicativeofDephosphorylation,notPhosphorylation\.Examplesare’dephosphorylationofMcl\-1’,’phosphataseactivity’,’removalofphosphategroup’\."""

4mention:str

5Site:List

6Theme:List

Figure 19:MLEEevent schema python class example1@dataclass

2classJustice\_arrestJail:

3"""Theeventistriggeredbythearrestordetentionofapersonorpersons,whichcanbeinitiatedbyauthoritiesorlawenforcement\.Thiseventtypecoversvariouscontexts,includingbutnotlimitedto,arrests,detentions,andimprisonments\.Unlikeothereventtypes,suchasConflict\_Attack,thiseventdoesnotinvolvephysicalharmorviolence\.Triggerssuchas’arrested’,’jailed’,’detained’,and’imprisoned’areindicativeofthiseventtype\.Examplesinclude’He’sbeenarrestedanddetailswillsoonbereleased’,’Localmediacirculatedaphotoofwhattheydescribedasthemomenthewasarrested’,and’Awebsearchfoundreportsgoingbackfiveyearswhereauthoritieshadarrestedorchargedteenagersforrecruitingotherteengirlsforprostitution’\."""

4mention:str

5Agent:List

6Person:List

7Place:List

Figure 20:M2E2event schema python class example

Similar Articles

Incentivizing Parametric Knowledge via Reinforcement Learning with Verifiable Rewards for Cross-Cultural Entity Translation

arXiv cs.CL

Researchers from Tianjin University and Alibaba Group propose EA-RLVR, a reinforcement learning framework with verifiable rewards that improves cross-cultural entity translation in LLMs by activating parametric knowledge already encoded during pre-training, without relying on external knowledge bases. Training on 7k samples boosts Qwen3-14B's entity translation accuracy from 23.66% to 31.87% on unseen entities.

CogEvol: Towards Efficient and Reliable Learning Environment Generation

Hugging Face Daily Papers

CogEvol is a family of models that efficiently generate structured learning artifacts like slides and interactive HTML pages in a single pass using supervised fine-tuning and reinforcement learning with vision-language rewards, reducing cost and improving reliability.