From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

arXiv cs.CL Papers

Summary

This paper introduces MiGUE-Bench, a systematic benchmark for evaluating LLMs on multi-granularity event analysis, spanning single- to cross-document tasks including event detection, relation reasoning, structure induction, and future prediction.

arXiv:2607.27654v1 Announce Type: new Abstract: Event analysis is an essential and fundamental direction of information extraction, involving various event-centric tasks at different granularity of documents. While large language models (LLMs) have preliminarily achieved promising performance in part of these tasks individually, their capability in event analysis still lacks comprehensive understanding due to restricted document granularity, task designs, and data source of existing benchmarks. To address these limitations, we introduce MiGUE-Bench, a systematic benchmark for assessing the performance of LLMs in multi-granularity event analysis. To support large-scale evaluation, we first develop an LLM-driven self-correcting annotation framework called MiGUE-Pipeline, enabling scalable acquisition of high-quality source data of events with automatic labels. Then, we design four core tasks in our benchmark, i.e., event detection, relation reasoning, structure induction, and future prediction, to probe model competence at different levels, from atomic event details to complex cross-document narratives. Extensive experiments on state-of-the-art LLMs and retrieval-augmented generation (RAG) methods delineate the current capability boundary and identify critical deficiencies, providing insights into the future improvement of LLMs in challenging event analysis tasks.
Original Article
View Cached Full Text

Cached at: 07/31/26, 10:02 AM

# From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models
Source: [https://arxiv.org/html/2607.27654](https://arxiv.org/html/2607.27654)
Tao WenLaboratory of Intelligent Collaborative ComputingUniversity of Electronic Science and Technology of ChinaChengduChina[enril\.wentao@std\.uestc\.edu\.cn](https://arxiv.org/html/2607.27654v1/mailto:[email protected])Shuai ShaoLaboratory of Intelligent Collaborative ComputingUniversity of Electronic Science and Technology of ChinaChengduChina[202221080331@std\.uestc\.edu\.cn](https://arxiv.org/html/2607.27654v1/mailto:[email protected]),Pei KeLaboratory of Intelligent Collaborative ComputingUniversity of Electronic Science and Technology of ChinaChengduChina[kepei@uestc\.edu\.cn](https://arxiv.org/html/2607.27654v1/mailto:[email protected]),Xu HanDepartment of Computer Science and TechnologyTsinghua UniversityBeijingChina[han\-xu@mail\.tsinghua\.edu\.cn](https://arxiv.org/html/2607.27654v1/mailto:[email protected]),Jie ZouSchool of Computer Science and EngineeringUniversity of Electronic Science and Technology of ChinaChengduChina[jie\.zou@uestc\.edu\.cn](https://arxiv.org/html/2607.27654v1/mailto:[email protected]),Guannan LiLaboratory of Intelligent Collaborative ComputingUniversity of Electronic Science and Technology of ChinaChengduChina[lgn4sci@std\.uestc\.edu\.cn](https://arxiv.org/html/2607.27654v1/mailto:[email protected]),Tao Tian,Jinjie QiuLaboratory of Intelligent Collaborative ComputingUniversity of Electronic Science and Technology of ChinaChengduChina[taotian@std\.uestc\.edu\.cn](https://arxiv.org/html/2607.27654v1/mailto:[email protected])[202522900127@std\.uestc\.edu\.cn](https://arxiv.org/html/2607.27654v1/mailto:[email protected]),Lan WangSchool of Information and Software EngineeringUniversity of Electronic Science and Technology of ChinaChengduChina[202521090303@std\.uestc\.edu\.cn](https://arxiv.org/html/2607.27654v1/mailto:[email protected])andKe QinLaboratory of Intelligent Collaborative ComputingUniversity of Electronic Science and Technology of ChinaChengduChina[qinke@uestc\.edu\.cn](https://arxiv.org/html/2607.27654v1/mailto:[email protected])

\(2026\)

###### Abstract\.

Event analysis is an essential and fundamental direction of information extraction, involving various event\-centric tasks at different granularity of documents\. While large language models \(LLMs\) have preliminarily achieved promising performance in part of these tasks individually, their capability in event analysis still lacks comprehensive understanding due to restricted document granularity, task designs, and data source of existing benchmarks\. To address these limitations, we introduce MiGUE\-Bench, a systematic benchmark for assessing the performance of LLMs in multi\-granularity event analysis\. To support large\-scale evaluation, we first develop an LLM\-driven self\-correcting annotation framework called MiGUE\-Pipeline, enabling scalable acquisition of high\-quality source data of events with automatic labels\. Then, we design four core tasks in our benchmark, i\.e\., event detection, relation reasoning, structure induction, and future prediction, to probe model competence at different levels, from atomic event details to complex cross\-document narratives\. Extensive experiments on state\-of\-the\-art LLMs and retrieval\-augmented generation \(RAG\) methods delineate the current capability boundary and identify critical deficiencies, providing insights into the future improvement of LLMs in challenging event analysis tasks\.

Large Language Model, Event Analysis, Automatic Evaluation

††journalyear:2026††copyright:cc††conference:Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval; July 20–24, 2026; Melbourne, VIC, Australia\.††booktitle:Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval \(SIGIR ’26\), July 20–24, 2026, Melbourne, VIC, Australia††isbn:979\-8\-4007\-2599\-9/2026/07††doi:10\.1145/3805712\.3808607††ccs:Computing methodologies Information extraction## 1\.Introduction

Event analysis is a foundational pillar of information extraction, inherently requiring to extract the structure of events from raw textual data at different granularity levels\(Chenet al\.,[2021](https://arxiv.org/html/2607.27654#bib.bib25)\)\. This essential direction encompasses a hierarchical task taxonomy from the fine\-grained event detection of individual occurrences to the induction of global event structures across different documents\(Minardet al\.,[2015](https://arxiv.org/html/2607.27654#bib.bib1); Liet al\.,[2021](https://arxiv.org/html/2607.27654#bib.bib22); Yanget al\.,[2025](https://arxiv.org/html/2607.27654#bib.bib2); Foleyet al\.,[2015](https://arxiv.org/html/2607.27654#bib.bib48); Fanet al\.,[2022](https://arxiv.org/html/2607.27654#bib.bib49); Sankepally,[2019](https://arxiv.org/html/2607.27654#bib.bib50); Louet al\.,[2022](https://arxiv.org/html/2607.27654#bib.bib51); Liaoet al\.,[2021](https://arxiv.org/html/2607.27654#bib.bib52)\)\. Mastering the capabilities of analyzing events from intra\-document event mentions to inter\-document event relations is indispensable for the application of current information extraction systems in real\-world scenarios, moving beyond local semantic matching towards global relation reasoning\.

Early works on event analysis have extensively explored task\-specific deep learning paradigms, ranging from sequence labeling architectures\(Chenet al\.,[2015](https://arxiv.org/html/2607.27654#bib.bib3); Nguyenet al\.,[2016](https://arxiv.org/html/2607.27654#bib.bib4); Shaet al\.,[2018](https://arxiv.org/html/2607.27654#bib.bib6)\)to graph\-based reasoning frameworks\(Liet al\.,[2021](https://arxiv.org/html/2607.27654#bib.bib22); Nguyen and Grishman,[2018](https://arxiv.org/html/2607.27654#bib.bib7)\)\. Despite their advanced performance on specific datasets, the generalization ability of these models is rather limited\. Thus, recent works resort to large language models \(LLMs\) and preliminarily show their promising performance in part of event analysis tasks, such as event extraction\(Liet al\.,[2025](https://arxiv.org/html/2607.27654#bib.bib26)\), event relation prediction\(Chenet al\.,[2024](https://arxiv.org/html/2607.27654#bib.bib9); Huet al\.,[2025](https://arxiv.org/html/2607.27654#bib.bib27)\), and event reasoning\(Nakshatriet al\.,[2023](https://arxiv.org/html/2607.27654#bib.bib14); Taoet al\.,[2025](https://arxiv.org/html/2607.27654#bib.bib28)\)\. Equipped with strong abilities of context understanding and knowledge utilization, LLMs have shown great potential in dealing with complex event analysis tasks in a zero\-shot manner, gradually becoming a research focus in this field\.

However, we argue that there still lacks a comprehensive benchmark for assessing LLMs’ capabilities in event analysis\. Existing benchmarks mostly suffer from restricted granularity, lacking data source, and homogeneous task design, hindering a systematic understanding of LLMs’ deficiencies at different event\-centric tasks:

- •Restricted Granularity: Most of the current benchmarks only fall into the assessment at constrained granularity of documents\. While some of the datasets aim to measure the model performance in understanding events within a document\(Taoet al\.,[2025](https://arxiv.org/html/2607.27654#bib.bib28); Wanget al\.,[2022](https://arxiv.org/html/2607.27654#bib.bib17)\), others focus on the event relations across multiple documents\(Honget al\.,[2016](https://arxiv.org/html/2607.27654#bib.bib29); Bugertet al\.,[2021](https://arxiv.org/html/2607.27654#bib.bib30); Zhanget al\.,[2024](https://arxiv.org/html/2607.27654#bib.bib18)\), both of which fail to provide an entire perspective of LLMs’ capabilities at different granularity levels\.
- •Lacking Data Source: Most existing benchmarks are retrofitted from legacy datasets\(Gonget al\.,[2025](https://arxiv.org/html/2607.27654#bib.bib19)\), which restrict the scalability to diverse and contemporary corpora\. The lack of data sources also causes missing cross\-document dependencies, making it improper to measure LLMs’ capabilities to analyze events across multiple documents\.
- •Homogeneous Task Design: Existing benchmarks are mostly targeting at isolated evaluation dimensions and objects, analyzing the ability of LLMs to deal with specific event relations\(Zhanget al\.,[2024](https://arxiv.org/html/2607.27654#bib.bib18)\)or tasks\(Chenet al\.,[2024](https://arxiv.org/html/2607.27654#bib.bib9); Nakshatriet al\.,[2023](https://arxiv.org/html/2607.27654#bib.bib14); Huet al\.,[2025](https://arxiv.org/html/2607.27654#bib.bib27); Taoet al\.,[2025](https://arxiv.org/html/2607.27654#bib.bib28)\)\. Such homogeneous and narrow task design may be unable to reveal the capability boundary of general LLMs, exaggerating their performance in the limited task scope\.

To address these limitations, we introduceMiGUE\-Bench, a comprehensive benchmark designed forMulti\-GranUlarEvent analysis\. MiGUE\-Bench aims to cover the evolving process of event analysis from single\-document event detection and relation reasoning to cross\-document event structure induction and future prediction \(as shown in Table[1](https://arxiv.org/html/2607.27654#S1.T1)\), enabling a holistic assessment of LLMs’ capabilities in different granularity levels of documents\.Firstly, to deal with the limitation of document granularity and data source, we develop an LLM\-driven automatic pipeline named MiGUE\-Pipeline that includes document filtering, event / relation annotation, and document clustering, enabling scalable data construction from raw corpora to support various event\-centric tasks\.Secondly, to solve the challenge of homogeneous task design, we devise four task types to cover different stages of event analysis, Specifically,MiGUE\-Detectionassesses the performance in recognizing and extracting triggers from event mentions, whileMiGUE\-Reasoningaims to evaluate the capability to reason within and across fragmented document contexts to extract implicit event relations\. To further obtain a global understanding of event dynamics across different documents,MiGUE\-Inductionis targeted at evaluating the ability to acquire the global topological structure of events\. Finally,MiGUE\-Predictiontests the capability to infer the future development of events based on the understanding of event dynamics across different documents\. Our benchmark is expected to reveal the deficiencies of LLMs via comprehensive and challenging task design\.

- •We develop the first open\-source automatic pipeline named MiGUE\-Pipeline for multi\-granularity event data generation, alleviating the dependence on manual annotation and flexibly supporting various event\-centric tasks\.
- •We introduce MiGUE\-Bench, a comprehensive benchmark that provides systematic assessment for LLMs’ capabilities in the entire spectrum of event analysis\.
- •We conduct extensive experiments on state\-of\-the\-art LLMs and retrieval\-augmented generation frameworks, uncovering critical deficiencies of LLMs on event analysis and providing insights for future improvement\.

Table 1\.Comparison of different event analysis benchmarks\.BenchmarkACE05\-EN\(Doddingtonet al\.,[2004](https://arxiv.org/html/2607.27654#bib.bib55)\)✓✗✗✗Causal\-TimeBank\(Mirza and Tonelli,[2014](https://arxiv.org/html/2607.27654#bib.bib54)\)✓✗✓✗MAVEN\(Wanget al\.,[2020](https://arxiv.org/html/2607.27654#bib.bib53)\)✓✗✗✗MAVEN\-ERE\(Wanget al\.,[2022](https://arxiv.org/html/2607.27654#bib.bib17)\)✓✗✓✗TCELongBench\(Zhanget al\.,[2024](https://arxiv.org/html/2607.27654#bib.bib18)\)✗✓✗✗EventRelBench\(Gonget al\.,[2025](https://arxiv.org/html/2607.27654#bib.bib19)\)✓✗✓✗MiGUE\-Bench \(Ours\)✓✓✓✓
## 2\.Related Work

LLMs for Event Analysis\.Recent advances in large language models \(LLMs\) have enabled a prompt\-based paradigm for event\-centric understanding and reasoning in a zero\-/few\-shot setting\. Existing studies show that LLMs can support core event analysis capabilities, spanning event extraction\(Liet al\.,[2025](https://arxiv.org/html/2607.27654#bib.bib26); Choudhary and Du,[2024](https://arxiv.org/html/2607.27654#bib.bib44); Huanget al\.,[2024a](https://arxiv.org/html/2607.27654#bib.bib45); Liet al\.,[2022](https://arxiv.org/html/2607.27654#bib.bib47)\)and event relation prediction\(Chenet al\.,[2024](https://arxiv.org/html/2607.27654#bib.bib9); Huet al\.,[2025](https://arxiv.org/html/2607.27654#bib.bib27); Chanet al\.,[2024](https://arxiv.org/html/2607.27654#bib.bib46); Taoet al\.,[2025](https://arxiv.org/html/2607.27654#bib.bib28)\)\. Beyond event understanding in a single document, LLMs have been utilized for temporal/causal event reasoning across documents, commonly with retrieval modules to support more coherent information extraction\(Nakshatriet al\.,[2023](https://arxiv.org/html/2607.27654#bib.bib14); Taoet al\.,[2025](https://arxiv.org/html/2607.27654#bib.bib28); Yanget al\.,[2025](https://arxiv.org/html/2607.27654#bib.bib2)\)\. To summarize, existing studies suggest that LLMs hold strong potential for bridging multiple granularity of event analysis, ranging from single\-document understanding to cross\-document reasoning\.

Despite this advance, current work on LLM\-based event analysis typically focuses on isolated task settings, making it hard to delineate the strengths and weaknesses of LLMs along the whole process of event analysis, moving from local semantics to global structures\. Therefore, a more systematic benchmark is necessary to comprehensively reflect the LLMs’ performance at multiple granularity levels of event analysis\.

Benchmarks for LLMs in Event Analysis\.Motivated by the recent progress of LLMs on event\-centric tasks, several benchmarks have been proposed to probe LLM capabilities along different facets of event analysis\. Existing works typically formulate evaluation as instruction\-following tasks with structured outputs, aiming to measure whether LLMs can \(i\) recover event\-centric descriptions from long and temporally dense narratives\(Taoet al\.,[2025](https://arxiv.org/html/2607.27654#bib.bib28); Wanget al\.,[2022](https://arxiv.org/html/2607.27654#bib.bib17)\), \(ii\) reason about temporal/causal dependencies among events\(Honget al\.,[2016](https://arxiv.org/html/2607.27654#bib.bib29); Bugertet al\.,[2021](https://arxiv.org/html/2607.27654#bib.bib30); Zhanget al\.,[2024](https://arxiv.org/html/2607.27654#bib.bib18)\), and \(iii\) infer event relations under diverse types and conduct event reasoning with different formats\(Zhanget al\.,[2024](https://arxiv.org/html/2607.27654#bib.bib18); Gonget al\.,[2025](https://arxiv.org/html/2607.27654#bib.bib19)\)\.

Though existing benchmarks offer a preliminary view of LLMs’ event analysis capabilities, they mostly focus on specific event\-centric tasks with restricted granularity, data sources, and task design\. For comparison, our work aims to provide an complete picture of LLM performance involving different stages in the whole process of event analysis, with multiple granularity levels of documents, diverse data sources, and broad task scopes\.

## 3\.MiGUE\-Pipeline

To support scalable data construction of MiGUE\-Bench, we present MiGUE\-Pipeline, a four\-stage framework comprisingdocument filtering,event annotation,relation annotation, andcluster generation, as shown in Figure[1](https://arxiv.org/html/2607.27654#S3.F1)\. Transforming raw corpora into high\-quality multi\-granularity event data resources can serve as the foundation for subsequent evaluation\.

### 3\.1\.Document Filtering

We firstly employ an LLM\-driven filtering mechanism that preliminarily evaluates each candidate document across three critical dimensions:factual informativeness\(prioritizing objective reporting over subjective commentary or fragmentary discourse\),text quality\(ensuring the documents’ grammatical fluency and structural coherence\),event density\(excluding redundant or non\-informative contents\)\. A document will be retained only if it satisfies all the aforementioned criteria based on LLM\-as\-a\-Judge with GPT\-4o\(Liuet al\.,[2023](https://arxiv.org/html/2607.27654#bib.bib31)\)\. The multi\-dimensional filtering ensures that the resulting documents are both grammatically fluent and rich in event dynamics\.

### 3\.2\.Event Annotation

Since identifying events in open documents is challenging, we conduct a pilot study on part of documents with GPT\-5\.2\(OpenAI,[2026](https://arxiv.org/html/2607.27654#bib.bib32)\)followed by manual check, finding six typical error types, i\.e\.,No\-Occurrence, Negated\-Claim, Assumption, Abstraction, Named\-Entity,andNarrative\. Detailed descriptions of all these error types are provided in Table[2](https://arxiv.org/html/2607.27654#S3.T2)\. Based on these findings, we propose a two\-stage protocol for event annotation\. Firstly, we make an effective LLM \(i\.e\., DeepSeek\-V3\.2\(DeepSeek\-AI,[2025](https://arxiv.org/html/2607.27654#bib.bib36)\)\) generate a candidate set of event triggers\. Then, we devise a retrieval\-augmented reflection method to select semantically similar examples in the data pool of each error type from our pilot study, respectively\. This method requires the LLM to self\-correct the results by reasoning against similar historical failures\. Manual validation on a subset of 200 data samples shows that our automatic protocol yields a precision of 0\.87 during event annotation\.

Table 2\.Description of Event Annotation Error Types\.Error TypeDescriptionNo\-OccurrenceNo\-Occurrenceerrors arise when the model labels the expressions describing future plans, intentions, or predictions rather than events that have actually occurred as triggers\. Although such expressions are event\-like in form, they are not realized on the current timeline and should not be extracted as events\.Negated\-ClaimNegated\-Claimerrors arise when the model fails to distinguish between the statements about events that did not actually occur and genuinely occurring events that express negation \(e\.g\., denying\)\. Negation operators themselves \(e\.g\., not, no, never\) do not introduce new events, whereas verbs such as deny, reject, and refute denote individual events and should be identified as triggers\.AssumptionAssumptionerrors arise when the model identifies as triggers those verbs or action expressions that carry event semantics but occur only in conditional, hypothetical, intentional, advisory, or exhortative contexts, rather than being asserted as factual events on the real timeline\.AbstractionAbstractionerrors arise when the model labels the expressions that are event\-like in form but function rhetorically or metaphorically rather than denote a specific real\-world event as triggers\. Such expressions lack clear temporal anchoring, identifiable participants, and concrete boundaries\. Thus, they are not extractable factual events\.Named\-EntityNamed\-Entityerrors arise when the model incorrectly labels nominal elements of an event \(e\.g\., participants, carriers, results, named entities, event\-denoting nouns, or state descriptions\) as triggers\.NarrativeNarrativeerrors arise when the model labels the expressions that narrate, explain, modify, or evaluate an event rather than denote the core action as event triggers\. Such expressions \(e\.g\., stance and manner\) do not introduce new events but only supplement or interpret existing ones\. Thereby, they should not be annotated as events\.

### 3\.3\.Relation Annotation

To further acquire the relation between the events within a single document or among different ones, we devise different strategies to ensure the quality of automatic relation annotation\.

Single\-Document Relation Annotation\.Inspired by existing works\(Chenet al\.,[2024](https://arxiv.org/html/2607.27654#bib.bib9)\), we design a constraint\-aware multi\-step reasoning strategy to capture event dependencies within a document\. For each event pair, LLMs iteratively select the most confident relation type\. Upon selecting a candidate label at each iteration, LLMs are prompted to validate its decision against a set of explicit logical constraints \(e\.g\., a coreference label inherently precludes causal or sub\-event relations\)\. Based on self\-evaluation, the LLM should select one of the following operations, i\.e\.,confirmingthe current label to proceed to the next most confident relation type,revisingthe current label to select an alternative one, andrestartingthe entire decision process due to the conflict during reasoning\. Once a label is confirmed, we dynamically prune the search space by removing logically incompatible relation types\. To further enhance reliability, we employ majority voting with three LLMs \(i\.e\., GPT\-4o, DeepSeek\-V3\.2\(DeepSeek\-AI,[2025](https://arxiv.org/html/2607.27654#bib.bib36)\), and Qwen3\-Max\(Qwen Team,[2025](https://arxiv.org/html/2607.27654#bib.bib37)\)\)\. If all the models yield identical predictions, the result is accepted; otherwise, the instance is escalated to a high\-capacity LLM \(e\.g\., Claude\-4\.5\-Opus\) for final results\. This approach achieves an accuracy rate of 0\.82 on 100 data samples during manual check, which shows its effectiveness\.

![Refer to caption](https://arxiv.org/html/2607.27654v1/x1.png)Figure 1\.Overview of MiGUE\-Pipeline and MiGUE\-Bench\.Cross\-Document Relation Annotation\.For comparison, cross\-document relation annotation faces the challenges of relational sparsity and computational intractability\. To address these challenges, we propose a propagate\-via\-coreference strategy\. For preparation, we constrain the search space using three proximity heuristics, includingtemporal proximity\(where documents must fall within a one\-month timestamp window\),semantic affinity\(where event mentions should exceed a semantic similarity threshold\), andentity overlap\(where events should share common participants\)\. Potentially coreferential pairs are then validated using the same protocol as in the single\-document part\. In the propagate\-via\-coreference strategy, we densify the relation network by treating coreferential events as bridges for propagation\. By applying formal transitivity rules \(e\.g\., ife1\\ext@arrow0099\\arrowfill@===corefe2e\_\{1\}\\ext@arrow 0099\\arrowfill@\\Relbar\\Relbar\\Relbar\{\}\{coref\}e\_\{2\},e1→c​a​u​s​ee3e\_\{1\}\\xrightarrow\{cause\}e\_\{3\}, thene2→c​a​u​s​ee3e\_\{2\}\\xrightarrow\{cause\}e\_\{3\}\), we recover implicit dependencies across different documents\. This systematic approach circumvents exhaustive pairing while improving the global structural integrity of the event relation network\.

### 3\.4\.Cluster Generation

Since documents that share identical real\-world events are potentially linked by their underlying narrative contents, we utilize event coreference as a primary signal for document cluster generation, supporting the data construction of downstream tasks that require sufficient cross\-document relations\. Specifically, we construct a document affinity graph whose vertices represent individual documents and edges denote the presence of coreferential event pairs\. To partition this graph into dense subgraphs, we employ the Leiden community detection algorithm\(Traaget al\.,[2019](https://arxiv.org/html/2607.27654#bib.bib40)\), which isolates clusters with maximum narrative coherence\. By operating on these localized subgraphs whose relational density is highest, we streamline the discovery of closely\-connected documents for more efficient data construction of cross\-document event\-centric tasks\.

## 4\.MiGUE\-Benchmark

In this section, we systematically delineate the construction of MiGUE\-Bench, including four core tasks\. For each task , we provide a rigorous definition involving task inputs and outputs, followed by its data construction strategy with MiGUE\-Pipeline\.

### 4\.1\.MiGUE\-Detection

Since the ability of precisely detecting events is the foundational skill of event analysis, we devise the MiGUE\-Detection task to measure LLMs’ performance on distinguishing factual event triggers\.

Task Definition\.Given a task descriptionD​e​sd​e​tDes\_\{det\}, a sentenceSS, and a candidate set of event triggers𝒞=\{c1,c2,…,cr∣3⩽r⩽6\}\\mathcal\{C\}=\\left\\\{c\_\{1\},c\_\{2\},\.\.\.,c\_\{r\}\\mid 3\\leqslant r\\leqslant 6\\right\\\}whererrindicates the size of𝒞\\mathcal\{C\}, the model is required to provide all the valid triggers that appear inSSwithin the candidate set𝒞\\mathcal\{C\}\.

Data Construction\.The data instances for this task are directly acquired from the event annotation stage \(Section[3\.2](https://arxiv.org/html/2607.27654#S3.SS2)\)\. To further improve the difficulty of this fundamental task, we incorporate adversarial distractors into the options\. Specifically, we inject typical error types identified during event annotation of MiGUE\-Pipeline into candidate options for each instance\. By integrating these failures modes as distractors, we aim to assess the fine\-grained semantic understanding ability of LLMs\.

Table 3\.Statistics of MiGUE\-Bench including the number of data instances \(\#Instance\) and the average number of input tokens \(\#Token\) / options \(\#Option\)\.TaskSubtask\#Instance\#Token\#OptionDetection\-33055\.184\.56ReasoningIntraTemporal300796\.844\.00Causal300937\.314\.00Coreference300890\.002\.00Subevent300770\.583\.00CrossTemporal3001209\.374\.00Causal3001202\.184\.00Coreference3001202\.242\.00Subevent3001141\.263\.00InductionTemporal Order150671\.77\-Causal Graph1101431\.274\.48Prediction\-3002112\.034\.50Overall\-3,2901061\.243\.85
### 4\.2\.MiGUE\-Reasoning

Inferring relationships between events within / across documents is essential for event analysis\. MiGUE\-Reasoning focuses on evaluating LLMs’ capacity of reasoning across different events\.

Task Definition\.Given the task descriptionD​e​sr​e​aDes\_\{rea\}, the event pair\(e,e′\)\(e,e^\{\{\}^\{\\prime\}\}\)which comes from the documentD/D′D/D^\{\{\}^\{\\prime\}\}, respectively, the target relational dimensionℛ∈\{T​e​m​p​o​r​a​l,C​a​u​s​a​l,S​u​b​e​v​e​n​t,C​o−r​e​f​e​r​e​n​c​e\}\\mathcal\{R\}\\in\\\{Temporal,Causal,Subevent,Co\-reference\\\}, and a candidate set of relation labels𝒞=\{c1,c2,…,cr∣2⩽r⩽4\}\\mathcal\{C\}=\\\{c\_\{1\},c\_\{2\},\.\.\.,c\_\{r\}\\mid 2\\leqslant r\\leqslant 4\\\}, the model is required to select the a label from𝒞\\mathcal\{C\}that precisely characterizes the relationship within the corresponding dimension\. According to whetherDDandD′D^\{\{\}^\{\\prime\}\}are the same document, this task can be further categorized into two subtasks, i\.e\.,Reasoning\-IntraandReasoning\-Cross\.

Data Construction\.To ensure that our benchmark probes deep reasoning over events rather than surface\-level pattern matching, we explicitly exclude pairs linked by overt linguistic cues, such as direct temporal markers \(afterward\) or causal connectives \(because\)\. We retain only those relations that necessitate contextual understanding and multi\-hop reasoning\. We also prioritize event pairs that the effective LLMs \(such as GPT\-4o\) initially misclassify to further enhance the difficulty\.

### 4\.3\.MiGUE\-Induction

Global comprehension of event development necessitates inducing the topological structure of event across documents\. MiGUE\-Induction aims to measure the capacities of LLMs for structure induction through two sub\-tasks:Interval\-based Temporal OrderingandCausal Graph Construction\.

#### 4\.3\.1\.Interval\-based Temporal Ordering

Traditional temporal ordering often simplifies events into instantaneous points, relying on binary before/after comparisons\(Zhanget al\.,[2024](https://arxiv.org/html/2607.27654#bib.bib18)\)\. This task design fails to capture the temporal duration and interval logic \(e\.g\., overlap and containment\) inherently in real\-world narratives\. To address this limitation, we decouple each event into itsstartandendtemporal anchors for fine\-grained ordering\.

Task Definition\.Given the task descriptionD​e​so​r​d​e​rDes\_\{order\}, a set of event triggersℰ=\{e1,…,en∣n≥3\}\\mathcal\{E\}=\\\{e\_\{1\},\\dots,e\_\{n\}\\mid n\\geq 3\\\}, and a set of documents𝒟=\{D1,…,Dm∣m≥2\}\\mathcal\{D\}=\\\{D\_\{1\},\\dots,D\_\{m\}\\mid m\\geq 2\\\}where these events occur, the model should output two distinct permutations ofℰ\\mathcal\{E\}, representing the chronological order of event starts and event ends, respectively\.

Data Construction\.The data instances are derived from cross\-document temporal relations \(Section[3\.3](https://arxiv.org/html/2607.27654#S3.SS3)\) and document clusters \(Section[3\.4](https://arxiv.org/html/2607.27654#S3.SS4)\) generated by MiGUE\-Pipeline\. We specifically leverage the results of relation propagation, which yields dense temporal networks across different documents\. we prioritize clusters with high relation density, ensuring that the selected events form an interconnected network rather than isolated chains\. Furthermore, to prevent the models from relying on superficial pattern matching, we remove absolute temporal markers \(e\.g\., April 14, 2022\) in the data instances\. This forces the model to induct the relative temporal order from the narrative contents and implicit cues within the provided documents\.

#### 4\.3\.2\.Causal Graph Construction

Causality among events of different documents commonly manifests as a directed acyclic graph \(DAG\), where events may have multiple preconditions or cause several subsequent consequences\. Since causality is essential for structure induction of events, we design a task to assess the capability of constructing causal graphs\.

Task Definition\.Given the task descriptionD​e​sc​a​uDes\_\{cau\}, a set of event triggersℰ=\{e1,…,en∣n≥3\}\\mathcal\{E\}=\\\{e\_\{1\},\.\.\.,e\_\{n\}\\mid n\\geq 3\\\}, a set of documents𝒟=\{D1,…,Dm∣m≥2\}\\mathcal\{D\}=\\\{D\_\{1\},\.\.\.,D\_\{m\}\\mid m\\geq 2\\\}where these events occur, and a set of candidate options𝒞si=\{c1,c2,…,cr∣4⩽r⩽6\}\\mathcal\{C\}\_\{s\_\{i\}\}=\\\{c\_\{1\},c\_\{2\},\\dots,c\_\{r\}\\mid 4\\leqslant r\\leqslant 6\\\}, where each option corresponds to a causal path⟨e1→e2→…→en⟩\\langle e\_\{1\}\\rightarrow e\_\{2\}\\rightarrow\\dots\\rightarrow e\_\{n\}\\rangle, the model is required to identify all the longest valid causal chains within the causal DAG\.

Data Construction\.The data instances are curated from causal networks that are built based on causal relations \(Section[3\.3](https://arxiv.org/html/2607.27654#S3.SS3)\) and cluster generation \(Section[3\.4](https://arxiv.org/html/2607.27654#S3.SS4)\)\. To acquire challenging evaluation instances, we treat each maximal path in causal DAGs as a correct option and design three types of distractors as disturbing options, includingirrelevant links\(i\.e\., chains containing event pairs with no logical dependency\),incomplete path\(i\.e\., sub\-paths that are causally valid but fail to include all possible intermediate or terminal events\) , andincorrect order\(i\.e\., chains where events are causally related but presented in an incorrect logical order\)\.

Table 4\.Accuracy \(Acc\.\) and Micro\-F1 of different LLMs and RAG methods on MiGUE\-Benchmark\.GranularitySingle\-DocumentCross\-DocumentTaskDetectionReasoningInductionPredictionSubtask\-Reasoning\-IntraReasoning\-CrossTemporal OrderCausal Graph\-Temp\.Cau\.Coref\.Sub\.Temp\.Cau\.Coref\.Sub\.StartEndMetricMicro\-F1Acc\.Acc\.Acc\.Acc\.Acc\.Acc\.Acc\.Acc\.Acc\.Acc\.Acc\.Acc\.Closed\-Source LLMsGPT\-5\.2\-Pro0\.83980\.66000\.63730\.97350\.80330\.67860\.64080\.96360\.81880\.84120\.83330\.43750\.4231Gemini\-3\-Pro0\.71510\.68670\.62960\.95000\.78010\.79670\.55710\.91000\.80340\.66230\.37750\.21590\.5814Claude\-4\.5\-Opus0\.85520\.71000\.67780\.97000\.82890\.77810\.70480\.93330\.84210\.87420\.80790\.58140\.5544Claude\-4\.5\-Haiku0\.61010\.67670\.60000\.93000\.70000\.46330\.42380\.74000\.71640\.44370\.17220\.19320\.4452Qwen3\-Max0\.65430\.69670\.60370\.90500\.80220\.70670\.60000\.77000\.80990\.44370\.19870\.27270\.4585Open\-Source LLMsDeepSeek\-V3\.20\.81680\.66000\.66300\.96250\.79560\.67320\.62860\.85670\.84800\.82110\.63000\.28410\.5210GLM\-4\.70\.81330\.69330\.61670\.96330\.77660\.70660\.53330\.86000\.83000\.80670\.56000\.27270\.4867Kimi\-K20\.58940\.59000\.56300\.92000\.70440\.61670\.51430\.84000\.78070\.35760\.19870\.26140\.5150Qwen3\-8B0\.61380\.40000\.45990\.86770\.51330\.49000\.33900\.73330\.61700\.13910\.07950\.13640\.3355Qwen3\-30B0\.62980\.41000\.49260\.62110\.60220\.53720\.50950\.62000\.78950\.17220\.05300\.25000\.3522Qwen3\-235B0\.61150\.56000\.54810\.87250\.61330\.55670\.52860\.76670\.83620\.35100\.12580\.17050\.4352Llama\-3\.1\-7B0\.49590\.23260\.21850\.32960\.22380\.21670\.16670\.18000\.13160\.05960\.01320\.03410\.3654RAG Methods based on LLMsUltraRAG \(w/ Gemini\-3\-Pro\)0\.74570\.69390\.61480\.95250\.82630\.80000\.53330\.89000\.83040\.69540\.38410\.20450\.6246UltraRAG \(w/ Qwen3\-Max\)0\.69360\.70000\.60740\.91000\.79110\.70670\.61900\.79000\.79530\.47680\.21850\.29550\.4751UltraRAG \(w/ Qwen3\-8B\)0\.62870\.36670\.43330\.86900\.51780\.50470\.33010\.69000\.54090\.14570\.07280\.11360\.3522LightRAG \(w/ Gemini\-3\-Pro\)0\.71370\.71000\.61810\.95500\.73280\.75000\.55710\.84670\.78950\.70330\.41960\.20670\.6185LightRAG \(w/ Qwen3\-Max\)0\.64650\.71810\.61320\.89220\.74970\.69670\.62380\.79330\.79240\.48340\.23180\.29550\.4950LightRAG \(w/ Qwen3\-8B\)0\.58510\.38930\.44810\.86470\.51450\.48670\.40120\.64670\.53800\.11260\.03970\.12770\.3445

### 4\.4\.MiGUE\-Prediction

The purpose of event analysis is not merely the retrospective understanding of what has occurred, but also the prediction of what will follow\. Thus, we devise the MiGUE\-Prediction task as follows\.

Task Definition\.Given the the task descriptionD​e​sp​r​e​dDes\_\{pred\}, a temporally ordered document sequence𝒟=\{D1,…,Dm−1\|3⩽m⩽6\}\\mathcal\{D\}=\\\{D\_\{1\},\.\.\.,D\_\{m\-1\}\|3\\leqslant m\\leqslant 6\\\}with their timestamps, and a candidate set of events𝒞=\{c1,…,cr∣4⩽r⩽5\}\\mathcal\{C\}=\\\{c\_\{1\},\\dots,c\_\{r\}\\mid 4\\leqslant r\\leqslant 5\\\}, the model is required to select the event that occurs in the correct terminal documentDmD\_\{m\}from𝒞\\mathcal\{C\}\.

Data Construction\.Inspired by existing works\(Huanget al\.,[2024b](https://arxiv.org/html/2607.27654#bib.bib21); Liet al\.,[2021](https://arxiv.org/html/2607.27654#bib.bib22); Maet al\.,[2023a](https://arxiv.org/html/2607.27654#bib.bib23)\), we extract the document sequences from the clusters \(Section[3\.4](https://arxiv.org/html/2607.27654#S3.SS4)\) that exhibit cross\-document causal chains \(Section[3\.3](https://arxiv.org/html/2607.27654#S3.SS3)\)\. We further increase difficulty by appending redundant context—chronologically consistent but irrelevant documents from the same cluster—to discourage shallow pattern matching\. For option design, we devise the correct option as a paraphrased version of the event occurring inDmD\_\{m\}with specific details removed to prevent trivial lexical shortcuts\. Following existing works\(Guanet al\.,[2024](https://arxiv.org/html/2607.27654#bib.bib24)\), we generate disturbing options using four adversarial strategies, includingcounterfactual\(i\.e\., events that are logically opposite to the true outcome\),overgeneralization\(i\.e\., plausible but exaggerated conclusions that transcend the evidence\),temporal trap\(i\.e\., past events from the document sequence\), andirrelevant distraction\(i\.e\., fabricated but contextually consistent events unsupported by the document sequence\)\.

### 4\.5\.Quality Control

To ensure the quality of MiGUE\-Bench, we implement a multi\-stage validation protocol\. Firstly, we automatically filter all the instances to maintain a balanced distribution across document sources, event densities, and difficulties\. Then, we thoroughly detect and exclude offensive, biased, or harmful content, ensuring alignment with ethical standards\. Finally, we conduct manual review on all the instances to verify the correctness of answers and eliminate ambiguity in all the options\. The statistics of the final dataset are shown in Table[3](https://arxiv.org/html/2607.27654#S4.T3)\.

## 5\.Experiment

### 5\.1\.Setting

Source Data\.As MiGUE\-Pipeline is a scalable event data construction pipeline, our corpus is mainly drawn from two sources: \(1\) high\-quality open\-source news document datasets\(Maet al\.,[2023b](https://arxiv.org/html/2607.27654#bib.bib56)\), and \(2\) more than 9,000 news reports collected from Chinese official and mainstream media outlets between 2022 and 2024\. During data collection, we performed a preliminary filtering based on news tags to ensure that the corpus covers major events from this period as comprehensively as possible\. All the corpora are subsequently processed by our MiGUE\-Pipeline to automatically construct event data222Since the source data collected is initially Chinese, we use MiGUE\-Pipeline to construct the Chinese dataset\. Nevertheless, MiGUE\-Pipeline is language\-agnostic and can be also applied to other languages with corresponding prompts\.\.

Base Model Selection\.To comprehensively evaluate the performance of LLMs, we select two lines of work, i\.e\., LLMs and RAG methods\. For LLMs, we involve mainstream closed\-source \(including GPT\-5\.2\-Pro\(OpenAI,[2026](https://arxiv.org/html/2607.27654#bib.bib32)\), Gemini\-3\-Pro\(Google,[2025](https://arxiv.org/html/2607.27654#bib.bib35)\), Claude\-4\.5\-Opus\(Anthropic,[2025b](https://arxiv.org/html/2607.27654#bib.bib33)\), Claude\-4\.5\-Haiku\(Anthropic,[2025a](https://arxiv.org/html/2607.27654#bib.bib34)\), and Qwen3\-Max\(Qwen Team,[2025](https://arxiv.org/html/2607.27654#bib.bib37)\)\) and open\-source LLMs \(including DeepSeek\-V3\.2\(DeepSeek\-AI,[2025](https://arxiv.org/html/2607.27654#bib.bib36)\), GLM\-4\.7\(Team GLM,[2024](https://arxiv.org/html/2607.27654#bib.bib42)\), Kimi\-K2\(Kimi Team,[2025](https://arxiv.org/html/2607.27654#bib.bib41)\), Qwen\-3\-8B/30B/235B\(Qwen Team,[2025](https://arxiv.org/html/2607.27654#bib.bib37)\), and Llama\-3\.1\-7B\(Llama Team,[2024](https://arxiv.org/html/2607.27654#bib.bib43)\)\) as representatives, covering the models with different scales, families, and capacities\. Furthermore, we also adopt two mainstream RAG frameworks including LightRAG\(Guoet al\.,[2025](https://arxiv.org/html/2607.27654#bib.bib38)\)and UltraRAG\(Chenet al\.,[2025](https://arxiv.org/html/2607.27654#bib.bib39)\)\. The retrieval corpora contain 5,000 documents from the document filtering stage \(Section[3\.1](https://arxiv.org/html/2607.27654#S3.SS1)\) of MiGUE\-Pipeline333In the task of MiGUE\-Prediction, we restrict the documents to those occurring prior to the predicted event, preventing data leakage\.\. In our experiment, we utilize Gemini\-3\-Pro, Qwen3\-Max, and Qwen3\-8B as base models for RAG methods\.

Evaluation Metric\.For MiGUE\-Detection, we adopt Micro\-F1 to jointly reflect the precision and recall of this multi\-choice selection task\. As for the other tasks, accuracy is adopted for measurement\.

### 5\.2\.Main Result

The results in[Table 4](https://arxiv.org/html/2607.27654#S4.T4)show that closed\-source LLMs mostly outperform open\-source ones, especially on cross\-document event analysis tasks, demonstrating the effectiveness of state\-of\-the\-art proprietary models\. Among them, GPT\-5\.2\-Pro and Claude\-4\.5\-Opus achieve the best performance across nearly all the tasks, while Deepseek\-V3\.2 and GLM\-4\.7 rank highest among open\-source LLMs\. We also have other interesting findings:

RAG can generally improve performance across the four tasks, but the gains are unstable\.Specifically, RAG benefits strong LLMs such as Gemini\-3\-Pro on most of the tasks\. However, for weaker LLMs like Qwen\-3\-8B, the noisy retrieved documents may degrade the original performance in turn\. How to design specific RAG strategies for event analysis still needs further study\.

At the MiGUE\-Reasoning task, temporal and subevent relations generally benefit from cross\-document settings, whereas coreference and causal reasoning exhibit the opposite trend\.We conjecture that multi\-document contexts may introduce additional temporal signals \(e\.g\., explicit timestamps\), which can be easily captured by LLMs\. Consequently, temporal\-related relations benefit from these additional signals\. In contrast, for coreference and causal reasoning, different documents may describe the same event with distinct perspectives\. Also, temporally successive events are not necessarily causally related\. Such heterogeneity may introduces noise and ambiguity, leading to degraded performance in cross\-document settings\.

At the MiGUE\-Induction task, LLMs are sensitive to start times of events but exhibit limited understanding of their time spans\.We observe from the subtask of interval\-based temporal ordering that most LLMs perform significantly worse on end\-time ordering than on start\-time one, with the accuracy often dropping by nearly half\. This suggests that identifying relative event start times is much simpler for current LLMs, whereas reasoning about event durations remains challenging\.

Causal tasks remain challenging for current LLMs\.Most of these LLMs achieve unsatisfactory performance on the tasks of causal reasoning and causal graph induction\. We conjecture that one reason is the misalignment between model\-internal causal representations and human causal understanding\. Moreover, current causal definitions in event analysis are often confined to language\-based descriptions and lack rigorous mathematical formalization\. We leave the exploration of more fine\-grained probabilistic causal modeling at the schema level instead of merely textual logical formulations in event analysis as important future work\.

![Refer to caption](https://arxiv.org/html/2607.27654v1/x2.png)Figure 2\.Accuracy on MiGUE\-Induction \(Temporal Order\) with different numbers of retrieved documents\.![Refer to caption](https://arxiv.org/html/2607.27654v1/x3.png)Figure 3\.Accuracy on MiGUE\-Reasoning \(Subevent\) with different lengths of input instructions\.
### 5\.3\.Detailed Analysis

Analysis on the Number of Retrieved Documents\.To further analyze this effect of additional relevant documents on temporal reasoning, we evaluate Qwen3\-Max and Kimi\-K2 under UltraRAG with different numbers of retrieved documents\. Figure[2](https://arxiv.org/html/2607.27654#S5.F2)shows that both models exhibit a rise–fall–stabilization pattern for start\-ordering and end\-ordering accuracies as the number of retrieved documents increases\. Moderate retrieved documents improve performance, whereas excessive documents degrade accuracy, possibly due to noise accumulation and context dilution\. End\-ordering performance consistently remains lower than start\-ordering, revealing the deficiencies of LLMs in modeling event time spans\.

Analysis on the Length of Input Instructions\.We further analyze the scaling behavior of LLMs with different lengths of input instructions on subevent reasoning\. As shown in[Figure 3](https://arxiv.org/html/2607.27654#S5.F3), Qwen3\-235B exhibits steady performance gains as the instruction length increases, indicating strong capabilities of long\-context understanding\. For comparison, Qwen3\-30B/8B also benefits from longer inputs but reaches its peak in the 1500–2100 / 900–1500 range, followed by a slight decline\. This suggest that larger models benefit from richer contextual evidence, whereas smaller models saturate earlier and degrade under extended inputs\.

## 6\.Conclusion

In this work, we present a comprehensive benchmark called MiGUE\-Bench for multi\-granularity event analysis, which covers the tasks ranging from fine\-grained event detection and reasoning to cross\-document event structure induction and forecasting, together with an LLM\-driven automatic pipeline named MiGUE\-Pipeline for scalable event data construction\. Through extensive experiments on state\-of\-the\-art LLMs and RAG methods, we show that despite their strong general capabilities, these models still exhibit notable limitations in multi\-granularity event analysis\. We hope that MiGUE\-Bench and MiGUE\-Pipeline can serve as standardized testbeds for diagnosing the capability boundaries of LLMs in event analysis and inspire future advances in this research field\.

###### Acknowledgements\.

This work was supported by Noncommunicable Chronic Diseases\-National Science and Technology Major Project \(No\. 2023ZD0501806\), Sichuan Science and Technology Program \(No\. 2025ZNSFSC1488\) , Fundamental Research Funds for the Central Universities \(No\. ZYGX2025XJ041\), and CIPS\-SMP\-Zhipu Large Model Fund \(No\. CIPS\-SMP20250314\)\.

## References

- Anthropic \(2025a\)Introducing claude haiku 4\.5\.Note:[https://www\.anthropic\.com/news/claude\-haiku\-4\-5](https://www.anthropic.com/news/claude-haiku-4-5)Accessed: 2026\-04\-25Cited by:[§5\.1](https://arxiv.org/html/2607.27654#S5.SS1.p2.1)\.
- Anthropic \(2025b\)Introducing claude opus 4\.5\.Note:[https://www\.anthropic\.com/news/claude\-opus\-4\-5](https://www.anthropic.com/news/claude-opus-4-5)Accessed: 2026\-04\-25Cited by:[§5\.1](https://arxiv.org/html/2607.27654#S5.SS1.p2.1)\.
- M\. Bugert, N\. Reimers, and I\. Gurevych \(2021\)Generalizing cross\-document event coreference resolution across multiple corpora\.Computational Linguistics47\(3\),pp\. 575–614\.Cited by:[1st item](https://arxiv.org/html/2607.27654#S1.I1.i1.p1.1),[§2](https://arxiv.org/html/2607.27654#S2.p3.1)\.
- C\. Chan, C\. Jiayang, W\. Wang, Y\. Jiang, T\. Fang, X\. Liu, and Y\. Song \(2024\)Exploring the potential of ChatGPT on sentence level relations: a focus on temporal, causal, and discourse relations\.InFindings of the Association for Computational Linguistics: EACL 2024,pp\. 684–721\.Cited by:[§2](https://arxiv.org/html/2607.27654#S2.p1.1)\.
- M\. Chen, Y\. Ma, K\. Song, Y\. Cao, Y\. Zhang, and D\. Li \(2024\)Improving large language models in event relation logical prediction\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 9451–9478\.Cited by:[3rd item](https://arxiv.org/html/2607.27654#S1.I1.i3.p1.1),[§1](https://arxiv.org/html/2607.27654#S1.p2.1),[§2](https://arxiv.org/html/2607.27654#S2.p1.1),[§3\.3](https://arxiv.org/html/2607.27654#S3.SS3.p2.1)\.
- M\. Chen, H\. Zhang, Q\. Ning, M\. Li, H\. Ji, K\. McKeown, and D\. Roth \(2021\)Event\-centric natural language processing\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: Tutorial Abstracts,pp\. 6–14\.Cited by:[§1](https://arxiv.org/html/2607.27654#S1.p1.1)\.
- Y\. Chen, L\. Xu, K\. Liu, D\. Zeng, and J\. Zhao \(2015\)Event extraction via dynamic multi\-pooling convolutional neural networks\.InProceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),pp\. 167–176\.Cited by:[§1](https://arxiv.org/html/2607.27654#S1.p2.1)\.
- Y\. Chen, D\. Guo, S\. Mei, X\. Li, H\. Chen, Y\. Li, Y\. Wang, C\. Tang, R\. Wang, D\. Wu, Y\. Yan, Z\. Liu, S\. Yu, Z\. Liu, and M\. Sun \(2025\)UltraRAG: A modular and automated toolkit for adaptive retrieval\-augmented generation\.CoRRabs/2504\.08761\.Cited by:[§5\.1](https://arxiv.org/html/2607.27654#S5.SS1.p2.1)\.
- M\. Choudhary and X\. Du \(2024\)QAEVENT: event extraction as question\-answer pairs generation\.InFindings of the Association for Computational Linguistics: EACL 2024,pp\. 1860–1873\.Cited by:[§2](https://arxiv.org/html/2607.27654#S2.p1.1)\.
- DeepSeek\-AI \(2025\)DeepSeek\-v3\.2: pushing the frontier of open large language models\.CoRRabs/2512\.02556\.Cited by:[§3\.2](https://arxiv.org/html/2607.27654#S3.SS2.p1.1),[§3\.3](https://arxiv.org/html/2607.27654#S3.SS3.p2.1),[§5\.1](https://arxiv.org/html/2607.27654#S5.SS1.p2.1)\.
- G\. R\. Doddington, A\. Mitchell, M\. A\. Przybocki, L\. A\. Ramshaw, S\. M\. Strassel, and R\. M\. Weischedel \(2004\)The automatic content extraction \(ACE\) program \- tasks, data, and evaluation\.InProceedings of the Fourth International Conference on Language Resources and Evaluation,pp\. 837–840\.Cited by:[Table 1](https://arxiv.org/html/2607.27654#S1.T1.4.2.1)\.
- C\. Fan, D\. Liu, L\. Qin, Y\. Zhang, and R\. Xu \(2022\)Towards event\-level causal relation identification\.InSIGIR ’22: The 45th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 1828–1833\.Cited by:[§1](https://arxiv.org/html/2607.27654#S1.p1.1)\.
- J\. Foley, M\. Bendersky, and V\. Josifovski \(2015\)Learning to extract local events from the web\.InProceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 423–432\.Cited by:[§1](https://arxiv.org/html/2607.27654#S1.p1.1)\.
- J\. Gong, B\. Zheng, and Q\. Hu \(2025\)EventRelBench: a comprehensive benchmark for evaluating event relation understanding in large language models\.InFindings of the Association for Computational Linguistics: EMNLP 2025,pp\. 9084–9099\.Cited by:[2nd item](https://arxiv.org/html/2607.27654#S1.I1.i2.p1.1),[Table 1](https://arxiv.org/html/2607.27654#S1.T1.4.7.1),[§2](https://arxiv.org/html/2607.27654#S2.p3.1)\.
- Google \(2025\)Gemini 3: a new era of intelligence\.Note:[https://blog\.google/products\-and\-platforms/products/gemini/gemini\-3/](https://blog.google/products-and-platforms/products/gemini/gemini-3/)Accessed: 2026\-04\-25Cited by:[§5\.1](https://arxiv.org/html/2607.27654#S5.SS1.p2.1)\.
- Y\. Guan, H\. Peng, X\. Wang, L\. Hou, and J\. Li \(2024\)OpenEP: open\-ended future event prediction\.CoRRabs/2408\.06578\.Cited by:[§4\.4](https://arxiv.org/html/2607.27654#S4.SS4.p3.1)\.
- Z\. Guo, L\. Xia, Y\. Yu, T\. Ao, and C\. Huang \(2025\)LightRAG: simple and fast retrieval\-augmented generation\.InFindings of the Association for Computational Linguistics: EMNLP 2025,pp\. 10746–10761\.Cited by:[§5\.1](https://arxiv.org/html/2607.27654#S5.SS1.p2.1)\.
- Y\. Hong, T\. Zhang, T\. O’Gorman, S\. Horowit\-Hendler, H\. Ji, and M\. Palmer \(2016\)Building a cross\-document event\-event relation corpus\.InProceedings of the 10th Linguistic Annotation Workshop held in conjunction with ACL 2016,pp\. 1–6\.Cited by:[1st item](https://arxiv.org/html/2607.27654#S1.I1.i1.p1.1),[§2](https://arxiv.org/html/2607.27654#S2.p3.1)\.
- Z\. Hu, Z\. Li, X\. Jin, L\. Bai, J\. Guo, and X\. Cheng \(2025\)Large language model\-based event relation extraction with rationales\.InProceedings of the 31st International Conference on Computational Linguistics,pp\. 7484–7496\.Cited by:[3rd item](https://arxiv.org/html/2607.27654#S1.I1.i3.p1.1),[§1](https://arxiv.org/html/2607.27654#S1.p2.1),[§2](https://arxiv.org/html/2607.27654#S2.p1.1)\.
- K\. Huang, I\. Hsu, T\. Parekh, Z\. Xie, Z\. Zhang, P\. Natarajan, K\. Chang, N\. Peng, and H\. Ji \(2024a\)TextEE: benchmark, reevaluation, reflections, and future challenges in event extraction\.InFindings of the Association for Computational Linguistics: ACL 2024,pp\. 12804–12825\.Cited by:[§2](https://arxiv.org/html/2607.27654#S2.p1.1)\.
- Z\. Huang, J\. Hwang, J\. Zhang, J\. Baik, W\. Zhang, D\. Wodarz, Y\. Sun, Q\. Gu, and W\. Wang \(2024b\)Causal graph ode: continuous treatment effect modeling in multi\-agent dynamical systems\.InProceedings of the ACM Web Conference 2024,pp\. 4607–4617\.Cited by:[§4\.4](https://arxiv.org/html/2607.27654#S4.SS4.p3.1)\.
- Kimi Team \(2025\)Kimi K2: open agentic intelligence\.CoRRabs/2507\.20534\.Cited by:[§5\.1](https://arxiv.org/html/2607.27654#S5.SS1.p2.1)\.
- B\. Li, X\. Han, J\. Liu, Y\. Ding, L\. Jing, Z\. Zhang, J\. Li, X\. Du, F\. Li, M\. Zhang, M\. Zhang, A\. Sun, P\. S\. Yu, and H\. Fei \(2025\)Event extraction in large language model\.CoRRabs/2512\.19537\.Cited by:[§1](https://arxiv.org/html/2607.27654#S1.p2.1),[§2](https://arxiv.org/html/2607.27654#S2.p1.1)\.
- M\. Li, S\. Li, Z\. Wang, L\. Huang, K\. Cho, H\. Ji, J\. Han, and C\. Voss \(2021\)The future is not one\-dimensional: complex event schema induction by graph modeling for event prediction\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,pp\. 5203–5215\.Cited by:[§1](https://arxiv.org/html/2607.27654#S1.p1.1),[§1](https://arxiv.org/html/2607.27654#S1.p2.1),[§4\.4](https://arxiv.org/html/2607.27654#S4.SS4.p3.1)\.
- R\. Li, W\. Zhao, C\. Yang, and S\. Su \(2022\)A dual\-expert framework for event argument extraction\.InThe 45th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 1110–1121\.Cited by:[§2](https://arxiv.org/html/2607.27654#S2.p1.1)\.
- J\. Liao, X\. Zhao, X\. Li, L\. Zhang, and J\. Tang \(2021\)Learning discriminative neural representations for event detection\.InThe 44th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 644–653\.Cited by:[§1](https://arxiv.org/html/2607.27654#S1.p1.1)\.
- Y\. Liu, D\. Iter, Y\. Xu, S\. Wang, R\. Xu, and C\. Zhu \(2023\)G\-eval: NLG evaluation using gpt\-4 with better human alignment\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing,pp\. 2511–2522\.Cited by:[§3\.1](https://arxiv.org/html/2607.27654#S3.SS1.p1.1)\.
- Llama Team \(2024\)The llama 3 herd of models\.CoRRabs/2407\.21783\.Cited by:[§5\.1](https://arxiv.org/html/2607.27654#S5.SS1.p2.1)\.
- C\. Lou, J\. Gao, C\. Yu, W\. Wang, H\. Zhao, W\. Tu, and R\. Xu \(2022\)Translation\-based implicit annotation projection for zero\-shot cross\-lingual event argument extraction\.InThe 45th International ACM SIGIR Conference on Research and Development in Information Retrieval,pp\. 2076–2081\.Cited by:[§1](https://arxiv.org/html/2607.27654#S1.p1.1)\.
- Y\. Ma, C\. Ye, Z\. Wu, X\. Wang, Y\. Cao, and T\. Chua \(2023a\)Context\-aware event forecasting via graph disentanglement\.InProceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining,pp\. 1643–1652\.Cited by:[§4\.4](https://arxiv.org/html/2607.27654#S4.SS4.p3.1)\.
- Y\. Ma, C\. Ye, Z\. Wu, X\. Wang, Y\. Cao, L\. Pang, and T\. Chua \(2023b\)Structured, complex and time\-complete temporal event forecasting\.CoRRabs/2312\.01052\.Cited by:[§5\.1](https://arxiv.org/html/2607.27654#S5.SS1.p1.1)\.
- A\. Minard, M\. Speranza, E\. Agirre, I\. Aldabe, M\. Van Erp, B\. Magnini, G\. Rigau, and R\. Urizar \(2015\)Semeval\-2015 task 4: timeline: cross\-document event ordering\.Inproceedings of the 9th International Workshop on Semantic Evaluation \(SemEval 2015\),pp\. 778–786\.Cited by:[§1](https://arxiv.org/html/2607.27654#S1.p1.1)\.
- P\. Mirza and S\. Tonelli \(2014\)An analysis of causality between events and its relation to temporal information\.InProceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers,pp\. 2097–2106\.Cited by:[Table 1](https://arxiv.org/html/2607.27654#S1.T1.4.3.1)\.
- N\. Nakshatri, S\. Liu, S\. Chen, D\. Roth, D\. Goldwasser, and D\. Hopkins \(2023\)Using llm for improving key event discovery: temporal\-guided news stream clustering with event summaries\.InFindings of the Association for Computational Linguistics: EMNLP 2023,pp\. 4162–4173\.Cited by:[3rd item](https://arxiv.org/html/2607.27654#S1.I1.i3.p1.1),[§1](https://arxiv.org/html/2607.27654#S1.p2.1),[§2](https://arxiv.org/html/2607.27654#S2.p1.1)\.
- T\. H\. Nguyen, K\. Cho, and R\. Grishman \(2016\)Joint event extraction via recurrent neural networks\.InProceedings of the 2016 conference of the North American chapter of the association for computational linguistics: human language technologies,pp\. 300–309\.Cited by:[§1](https://arxiv.org/html/2607.27654#S1.p2.1)\.
- T\. H\. Nguyen and R\. Grishman \(2018\)Graph convolutional networks with argument\-aware pooling for event detection\.InProceedings of the Thirty\-Second AAAI Conference on Artificial Intelligence,pp\. 5900–5907\.Cited by:[§1](https://arxiv.org/html/2607.27654#S1.p2.1)\.
- OpenAI \(2026\)OpenAI GPT\-5 system card\.CoRRabs/2601\.03267\.Cited by:[§3\.2](https://arxiv.org/html/2607.27654#S3.SS2.p1.1),[§5\.1](https://arxiv.org/html/2607.27654#S5.SS1.p2.1)\.
- Qwen Team \(2025\)Qwen3 technical report\.CoRRabs/2505\.09388\.Cited by:[§3\.3](https://arxiv.org/html/2607.27654#S3.SS3.p2.1),[§5\.1](https://arxiv.org/html/2607.27654#S5.SS1.p2.1)\.
- R\. Sankepally \(2019\)Event information retrieval from text\.InProceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval,SIGIR’19,pp\. 1447\.Cited by:[§1](https://arxiv.org/html/2607.27654#S1.p1.1)\.
- L\. Sha, F\. Qian, B\. Chang, and Z\. Sui \(2018\)Jointly extracting event triggers and arguments by dependency\-bridge RNN and tensor\-based argument interaction\.InProceedings of the Thirty\-Second AAAI Conference on Artificial Intelligence, \(AAAI\-18\), the 30th innovative Applications of Artificial Intelligence \(IAAI\-18\), and the 8th AAAI Symposium on Educational Advances in Artificial Intelligence \(EAAI\-18\), New Orleans, Louisiana, USA, February 2\-7, 2018,pp\. 5916–5923\.Cited by:[§1](https://arxiv.org/html/2607.27654#S1.p2.1)\.
- Z\. Tao, Z\. Jin, Y\. Zhang, X\. Chen, H\. Zhao, J\. Li, B\. Liang, C\. Tao, Q\. Liu, and K\. Wong \(2025\)A comprehensive evaluation on event reasoning of large language models\.InThe Thirty\-Ninth AAAI Conference on Artificial Intelligence,pp\. 25273–25281\.Cited by:[1st item](https://arxiv.org/html/2607.27654#S1.I1.i1.p1.1),[3rd item](https://arxiv.org/html/2607.27654#S1.I1.i3.p1.1),[§1](https://arxiv.org/html/2607.27654#S1.p2.1),[§2](https://arxiv.org/html/2607.27654#S2.p1.1),[§2](https://arxiv.org/html/2607.27654#S2.p3.1)\.
- Team GLM \(2024\)ChatGLM: A family of large language models from GLM\-130B to GLM\-4 all tools\.CoRRabs/2406\.12793\.Cited by:[§5\.1](https://arxiv.org/html/2607.27654#S5.SS1.p2.1)\.
- V\. A\. Traag, L\. Waltman, and N\. J\. Van Eck \(2019\)From louvain to leiden: guaranteeing well\-connected communities\.Scientific reports9\(1\),pp\. 1–12\.Cited by:[§3\.4](https://arxiv.org/html/2607.27654#S3.SS4.p1.1)\.
- X\. Wang, Y\. Chen, N\. Ding, H\. Peng, Z\. Wang, Y\. Lin, X\. Han, L\. Hou, J\. Li, Z\. Liu, P\. Li, and J\. Zhou \(2022\)MAVEN\-ERE: A unified large\-scale dataset for event coreference, temporal, causal, and subevent relation extraction\.InProceedings of the 2022 Conference on Empirical Methods in Natural Language Processing,pp\. 926–941\.Cited by:[1st item](https://arxiv.org/html/2607.27654#S1.I1.i1.p1.1),[Table 1](https://arxiv.org/html/2607.27654#S1.T1.4.5.1),[§2](https://arxiv.org/html/2607.27654#S2.p3.1)\.
- X\. Wang, Z\. Wang, X\. Han, W\. Jiang, R\. Han, Z\. Liu, J\. Li, P\. Li, Y\. Lin, and J\. Zhou \(2020\)MAVEN: A Massive General Domain Event Detection Dataset\.InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 1652–1671\.Cited by:[Table 1](https://arxiv.org/html/2607.27654#S1.T1.4.4.1)\.
- Z\. Yang, Y\. Wang, Z\. Shi, Y\. Yao, L\. Liang, K\. Ding, E\. Yilmaz, H\. Chen, and Q\. Zhang \(2025\)EventRAG: enhancing LLM generation with event knowledge graphs\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 16967–16979\.Cited by:[§1](https://arxiv.org/html/2607.27654#S1.p1.1),[§2](https://arxiv.org/html/2607.27654#S2.p1.1)\.
- Z\. Zhang, Y\. Cao, C\. Ye, Y\. Ma, L\. Liao, and T\. Chua \(2024\)Analyzing temporal complex events with large language models? A benchmark towards temporal, long context understanding\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),pp\. 1588–1606\.Cited by:[1st item](https://arxiv.org/html/2607.27654#S1.I1.i1.p1.1),[3rd item](https://arxiv.org/html/2607.27654#S1.I1.i3.p1.1),[Table 1](https://arxiv.org/html/2607.27654#S1.T1.4.6.1),[§2](https://arxiv.org/html/2607.27654#S2.p3.1),[§4\.3\.1](https://arxiv.org/html/2607.27654#S4.SS3.SSS1.p1.1)\.

Similar Articles

Benchmarking LLMs

Reddit r/AI_Agents

A study or report on benchmarking large language models, likely comparing performance across various tasks.

MLUBench: A Benchmark for Lifelong Unlearning Evaluation in MLLMs

arXiv cs.AI

MLUBench is a large-scale benchmark for lifelong unlearning in multimodal large language models (MLLMs), featuring 127 entities across 9 classes. The paper identifies that existing unlearning methods suffer from cumulative degradation and proposes LUMoE to mitigate this, showing significant improvements.

ModelEquivBench: Certifying Multi-Relational Evaluation of LLM-Generated Optimization Models

arXiv cs.AI

ModelEquivBench is a certifying multi-relational evaluation system for LLM-generated optimization models, reporting per-pair semantic profiles across seven equivalence relations instead of a single accuracy score. It evaluates GPT-5.4, Claude Sonnet 4.6, and Qwen3.5-397B-A17B on a fixed benchmark, revealing stage-wise failures that coarse baselines miss.