Tag
研究人员提出了一种结合深度学习与图像预处理技术的OCR模型,用于自动提取韩文发票中的关键信息,在自建数据集上达到了87%的F1分数,且处理耗时极低。
This paper evaluates the use of large language models as AI respondents to generate structured survey responses from policy documents, demonstrating high agreement in structured indicators and potential for hybrid human-AI workflows in policy monitoring.
EAGER is a reinforcement learning framework that improves generative event extraction through fine-grained verifiable rewards and schema-contrastive advantage estimation, outperforming prompting, fine-tuning, and prior RL methods on seven benchmark datasets.
This paper presents WaterBERT, a domain-adapted BERT model for water treatment literature mining, enhancing semantic representation and enabling large-scale structured information extraction and knowledge graph construction.
Schematize is an open-source multi-agent system that interactively generates and refines information-extraction schemas for legal research, achieving top performance in human evaluations.
GLiNER2.5 is a multilingual boundary checkpoint for unified schema-based information extraction, supporting entities, classification, structured records, relations, and span attributes in one model.
This paper introduces Self-Meta-Evolve, a hierarchical framework that personalizes prompts for each user in enterprise information extraction tasks, improving performance through continuous refinement based on interaction feedback.
The paper introduces a weakly supervised framework using large language models to extract dataset mentions in forced displacement and FCV documents, achieving high accuracy with limited labeled data.
CMNIE is a new information extraction benchmark for Chinese military news that jointly annotates events, entities, and relations to evaluate models on schema adherence and exact span matching.
PiPMRE is a novel pipeline framework for medical relation extraction that uses a relation generator and filter to enhance performance, surpassing previous state-of-the-art methods on public datasets.
This paper benchmarks open-source OCR, LLM, and VLM systems for structured information extraction in a high-risk public sector application, finding that VLMs generally outperform OCR+LLM pipelines but most configurations struggle in zero-shot settings, emphasizing the critical role of input quality.
This paper introduces the ALD/E-ImageMiner benchmark for multimodal comprehension of scientific images and discusses future research directions for general-purpose scientific AI, based on insights from the ICDAR 2026 competition.
Jerry Liu introduces ExtractBench and highlights the 'agentic plus' extractor in LlamaParse for handling massive volumes of fields in long documents, with benchmark results available on ExtractBench.
The article explores the vast, untapped potential in human-AI conversations, suggesting that systems could be built to extract valuable insights and novel ideas from billions of interactions, potentially revolutionizing how we understand human reasoning.
This paper introduces MUSE, a full-text, multi-domain scientific knowledge base of Problem-Solution-Rationale triplets mined from research papers, along with an extraction pipeline and initial experiments on rationale-supervised LLM fine-tuning.
This paper proposes TdSciNER, a type-driven multi-task learning approach that leverages LLMs to improve scientific named entity recognition by filtering entity types, adding an auxiliary typing task, and using a demonstration selection strategy. Experiments on three datasets show performance comparable to fully supervised models.
Presents Doc2DB-Bench, a benchmark for evaluating LLM-based extraction of relational databases from long documents, with 203 instances across 42 schemas and seven domains.
This paper introduces ConstructCIE, a manually annotated dataset for extracting causal information from OSHA construction accident narratives, and evaluates supervised sequence taggers and instruction-tuned LLMs on end-to-end hierarchical extraction.
This paper presents a two-stage LLM pipeline for extracting and triangulating causal evidence from humanitarian crisis reports, achieving strong F1 scores on a ReliefWeb dataset and proposing a Level-of-Evidence score for cross-context convergence.
This paper introduces MiGUE-Bench, a systematic benchmark for evaluating LLMs on multi-granularity event analysis, spanning single- to cross-document tasks including event detection, relation reasoning, structure induction, and future prediction.