information-extraction

Tag

Cards List
#information-extraction

Developing an OCR model for Extracting Information from Invoices with Korean Language

arXiv cs.CL ↗ · 8h ago Cached

研究人员提出了一种结合深度学习与图像预处理技术的OCR模型,用于自动提取韩文发票中的关键信息,在自建数据集上达到了87%的F1分数,且处理耗时极低。

0 favorites 0 likes
#information-extraction

From Policy Documents to Structured Survey Responses: Evaluating Large Language Models for Policy Monitoring

arXiv cs.CL ↗ · 5d ago Cached

This paper evaluates the use of large language models as AI respondents to generate structured survey responses from policy documents, demonstrating high agreement in structured indicators and potential for hybrid human-AI workflows in policy monitoring.

0 favorites 0 likes
#information-extraction

EAGER: Enhancing Generative Event Extraction via Reinforcement Learning with Verifiable Rewards

arXiv cs.CL ↗ · 5d ago Cached

EAGER is a reinforcement learning framework that improves generative event extraction through fine-grained verifiable rewards and schema-contrastive advantage estimation, outperforming prompting, fine-tuning, and prior RL methods on seven benchmark datasets.

0 favorites 0 likes
#information-extraction

Domain-Adaptive Pretraining Enhances Water Treatment Semantic Representation for Large-Scale Structured Literature Mining

arXiv cs.CL ↗ · 2026-09-23 Cached

This paper presents WaterBERT, a domain-adapted BERT model for water treatment literature mining, enhancing semantic representation and enabling large-scale structured information extraction and knowledge graph construction.

0 favorites 0 likes
#information-extraction

Schematize: An Agentic System for Generating and Refining Information-Extraction Schemas for Legal Research

arXiv cs.CL ↗ · 2026-09-22 Cached

Schematize is an open-source multi-agent system that interactively generates and refines information-extraction schemas for legal research, achieving top performance in human evaluations.

0 favorites 0 likes
#information-extraction

@george_onx: Seeing a lot of comparisons of Jev w/ GLiNER models lately. If you're benchmarking, GLiNER2.5 is the latest version we'…

X AI KOLs Timeline ↗ · 2026-09-21 Cached

GLiNER2.5 is a multilingual boundary checkpoint for unified schema-based information extraction, supporting entities, classification, structured records, relations, and span attributes in one model.

0 favorites 0 likes
#information-extraction

One Prompt Does Not Fit All: Self-Meta-Evolve for Personalized Information Extraction

arXiv cs.AI ↗ · 2026-09-21 Cached

This paper introduces Self-Meta-Evolve, a hierarchical framework that personalizes prompts for each user in enterprise information extraction tasks, improving performance through continuous refinement based on interaction feedback.

0 favorites 0 likes
#information-extraction

Extracting Dataset Mentions in Forced Displacement and FCV Documents: A Weakly Supervised Framework with LLM-Based Label Refinement

arXiv cs.CL ↗ · 2026-09-14 Cached

The paper introduces a weakly supervised framework using large language models to extract dataset mentions in forced displacement and FCV documents, achieving high accuracy with limited labeled data.

0 favorites 0 likes
#information-extraction

CMNIE: An Information Extraction Benchmark for Chinese Military News

arXiv cs.CL ↗ · 2026-09-11 Cached

CMNIE is a new information extraction benchmark for Chinese military news that jointly annotates events, entities, and relations to evaluate models on schema adherence and exact span matching.

0 favorites 0 likes
#information-extraction

PiPMRE: A Pipeline Based on Language Model for Medical Relation Extraction

arXiv cs.CL ↗ · 2026-09-04 Cached

PiPMRE is a novel pipeline framework for medical relation extraction that uses a relation generator and filter to enhance performance, surpassing previous state-of-the-art methods on public datasets.

0 favorites 0 likes
#information-extraction

Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector Application

arXiv cs.AI ↗ · 2026-08-20 Cached

This paper benchmarks open-source OCR, LLM, and VLM systems for structured information extraction in a high-risk public sector application, finding that VLMs generally outperform OCR+LLM pipelines but most configurations struggle in zero-shot settings, emphasizing the critical role of input quality.

0 favorites 0 likes
#information-extraction

A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

arXiv cs.AI ↗ · 2026-08-17 Cached

This paper introduces the ALD/E-ImageMiner benchmark for multimodal comprehension of scientific images and discusses future research directions for general-purpose scientific AI, based on insights from the ICDAR 2026 competition.

0 favorites 0 likes
#information-extraction

@jerryjliu0: Our "agentic plus" extractor in LlamaParse is great for extracting out massive volumes of fields (e.g. 10k-100k+ fields…

X AI KOLs Following ↗ · 2026-08-15 Cached

Jerry Liu introduces ExtractBench and highlights the 'agentic plus' extractor in LlamaParse for handling massive volumes of fields in long documents, with benchmark results available on ExtractBench.

0 favorites 0 likes
#information-extraction

The profound and untapped potential within human-AI conversations. I think this will eventually change everything.

Reddit r/ArtificialInteligence ↗ · 2026-08-14

The article explores the vast, untapped potential in human-AI conversations, suggesting that systems could be built to extract valuable insights and novel ideas from billions of interactions, potentially revolutionizing how we understand human reasoning.

0 favorites 0 likes
#information-extraction

MUSE: A Full-Text Cross-Domain Knowledge Base of Scientific Problems, Solutions, and Rationales

arXiv cs.CL ↗ · 2026-08-12 Cached

This paper introduces MUSE, a full-text, multi-domain scientific knowledge base of Problem-Solution-Rationale triplets mined from research papers, along with an extraction pipeline and initial experiments on rationale-supervised LLM fine-tuning.

0 favorites 0 likes
#information-extraction

Enhancing Scientific Named Entity Recognition via Large Language Models: A Type-driven Multi-task Learning Approach

arXiv cs.CL ↗ · 2026-08-11 Cached

This paper proposes TdSciNER, a type-driven multi-task learning approach that leverages LLMs to improve scientific named entity recognition by filtering entity types, adding an auxiliary typing task, and using a demonstration selection strategy. Experiments on three datasets show performance comparable to fully supervised models.

0 favorites 0 likes
#information-extraction

Beyond Tables: Doc2DB-Bench for Relationally Faithful Document-to-Database Construction

arXiv cs.CL ↗ · 2026-08-11 Cached

Presents Doc2DB-Bench, a benchmark for evaluating LLM-based extraction of relational databases from long documents, with 203 instances across 42 schemas and seven domains.

0 favorites 0 likes
#information-extraction

ConstructCIE: A Dataset for Extracting Causal Information from Construction Accident Narratives

arXiv cs.CL ↗ · 2026-08-10 Cached

This paper introduces ConstructCIE, a manually annotated dataset for extracting causal information from OSHA construction accident narratives, and evaluates supervised sequence taggers and instruction-tuned LLMs on end-to-end hierarchical extraction.

0 favorites 0 likes
#information-extraction

Causal Evidence Extraction and Triangulation in Crisis Reports using Large Language Models: A ReliefWeb-based Study

arXiv cs.CL ↗ · 2026-08-06 Cached

This paper presents a two-stage LLM pipeline for extracting and triangulating causal evidence from humanitarian crisis reports, achieving strong F1 scores on a ReliefWeb dataset and proposing a Level-of-Evidence score for cross-context convergence.

0 favorites 0 likes
#information-extraction

From Single- to Cross-Document: Benchmarking Multi-Granularity Event Analysis of Large Language Models

arXiv cs.CL ↗ · 2026-07-31 Cached

This paper introduces MiGUE-Bench, a systematic benchmark for evaluating LLMs on multi-granularity event analysis, spanning single- to cross-document tasks including event detection, relation reasoning, structure induction, and future prediction.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback