information-extraction

Tag

Cards List
#information-extraction

DIRECT: Direct Decoding for Efficient and Aligned Sequence Labeling with Large Language Models

arXiv cs.CL ↗ · 2026-07-30 Cached

DIRECT is a framework for sequence labeling using large language models that improves domain alignment through Direct Preference Optimization (DPO) after supervised fine-tuning and increases inference efficiency via controlled decoding with template-filling and KV cache reuse.

0 favorites 0 likes
#information-extraction

Enhancing Generative Information Extraction with Two-step Validation: A Product Attribute Use Case

arXiv cs.CL ↗ · 2026-07-30 Cached

This paper proposes a two-step validation method for generative information extraction, integrating a PLM block into the pipeline to enhance LLM performance, particularly for weakly expressed entities in product attribute extraction for digital product passports.

0 favorites 0 likes
#information-extraction

ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures

Hugging Face Daily Papers ↗ · 2026-07-29 Cached

This paper presents the ICDAR 2026 Competition on Information Extraction from ALD/E Scientific Figures, introducing the Sci-ImageMiner benchmark with four complementary tasks. Results show SOTA multimodal models perform well on classification and summarization but struggle with data extraction and scientific reasoning, especially visual question answering.

0 favorites 0 likes
#information-extraction

LA-RL: Label-Aware Self-Reflection for Reinforcement Learning in Information Extraction

arXiv cs.CL ↗ · 2026-07-28 Cached

LA-RL introduces a label-aware self-reflection framework for reinforcement learning in information extraction, achieving consistent improvements on named entity recognition, relation extraction, and event extraction tasks with gains of up to 20 F1 on out-of-distribution benchmarks.

0 favorites 0 likes
#information-extraction

Scope3Trace: Evidence-Based Identification and Extraction of Scope 3 GHG Emissions from Sustainability Reports

arXiv cs.CL ↗ · 2026-07-21 Cached

Scope3Trace proposes an evidence-grounded information extraction framework that leverages large language models to identify and extract Scope 3 greenhouse gas emissions from sustainability reports, contributing a dual-level multimodal dataset and achieving high extraction accuracy.

0 favorites 0 likes
#information-extraction

Schema-Constrained Document-Level Event Argument Extraction with Lightweight LLM Fine-Tuning

arXiv cs.CL ↗ · 2026-07-21 Cached

This paper proposes a schema-constrained document-level event argument extraction method using fine-tuned mid-sized open LLMs with LoRA, role-set injection, and deterministic decoding, achieving state-of-the-art results on MAVEN-ARG.

0 favorites 0 likes
#information-extraction

Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents

arXiv cs.AI ↗ · 2026-07-20 Cached

This paper investigates whether agentic mechanisms such as reflection and memory lead to controllable improvements over fixed LLM workflows for information extraction from scholarly PDFs, using conference-paper dataset extraction as a testbed.

0 favorites 0 likes
#information-extraction

Relation Extraction Model Based on Semantic Enhancement Mechanism

arXiv cs.CL ↗ · 2026-07-13 Cached

The paper proposes CasAug, a relation extraction model based on the CasRel framework with a semantic enhancement mechanism to address the triple overlap problem, showing improved performance over baseline models.

0 favorites 0 likes
#information-extraction

Do It Right! A Methodology for Successful NLP System Development

arXiv cs.CL ↗ · 2026-07-08 Cached

This paper presents a stepwise methodology for developing NLP systems in the clinical domain, applying the Systems Development Life Cycle approach, and discusses the challenges of using large language models for information extraction from electronic medical records.

0 favorites 0 likes
#information-extraction

@TheAhmadOsman: Everyone is talking about small and specialized models finally Tweet below is from 17 months ago

X AI KOLs Timeline ↗ · 2026-07-07 Cached

A tweet highlights Jina AI's ReaderLM-v2, a small 4GB model that achieves high accuracy in extracting information from messy DOM elements, exemplifying the trend toward specialized small language models.

0 favorites 0 likes
#information-extraction

SchemaRAG: Dynamic Large Schema Reduction for LLM-driven Structured Information Extraction

arXiv cs.AI ↗ · 2026-07-02 Cached

SchemaRAG is a retrieval-augmented generation framework that dynamically reduces the output schema space for LLM-driven structured information extraction, achieving improved performance and efficiency on healthcare and e-commerce datasets.

0 favorites 0 likes
#information-extraction

LC-ICL: Label-Guided Contrastive In-Context Learning for Robust Information Extraction

arXiv cs.CL ↗ · 2026-06-30 Cached

This paper proposes LC-ICL, a novel few-shot technique that uses both correct and incorrect examples with error-cause labels to improve large language models' performance on information extraction tasks like named entity recognition and relation extraction.

0 favorites 0 likes
#information-extraction

Extracting Knowledge from an Arabic-English Machine-Readable Dictionary Using Information Extraction

arXiv cs.CL ↗ · 2026-06-30 Cached

This paper presents a method for automatically extracting lexical knowledge from the Arabic-English Al-Mawrid dictionary using n-gram analysis, keyword-in-context analysis, and rule-based information extraction.

0 favorites 0 likes
#information-extraction

ReaORE: Reasoning-Guided Progressive Open Relation Extraction Empowered by Large Reasoning Models

arXiv cs.CL ↗ · 2026-06-26 Cached

Proposes ReaORE, a reasoning-guided framework for open relation extraction that progressively filters and predicts relations via coarse-to-fine reasoning, outperforming existing baselines on two datasets.

0 favorites 0 likes
#information-extraction

Extracting Problem and Method Sentence from Scientific Papers: A Context-enhanced Transformer Using Formulaic Expression Desensitization

arXiv cs.CL ↗ · 2026-06-26 Cached

This paper proposes a context-enhanced transformer using formulaic expression desensitization for extracting problem and method sentences from scientific papers, achieving improvements of 3.71% and 2.67% in macro F1 score on two datasets.

0 favorites 0 likes
#information-extraction

BCL: Bayesian In-Context Learning Framework for Information Extraction

arXiv cs.CL ↗ · 2026-06-18 Cached

BCL is the first optimization framework that uses particle filtering with Bayesian updates to systematically refine label representations for information extraction tasks, showing consistent improvements over existing methods.

0 favorites 0 likes
#information-extraction

Configurable Clinical Information Extraction with Agentic RAG: What Works, What Breaks, and Why

Hugging Face Daily Papers ↗ · 2026-06-17 Cached

ACIE, an agentic RAG system for clinical information extraction, achieves 96.5% acceptance rate in nuclear-medicine physicians' judgments across 7,326 instances, addressing challenges of heterogeneous patient contexts and missing metadata.

0 favorites 0 likes
#information-extraction

AAbAAC: An Annotated Corpus for Autoimmunity Information Extraction

arXiv cs.AI ↗ · 2026-06-12 Cached

AAbAAC is a manually annotated corpus of 115 PubMed abstracts for autoimmunity information extraction, focusing on entities like autoimmune diseases and autoantibodies. The study demonstrates improved NER performance after fine-tuning on this corpus.

0 favorites 0 likes
#information-extraction

sebis at CRF Filling 2026: A Two-Stage Local LLM Pipeline for Medical CRF Filling

arXiv cs.CL ↗ · 2026-06-12 Cached

This paper presents a fully local, two-stage LLM pipeline using MedGemma-27B for filling Case Report Forms from clinical notes, achieving a macro-F1 of 0.55 on the English test track and securing second place among local open-source submissions.

0 favorites 0 likes
#information-extraction

Benchmarking Large Language Models for Safety Data Extraction

arXiv cs.CL ↗ · 2026-06-11 Cached

This paper benchmarks four large language models (Gemini 1.5 Pro, GPT-4o, Claude 3.7 Sonnet, Llama 3.1-70B) for extracting structured information from Safety Data Sheets, finding that text-based extraction with chain-of-thought prompting yields the highest accuracy (84% by Gemini 1.5 Pro) but no model surpasses the 90% threshold required for reliable industrial deployment.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback