Tag
DIRECT is a framework for sequence labeling using large language models that improves domain alignment through Direct Preference Optimization (DPO) after supervised fine-tuning and increases inference efficiency via controlled decoding with template-filling and KV cache reuse.
This paper proposes a two-step validation method for generative information extraction, integrating a PLM block into the pipeline to enhance LLM performance, particularly for weakly expressed entities in product attribute extraction for digital product passports.
This paper presents the ICDAR 2026 Competition on Information Extraction from ALD/E Scientific Figures, introducing the Sci-ImageMiner benchmark with four complementary tasks. Results show SOTA multimodal models perform well on classification and summarization but struggle with data extraction and scientific reasoning, especially visual question answering.
LA-RL introduces a label-aware self-reflection framework for reinforcement learning in information extraction, achieving consistent improvements on named entity recognition, relation extraction, and event extraction tasks with gains of up to 20 F1 on out-of-distribution benchmarks.
Scope3Trace proposes an evidence-grounded information extraction framework that leverages large language models to identify and extract Scope 3 greenhouse gas emissions from sustainability reports, contributing a dual-level multimodal dataset and achieving high extraction accuracy.
This paper proposes a schema-constrained document-level event argument extraction method using fine-tuned mid-sized open LLMs with LoRA, role-set injection, and deterministic decoding, achieving state-of-the-art results on MAVEN-ARG.
This paper investigates whether agentic mechanisms such as reflection and memory lead to controllable improvements over fixed LLM workflows for information extraction from scholarly PDFs, using conference-paper dataset extraction as a testbed.
The paper proposes CasAug, a relation extraction model based on the CasRel framework with a semantic enhancement mechanism to address the triple overlap problem, showing improved performance over baseline models.
This paper presents a stepwise methodology for developing NLP systems in the clinical domain, applying the Systems Development Life Cycle approach, and discusses the challenges of using large language models for information extraction from electronic medical records.
A tweet highlights Jina AI's ReaderLM-v2, a small 4GB model that achieves high accuracy in extracting information from messy DOM elements, exemplifying the trend toward specialized small language models.
SchemaRAG is a retrieval-augmented generation framework that dynamically reduces the output schema space for LLM-driven structured information extraction, achieving improved performance and efficiency on healthcare and e-commerce datasets.
This paper proposes LC-ICL, a novel few-shot technique that uses both correct and incorrect examples with error-cause labels to improve large language models' performance on information extraction tasks like named entity recognition and relation extraction.
This paper presents a method for automatically extracting lexical knowledge from the Arabic-English Al-Mawrid dictionary using n-gram analysis, keyword-in-context analysis, and rule-based information extraction.
Proposes ReaORE, a reasoning-guided framework for open relation extraction that progressively filters and predicts relations via coarse-to-fine reasoning, outperforming existing baselines on two datasets.
This paper proposes a context-enhanced transformer using formulaic expression desensitization for extracting problem and method sentences from scientific papers, achieving improvements of 3.71% and 2.67% in macro F1 score on two datasets.
BCL is the first optimization framework that uses particle filtering with Bayesian updates to systematically refine label representations for information extraction tasks, showing consistent improvements over existing methods.
ACIE, an agentic RAG system for clinical information extraction, achieves 96.5% acceptance rate in nuclear-medicine physicians' judgments across 7,326 instances, addressing challenges of heterogeneous patient contexts and missing metadata.
AAbAAC is a manually annotated corpus of 115 PubMed abstracts for autoimmunity information extraction, focusing on entities like autoimmune diseases and autoantibodies. The study demonstrates improved NER performance after fine-tuning on this corpus.
This paper presents a fully local, two-stage LLM pipeline using MedGemma-27B for filling Case Report Forms from clinical notes, achieving a macro-F1 of 0.55 on the English test track and securing second place among local open-source submissions.
This paper benchmarks four large language models (Gemini 1.5 Pro, GPT-4o, Claude 3.7 Sonnet, Llama 3.1-70B) for extracting structured information from Safety Data Sheets, finding that text-based extraction with chain-of-thought prompting yields the highest accuracy (84% by Gemini 1.5 Pro) but no model surpasses the 90% threshold required for reliable industrial deployment.