Enhancing Generative Information Extraction with Two-step Validation: A Product Attribute Use Case

arXiv cs.CL Papers

Summary

This paper proposes a two-step validation method for generative information extraction, integrating a PLM block into the pipeline to enhance LLM performance, particularly for weakly expressed entities in product attribute extraction for digital product passports.

arXiv:2607.26780v1 Announce Type: new Abstract: The ability of large language models (LLMs) to process and generate text has introduced potential for applications in information extraction (IE). While it's debated whether LLMs outperform smaller fine-tuned models for classification tasks, their strong generalization capability makes them promising for domains with limited labeled data available for fine-tuning. This advantage is particularly relevant for the emerging application of the digital product passport (DPP), where the problem space is broad but domain-specific data remains scarce. Motivated by this use case, we apply generative IE to the product domain, explicitly addressing efficiency, generalizability, and data privacy constraints. We propose a two-step validation method that integrates a PLM block into the generative IE pipeline and thereby leverages LLMs' correction capability. We discover that such a validation task enhances LLM performance, particularly on the extraction of weakly expressed, low-salience entities that appear sparsely throughout the text. For certain entities, the performance of mid-size models can even reach levels comparable to larger models, and the improvement of first-step PLM predictions also enhance the final LLM output. Nevertheless, the effects on the smallest open-source LLMs (e.g., Llama-3.2 3B) is limited. Based on the findings, we develop a demo application for product information extraction that utilizes locally deployed LLMs, targeting further adaptations to real-world DPP use cases.
Original Article
View Cached Full Text

Cached at: 07/30/26, 09:59 AM

# Enhancing Generative Information Extraction with Two-step Validation: A Product Attribute Use Case
Source: [https://arxiv.org/html/2607.26780](https://arxiv.org/html/2607.26780)
Yi\-Sheng Hsu Nermeen Abou Baker Uwe Handmann Computer Science Institute, Ruhr West University of Applied Sciences Bottrop, Germany \{firstname\.lastname\}@hs\-ruhrwest\.de

###### Abstract

The ability of large language models \(LLMs\) to process and generate text has introduced potential for applications in information extraction \(IE\)\. While it’s debated whether LLMs outperform smaller fine\-tuned models for classification tasks, their strong generalization capability makes them promising for domains with limited labeled data available for fine\-tuning\. This advantage is particularly relevant for the emerging application of the digital product passport \(DPP\), where the problem space is broad but domain\-specific data remains scarce\. Motivated by this use case, we apply generative IE to the product domain, explicitly addressing efficiency, generalizability, and data privacy constraints\. We propose a two\-step validation method that integrates a PLM block into the generative IE pipeline and thereby leverages LLMs’ correction capability\. We discover that such a validation task enhances LLM performance, particularly on the extraction of weakly expressed, low\-salience entities that appear sparsely throughout the text\. For certain entities, the performance of mid\-size models can even reach levels comparable to larger models, and the improvement of first\-step PLM predictions also enhance the final LLM output\. Nevertheless, the effects on the smallest open\-source LLMs \(e\.g\.,Llama\-3\.2 3B\) is limited\. Based on the findings, we develop a demo application for product information extraction that utilizes locally deployed LLMs, targeting further adaptations to real\-world DPP use cases\.111The repo for experiment is available at[https://github\.com/doyouwantsometea/pie\_paper](https://github.com/doyouwantsometea/pie_paper)\. The demo app is available at[https://github\.com/hrw\-neurolab/transferhub\_pie\_demo](https://github.com/hrw-neurolab/transferhub_pie_demo)\.

Enhancing Generative Information Extraction with Two\-step Validation: A Product Attribute Use Case

Yi\-Sheng Hsu Nermeen Abou Baker Uwe HandmannComputer Science Institute, Ruhr West University of Applied SciencesBottrop, Germany\{firstname\.lastname\}@hs\-ruhrwest\.de

## 1Introduction

As a long\-established NLP task, information extraction \(IE\) remains challenging owing to the various and continually changing sources of information and task requirementsXuet al\.\([2024](https://arxiv.org/html/2607.26780#bib.bib7)\)\. Earlier works perform IE through fine\-tuning pretrained language models \(PLMs\) such asBERTJianget al\.\([2020](https://arxiv.org/html/2607.26780#bib.bib30)\); Shinet al\.\([2020](https://arxiv.org/html/2607.26780#bib.bib29)\), while more recent studies attempt to leverage the generation capability of LLMs to directly output structured dataZhanget al\.\([2025](https://arxiv.org/html/2607.26780#bib.bib9)\)\. Both mainstream approaches have their advantages and drawbacks: Fine\-tuned PLMs can often surpass LLM performance owing to the task’s classification\-oriented natureWanget al\.\([2025](https://arxiv.org/html/2607.26780#bib.bib6)\)\. By contrast, LLMs depend less on domain\-specific data and generalize better than PLMs when such training data are scarce\. In real\-world use cases, the trade\-off often imposes challenges to deploying robust IE applications that are both effective and well\-integrated with domain\-specific knowledgeMaet al\.\([2023](https://arxiv.org/html/2607.26780#bib.bib14)\); Blumeet al\.\([2023](https://arxiv.org/html/2607.26780#bib.bib34)\); Kimet al\.\([2024](https://arxiv.org/html/2607.26780#bib.bib36)\)\.

![Refer to caption](https://arxiv.org/html/2607.26780v1/graphs/workflow.png)Figure 1:The workflow of the two\-step validation task, where LLM is instructed to correct a PLM output instead of extracting information directly from the text\.This challenge is especially relevant to the product domainSchönet al\.\([2018](https://arxiv.org/html/2607.26780#bib.bib33)\); Brinkmannet al\.\([2024](https://arxiv.org/html/2607.26780#bib.bib1)\); Hättyet al\.\([2024](https://arxiv.org/html/2607.26780#bib.bib35)\): Structured, domain\-specific annotations are costly and limited, while product information is highly heterogeneous and frequently subject to confidentiality constraints\. These challenges draw our attention to this domain, particularly under the context of the European Commission’s Digital Product Passport \(DPP\)Petriket al\.\([2025](https://arxiv.org/html/2607.26780#bib.bib32)\), which requires the structured disclosure of product attributes including materials, components, and manufacturing details\. The implementation of the DPP is planned from 2027, starting from batteries and gradually extending to a broader coverage\. Although the establishment of the DPP can largely benefit from IE applications in creating and processing the required structured data, non\-existing data and the broad scope introduce bottlenecks to addressing such a use case\. Furthermore, while handling product specifications or confidential information, companies often avoid commercial LLMs \(e\.g\. ChatGPT or Gemini\) owing to data privacy concerns, hindering both the generation and use of structured data required by these new regulations\.

Motivated by the DPP use case, we focus on product information extraction using open\-source LLMs in this study, eventually building a demo application that can be further extended to broaden use cases\. Aiming at enhancing generalizability across underexplored product domains while reducing model size to allow local deployment, we reformulate the conventional generative IE approach as a validation task and instruct LLMs to correct first\-step predictions from a fine\-tuned PLM \(Figure[1](https://arxiv.org/html/2607.26780#S1.F1)\)\. The twist strategically exploits LLMs’ text\-refinement strengths to boost output quality without increasing model size, thereby achieving a more efficient performance\-to\-size ratio\. Our study contributes the following:

- •We introduce a hybrid, two\-step generative IE method, bringing forward the validation task by integrating a PLM block into the generative IE pipeline\.
- •Our method enhances generative IE performance particularly on low\-salience, weakly expressed entity classes\. This enables smaller LLMs to deliver performance comparable to larger models and thereby benefits local deployment and data privacy\.
- •Based on the proposed validation task reformulation, we prototype a demo application \(§[5](https://arxiv.org/html/2607.26780#S5)\) that can be further integrated to real\-world applications, particularly under the context of the DPP\.

EntityAmazonE\-commerceSize570349Weight90771Product number710133Component993718Material637400Manufacturer722447Data Instances426512Table 1:Dataset size, entity classes and their counts in each dataset\. While the first three entities are more explicit, the later ones tend to be weakly expressed in the text and are therefore considered more challenging\.
## 2Related Work

#### Conventional and generative IE\.

Information extraction \(IE\) transforms plain text into structured information and is commonly categorized into named entity recognition \(NER\), relation extraction \(RE\), and event extraction \(EE\)Luet al\.\([2022](https://arxiv.org/html/2607.26780#bib.bib12)\); Xuet al\.\([2024](https://arxiv.org/html/2607.26780#bib.bib7)\)\. Conventional approaches for NER includeBiLSTM\-CRFLampleet al\.\([2016](https://arxiv.org/html/2607.26780#bib.bib19)\)and later on PLMs such asBERTDevlinet al\.\([2019](https://arxiv.org/html/2607.26780#bib.bib21)\)andRoBERTaLiuet al\.\([2019](https://arxiv.org/html/2607.26780#bib.bib20)\); following the recent development of LLMs, their widespread use has shifted the task from its original classification\-based nature to a generative oneHsu and Roberts \([2024](https://arxiv.org/html/2607.26780#bib.bib4)\); Xuet al\.\([2024](https://arxiv.org/html/2607.26780#bib.bib7)\); Zhanget al\.\([2025](https://arxiv.org/html/2607.26780#bib.bib9)\)\.

Whether LLMs’ generative capability suffices for IE tasks nevertheless remains a debatable topic\. LLMs were often found outperformed by fine\-tuned PLMs on IE tasksGaoet al\.\([2024](https://arxiv.org/html/2607.26780#bib.bib17)\); Penget al\.\([2024](https://arxiv.org/html/2607.26780#bib.bib18)\)\. Several studies highlighted the notable gap remaining between PLMs and LLMsHanet al\.\([2023](https://arxiv.org/html/2607.26780#bib.bib3)\); Liet al\.\([2023](https://arxiv.org/html/2607.26780#bib.bib11)\); Chenet al\.\([2024a](https://arxiv.org/html/2607.26780#bib.bib15)\); Liaoet al\.\([2025](https://arxiv.org/html/2607.26780#bib.bib16)\); LLMs were not considered sufficient on few\-shot IE, with their performance possibly worsening on easy samples owing to hallucination or span boundary mismatchMaet al\.\([2023](https://arxiv.org/html/2607.26780#bib.bib14)\)\.Hanet al\.\([2023](https://arxiv.org/html/2607.26780#bib.bib3)\)highlighted existing challenges LLMs, such as extending span lengths and being oversensitive to irrelevant context\. According toWanget al\.\([2025](https://arxiv.org/html/2607.26780#bib.bib6)\), LLMs without fine\-tuning tended to fall behind supervised BERT\-based models on NER task and may further suffer from known challenges such as hallucination\.

#### Enhancing performance\.

Previous studies explored several approaches in pursuit of increasing the performance of generative IE\. Generally, LLMs have been known to be capable of revising self\-generated outputs for better resultsMadaanet al\.\([2023](https://arxiv.org/html/2607.26780#bib.bib28)\); Kamoiet al\.\([2024](https://arxiv.org/html/2607.26780#bib.bib39)\)\. When it comes to IE tasks, fine\-tuning with fewer than 1k data instances could boost performanceDunnet al\.\([2022](https://arxiv.org/html/2607.26780#bib.bib2)\), and LLMs could also improve in NER when provided with self\-annotated entitiesXieet al\.\([2024](https://arxiv.org/html/2607.26780#bib.bib5)\)\. Alternatively, modifying output structureSainzet al\.\([2024](https://arxiv.org/html/2607.26780#bib.bib25)\)such as applying code\-styled formulationLiet al\.\([2025a](https://arxiv.org/html/2607.26780#bib.bib24)\)has been found beneficial\.

LLMs were also often applied as a component in an IE workflow to achieve better overall performance\.Weiet al\.\([2023](https://arxiv.org/html/2607.26780#bib.bib8)\)proposed a two\-step QA process to further break down zero\-shot IE task\. Similarly,Liet al\.\([2024](https://arxiv.org/html/2607.26780#bib.bib13)\)split the NER task into identifying named entities and formatting structured output\. LLMs may also be added upon smaller models to perform post\-hoc verificationKimet al\.\([2024](https://arxiv.org/html/2607.26780#bib.bib36)\), classificationZhanget al\.\([2024](https://arxiv.org/html/2607.26780#bib.bib37)\), or selectionChenet al\.\([2024b](https://arxiv.org/html/2607.26780#bib.bib38)\)\. AsZhanget al\.\([2025](https://arxiv.org/html/2607.26780#bib.bib9)\)highlighted the increased cost and computational demand with LLMs,Kimet al\.\([2025](https://arxiv.org/html/2607.26780#bib.bib10)\)introduced a plug\-in architecture that utilized both PLM and LLM in a pipeline to take into account efficiency and generalizability at once\.

#### IE in the product domain\.

Although IE tasks are widely applicable to a range of real\-world scenarios, when it comes to product information, the scarcity of training data often hinders the usage of NER applications with conventional methodsSchönet al\.\([2018](https://arxiv.org/html/2607.26780#bib.bib33)\)\. Nevertheless, the development of LLMs introduced a breakthrough to such a bottleneck\.Hättyet al\.\([2024](https://arxiv.org/html/2607.26780#bib.bib35)\)highlighted LLMs’ outstanding capability of extracting explicit labels; moreover, LLMs could extract implicit labels, which was hardly achievable with the conventional token classification approach\.Brinkmannet al\.\([2024](https://arxiv.org/html/2607.26780#bib.bib1)\)investigated product IE using GPT models and hinted at the potential of normalizing numeric attribute values\. Similarly,Liet al\.\([2025b](https://arxiv.org/html/2607.26780#bib.bib26)\)probed into attribute mining in product description, introducing a framework adopting the chain\-of\-thought reasoning methodWeiet al\.\([2022](https://arxiv.org/html/2607.26780#bib.bib27)\)\.

## 3Generative IE with Validation Task

### 3\.1Task reformulation

In our experiments, we adapt the conventional generative IE into a validation task, running LLMs on two task formulations:

1. 1\.Extraction, where the model is asked to identify entities from plain text and directly generate structured output\.
2. 2\.Validation, which features two\-step generation: An initial structured prediction is first made by the PLM\. The prediction is then provided to the LLM, which validates and corrects the initial output \(Figure[1](https://arxiv.org/html/2607.26780#S1.F1)\)\.

The validation task imitates the self\-refine frameworkMadaanet al\.\([2023](https://arxiv.org/html/2607.26780#bib.bib28)\)without recursive prompting\. The method also echoes the usage of LLMs for post\-hoc processingKimet al\.\([2024](https://arxiv.org/html/2607.26780#bib.bib36)\); Chenet al\.\([2024b](https://arxiv.org/html/2607.26780#bib.bib38)\)in NER workflow to leverage LLMs’ general knowledge to the task\.

In the validation task, we fine\-tune oneRoBERTaand oneDeBERTamodel per dataset, yielding four setups for first\-step PLM predictions\. As a baseline, LLMs are tasked with correcting an empty dictionary formatted as PLM output, adding up to a total of five setups for the validation task\.

### 3\.2Experiments

#### Data\.

Focusing on product domain for further DPP applications, we use two open\-source datasets as the main materials: Hugging FaceAmazon Product Description222[https://huggingface\.co/datasets/philschmid/amazon\-product\-descriptions\-vlm](https://huggingface.co/datasets/philschmid/amazon-product-descriptions-vlm)\(hereinafterAmazon/AZ\) and KaggleE\-commerce Text Classification333[https://www\.kaggle\.com/datasets/saurabhshahane/ecommerce\-text\-classification](https://www.kaggle.com/datasets/saurabhshahane/ecommerce-text-classification)\(hereinafterE\-commerce/EC\)\. Both datasets are resampled to roughly 500 instances and then annotated with six named\-entity labels \(Table[1](https://arxiv.org/html/2607.26780#S1.T1)\) by a single annotator usingLabel StudioTkachenkoet al\.\([2020](https://arxiv.org/html/2607.26780#bib.bib31)\)\. The definition of the labels is provided in Table[5](https://arxiv.org/html/2607.26780#A1.T5)in the Appendix\.

ModelExtractionValidationbaselineRoBERTa\-AZRoBERTa\-ECDeBERTa\-AZDeBERTa\-ECAmazonLlama\-3\.2 3B53\.0753\.2046\.4943\.7646\.8144\.46Llama\-3\.1 8B54\.1758\.91\\cellcolorgreen\!1061\.45\*\\cellcolorgreen\!1058\.06\\cellcolorgreen\!1060\.90\\cellcolorgreen\!1058\.28Llama\-3\.3 70B70\.0170\.71\\cellcolorgreen\!1072\.94\\cellcolorgreen\!1071\.44\\cellcolorgreen\!1072\.84\\cellcolorgreen\!1071\.33Mistral\-0\.3 7B46\.4850\.6839\.13\*36\.8641\.3238\.66Mistral\-small\-3\.1 24B66\.4962\.89\\cellcolorgreen\!1069\.36\\cellcolorgreen\!1067\.20\\cellcolorgreen\!1069\.72\\cellcoloryellow\!1566\.45Gemma\-3 4B53\.2452\.09\\cellcolorgreen\!1055\.1950\.88\\cellcolorgreen\!1056\.69\\cellcoloryellow\!1552\.87Gemma\-3 27B65\.0366\.50\\cellcolorgreen\!1072\.64\\cellcolorgreen\!1067\.63\\cellcolorgreen\!1072\.39\\cellcolorgreen\!1067\.74E\-commerceLlama\-3\.2 3B30\.2727\.1623\.82\\cellcoloryellow\!1529\.3926\.12\\cellcoloryellow\!1530\.04Llama\-3\.1 8B35\.7938\.75\\cellcoloryellow\!1537\.12\\cellcolorgreen\!1042\.24\*\\cellcoloryellow\!1537\.50\\cellcolorgreen\!1042\.80Llama\-3\.3 70B44\.4844\.89\\cellcolorgreen\!1046\.32\\cellcolorgreen\!1046\.49\\cellcolorgreen\!1046\.03\\cellcolorgreen\!1046\.74Mistral\-0\.3 7B30\.4833\.08\\cellcoloryellow\!1532\.74\\cellcolorgreen\!1034\.75\*\\cellcolorgreen\!1033\.52\\cellcolorgreen\!1034\.99Mistral\-small\-3\.1 24B44\.5743\.23\\cellcoloryellow\!1544\.23\\cellcolorgreen\!1049\.64\\cellcolorgreen\!1046\.57\\cellcolorgreen\!1050\.19Gemma\-3 4B37\.5437\.8937\.50\\cellcolorgreen\!1041\.35\\cellcoloryellow\!1537\.62\\cellcolorgreen\!1041\.89Gemma\-3 27B46\.2045\.5245\.50\\cellcolorgreen\!1049\.20\\cellcoloryellow\!1545\.63\\cellcolorgreen\!1048\.67Table 2:F1score \(%\) of generative IE across two task variations and different validation setups\. The highest score per model is highlighted in bold face\. Scores higher than both the extraction task and the validation baseline are marked in green, and those higher than either of them are in yellow\. The starred scores denote the average over five runs to assess robustness\.
#### Models\.

We include seven LLMs from three model families:Llama,Mistral, andGemmawith different sizes, using open\-source models to support local deployment for the DPP use cases\. All the models are instruction\-tuned variants downloaded from Hugging Face\. The task is conducted through one\-shot prompting with few\-shot training\. The prompt is provided in Figure[6](https://arxiv.org/html/2607.26780#A2.F6)in the Appendix\.

#### Evaluation\.

The outcome of generative IE is assessed through the F1score\. Considering that gold labels might contain heterogeneous formats of a label \(e\.g\. “120cm” and “120 cm” should be considered identical\) and that LLMs tend to correct minor grammatical errors, we first eliminate spaces and special characters in both gold labels and predictions, and then remove duplicated entities and cases\. The evaluation is conducted on an entity\-level: A predicted entity is considered correct only while fully mapping a gold label of the same class\. For example, “wooden frame” as acomponentwould be considered false because the gold labels mark “frame” as acomponentand “wood” as amaterial\.

## 4Results and Discussion

#### IE performance using validation method\.

We explore two generative IE task formulations and report the F1score in Table[2](https://arxiv.org/html/2607.26780#S3.T2)\. To assess the robustness of the results, we repeated experiments withLlama\-3\.1 8BandMistral\-0\.3 7B, together with fine\-tunedRoBERTaon both datasets, across five independent runs\. The standard deviations range from5\.7×10−35\.7\\times 10^\{\-3\}to1\.2×10−21\.2\\times 10^\{\-2\}, while the 95% confidence intervals fall between7\.1×10−37\.1\\times 10^\{\-3\}and1\.5×10−21\.5\\times 10^\{\-2\}\. The low variance across different runs suggests that the observed performance remains stable\.

Through the validation task, larger LLMs from the same model family tend to deliver better results, while the three model families do not substantially differ in performance\. All models tend to score higher onAmazonacross all settings; this may result from companies complying with certain format requirements in providing information on the Amazon website\. Such a reason also leads to PLMs’ outstanding capability of extractingsize,weight, andproduct numberon theAmazondataset \(Appendix[A](https://arxiv.org/html/2607.26780#A1)\)\.

While applying the two\-step validation method, we find that mid\-size and large models frequently benefit from the task reformulation, which echoes LLMs’ capability of self\-improvementMadaanet al\.\([2023](https://arxiv.org/html/2607.26780#bib.bib28)\)\. Most models score better F1in at least one scenario, with the improvement being the most consistent withLlama\-3\.3 70B\. Notably, even empty PLM predictions i\.e\. the baseline setup can already improve IE performance in multiple cases\. This suggests that the gains are not solely driven by the good PLM predictions; since LLMs are proven capable of revising empty predictions, error propagation from incorrect PLM predictions to the final LLM output also becomes limited\. However, the smallest LLMs such asLlama\-3\.2 3BandGemma\-3 4Boften suffer from a drastic performance drop, potentially because the longer prompted context increases difficulty of the task\.

Precision and recall scores provide further details on how the validation task affects model performance\. Although the smallest models could already achieve good precision with the intuitive extraction task, they suffer from significantly lower recall, which can be improved through the validation task\. In comparison, the enhancement of precision score contributes more to the overall improvement with mid\-sized and larger LLMs\. Such an observation suggests that the two\-step method could take advantages of LLMs’ generalizability across models and validation settings in reducing ignored, undetected labels\.

Among the four tested PLM blocks, predictions from a PLM fine\-tuned on the same dataset consistently outperform other setups, boosting performance by up to over 7% \(Gemma\-3 27BonAmazon\)\. Although the baseline setting is found already beneficial for the two\-step method, this trend highlights the importance of PLM prediction quality, indicating that LLMs can directly take advantages of the improvements on the first\-step prediction and thereby generate better final output\.

![Refer to caption](https://arxiv.org/html/2607.26780v1/graphs/llama_f1.png)

![Refer to caption](https://arxiv.org/html/2607.26780v1/graphs/mistral_f1.png)

![Refer to caption](https://arxiv.org/html/2607.26780v1/graphs/gemma_f1.png)

Figure 2:F1score \(%\) per label forLlama\(top\),Mistral\(middle\), andGemma\(bottom\) model family\. The black solid lines denote the extraction task, while the other line styles represent different validation setups\.
#### Enhancement on weakly expressed labels\.

Although the overall F1score does not fluctuate massively, performance varies across labels\. Figure[2](https://arxiv.org/html/2607.26780#S4.F2)reveals that task reformulation introduces improvement more notably tocomponentandmanufacturer, whereas the original extraction task outperforms most validation setups onsizeandweight\. This pattern holds across all three model families, with improvements reflected in both precision and recall \(Appendix[B](https://arxiv.org/html/2607.26780#A2)\)\. Furthermore, on labels such ascomponentandmanufacturer, applying PLM block fine\-tuned on the same dataset could enhance the performance of smaller models, enabling them to approximate and even occasionally outperform larger models\.

This discrepancy highlights the advantages of the validation method on labels that are vague or appear sparsely in the text\. Applications of product IE such as recommendation, decision\-makingBrinkmannet al\.\([2024](https://arxiv.org/html/2607.26780#bib.bib1)\), and the DPPPetriket al\.\([2025](https://arxiv.org/html/2607.26780#bib.bib32)\)can often depend on the weakly expressed labels\. Moreover, the enhancement of lightweight LLMs enables local deployment even on limited hardware, making the method practical in use cases that have to comply with data privacy requirements\. Overall, by reformulating generative IE and integrating the validation method, the approach offers strong potential for broader adoption and practical application in real\-world scenarios\.

## 5Demo Application

We build a demo application upon the findings from the experiments to performs product information extraction, ultimately aiming to accelerate or even automate the process of implementing the DPP\. The application adopts our proposed two\-step method to generate structured information out of product descriptions in plain text, which we assume companies maintain as part of their product launch processes in the modern market\. Considering the structured information demanded by the DPP remains unclear under the ongoing legislative process, we currently stick to the labels explored in the experiments for the application\.

Using examples from a web UI screenshot, Figure[3](https://arxiv.org/html/2607.26780#S5.F3)visualizes the architecture of the application\. In light of the divergent enhancement across the labels, we categorize the six labels into the explicit entities \(size,weight,product number\) and the weakly expressed ones \(component,material,manufacturer\), assigning the two groups with different tasks: The explicit entities are directly extracted by the LLM, while the weakly expressed entities undergo the two\-step validation process\. Among the labels, we considercomponentandmaterialthe most relevant to a DPP use case\. For the validation task, we use fine\-tunedRoBERTaas the PLM block despite its slightly poorer performance in comparison toDeBERTa, since the inference time is significantly shorter\. Based on the respectively better performing method, the label split enhances the overall extraction quality\.

Although the F1scores \(Figure[2](https://arxiv.org/html/2607.26780#S4.F2)\) suggest that the smallest LLMs \(e\.g\.,Llama\-3\.2 3BandGemma\-3 4B\) suffer from performance drop under the validation task, the small size makes them to some extent usable while facing strict hardware constraints, particularly considering the increased recall: Extraction task typically yield lower recall for weakly expressed labels, especiallycomponent\. From this point of view, improvements are already noticeable with even the smallest LLMs \(Figure[5](https://arxiv.org/html/2607.26780#A1.F5)\)\. Although the task reformulation has limited impact on precision, it helps reduce false negative predictions, which is already beneficial in our use case\. A possible minor factor is the task split reduces the label amount from 6 to 3, which potentially simplifies the task with one\-shot prompting\. With the smallest LLMs, the application can be efficiently deployed on a local device without the need for a traditional GPU\. For example, in one of our tested environments with an M1 MacBookPro \(produced 2021\), processing a one\-page\-long product description usually takes less than 5 seconds withLlama\-3\.2, highlighting the advantage in lowering processing time\.

The application provides companies, especially small and medium\-sized enterprises \(SMEs\), a solution to utilize existing data \(e\.g\. product description\) to simplify the establishment of DPP while conforming to the upcoming EU regulations\. Leveraging the validation task, the application improves the performance\-to\-size ratio of generative IE workflow and provides more flexibility with local setups\. This also prevents data leakage and guarantees privacy, which is a huge concern for SMEs to process sensitive data with commercial models such as GhatGPT\. Through further adaptation to the updated regulations, the application demonstrates the potential for efficient deployment and integration into other systems\.

![Refer to caption](https://arxiv.org/html/2607.26780v1/graphs/demo_anonymous_2.png)Figure 3:The architecture of the product information extractor demo tool, using the screenshot of processing an Amazon page of a vacuum cleaner as an example\. In the demo, we split the labels into two groups: explicit and non\-explicit, running both extraction and validation tasks in parallel to deliver better performance respectively on each group\.
## 6Conclusion

In this study, we examined generative IE methods in the product domain to support applications regarding the DPP\. We introduced a two\-step validation method to reformulate generative IE task and discovered that open\-source LLMs can enhance IE performance in most cases\. Furthermore, the F1score enhancement was proportional to the quality of PLM predictions given as the source to correct\. The improvement was most evident for weakly expressed entities with higher semantic complexity; more explicitly stated entities were already well handled by standard extraction methods\. Based on the findings, we further implemented a demo application that performs IE on product description, aiming to accelerate the establishment of the DPP\. The enhancement of lightweight LLMs in the IE pipeline enables local deployment, ensuring stronger data privacy for companies and eventually paving the way for broader real\-world adoption\.

## Limitations

The datasets are annotated by only one annotator because of the dataset size and the clear entity span; although the labels are straightforward as product information tend to prevent semantic ambiguity, mislabeling may occur and induce biases into the results\. The annotated spans can sometimes be affected by tokenization\. In particular, we find the mixture of alphabet and numbers \(e\.g\. as “1x1\.2 m” or “ABC1000TUV”\) occasionally incorrectly tokenized, which propagates to the following evaluation\. While evaluating LLM output, strict matching does not fully capture how practical the results are from an application\-oriented perspective\. When the gold label marks “gloves” as a component, “white gloves” would be considered false, whereas the later may seem slightly better considering the DPP use case\. Furthermore, we employ token\-level F1to evaluate fine\-tuned PLMs and yet entity\-level F1for LLMs\. It is therefore challenging to make direct comparisons between scores from the two approaches\.

## Ethical considerations

Early stage of framing the research directions involve using ChatGPT \(GPT\-5\.1\) to review feasibility of new ideas\. Regarding contents, we do not see immediate ethical concerns in terms of research and development\.

## Acknowledgments

This work has been funded by the Ministry of Economy, Innovation, Digitalization, and Energy of the State of North Rhine\-Westphalia, Germany, and the European Union within the projectDer Transferhub Digitalisierung & Circular Economy im Prosperkolleg\. We thank Nils Feldhus for his careful review and valuable feedback of the paper draft\.

## References

- Generative models for product attribute extraction\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: EMNLP 2023 \- Industry Track, Singapore, December 6\-10, 2023,M\. Wang and I\. Zitouni \(Eds\.\),pp\. 575–585\.External Links:[Link](https://doi.org/10.18653/v1/2023.emnlp-industry.55),[Document](https://dx.doi.org/10.18653/V1/2023.EMNLP-INDUSTRY.55)Cited by:[§1](https://arxiv.org/html/2607.26780#S1.p1.1)\.
- A\. Brinkmann, N\. Baumann, and C\. Bizer \(2024\)Using llms for the extraction and normalization of product attribute values\.InAdvances in Databases and Information Systems \- 28th European Conference, ADBIS 2024, Bayonne, France, August 28\-31, 2024, Proceedings,J\. Tekli, J\. Gamper, R\. Chbeir, and Y\. Manolopoulos \(Eds\.\),Lecture Notes in Computer Science, Vol\.14918,pp\. 217–230\.External Links:[Link](https://doi.org/10.1007/978-3-031-70626-4%5C_15),[Document](https://dx.doi.org/10.1007/978-3-031-70626-4%5F15)Cited by:[§1](https://arxiv.org/html/2607.26780#S1.p2.1),[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px3.p1.1),[§4](https://arxiv.org/html/2607.26780#S4.SS0.SSS0.Px2.p2.1)\.
- R\. Chen, C\. Qin, W\. Jiang, and D\. Choi \(2024a\)Is a large language model a good annotator for event extraction?\.InThirty\-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty\-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2014, February 20\-27, 2024, Vancouver, Canada,M\. J\. Wooldridge, J\. G\. Dy, and S\. Natarajan \(Eds\.\),pp\. 17772–17780\.External Links:[Link](https://doi.org/10.1609/aaai.v38i16.29730),[Document](https://dx.doi.org/10.1609/AAAI.V38I16.29730)Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px1.p2.1)\.
- W\. Chen, L\. Zhao, Z\. Zheng, T\. Xu, Y\. Wang, and E\. Chen \(2024b\)Double\-checker: large language model as a checker for few\-shot named entity recognition\.InFindings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12\-16, 2024,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Findings of ACL, Vol\.EMNLP 2024,pp\. 3172–3181\.External Links:[Link](https://doi.org/10.18653/v1/2024.findings-emnlp.180),[Document](https://dx.doi.org/10.18653/V1/2024.FINDINGS-EMNLP.180)Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px2.p2.1),[§3\.1](https://arxiv.org/html/2607.26780#S3.SS1.p3.1)\.
- J\. Devlin, M\. Chang, K\. Lee, and K\. Toutanova \(2019\)BERT: pre\-training of deep bidirectional transformers for language understanding\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL\-HLT 2019, Minneapolis, MN, USA, June 2\-7, 2019, Volume 1 \(Long and Short Papers\),J\. Burstein, C\. Doran, and T\. Solorio \(Eds\.\),pp\. 4171–4186\.External Links:[Link](https://doi.org/10.18653/v1/n19-1423),[Document](https://dx.doi.org/10.18653/V1/N19-1423)Cited by:[Appendix A](https://arxiv.org/html/2607.26780#A1.p1.1),[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Dunn, J\. Dagdelen, N\. Walker, S\. Lee, A\. S\. Rosen, G\. Ceder, K\. A\. Persson, and A\. Jain \(2022\)Structured information extraction from complex scientific text with fine\-tuned large language models\.CoRRabs/2212\.05238\.External Links:[Link](https://doi.org/10.48550/arXiv.2212.05238),[Document](https://dx.doi.org/10.48550/ARXIV.2212.05238),2212\.05238Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px2.p1.1)\.
- J\. Gao, H\. Zhao, W\. Wang, C\. Yu, and R\. Xu \(2024\)EventRL: enhancing event extraction with outcome supervision for large language models\.CoRRabs/2402\.11430\.External Links:[Link](https://doi.org/10.48550/arXiv.2402.11430),[Document](https://dx.doi.org/10.48550/ARXIV.2402.11430),2402\.11430Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px1.p2.1)\.
- R\. Han, T\. Peng, C\. Yang, B\. Wang, L\. Liu, and X\. Wan \(2023\)Is information extraction solved by chatgpt? an analysis of performance, evaluation criteria, robustness and errors\.CoRRabs/2305\.14450\.External Links:[Link](https://doi.org/10.48550/arXiv.2305.14450),[Document](https://dx.doi.org/10.48550/ARXIV.2305.14450),2305\.14450Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px1.p2.1)\.
- A\. Hätty, D\. Milchevski, K\. Döring, M\. Putnikovic, M\. Mesgar, F\. Novovic, M\. Braun, K\. Borimann, and I\. Stranjanac \(2024\)A cost\-efficient modular sieve for extracting product information from company websites\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: EMNLP 2024 \- Industry Track, Miami, Florida, USA, November 12\-16, 2024,F\. Dernoncourt, D\. Preotiuc\-Pietro, and A\. Shimorina \(Eds\.\),pp\. 1444–1456\.External Links:[Link](https://doi.org/10.18653/v1/2024.emnlp-industry.106),[Document](https://dx.doi.org/10.18653/V1/2024.EMNLP-INDUSTRY.106)Cited by:[§1](https://arxiv.org/html/2607.26780#S1.p2.1),[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px3.p1.1)\.
- P\. He, X\. Liu, J\. Gao, and W\. Chen \(2021\)Deberta: decoding\-enhanced bert with disentangled attention\.In9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3\-7, 2021,External Links:[Link](https://openreview.net/forum?id=XPZIaotutsD)Cited by:[Appendix A](https://arxiv.org/html/2607.26780#A1.p1.1)\.
- E\. Hsu and K\. Roberts \(2024\)LLM\-IE: A python package for generative information extraction with large language models\.CoRRabs/2411\.11779\.External Links:[Link](https://doi.org/10.48550/arXiv.2411.11779),[Document](https://dx.doi.org/10.48550/ARXIV.2411.11779),2411\.11779Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px1.p1.1)\.
- Z\. Jiang, A\. El\-Jaroudi, W\. Hartmann, D\. Karakos, and L\. Zhao \(2020\)Cross\-lingual information retrieval with BERT\.InProceedings of the workshop on Cross\-Language Search and Summarization of Text and Speech \(CLSSTS2020\),K\. McKeown, D\. W\. Oard, Elizabeth, and R\. Schwartz \(Eds\.\),Marseille, France,pp\. 26–31\(eng\)\.External Links:[Link](https://aclanthology.org/2020.clssts-1.5/),ISBN 979\-10\-95546\-55\-9Cited by:[§1](https://arxiv.org/html/2607.26780#S1.p1.1)\.
- R\. Kamoi, Y\. Zhang, N\. Zhang, J\. Han, and R\. Zhang \(2024\)When can llms*Actually*correct their own mistakes? A critical survey of self\-correction of llms\.Trans\. Assoc\. Comput\. Linguistics12,pp\. 1417–1440\.External Links:[Link](https://doi.org/10.1162/tacl%5C_a%5C_00713),[Document](https://dx.doi.org/10.1162/TACL%5FA%5F00713)Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px2.p1.1)\.
- K\. Kim, J\. Jang, J\. Choi, Y\. Lee, K\. Jin, and Y\. Kim \(2025\)Plug\-in and fine\-tuning: bridging the gap between small language models and large language models\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), ACL 2025, Vienna, Austria, July 27 \- August 1, 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),pp\. 5434–5452\.External Links:[Link](https://aclanthology.org/2025.acl-long.271/)Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px2.p2.1)\.
- S\. Kim, K\. Seo, H\. Chae, J\. Yeo, and D\. Lee \(2024\)VerifiNER: verification\-augmented NER via knowledge\-grounded reasoning with large language models\.InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), ACL 2024, Bangkok, Thailand, August 11\-16, 2024,L\. Ku, A\. Martins, and V\. Srikumar \(Eds\.\),pp\. 2441–2461\.External Links:[Link](https://doi.org/10.18653/v1/2024.acl-long.134),[Document](https://dx.doi.org/10.18653/V1/2024.ACL-LONG.134)Cited by:[§1](https://arxiv.org/html/2607.26780#S1.p1.1),[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px2.p2.1),[§3\.1](https://arxiv.org/html/2607.26780#S3.SS1.p3.1)\.
- G\. Lample, M\. Ballesteros, S\. Subramanian, K\. Kawakami, and C\. Dyer \(2016\)Neural architectures for named entity recognition\.InNAACL HLT 2016, The 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, San Diego California, USA, June 12\-17, 2016,K\. Knight, A\. Nenkova, and O\. Rambow \(Eds\.\),pp\. 260–270\.External Links:[Link](https://doi.org/10.18653/v1/n16-1030),[Document](https://dx.doi.org/10.18653/V1/N16-1030)Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px1.p1.1)\.
- B\. Li, G\. Fang, Y\. Yang, Q\. Wang, W\. Ye, W\. Zhao, and S\. Zhang \(2023\)Evaluating chatgpt’s information extraction capabilities: an assessment of performance, explainability, calibration, and faithfulness\.CoRRabs/2304\.11633\.External Links:[Link](https://doi.org/10.48550/arXiv.2304.11633),[Document](https://dx.doi.org/10.48550/ARXIV.2304.11633),2304\.11633Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px1.p2.1)\.
- B\. Li, G\. Fang, W\. Ye, Z\. Xu, J\. Zhang, H\. Cheng, and S\. Zhang \(2025a\)MPL: multiple programming languages with large language models for information extraction\.InFindings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 \- August 1, 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),pp\. 2403–2414\.External Links:[Link](https://aclanthology.org/2025.findings-acl.122/)Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px2.p1.1)\.
- J\. Li, Y\. Li, X\. Shen, C\. Zhang, G\. Qi, and S\. Bi \(2025b\)Open\-world attribute mining for e\-commerce products with multimodal self\-correction instruction tuning\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), ACL 2025, Vienna, Austria, July 27 \- August 1, 2025,W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),pp\. 1702–1714\.External Links:[Link](https://aclanthology.org/2025.acl-long.85/)Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px3.p1.1)\.
- Y\. Li, R\. Ramprasad, and C\. Zhang \(2024\)A simple but effective approach to improve structured language model output for information extraction\.InFindings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12\-16, 2024,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),pp\. 5133–5148\.External Links:[Link](https://doi.org/10.18653/v1/2024.findings-emnlp.295),[Document](https://dx.doi.org/10.18653/V1/2024.FINDINGS-EMNLP.295)Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px2.p2.1)\.
- X\. Liao, J\. Duan, Y\. Huang, and J\. Wang \(2025\)RUIE: retrieval\-based unified information extraction using large language model\.InProceedings of the 31st International Conference on Computational Linguistics, COLING 2025, Abu Dhabi, UAE, January 19\-24, 2025,O\. Rambow, L\. Wanner, M\. Apidianaki, H\. Al\-Khalifa, B\. D\. Eugenio, and S\. Schockaert \(Eds\.\),pp\. 9640–9655\.External Links:[Link](https://aclanthology.org/2025.coling-main.645/)Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px1.p2.1)\.
- Y\. Liu, M\. Ott, N\. Goyal, J\. Du, M\. Joshi, D\. Chen, O\. Levy, M\. Lewis, L\. Zettlemoyer, and V\. Stoyanov \(2019\)RoBERTa: A robustly optimized BERT pretraining approach\.CoRRabs/1907\.11692\.External Links:[Link](http://arxiv.org/abs/1907.11692),1907\.11692Cited by:[Appendix A](https://arxiv.org/html/2607.26780#A1.p1.1),[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px1.p1.1)\.
- Y\. Lu, Q\. Liu, D\. Dai, X\. Xiao, H\. Lin, X\. Han, L\. Sun, and H\. Wu \(2022\)Unified structure generation for universal information extraction\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\), ACL 2022, Dublin, Ireland, May 22\-27, 2022,S\. Muresan, P\. Nakov, and A\. Villavicencio \(Eds\.\),pp\. 5755–5772\.External Links:[Link](https://doi.org/10.18653/v1/2022.acl-long.395),[Document](https://dx.doi.org/10.18653/V1/2022.ACL-LONG.395)Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px1.p1.1)\.
- Y\. Ma, Y\. Cao, Y\. Hong, and A\. Sun \(2023\)Large language model is not a good few\-shot information extractor, but a good reranker for hard samples\!\.InFindings of the Association for Computational Linguistics: EMNLP 2023, Singapore, December 6\-10, 2023,H\. Bouamor, J\. Pino, and K\. Bali \(Eds\.\),pp\. 10572–10601\.External Links:[Link](https://doi.org/10.18653/v1/2023.findings-emnlp.710),[Document](https://dx.doi.org/10.18653/V1/2023.FINDINGS-EMNLP.710)Cited by:[§1](https://arxiv.org/html/2607.26780#S1.p1.1),[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px1.p2.1)\.
- A\. Madaan, N\. Tandon, P\. Gupta, S\. Hallinan, L\. Gao, S\. Wiegreffe, U\. Alon, N\. Dziri, S\. Prabhumoye, Y\. Yang, S\. Gupta, B\. P\. Majumder, K\. Hermann, S\. Welleck, A\. Yazdanbakhsh, and P\. Clark \(2023\)Self\-refine: iterative refinement with self\-feedback\.InAdvances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 \- 16, 2023,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),External Links:[Link](http://papers.nips.cc/paper%5C_files/paper/2023/hash/91edff07232fb1b55a505a9e9f6c0ff3-Abstract-Conference.html)Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px2.p1.1),[§3\.1](https://arxiv.org/html/2607.26780#S3.SS1.p3.1),[§4](https://arxiv.org/html/2607.26780#S4.SS0.SSS0.Px1.p3.1)\.
- L\. Peng, Z\. Wang, F\. Yao, Z\. Wang, and J\. Shang \(2024\)MetaIE: distilling a meta model from LLM for all kinds of information extraction tasks\.CoRRabs/2404\.00457\.External Links:[Link](https://doi.org/10.48550/arXiv.2404.00457),[Document](https://dx.doi.org/10.48550/ARXIV.2404.00457),2404\.00457Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px1.p2.1)\.
- D\. Petrik, F\. Dzierzawa, and K\. Warthmann \(2025\)A maturity model for digital product passports: A design science study\.IEEE Access13,pp\. 114575–114594\.External Links:[Link](https://doi.org/10.1109/ACCESS.2025.3584842),[Document](https://dx.doi.org/10.1109/ACCESS.2025.3584842)Cited by:[§1](https://arxiv.org/html/2607.26780#S1.p2.1),[§4](https://arxiv.org/html/2607.26780#S4.SS0.SSS0.Px2.p2.1)\.
- O\. Sainz, I\. García\-Ferrero, R\. Agerri, O\. L\. de Lacalle, G\. Rigau, and E\. Agirre \(2024\)GoLLIE: annotation guidelines improve zero\-shot information\-extraction\.InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7\-11, 2024,External Links:[Link](https://openreview.net/forum?id=Y3wpuxd7u9)Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px2.p1.1)\.
- V\. Sanh, L\. Debut, J\. Chaumond, and T\. Wolf \(2019\)DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter\.CoRRabs/1910\.01108\.External Links:[Link](http://arxiv.org/abs/1910.01108),1910\.01108Cited by:[Appendix A](https://arxiv.org/html/2607.26780#A1.p1.1)\.
- S\. Schön, V\. Mironova, A\. Gabryszak, and L\. Hennig \(2018\)A corpus study and annotation schema for named entity recognition and relation extraction of business products\.InProceedings of the Eleventh International Conference on Language Resources and Evaluation, LREC 2018, Miyazaki, Japan, May 7\-12, 2018,N\. Calzolari, K\. Choukri, C\. Cieri, T\. Declerck, S\. Goggi, K\. Hasida, H\. Isahara, B\. Maegaard, J\. Mariani, H\. Mazo, A\. Moreno, J\. Odijk, S\. Piperidis, and T\. Tokunaga \(Eds\.\),External Links:[Link](http://www.lrec-conf.org/proceedings/lrec2018/summaries/88.html)Cited by:[§1](https://arxiv.org/html/2607.26780#S1.p2.1),[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px3.p1.1)\.
- H\. J\. Shin, J\. Y\. Park, D\. B\. Yuk, and J\. S\. Lee \(2020\)BERT\-based spatial information extraction\.InProceedings of the Third International Workshop on Spatial Language Understanding,P\. Kordjamshidi, A\. Bhatia, M\. Alikhani, J\. Baldridge, M\. Bansal, and M\. Moens \(Eds\.\),Online,pp\. 10–17\.External Links:[Link](https://aclanthology.org/2020.splu-1.2/),[Document](https://dx.doi.org/10.18653/v1/2020.splu-1.2)Cited by:[§1](https://arxiv.org/html/2607.26780#S1.p1.1)\.
- M\. Tkachenko, M\. Malyuk, A\. Holmanyuk, and N\. Liubimov \(2020\)Label Studio: data labeling software\.Note:Open source software available from https://github\.com/HumanSignal/label\-studioExternal Links:[Link](https://github.com/HumanSignal/label-studio)Cited by:[§3\.2](https://arxiv.org/html/2607.26780#S3.SS2.SSS0.Px1.p1.1)\.
- S\. Wang, X\. Sun, X\. Li, R\. Ouyang, F\. Wu, T\. Zhang, J\. Li, G\. Wang, and C\. Guo \(2025\)GPT\-NER: named entity recognition via large language models\.InFindings of the Association for Computational Linguistics: NAACL 2025, Albuquerque, New Mexico, USA, April 29 \- May 4, 2025,L\. Chiruzzo, A\. Ritter, and L\. Wang \(Eds\.\),pp\. 4257–4275\.External Links:[Link](https://doi.org/10.18653/v1/2025.findings-naacl.239),[Document](https://dx.doi.org/10.18653/V1/2025.FINDINGS-NAACL.239)Cited by:[§1](https://arxiv.org/html/2607.26780#S1.p1.1),[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px1.p2.1)\.
- J\. Wei, X\. Wang, D\. Schuurmans, M\. Bosma, B\. Ichter, F\. Xia, E\. H\. Chi, Q\. V\. Le, and D\. Zhou \(2022\)Chain\-of\-thought prompting elicits reasoning in large language models\.InAdvances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 \- December 9, 2022,S\. Koyejo, S\. Mohamed, A\. Agarwal, D\. Belgrave, K\. Cho, and A\. Oh \(Eds\.\),External Links:[Link](http://papers.nips.cc/paper%5C_files/paper/2022/hash/9d5609613524ecf4f15af0f7b31abca4-Abstract-Conference.html)Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px3.p1.1)\.
- X\. Wei, X\. Cui, N\. Cheng, X\. Wang, X\. Zhang, S\. Huang, P\. Xie, J\. Xu, Y\. Chen, M\. Zhang, Y\. Jiang, and W\. Han \(2023\)Zero\-shot information extraction via chatting with chatgpt\.CoRRabs/2302\.10205\.External Links:[Link](https://doi.org/10.48550/arXiv.2302.10205),[Document](https://dx.doi.org/10.48550/ARXIV.2302.10205),2302\.10205Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px2.p2.1)\.
- T\. Xie, Q\. Li, Y\. Zhang, Z\. Liu, and H\. Wang \(2024\)Self\-improving for zero\-shot named entity recognition with large language models\.InProceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies: Short Papers, NAACL 2024, Mexico City, Mexico, June 16\-21, 2024,K\. Duh, H\. Gómez\-Adorno, and S\. Bethard \(Eds\.\),pp\. 583–593\.External Links:[Link](https://doi.org/10.18653/v1/2024.naacl-short.49),[Document](https://dx.doi.org/10.18653/V1/2024.NAACL-SHORT.49)Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px2.p1.1)\.
- D\. Xu, W\. Chen, W\. Peng, C\. Zhang, T\. Xu, X\. Zhao, X\. Wu, Y\. Zheng, Y\. Wang, and E\. Chen \(2024\)Large language models for generative information extraction: a survey\.Frontiers Comput\. Sci\.18\(6\),pp\. 186357\.External Links:[Link](https://doi.org/10.1007/s11704-024-40555-y),[Document](https://dx.doi.org/10.1007/S11704-024-40555-Y)Cited by:[§1](https://arxiv.org/html/2607.26780#S1.p1.1),[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px1.p1.1)\.
- Z\. Zhang, Y\. Zhao, H\. Gao, and M\. Hu \(2024\)LinkNER: linking local named entity recognition models to large language models using uncertainty\.InProceedings of the ACM on Web Conference 2024, WWW 2024, Singapore, May 13\-17, 2024,T\. Chua, C\. Ngo, R\. Kumar, H\. W\. Lauw, and R\. K\. Lee \(Eds\.\),pp\. 4047–4058\.External Links:[Link](https://doi.org/10.1145/3589334.3645414),[Document](https://dx.doi.org/10.1145/3589334.3645414)Cited by:[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px2.p2.1)\.
- Z\. Zhang, W\. You, T\. Wu, X\. Wang, J\. Li, and M\. Zhang \(2025\)A survey of generative information extraction\.InProceedings of the 31st International Conference on Computational Linguistics, COLING 2025, Abu Dhabi, UAE, January 19\-24, 2025,O\. Rambow, L\. Wanner, M\. Apidianaki, H\. Al\-Khalifa, B\. D\. Eugenio, and S\. Schockaert \(Eds\.\),pp\. 4840–4870\.External Links:[Link](https://aclanthology.org/2025.coling-main.324/)Cited by:[§1](https://arxiv.org/html/2607.26780#S1.p1.1),[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2607.26780#S2.SS0.SSS0.Px2.p2.1)\.

## Appendix AIE with PLMs

DataModelOSizeWeightPNComponentMaterialManufacturerOverallCrossevalBIBIBIBIBIBIAZRoBERTa98\.4090\.3291\.5896\.0095\.2693\.7693\.1854\.6052\.1175\.9377\.1772\.2071\.3396\.9993\.72DeBERTa98\.3591\.9492\.1296\.8796\.2093\.8093\.4855\.1146\.2577\.0873\.9471\.6271\.9896\.9294\.19ECRoBERTa97\.3769\.1777\.9573\.5075\.0064\.6175\.0246\.5352\.2973\.5567\.5367\.5766\.6694\.8394\.57DeBERTa97\.2970\.0180\.3688\.8585\.3665\.3066\.3139\.4535\.4275\.7468\.6168\.8869\.9494\.4394\.72Table 3:F1score \(%\) of fine\-tuned PLMs on NER under the BIO schema\. The overall score is weighted by the number of tokens\. The scores denote the mean value of the three runs, and the highest score per label is highlighted in bold face\.DataModelOSizeWeightPNComponentMaterialManufacturerOverallBIBIBIBIBIBIAZRoBERTa96\.9156\.8172\.7338\.3835\.3638\.2057\.1423\.0522\.1262\.8054\.1426\.6741\.3293\.72DeBERTa97\.1360\.4475\.5952\.9248\.8058\.4466\.1325\.8724\.8062\.5150\.9242\.6943\.3494\.19ECRoBERTa97\.3072\.5485\.0087\.1089\.6980\.7687\.9330\.8624\.1057\.7254\.7441\.9448\.3094\.57DeBERTa97\.3878\.1487\.6195\.3796\.0358\.1582\.5236\.0228\.7559\.2955\.1345\.6448\.2694\.72Table 4:F1score \(%\) of cross\-evaluation under the BIO schema\. The overall score is weighted by the number of tokens\. The scores denote the mean value of the three runs, and the highest score per label is highlighted in bold face\.LabelDescriptionSizeSize measurement of the product, e\.g\. 120 cmWeightWeight measurement of the product, e\.g\. 1 kgProduct numberNumeric or alphabetic identifier of the product, e\.g\. AM12345678ComponentProduct composition and removable segments, e\.g\. frame, cable, batteryMaterialMaterial of the product or its components, e\.g\. polyesterManufacturerName of the Manufacturer, e\.g\. IKEATable 5:Labels and the corresponding descriptions provided as the annotation guideline\.We fine\-tune BERT\-based models as the PLM block in the two\-step validation method\. We first perform hyperparameter search withBERTDevlinet al\.\([2019](https://arxiv.org/html/2607.26780#bib.bib21)\),DistilBERTSanhet al\.\([2019](https://arxiv.org/html/2607.26780#bib.bib22)\),RoBERTaLiuet al\.\([2019](https://arxiv.org/html/2607.26780#bib.bib20)\), andDeBERTaHeet al\.\([2021](https://arxiv.org/html/2607.26780#bib.bib23)\), deciding to use the later two because of better performance\.

ModelExtractionValidationbaselineRoBERTa\-AZRoBERTa\-ECDeBERTa\-AZDeBERTa\-ECAmazonLlama\-3\.2 3B67\.6461\.1149\.1548\.0649\.5446\.16Llama\-3\.1 8B66\.9664\.97\\cellcoloryellow\!1566\.87\*62\.32\\cellcoloryellow\!1566\.3263\.13Llama\-3\.3 70B63\.9666\.04\\cellcolorgreen\!1068\.25\\cellcolorgreen\!1067\.24\\cellcolorgreen\!1068\.18\\cellcolorgreen\!1066\.95Mistral\-0\.3 7B61\.8765\.2643\.31\*41\.5445\.3942\.05Mistral\-small\-3\.1 24B64\.3663\.52\\cellcolorgreen\!1070\.95\\cellcolorgreen\!1069\.04\\cellcolorgreen\!1071\.60\\cellcolorgreen\!1068\.79Gemma\-3 4B58\.5754\.58\\cellcoloryellow\!1554\.8951\.31\\cellcoloryellow\!1556\.8552\.33Gemma\-3 27B61\.4465\.88\\cellcolorgreen\!1071\.13\\cellcolorgreen\!1066\.70\\cellcolorgreen\!1070\.95\\cellcolorgreen\!1065\.90E\-commerceLlama\-3\.2 3B34\.9623\.8321\.24\\cellcoloryellow\!1525\.7422\.99\\cellcoloryellow\!1526\.61Llama\-3\.1 8B36\.7136\.7834\.87\\cellcolorgreen\!1040\.19\*35\.55\\cellcolorgreen\!1040\.38Llama\-3\.3 70B37\.0236\.57\\cellcolorgreen\!1038\.62\\cellcolorgreen\!1038\.84\\cellcolorgreen\!1038\.50\\cellcolorgreen\!1039\.13Mistral\-0\.3 7B29\.3031\.48\\cellcoloryellow\!1530\.31\\cellcolorgreen\!1032\.72\*\\cellcolorgreen\!1031\.77\\cellcolorgreen\!1032\.77Mistral\-small\-3\.1 24B43\.7743\.26\\cellcolorgreen\!1044\.91\\cellcolorgreen\!1050\.38\\cellcolorgreen\!1047\.53\\cellcolorgreen\!1050\.84Gemma\-3 4B42\.9734\.24\\cellcoloryellow\!1535\.08\\cellcoloryellow\!1538\.09\\cellcoloryellow\!1535\.15\\cellcoloryellow\!1538\.37Gemma\-3 27B39\.6341\.92\\cellcoloryellow\!1541\.72\\cellcolorgreen\!1045\.75\\cellcolorgreen\!1042\.47\\cellcolorgreen\!1045\.03Table 6:Precision score \(%\) of generative IE across two task variations and different validation setups\. The highest score per model is highlighted in bold face\. Scores higher than both the extraction task and the validation baseline are marked in green, and those higher than either of them are in yellow\. The starred numbers denote the averages over five runs\. Across the four reported mean scores, standard deviations range from4\.1×10−34\.1\\times 10^\{\-3\}to1\.0×10−21\.0\\times 10^\{\-2\}, while 95% confidence intervals fall between5\.1×10−35\.1\\times 10^\{\-3\}and1\.3×10−21\.3\\times 10^\{\-2\}\.ModelExtractionValidationbaselineRoBERTa\-AZRoBERTa\-ECDeBERTa\-AZDeBERTa\-ECAmazonLlama\-3\.2 3B43\.6747\.11\\cellcoloryellow\!1544\.1040\.17\\cellcoloryellow\!1544\.3642\.88Llama\-3\.1 8B45\.4953\.88\\cellcolorgreen\!1056\.86\*\\cellcolorgreen\!1054\.35\\cellcolorgreen\!1056\.30\\cellcolorgreen\!1054\.12Llama\-3\.3 70B77\.3276\.10\\cellcolorgreen\!1078\.31\\cellcoloryellow\!1576\.20\\cellcolorgreen\!1078\.18\\cellcoloryellow\!1576\.33Mistral\-0\.3 7B37\.2241\.4235\.69\*33\.12\\cellcoloryellow\!1537\.9235\.77Mistral\-small\-3\.1 24B68\.7662\.28\\cellcoloryellow\!1567\.83\\cellcoloryellow\!1565\.45\\cellcoloryellow\!1567\.93\\cellcoloryellow\!1564\.26Gemma\-3 4B48\.7949\.82\\cellcolorgreen\!1055\.50\\cellcolorgreen\!1050\.45\\cellcolorgreen\!1056\.53\\cellcolorgreen\!1053\.42Gemma\-3 27B69\.0661\.74\\cellcolorgreen\!1074\.21\\cellcoloryellow\!1568\.60\\cellcolorgreen\!1073\.88\\cellcolorgreen\!1069\.69E\-commerceLlama\-3\.2 3B26\.6931\.57\\cellcoloryellow\!1527\.12\\cellcolorgreen\!1034\.2520\.23\\cellcolorgreen\!1034\.49Llama\-3\.1 8B34\.9240\.95\\cellcoloryellow\!1539\.67\\cellcolorgreen\!1044\.54\*\\cellcoloryellow\!1539\.67\\cellcolorgreen\!1045\.52Llama\-3\.3 70B55\.7058\.14\\cellcoloryellow\!1557\.83\\cellcoloryellow\!1557\.89\\cellcoloryellow\!1557\.22\\cellcoloryellow\!1558\.01Mistral\-0\.3 7B31\.7534\.86\\cellcolorgreen\!1035\.59\\cellcolorgreen\!1037\.04\*\\cellcolorgreen\!1035\.47\\cellcolorgreen\!1037\.54Mistral\-small\-3\.1 24B45\.4043\.21\\cellcoloryellow\!1543\.57\\cellcolorgreen\!1048\.93\\cellcolorgreen\!1045\.64\\cellcolorgreen\!1049\.54Gemma\-3 4B33\.3342\.41\\cellcoloryellow\!1540\.28\\cellcolorgreen\!1045\.22\\cellcoloryellow\!1540\.46\\cellcolorgreen\!1046\.13Gemma\-3 27B55\.3949\.79\\cellcoloryellow\!1550\.03\\cellcoloryellow\!1553\.2049\.30\\cellcoloryellow\!1552\.96Table 7:Reacll score \(%\) of generative IE across two task variations and different validation setups\. The highest score per model is highlighted in bold face\. Scores higher than both the extraction task and the validation baseline are marked in green, and those higher than either of them are in yellow\. The starred numbers denote the averages over five runs\. Across the four reported mean scores, standard deviations range from6\.1×10−36\.1\\times 10^\{\-3\}to1\.8×10−21\.8\\times 10^\{\-2\}, while the corresponding 95% confidence intervals fall between7\.6×10−37\.6\\times 10^\{\-3\}and2\.3×10−22\.3\\times 10^\{\-2\}\.### A\.1Experiments

#### Setup\.

We annotate the datasets with six entity classes, using the label description as guidelines \(Table[5](https://arxiv.org/html/2607.26780#A1.T5)\)\. The same annotation is also used in the experiments \(§[3](https://arxiv.org/html/2607.26780#S3)\)\. For each setup of dataset and PLM, we run the fine\-tuning for three times, using 5\-fold cross\-validation and training for 8 epochs with a batch size of 16 and a learning rate set to2×10−52\\times 10^\{\-5\}\. The experiments are conducted on 4 NVIDIA Quadro RTX8000 GPUs\.

#### Model selection\.

While performing the validation task reformulation, we select oneRoBERTaand oneDeBERTaper dataset, making four setups in total for generating the first\-step PLM prediction\. In fine\-tuning, the repeated experiment runs produce consistent model performance, with the overall F1scores deviating by less than 1%\. We thereby further compare the models through evaluating them with the dataset they arenotfine\-tuned on, selecting the model scoring the highest accuracy under this cross\-evaluation schema, which implies better generalizability over unseen data\.

![Refer to caption](https://arxiv.org/html/2607.26780v1/graphs/llama_precision.png)

![Refer to caption](https://arxiv.org/html/2607.26780v1/graphs/mistral_precision.png)

![Refer to caption](https://arxiv.org/html/2607.26780v1/graphs/gemma_precision.png)

Figure 4:Precision score \(%\) per label forLlama\(top\),Mistral\(middle\), andGemma\(bottom\) model family\. The black solid lines denote the extraction task, while the other line styles represent different validation setups\.![Refer to caption](https://arxiv.org/html/2607.26780v1/graphs/llama_recall.png)

![Refer to caption](https://arxiv.org/html/2607.26780v1/graphs/mistral_recall.png)

![Refer to caption](https://arxiv.org/html/2607.26780v1/graphs/gemma_recall.png)

Figure 5:Recall score \(%\) per label forLlama\(top\),Mistral\(middle\), andGemma\(bottom\) model family\. The black solid lines denote the extraction task, while the other line styles represent different validation setups\.

### A\.2Results

For each setting, we run the experiments three times and report the mean score in Table[3](https://arxiv.org/html/2607.26780#A1.T3)\. Across the models and datasets, the standard deviations of overall F1range from3\.0×10−43\.0\\times 10^\{\-4\}to3\.7×10−33\.7\\times 10^\{\-3\}, while the standard deviations of cross\-evaluation from9\.7×10−49\.7\\times 10^\{\-4\}to3\.8×10−33\.8\\times 10^\{\-3\}\.

PLMs achieve high overall F1scores, though this is largely driven by performance on the O\-label, masking struggles with specific entity types\. We find that PLMs extract information more accurately at entities that are more explicitly stated in the text, such asweightandproduct number, as they tend to show up in specifications and could often include numeric values\. In contrast,componentbecomes the most challenging entity for both models since it involves more complex semantic comprehension over a paragraph or even the entire text\.

Though performing cross\-evaluation using a different dataset, we further examine fine\-tuned PLMs’ generalizability \(Table[4](https://arxiv.org/html/2607.26780#A1.T4)\)\. While using a different dataset for evaluation, we find performance more significantly dropping on the weakly expressed entities such ascomponentandmanufacturer\. The trend highlights data collection and labeling as bottlenecks in applying conventional methods to real\-world IE tasks, as the model may still fail to generalize to unseen data instances in spite of the shared domain knowledge\.

## Appendix BSupplementary Results

Extending F1scores in Table[2](https://arxiv.org/html/2607.26780#S3.T2), we report precision \(Table[6](https://arxiv.org/html/2607.26780#A1.T6)\) and recall \(Table[7](https://arxiv.org/html/2607.26780#A1.T7)\) scores of generative IE tasks and setups\. For the scores on each entity class, we visualize the scores in Figure[4](https://arxiv.org/html/2607.26780#A1.F4)and[5](https://arxiv.org/html/2607.26780#A1.F5)\.

You are an expert in information extraction\. You are given a task to find information described in a free\-text product description, and put them into structured form\.The designated output format should be in JSON and contains these keys: \{"size":"", "weight": "", "product number": "", "composition": "", "material": "", "manufacturer": ""\}Follow these principles:1\. Be concise and avoid verbose explanations\.2\. Use a list as value to include all the information\. Each key may include zero, one, or multiple suitable values\.3\. Some information may be lacking in the input text\. If no proper information could be retrieved, put None as the value\.4\. Output only the designated JSON, and don’t add any keys to its format\.Here are some examples of a free\-text product description and the target output:\*\*Input:\*\* This \#VAGA\-198964 table from Ikea is composed of a wooden surface and a steel frame\. Package size: 2x2m \| Package weight: 5\.8kg \| Shipping policy: Free delivery from 59€\*\*Output:\*\* \{"size": \["2x2m"\], "weight": \["5\.8kg"\], "product number": \["\#VAGA\-198964"\], "composition": \["surface", "frame"\], "material": \["wood", "steel"\], "manufacturer": \["Ikea"\]\}\*\*Input:\*\* The ergonomic wireless mouse is designed specifically for right\-handed users\. It is also pleasant to the touch and encourages a natural hand position\. Thanks to the interference\-free 2\.4 GHz wireless technology, the mouse can be used flexibly within a range of up to 10 m \- without any cables at all\*\*Output:\*\* \{"size": \[\], "weight": \[\], "product number": \[\], "composition": \[\], "material": \[\], "manufacturer": \[\]\},\*\*Input:\*\* The bed suite from Zara Home features two pillows and one quilt\. Suitable for double bed \(140x200\), they are made of cotton fibre and are machine washable\.\*\*Output:\*\* \{"size": \["140x220"\], "weight": \[\], "product number": \[\], "composition": \["pillow", "quilt"\], "material": \["cotton"\], "manufacturer": \["Zara Home"\]\}Perform the given task on the following text: \{TEXT\}You must return valid JSON with the required keys at the top level without adding any text\. If product description or any information is missing, return an empty list for that field\. Do not return error messages or any other keys\.

You are an expert in processing product data\. Your task is to validate structured output from a language model on information extraction\. Given the input and model output, check whether the model\-generated answer is correct or should be revised\.Stick to this JSON format and do not add any keys: \{"size":"", "weight": "", "product number": "", "composition": "", "material": "", "manufacturer": ""\}Follow these principles:1\. Be concise and avoid verbose explanations\.2\. Use a list as value to include all the information\. Each key may include zero, one, or multiple suitable values\.3\. Make minimal corrections to the model\-generated answer\. If it already looks correct, return it as the output\.4\. Do not include information that is not mentioned in the text\.Here are some examples of a free\-text product description and the validated, corrected target output:\*\*Input:\*\* This \#VAGA\-198964 table from Ikea is composed of a wooden surface and a steel frame\. Package size: 2x2m \| Package weight: 5\.8kg \| Shipping policy: Free delivery from 59€\*\*Output:\*\* \{"size": \["2x2m"\], "weight": \["5\.8kg"\], "product number": \["\#VAGA\-198964"\], "composition": \["surface", "frame"\], "material": \["wood", "steel"\], "manufacturer": \["Ikea"\]\}\*\*Input:\*\* The ergonomic wireless mouse is designed specifically for right\-handed users\. It is also pleasant to the touch and encourages a natural hand position\. Thanks to the interference\-free 2\.4 GHz wireless technology, the mouse can be used flexibly within a range of up to 10 m \- without any cables at all\*\*Output:\*\* \{"size": \[\], "weight": \[\], "product number": \[\], "composition": \[\], "material": \[\], "manufacturer": \[\]\},\*\*Input:\*\* The bed suite from Zara Home features two pillows and one quilt\. Suitable for double bed \(140x200\), they are made of cotton fibre and are machine washable\.\*\*Output:\*\* \{"size": \["140x220"\], "weight": \[\], "product number": \[\], "composition": \["pillow", "quilt"\], "material": \["cotton"\], "manufacturer": \["Zara Home"\]\}Perform the given task on the following text: \{TEXT\}You must return valid JSON with the required keys at the top level without adding any text\. If product description or any information is missing, return an empty list for that field\. Do not return error messages or any other keys\.

Figure 6:Prompt structure of the extraction \(green\) and validation task \(blue\)\. Considering that the instances from the datasets feature long text and thereby sparse named entities, we provide three manually created few\-shot examples in the prompt, each containing no more than three sentences to ensure high entity density\. Throughout the experiments, we identify issues with LLM output such as additional keys and values, gradually optimizing the instructions to ensure the robustness of the prompt\.

Similar Articles