AutoSpecNER: A Fine-Grained Named Entity Recognition Dataset for Vehicle Specification Extraction
Summary
Introduces AutoSpecNER, an expert-annotated dataset for fine-grained named entity recognition in vehicle listings, with 659 advertisements annotated across 15 entity types. Benchmark results show DeBERTa achieves 90% micro-F1, outperforming rule-based and LLM approaches.
View Cached Full Text
Cached at: 06/24/26, 07:47 AM
# AutoSpecNER: A Fine-Grained Named Entity Recognition Dataset for Vehicle Specification Extraction
Source: [https://arxiv.org/html/2606.24387](https://arxiv.org/html/2606.24387)
Jordan Lee1,2,\*,Filippos Ventirozos1,2,Abdirahman Abdullahm1,Ioanna Nteka2, Peter Appleby2,Matthew Shardlow1 1Department of Computing and Mathematics, Manchester Metropolitan University, UK 2Autotrader Research Group, Autotrader UK \{f\.ventirozos,m\.shardlow\}@mmu\.ac\.uk \*Work conducted during an internship at Autotrader UK
###### Abstract
Vehicle advertisements contain rich specification information, but automotive NER resources remain limited\. We introduceAutoSpecNER, an expert\-annotated dataset for fine\-grained entity recognition in vehicle listings\. The dataset includes 659 advertisements from a popular car\-selling website, with over 10,000 entities annotated across 15 categories, includingMODEL,ENGINE\_SPEC, andBATTERY\_CAPACITY\. Annotation quality was validated through inter\-annotator agreement, achieving an average score of 91\.5%\. We benchmark rule\-based extraction, fine\-tuned transformer encoders, and large language models\. DeBERTa achieves the best performance with a 90% micro\-F1 score, outperforming the rule\-based baseline \(43%\) and the strongest large language model \(77\.8%\)\.
AutoSpecNER: A Fine\-Grained Named Entity Recognition Dataset for Vehicle Specification Extraction
Jordan Lee1,2,\*,Filippos Ventirozos1,2,Abdirahman Abdullahm1,Ioanna Nteka2,Peter Appleby2,Matthew Shardlow11Department of Computing and Mathematics, Manchester Metropolitan University, UK2Autotrader Research Group, Autotrader UK\{f\.ventirozos,m\.shardlow\}@mmu\.ac\.uk\*Work conducted during an internship at Autotrader UK\.
## 1Introduction
The automotive industry generates vast amounts of unstructured text through online vehicle advertisements\. These advertisements contain valuable specification information embedded in free\-form descriptions, but extracting this structured data manually is impractical at scale\. This challenge is compounded by the recent introduction of AI\-generated advertisement content, which can contain hallucinations—factually incorrect information that may mislead consumers\.
Named entity recognition \(NER\) offers a solution by automatically identifying and extracting specific information spans from text\. While NER has been successfully applied to various domains including biomedical texts\(Majidet al\.,[2024](https://arxiv.org/html/2606.24387#bib.bib10)\), news articles\(Tjong Kim Sang and De Meulder,[2003](https://arxiv.org/html/2606.24387#bib.bib1)\), and social media\(Derczynskiet al\.,[2017](https://arxiv.org/html/2606.24387#bib.bib2)\), the automotive advertisement domain presents unique challenges:
- •Domain\-specific terminology: Technical specifications like “2\.0L TDI”, “DSG transmission”, or “Santorini \[black colour\]” require specialised understanding\.
- •Fine\-grained distinctions: Differentiating between similar concepts \(e\.g\., exterior vs\. interior colour, battery capacity vs\. range\)\.
- •Multi\-word entities: Complex specifications often span multiple tokens \(e\.g\., "18 minutes with 350kW charger"\)\.
- •Mixed content sources: Advertisements include both user\-generated content with typos and informal language, and AI\-generated content with potential hallucinations\.
This paper presents two main contributions\. First, we introduce Automotive Specification NER \(AutoSpecNER\), a new, publicly available dataset111Available at[github\.com/FilipposVentirozos/AutoSpecNER](https://github.com/FilipposVentirozos/AutoSpecNER)\.for fine\-grained NER tailored to the automotive domain\. The dataset contains 659 vehicle advertisements annotated with 15 entity types critical for vehicle identification and comparison\.
Second, we provide a comprehensive benchmark evaluation on AutoSpecNER\. We compare the performance of three diverse approaches: a rules\-based system, transformer\-based encoder models, and large language models \(LLMs\) prompted using few\-shot and self\-verification techniques\(Wanget al\.,[2023](https://arxiv.org/html/2606.24387#bib.bib21)\)\. Our analysis confirms that while NER is a viable technique for this task, the choice of model has a profound impact on performance, with fine\-tuned encoders demonstrating superior capabilities\.
## 2Related Work
### 2\.1Domain\-Specific NER and Fine\-Grained Entity Recognition
Standard NER benchmarks like CoNLL\-2003\(Tjong Kim Sang and De Meulder,[2003](https://arxiv.org/html/2606.24387#bib.bib1)\)focus on coarse\-grained entities \(person, location, organisation\) inadequate for technical domains\. Fine\-grained entity recognition\(Ling and Weld,[2012](https://arxiv.org/html/2606.24387#bib.bib11)\)requires distinguishing between closely related entity types—a challenge particularly acute in automotive contexts where hierarchical relationships exist between entities \(e\.g\., “2024 Ford F\-150 Limited” contains YEAR, MAKE, MODEL, and TRIM entities\)\.
Recent work has explored product and attribute extraction\(Putthividhya and Hu,[2011](https://arxiv.org/html/2606.24387#bib.bib12); Chenet al\.,[2023](https://arxiv.org/html/2606.24387#bib.bib13)\), demonstrating unique challenges in technical NER: domain\-specific abbreviations, overlapping entity boundaries, and hierarchical entity relationships\. Fine\-grained annotation schemas have proven essential for capturing technical specifications in industrial domains\(Bikaunet al\.,[2024](https://arxiv.org/html/2606.24387#bib.bib14)\), yet automotive\-specific resources remain limited\.
### 2\.2Automotive NER Research
The automotive domain has received minimal attention in NER research\.Hu and Ma \([2024](https://arxiv.org/html/2606.24387#bib.bib15)\)addressed NER for automotive accessories in Chinese, while recent work byVentirozoset al\.\([2024](https://arxiv.org/html/2606.24387#bib.bib16)\)introduces the Auto\-AdvER approach for English vehicle advertisements to understand the condition, historic claims and sales options offered\. Recent datasets like FindVehicle\(Guanet al\.,[2024](https://arxiv.org/html/2606.24387#bib.bib17)\)target vehicle retrieval with entity types including vehicle colour, brand, model, and location, but lack the granularity needed for technical specification extraction\.
Parket al\.\([2023](https://arxiv.org/html/2606.24387#bib.bib18)\)introduced ADMit, combining adversarial training and multi\-task learning for automotive NER for a Korean and English\. Their work addresses domain adaptation between general and automotive\-specific terminology but focuses on FAQ systems rather than technical specifications\. Our work differs by targeting fine\-grained vehicle identifiables critical for specification verification and hallucination detection\.
### 2\.3Neural Approaches and Domain Adaptation
Transformer\-based models have become standard for NER, with BERT\(Devlinet al\.,[2019](https://arxiv.org/html/2606.24387#bib.bib3)\)and its variants achieving strong baseline performance\. Domain\-specific pre\-training significantly improves technical entity recognition, as demonstrated in manufacturing domains where morphological patterns guide entity recognition\(Liet al\.,[2024](https://arxiv.org/html/2606.24387#bib.bib19)\)\.
Few\-shot NER methods combining knowledge graphs and contrastive learning show promise for low\-resource domains\(Zhanget al\.,[2024](https://arxiv.org/html/2606.24387#bib.bib20)\), addressing the scarcity of labelled automotive data\. While LLMs demonstrate competitive performance through prompting\(Wanget al\.,[2023](https://arxiv.org/html/2606.24387#bib.bib21)\), recent studies\(Naguibet al\.,[2024](https://arxiv.org/html/2606.24387#bib.bib24)\)show that smaller, specialised models often outperform LLMs in low\-resource technical domains, making them more practical for deployment in automotive applications\.
## 3The AutoSpecNER Dataset
### 3\.1Data Collection and Composition
We were provided 659 vehicle advertisements from one of the UK’s biggest online vehicle\-selling websites\. The dataset comprises two distinct sources:
User\-generated advertisements \(350\): The dataset was a sample of an equal distribution of ads written by individual sellers \(stratified across different UK counties\) and dealerships \(stratified across the largest dealerships in the UK\)\. The advertisements contain natural language variation, informal descriptions, typographical errors, and inconsistent formatting\.
Example:“lovleyford focus 1\.8diesal, 2015 plate, full mot till next yr, greymetalicpaint”\.
AI\-generated advertisements \(309\): Created222These were generated by the team of the aforementioned UK website company\. These AI generated descriptions were shown to the users and the users decided whether they want to keep them or edit/delete them\.using Google Gemini333gemini\-2\.0\-flash\-001and Meta’s LLaMA3444llama\-3\.1\-8b\-instructby providing in the prompt the vehicle specifications\. These are grammatically correct and well\-structured but may contain hallucinations\. For example in this advertisement for an Audi RSQ8, where the generated advert incorrectly refers to the vehicle as an Audi RS6:
> “The AudiRS6is a high\-performance car that boasts a powerful 4\.0\-litre V8 engine\. This petrol engine is paired with an automatic transmission…”
Other hallucinations exist which are more subtle than this such as many adverts have additional information inserted not present in the specifications, but could be true\. An example of this is in the following advert in which a Volkswagen Polo Match is referred to as a Volkswagen Polo EVO Match:
> “With only 21,029 miles on the clock, this 2021 Volkswagen PoloEVOMatch is manufacturer approved…”
This dual\-source approach enables investigation of how NER models handle both human errors \(typos, informality\) and AI errors \(hallucinations, specification confusion\)\. Crucially, each advertisement is accompanied by structured meta\-data \(a fact table\) that lists the vehicle’s actual specifications\. This is a vital characteristic, as it allows for the verification of extracted entities and forms the basis for potential error/hallucination detection systems\.
### 3\.2Corpus Annotation and Inter\-Annotator Agreement
To ensure the reliability and interpretability of our proposed 15\-label schema, we undertook a multi\-stage annotation process\. The process was iterative, involving an initial schema refinement phase to produce clear guidelines555The annotation guidelines can be found at[github\.com/FilipposVentirozos/AutoSpecNER](https://github.com/FilipposVentirozos/AutoSpecNER)\., followed by a final validation phase to confirm their efficacy\.
Our initial phase involved two annotators independently labelling a pilot set of 50 advertisements, balanced between user\- and AI\-generated content\. While this yielded a promising micro F1\-score of 0\.81, a qualitative analysis of the disagreements revealed systematic ambiguities\. This analysis was crucial for developing a set of explicit annotation principles\.
The key principles that emerged from this refinement process were:
- •Prioritise Specificity over General Description:Labels are reserved for specific, quantifiable details, not general descriptive statements\. For instance, the phrase “low CO2 emissions”, while describing an engine’s characteristic, is not annotated asENGINE\_SPECbecause it lacks a specific value\.
- •Maintain Entity Separation for Adjacent, Distinct Items:When multiple entities of the same type appear adjacently but refer to distinct details, they must be annotated as separate entities\. For example, in the text “…TDCI 1\.6L ECOnetic…”, each component \(“TDCI”, “1\.6L”, “ECOnetic”\) is labelled as a separateENGINE\_SPECentity\.
- •Disambiguate Overlapping Concepts:Clear distinctions were established to prevent confusion\. TheMAKElabel, for example, is restricted to the primary vehicle manufacturer; in ‘…a Ford Focus with a Mercedes…engine‘, only “Ford” is labelled asMAKE\. Similarly, composite names like “e\-SKYACTIV G Centre\-Line” are split into “e\-SKYACTIV G” \(ENGINE\_SPEC\) and “Centre\-Line” \(TRIM\)\.
- •Enforce Strict Contextual Boundaries:Annotators were instructed to capture the full, self\-contained descriptive phrase\. ForBATTERY\_RANGE, the entire phrase “maximum range of 280 miles when new” is captured, preserving the qualifying context\.
- •Exclude Speculative and Ancillary Information:The guidelines were refined to ensure that only the direct attribute of the advertised vehicle is annotated\. In cases where an advert mentions other available models \(e\.g\., “The Golf is also available as a Station Wagon \(Estate\)”\), only the body type of the primary advertised vehicle \(e\.g\., “Hatchback”\) is to be labelled\.
- •Annotate the Text, Not the Structured Specification:Annotations should reflect the entity span as it appears in the advertisement, even when it differs from the accompanying structured specification data or appears factually incorrect\. For example, if the advert describes the trim as “220d Luxury”, annotators label the span in the text rather than normalising it to the official specification value, “Luxury”\. This ensures that models learn to extract the claims made in the free\-text advert, which can later be compared against structured specifications to detect inconsistencies or hallucinations\.
### 3\.3Annotation Schema Overview
The final schema comprises 15 entity labels, organised into four groups and defined with representative examples in Table[1](https://arxiv.org/html/2606.24387#S3.T1)\.
Table 1:The 15\-labelAutoSpecNERannotation schema, grouped by category with definitions and representative examples\.#### 3\.3\.1Final Validation and Agreement
Following the guideline refinement, a final inter\-annotator agreement study was conducted\. This involved three annotators \(the authors of this paper\) of mixed nationalities with English as either their first or second professional language and varying levels of prior exposure to NLP annotation tasks, who labelled a new, larger random sample of 100 advertisements \(50 user\-generated and 50 AI\-generated\)\. Using the finalised guidelines, this validation achieved a strict\-matching average micro F1\-score666In NER in most cases Kappa is equivalent to F1, hence, we just measure F1\(Richieet al\.,[2022](https://arxiv.org/html/2606.24387#bib.bib34); Hripcsak and Rothschild,[2005](https://arxiv.org/html/2606.24387#bib.bib35)\)of91\.5%across all three annotators \(standard deviation:3\.2%\)\. Precision consistently exceeded recall by approximately 2 percentage points, indicating strong agreement on positive entity identification while annotators exhibited greater conservatism when entity boundaries or labels were ambiguous\.
### 3\.4Dataset Statistics
The dataset contains a total of 11,117 labelled entities overall\. Entity frequencies approximate a power law distribution, from MODEL \(2,426 instances\) to NO\_SEATS \(10 instances\)\. This reflects natural occurrence patterns \- every advertisement mentions the model, but seat count is specified only when notable\. One can look into Figure[1](https://arxiv.org/html/2606.24387#S3.F1)to view the label distributions for the user and the AI advertisement generated text respectively\. In total, the dataset contains 44 different brands, 237 unique models and the vehicles range from older \(first registered in 2006\) to newest \(this year 2025\)\.
Figure 1:Proportion of labels from each source within the dataset\. In dark blue are the the text advertisements generated by either Gemini or Llama and in green are the ones written by users\.
## 4Experiments
To prepare the annotated corpus for training, we performed two key steps: pre\-processing the span\-level annotations into a token\-level format and partitioning the corpus into training \(~70%\), validation \(~15%\), and testing \(~15%\) sets, the distribution of the entities can be found in Table[2](https://arxiv.org/html/2606.24387#S4.T2)\.
Table 2:Distribution of entity counts across the final training, validation, and testing sets\.The three different ML approaches operate on different units: sub\-word tokens for the encoders, whitespace tokens for the rule\-based matcher, and LLM\-generated free\-form text\. We therefore evaluate every system on a single, tokeniser\-independent character\-level IOB2 representation\. Specifically, for each advert, the gold character\-offset spans are expanded into a per\-character IOB2 sequence in which every character carries exactly one tag\. Within an entity span, the first character is taggedB\-and each subsequent character—including whitespace within a multi\-word entity—is taggedI\-; every character that falls outside any entity span is taggedO\(“Outside”\)\. Each system’s predictions are mapped onto the same per\-character IOB2 representation, so all systems are scored on identical units; we then reconstruct entities and report entity\-level precision, recall, and micro\-F1 withseqevalNakayama \([2018](https://arxiv.org/html/2606.24387#bib.bib26)\)\.
### 4\.1Models
We evaluated three classes of models on the AutoSpecNER dataset to establish a robust performance benchmark\.
#### 4\.1\.1Rules\-Based Approach
To establish a performance baseline and explore a non\-AI alternative, we developed a hybrid rules\-based system for entity extraction\. This approach was designed to be computationally inexpensive at inference time and serves as a benchmark against which our neural models are compared\. The system combines two core methodologies: taxonomy\-based matching and regular expressions\.
##### Methodology
Our system employs a two\-pronged strategy, applying different techniques based on the nature of the target entity\.
##### Taxonomy\-Based Matching
For eight of the core vehicle attributes \(MAKE,MODEL,TRIM,FUEL\_TYPE,BODY\_TYPE,TRANSMISSION,INTERIOR\_COLOUR, andEXTERIOR\_COLOUR\), we leveraged an external taxonomy provided by the company that provided the data\. This taxonomy acts as a gazetteer of known values for each entity\. The system iterates through the advertisement text, comparing text spans to the values in the gazetteer\. We evaluated several matching strategies:
- •Strict Matching:Requires an exact, case\-insensitive string match\.
- •Partial Matching:To account for minor variations and typographical errors, we implemented partial matching using a normalised Levenshtein distance function\. Predictions were made only if the similarity score exceeded a pre\-determined threshold\.
- •Heuristic Filtering:To improve precision, we experimented with two disambiguation heuristics: \(1\) amost\-common\-valuefilter, which retains only the most frequently matched entity for labels expected to appear once \(e\.g\.,MAKE\); and \(2\)hierarchical filtering, which uses known make\-model\-trim relationships from the taxonomy to prune invalid predictions \(e\.g\., removing a predicted trim that does not belong to the predicted model\)\.
##### Regular Expressions
For entities with highly predictable syntactic patterns that were not present in the taxonomy, we employed regular expressions\. This approach was used for three specific labels:BATTERY\_CAPACITY,NUMBER\_OF\_SEATS, andYEAR\. The patterns, detailed in Table[3](https://arxiv.org/html/2606.24387#S4.T3), were designed to be robust enough to capture common formulations of these attributes\.
Table 3:Regular expression patterns used for specific entity extraction\.
#### 4\.1\.2Transformer Encoders
We fine\-tuned four widely\-used encoder models: BERT\-base\-cased\(Devlinet al\.,[2019](https://arxiv.org/html/2606.24387#bib.bib3)\), RoBERTa\-base\(Liuet al\.,[2019](https://arxiv.org/html/2606.24387#bib.bib4)\), ModernBERT\-base\(Warneret al\.,[2025](https://arxiv.org/html/2606.24387#bib.bib27)\)and DeBERTa\-v3\-base\(Heet al\.,[2023](https://arxiv.org/html/2606.24387#bib.bib8)\)\. To perform the training and inference, we utilised the HuggingFace library\(Wolfet al\.,[2020](https://arxiv.org/html/2606.24387#bib.bib28)\)for token classification\.
We perform grid search over learning rates\{5×10−5,3×10−5\}\\\{5\\times 10^\{\-5\},3\\times 10^\{\-5\}\\\}, batch sizes\{4,8,16\}\\\{4,8,16\\\}, and weight decay values\{0\.0,0\.01,0\.1\}\\\{0\.0,0\.01,0\.1\\\}to identify optimal hyperparameters for each encoder architecture\. We train all models for a maximum of 50 epochs with early stopping patience of 3 epochs, monitoring the weighted F1 score on the validation set\. The best configuration across all models uses a learning rate2×10−52\\times 10^\{\-5\}, batch size 8, weight decay 0\.01, with a warmup of 100 steps and a maximum sequence length of 512 tokens used in all experiments\.
#### 4\.1\.3Large Language Models
To provide a contemporary comparison to encoder\-based models, we evaluated four open sourced LLMs for the NER task: Qwen3\(Yanget al\.,[2025](https://arxiv.org/html/2606.24387#bib.bib29)\)Mistral\-7B\(Jianget al\.,[2023](https://arxiv.org/html/2606.24387#bib.bib5)\), LLaMA\-3\-8B\(Grattafioriet al\.,[2024](https://arxiv.org/html/2606.24387#bib.bib23)\), Gemma\-3\(Teamet al\.,[2025b](https://arxiv.org/html/2606.24387#bib.bib30)\)\. Additionally, we evaluated two closed\-sourced LLMs: Gemini\-2\.5\(Teamet al\.,[2025a](https://arxiv.org/html/2606.24387#bib.bib32)\)and GPT\-5\(Singhet al\.,[2026](https://arxiv.org/html/2606.24387#bib.bib33)\)\. These models were selected as representative of smaller and resource\-efficient architectures\. All the specific versions with their respective parameter sizes are listed in Table[4](https://arxiv.org/html/2606.24387#S5.T4)\.
##### Prompt Engineering
We adopted the GPT\-NER methodology\(Wanget al\.,[2023](https://arxiv.org/html/2606.24387#bib.bib21)\), implementing a single\-turn per\-label prompting strategy, where each advertisement was processed 15 times, once for each entity type\. The refined prompting approach incorporated three key components:
- •Task\-specific prompts: Each prompt focused on a single entity type with a concise task description and label definition
- •Few\-shot examples: 2 few\-shot in\-context learning examples were used\.
- •Simplified output format: The output would copy verbatim the example that we want labelled and the entities of interested would be marked with boundary delimiters \(@@…\|\|\) rather than full NER formatting\.
##### Self\-Verification
Following\(Wanget al\.,[2023](https://arxiv.org/html/2606.24387#bib.bib21)\), a self\-verification step was implemented to mitigate the high false positive rate characteristic of LLM\-based NER\. This secondary prompt presented the model’s initial predictions for validation, significantly improving precision whilst marginally reducing recall\. Preliminary experiments, found in Appendix[A](https://arxiv.org/html/2606.24387#A1), demonstrated the efficacy of including this step\.
Complete prompt templates are provided in Appendix[B](https://arxiv.org/html/2606.24387#A2)\.
## 5Results
### 5\.1Rules\-based Approach
The rules\-based approach utilised a combination of taxonomy\-based string matching and regular expressions\. Performance was highly dependent on the complexity of the entity and the matching method employed\. Methods involving partial matching via Fuzzy and Levenshtein algorithms, while more accurate for variable text, were significantly slower at inference time compared to exact matching\.
Table[5](https://arxiv.org/html/2606.24387#S5.T5)summarises the best F1\-scores achieved for each label\. As shown, simpler, more regular entities likeYEAR,MAKE, andMODELachieved strong performance \(0\.77–0\.79 F1\)\. In contrast, more complex and nuanced labels such asINTERIOR\_COLOURandEXTERIOR\_COLOURperformed poorly, with F1\-scores below 0\.03\.
An analysis of precision and recall scores reported in Appendix[C](https://arxiv.org/html/2606.24387#A3)reveal that regex\-based methods, where applicable \(e\.g\., forYEAR\), achieved high recall at the cost of precision, whereas taxonomy\-based matching often resulted in higher precision but lower recall, as it failed to capture variations not present in the pre\-defined lists\. Overall, when evaluated across all present labels, the best\-performing configuration of the rules\-based approach achieved a micro\-averaged F1\-score of 0\.43\.
### 5\.2Encoder\-based Models
We evaluated four transformer\-based encoders: BERT, RoBERTa, ModernBERT and DeBERTa\. Table[4](https://arxiv.org/html/2606.24387#S5.T4)summarises the character\-level performance of all evaluated models\. Fine\-tuned encoder models achieve higher performance, with micro\-F1 scores ranging from 82\.9% to 90\.1%, compared to LLMs which achieve only 77\.8% to 41\.7% micro\-F1\.
Table 4:Overall character\-level performance comparison\. Encoders consistently outperform LLMs with 2\-shot in\-context learning\. The micro\-F1 scores are shown in descending order\.The DeBERTa model demonstrated the best overall performance, achieving an F1 score of 90\.1%\. After resampling the data splits and repeating the experiments three times, the standard deviation was 2\.3%, indicating relatively stable performance across runs\. This strong performance may be attributable to DeBERTa’s disentangled attention mechanism, which more effectively captures positional information, a feature that is particularly relevant for the consistently structured AI\-generated adverts\.
Table[5](https://arxiv.org/html/2606.24387#S5.T5)shows per\-label F1 scores for all the family models\. The encoders perform best on numeric entities such asYEARandBATTERY\_CAPACITY, standardised attributes likeFUEL\_TYPEandTRANSMISSION, and vehicle identifiers includingMAKEandMODEL\. In contrast, they struggle with more subjective attributes such asINTERIOR\_COLOURandEXTERIOR\_COLOUR, likely due to the wide lexical variation in colour descriptions \(e\.g\., “midnight blue”, “pearl white”, “metallic silver”\)\.
Our analysis of the training data size indicated that the dataset was sufficient for the task, as model performance \(measured by evaluation loss\) began to plateau after approximately 300–350 training samples\. Further details and the corresponding graph are available in Appendix[D](https://arxiv.org/html/2606.24387#A4)\. The one exception is theNO\_SEATSwhich was scarcely reported in the text descriptions\.
The decoder models, which are much larger in parameter size—many times more than 100 times larger—than their encoder counterparts excelled in various labels, especially in those with less support, comparatively to the encoders\.
Table 5:Per\-label F1 scores comparing encoder, LLM, and rules\-based performance\. In bold are the top performants for each label\. Some model names are shortened for illustration purposes but are the same version with the ones presented in Table[4](https://arxiv.org/html/2606.24387#S5.T4)\.Gemini\-2\.5 scored the highest amongst the LLMs but still second to the encoders\. Its predictions across the test set reveal heterogeneous performance across entity types: strong \(F1\>\>75%\) on YEAR,MAKE, andMODEL; moderate \(F1 50–75%\) onENGINE\_SPEC,EXTERIOR\_COLOUR,TRIM,BATTERY\_CAPACITY, andBODY\_TYPE; and poor \(F1<<50%\) onINTERIOR\_COLOUR,TRANSMISSION, and notablyRECHARGE\_TIME\(0% F1 despite 10 test instances\)\.
LLMs exhibit three primary failure modes\. First,boundary detection errors: extracting “300 miles” instead of “approximately 300 miles”BATTERY\_RANGE\), or “408” instead of “408 horsepower” \(ENGINE\_SPEC\)\. Second,variant normalisation failures: achieving high recall on simpleFUEL\_TYPEvalues \(“petrol”, “diesel”, “electric”\) while missing hybrid variants \(“Petrol Hybrid”, “Petrol Plug\-in Hybrid” , “Diesel Hybrid”\)\.
We identify three primary factors contributing to these LLM failures\. First, our 2\-shot learning approach may not be indicative enough for the LLMs to perform inference\. Second, vehicle specifications involve domain\-specific terminology \(“kWh”, “bhp”, “T\-GDI”, “BiTurbo”\) and complex entity boundaries \(e\.g\., “2\.0L Turbocharged Inline\-4” asENGINE\_SPECentities\), for which pre\-trained LLMs lack adequate exposure during training, unlike general entities such as persons or organisations\. Third, the severe lexical variation within certain entity types \(e\.g\.,INTERIOR\_COLOUR: “Bengal red Nappa leather”, “Light Oyster and Ebony”, “Full Black Valcona Leather”\) exceeds what can be captured in 2\-shot demonstrations, while more standardised types \(YEAR,MAKE\) benefit from prior knowledge in the pre\-training corpus\.
Albeit, the LLMs proved resourceful, achieving the highest F1 score amongst the other approaches, when the support was quite low in the labelsNO\_SEATS,BATTERY\_RANGEandBOOT\_SIZE\.
### 5\.3Computational Resources
All encoder\-based transformer models were fine\-tuned on an Apple M4 MacBook\. Open\-source LLM inference was performed using vLLM\(Kwonet al\.,[2023](https://arxiv.org/html/2606.24387#bib.bib25)\)on NVIDIA RTX PRO 6000 \(96GB VRAM\) or NVIDIA H200 SXM \(141GB VRAM\) GPUs, depending on model size\. Closed\-source models were accessed through their respective provider APIs\. To ensure deterministic outputs, we set the temperature to 0\.0 where supported and used greedy decoding for all generations\. Additional hardware and runtime details are provided in Appendix[E](https://arxiv.org/html/2606.24387#A5)\.
## 6Discussion
The current investigation has implications for deploying NER systems in production automotive applications\. When the goal is extracting entity types and populating structured vehicle databases, fine\-tuned encoder models—despite their training requirement—remain the best solution overall, with DeBERTa leading when the supporting examples are sufficient, followed by the LLMs and last the rule\-based approaches\.
Looking beyond immediate deployment constraints and efficiency considerations, several avenues could improve LLM performance on this task\. Increasing the few\-shot to more examples per entity type could improve prediction accuracy\. Similarly, retrieval\-augmented prompting could dynamically retrieve relevant examples that could boost the score as well\. Finally, models further pre\-trained or post\-trained on automotive corpora may better handle specialised terminology and entity patterns\.
The high accuracy of the fine\-tuned encoder models enables several valuable downstream applications\. While this work focused on advertisements, the models could be evaluated for tracking vehicles from unstructured social media posts to monitor consumer trends, popular features, and brand perception\. Additionally, the process of extracting entities can be used to detect hallucinations in AI\-generated content by verifying whether a model that generated the content included additional or different specifications and whether it adhered to the structured vehicle specification data given, thereby enhancing platform integrity\. It was notable how on average we could spot an around 20% increase in F1 score when detecting entities in the AI generated text alone\. The fine\-tuned encoders could prove resourceful at tracking these mentions from AI generated text, raising flags when necessary\.
## 7Conclusions
In this paper, we addressed the high\-impact challenge of extracting fine\-grained vehicle specifications from the text descriptions of car advertisements\. Our dataset and annotation schemaAutoSpecNERcomprises 659 advertisements—written by real users and AI\-generated proprietary data acquired from the UK’s biggest online vehicle\-selling website—labelled with 15 specification types ranging fromMAKEandMODELtoRECHARGE\_TIME, with the schema validated at a commendable 91\.5% inter\-annotator F1 score\. Benchmarking rule\-based, fine\-tuned encoder, and few\-shot decoder approaches, we found the fine\-tuned encoders exceptionally well\-suited to this task, with DeBERTa achieving a top micro F1\-score of 0\.901, followed by the decoders \(led by Gemini\-2\.5\) and last the rule\-based approaches; the LLMs proved particularly resourceful for low\-support labels scarcely reported in the text\. Finally, this dataset can support market trend analysis and, crucially, the detection of factual hallucinations in AI\-generated content\.
## 8Limitations
The dataset was drawn exclusively from a proprietary database maintained by a single UK\-based company\. Although substantial, it is unlikely to capture the full diversity of linguistic usage, document structures, and formatting conventions observed across different countries, platforms, and automotive retailers\. In addition, the corpus is UK\-specific and monolingual \(English\), preventing assessment of cross\-lingual performance\. Improving generalisability would require a more heterogeneous, multi\-source, and multilingual dataset\.
## 9Ethical Considerations
The development of theAutoSpecNERdataset and its associated models has been guided by key ethical considerations\. The dataset was constructed from publicly accessible advertisements available on a major commercial platform\. A manual review of the user\-generated content was conducted, which confirmed the absence of Personally Identifiable Information \(PII\) in the source data\. While the data’s public availability and freedom from PII mitigate privacy concerns, its origin introduces a risk of socio\-economic bias\. The model may exhibit performance disparities across content from different seller types, and if the system is more accurate on professionally formatted advertisements than on informal text from private sellers, its application could inadvertently create an unfair marketplace\.
Consideration must also be given to the potential for misuse and downstream harms\. The same technology that extracts specifications can be repurposed for malicious ends\. Models trained on this data could be exploited to generate convincing but fraudulent vehicle advertisements, either for individual scams or at scale to manipulate market perceptions of a vehicle’s value and availability\. In a deployed application, over\-reliance on the system’s output could lead to automation bias, where users and platform administrators place undue trust in the automated extractions\. This could cause model errors to be systematically propagated, leading consumers to make purchasing decisions based on incorrect information\.
Finally, we address the environmental impact of this research\. Our experiments involved training and evaluating multiple large\-scale models, which consumed significant computational resources and energy\. However, our findings make a positive contribution in this area by demonstrating that smaller, fine\-tuned encoder models significantly outperform their much larger and more resource\-intensive LLM counterparts for this task\. This result advocates for a more sustainable approach to deploying NLP in production environments, showing that for specialised industrial tasks, more energy\-efficient models can also be the most effective\.
## References
- MaintIE: a fine\-grained annotation schema and benchmark for information extraction from maintenance short texts\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\),pp\. 10939–10951\.External Links:[Link](https://aclanthology.org/2024.lrec-main.954/)Cited by:[§2\.1](https://arxiv.org/html/2606.24387#S2.SS1.p2.1)\.
- W\. Chen, K\. Shinzato, N\. Yoshinaga, and Y\. Xia \(2023\)Does named entity recognition truly not scale up to real\-world product attribute extraction?\.InProceedings of the 2023 Conference on Empirical Methods in Natural Language Processing: Industry Track,pp\. 163–173\.External Links:[Link](https://aclanthology.org/2023.emnlp-industry.16/)Cited by:[§2\.1](https://arxiv.org/html/2606.24387#S2.SS1.p2.1)\.
- L\. Derczynski, E\. Nichols, M\. van Erp, and N\. Limsopatham \(2017\)Results of the WNUT2017 shared task on novel and emerging entity recognition\.InProceedings of the 3rd Workshop on Noisy User\-generated Text,L\. Derczynski, W\. Xu, A\. Ritter, and T\. Baldwin \(Eds\.\),Copenhagen, Denmark,pp\. 140–147\.External Links:[Link](https://aclanthology.org/W17-4418/),[Document](https://dx.doi.org/10.18653/v1/W17-4418)Cited by:[§1](https://arxiv.org/html/2606.24387#S1.p2.1)\.
- J\. Devlin, M\. Chang, K\. Lee, and K\. Toutanova \(2019\)BERT: pre\-training of deep bidirectional transformers for language understanding\.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 \(Long and Short Papers\),J\. Burstein, C\. Doran, and T\. Solorio \(Eds\.\),Minneapolis, Minnesota,pp\. 4171–4186\.External Links:[Link](https://aclanthology.org/N19-1423/),[Document](https://dx.doi.org/10.18653/v1/N19-1423)Cited by:[§2\.3](https://arxiv.org/html/2606.24387#S2.SS3.p1.1),[§4\.1\.2](https://arxiv.org/html/2606.24387#S4.SS1.SSS2.p1.1)\.
- A\. Grattafiori, A\. Dubey, A\. Jauhri, A\. Pandey, A\. Kadian, A\. Al\-Dahle, A\. Letman, A\. Mathur, A\. Schelten, A\. Vaughan, A\. Yang, A\. Fan, A\. Goyal, A\. Hartshorn, A\. Yang, A\. Mitra, A\. Sravankumar, A\. Korenev, A\. Hinsvark, A\. Rao, A\. Zhang, A\. Rodriguez, A\. Gregerson, A\. Spataru, B\. Roziere, B\. Biron, B\. Tang, B\. Chern, C\. Caucheteux, C\. Nayak, C\. Bi, C\. Marra, C\. McConnell, C\. Keller, C\. Touret, C\. Wu, C\. Wong, C\. C\. Ferrer, C\. Nikolaidis, D\. Allonsius, D\. Song, D\. Pintz, D\. Livshits, D\. Wyatt, D\. Esiobu, D\. Choudhary, D\. Mahajan, D\. Garcia\-Olano, D\. Perino, D\. Hupkes, E\. Lakomkin, E\. AlBadawy, E\. Lobanova, E\. Dinan, E\. M\. Smith, F\. Radenovic, F\. Guzmán, F\. Zhang, G\. Synnaeve, G\. Lee, G\. L\. Anderson, G\. Thattai, G\. Nail, G\. Mialon, G\. Pang, G\. Cucurell, H\. Nguyen, H\. Korevaar, H\. Xu, H\. Touvron, I\. Zarov, I\. A\. Ibarra, I\. Kloumann, I\. Misra, I\. Evtimov, J\. Zhang, J\. Copet, J\. Lee, J\. Geffert, J\. Vranes, J\. Park, J\. Mahadeokar, J\. Shah, J\. van der Linde, J\. Billock, J\. Hong, J\. Lee, J\. Fu, J\. Chi, J\. Huang, J\. Liu, J\. Wang, J\. Yu, J\. Bitton, J\. Spisak, J\. Park, J\. Rocca, J\. Johnstun, J\. Saxe, J\. Jia, K\. V\. Alwala, K\. Prasad, K\. Upasani, K\. Plawiak, K\. Li, K\. Heafield, K\. Stone, K\. El\-Arini, K\. Iyer, K\. Malik, K\. Chiu, K\. Bhalla, K\. Lakhotia, L\. Rantala\-Yeary, L\. van der Maaten, L\. Chen, L\. Tan, L\. Jenkins, L\. Martin, L\. Madaan, L\. Malo, L\. Blecher, L\. Landzaat, L\. de Oliveira, M\. Muzzi, M\. Pasupuleti, M\. Singh, M\. Paluri, M\. Kardas, M\. Tsimpoukelli, M\. Oldham, M\. Rita, M\. Pavlova, M\. Kambadur, M\. Lewis, M\. Si, M\. K\. Singh, M\. Hassan, N\. Goyal, N\. Torabi, N\. Bashlykov, N\. Bogoychev, N\. Chatterji, N\. Zhang, O\. Duchenne, O\. Çelebi, P\. Alrassy, P\. Zhang, P\. Li, P\. Vasic, P\. Weng, P\. Bhargava, P\. Dubal, P\. Krishnan, P\. S\. Koura, P\. Xu, Q\. He, Q\. Dong, R\. Srinivasan, R\. Ganapathy, R\. Calderer, R\. S\. Cabral, R\. Stojnic, R\. Raileanu, R\. Maheswari, R\. Girdhar, R\. Patel, R\. Sauvestre, R\. Polidoro, R\. Sumbaly, R\. Taylor, R\. Silva, R\. Hou, R\. Wang, S\. Hosseini, S\. Chennabasappa, S\. Singh, S\. Bell, S\. S\. Kim, S\. Edunov, S\. Nie, S\. Narang, S\. Raparthy, S\. Shen, S\. Wan, S\. Bhosale, S\. Zhang, S\. Vandenhende, S\. Batra, S\. Whitman, S\. Sootla, S\. Collot, S\. Gururangan, S\. Borodinsky, T\. Herman, T\. Fowler, T\. Sheasha, T\. Georgiou, T\. Scialom, T\. Speckbacher, T\. Mihaylov, T\. Xiao, U\. Karn, V\. Goswami, V\. Gupta, V\. Ramanathan, V\. Kerkez, V\. Gonguet, V\. Do, V\. Vogeti, V\. Albiero, V\. Petrovic, W\. Chu, W\. Xiong, W\. Fu, W\. Meers, X\. Martinet, X\. Wang, X\. Wang, X\. E\. Tan, X\. Xia, X\. Xie, X\. Jia, X\. Wang, Y\. Goldschlag, Y\. Gaur, Y\. Babaei, Y\. Wen, Y\. Song, Y\. Zhang, Y\. Li, Y\. Mao, Z\. D\. Coudert, Z\. Yan, Z\. Chen, Z\. Papakipos, A\. Singh, A\. Srivastava, A\. Jain, A\. Kelsey, A\. Shajnfeld, A\. Gangidi, A\. Victoria, A\. Goldstand, A\. Menon, A\. Sharma, A\. Boesenberg, A\. Baevski, A\. Feinstein, A\. Kallet, A\. Sangani, A\. Teo, A\. Yunus, A\. Lupu, A\. Alvarado, A\. Caples, A\. Gu, A\. Ho, A\. Poulton, A\. Ryan, A\. Ramchandani, A\. Dong, A\. Franco, A\. Goyal, A\. Saraf, A\. Chowdhury, A\. Gabriel, A\. Bharambe, A\. Eisenman, A\. Yazdan, B\. James, B\. Maurer, B\. Leonhardi, B\. Huang, B\. Loyd, B\. D\. Paola, B\. Paranjape, B\. Liu, B\. Wu, B\. Ni, B\. Hancock, B\. Wasti, B\. Spence, B\. Stojkovic, B\. Gamido, B\. Montalvo, C\. Parker, C\. Burton, C\. Mejia, C\. Liu, C\. Wang, C\. Kim, C\. Zhou, C\. Hu, C\. Chu, C\. Cai, C\. Tindal, C\. Feichtenhofer, C\. Gao, D\. Civin, D\. Beaty, D\. Kreymer, D\. Li, D\. Adkins, D\. Xu, D\. Testuggine, D\. David, D\. Parikh, D\. Liskovich, D\. Foss, D\. Wang, D\. Le, D\. Holland, E\. Dowling, E\. Jamil, E\. Montgomery, E\. Presani, E\. Hahn, E\. Wood, E\. Le, E\. Brinkman, E\. Arcaute, E\. Dunbar, E\. Smothers, F\. Sun, F\. Kreuk, F\. Tian, F\. Kokkinos, F\. Ozgenel, F\. Caggioni, F\. Kanayet, F\. Seide, G\. M\. Florez, G\. Schwarz, G\. Badeer, G\. Swee, G\. Halpern, G\. Herman, G\. Sizov, Guangyi, Zhang, G\. Lakshminarayanan, H\. Inan, H\. Shojanazeri, H\. Zou, H\. Wang, H\. Zha, H\. Habeeb, H\. Rudolph, H\. Suk, H\. Aspegren, H\. Goldman, H\. Zhan, I\. Damlaj, I\. Molybog, I\. Tufanov, I\. Leontiadis, I\. Veliche, I\. Gat, J\. Weissman, J\. Geboski, J\. Kohli, J\. Lam, J\. Asher, J\. Gaya, J\. Marcus, J\. Tang, J\. Chan, J\. Zhen, J\. Reizenstein, J\. Teboul, J\. Zhong, J\. Jin, J\. Yang, J\. Cummings, J\. Carvill, J\. Shepard, J\. McPhie, J\. Torres, J\. Ginsburg, J\. Wang, K\. Wu, K\. H\. U, K\. Saxena, K\. Khandelwal, K\. Zand, K\. Matosich, K\. Veeraraghavan, K\. Michelena, K\. Li, K\. Jagadeesh, K\. Huang, K\. Chawla, K\. Huang, L\. Chen, L\. Garg, L\. A, L\. Silva, L\. Bell, L\. Zhang, L\. Guo, L\. Yu, L\. Moshkovich, L\. Wehrstedt, M\. Khabsa, M\. Avalani, M\. Bhatt, M\. Mankus, M\. Hasson, M\. Lennie, M\. Reso, M\. Groshev, M\. Naumov, M\. Lathi, M\. Keneally, M\. Liu, M\. L\. Seltzer, M\. Valko, M\. Restrepo, M\. Patel, M\. Vyatskov, M\. Samvelyan, M\. Clark, M\. Macey, M\. Wang, M\. J\. Hermoso, M\. Metanat, M\. Rastegari, M\. Bansal, N\. Santhanam, N\. Parks, N\. White, N\. Bawa, N\. Singhal, N\. Egebo, N\. Usunier, N\. Mehta, N\. P\. Laptev, N\. Dong, N\. Cheng, O\. Chernoguz, O\. Hart, O\. Salpekar, O\. Kalinli, P\. Kent, P\. Parekh, P\. Saab, P\. Balaji, P\. Rittner, P\. Bontrager, P\. Roux, P\. Dollar, P\. Zvyagina, P\. Ratanchandani, P\. Yuvraj, Q\. Liang, R\. Alao, R\. Rodriguez, R\. Ayub, R\. Murthy, R\. Nayani, R\. Mitra, R\. Parthasarathy, R\. Li, R\. Hogan, R\. Battey, R\. Wang, R\. Howes, R\. Rinott, S\. Mehta, S\. Siby, S\. J\. Bondu, S\. Datta, S\. Chugh, S\. Hunt, S\. Dhillon, S\. Sidorov, S\. Pan, S\. Mahajan, S\. Verma, S\. Yamamoto, S\. Ramaswamy, S\. Lindsay, S\. Lindsay, S\. Feng, S\. Lin, S\. C\. Zha, S\. Patil, S\. Shankar, S\. Zhang, S\. Zhang, S\. Wang, S\. Agarwal, S\. Sajuyigbe, S\. Chintala, S\. Max, S\. Chen, S\. Kehoe, S\. Satterfield, S\. Govindaprasad, S\. Gupta, S\. Deng, S\. Cho, S\. Virk, S\. Subramanian, S\. Choudhury, S\. Goldman, T\. Remez, T\. Glaser, T\. Best, T\. Koehler, T\. Robinson, T\. Li, T\. Zhang, T\. Matthews, T\. Chou, T\. Shaked, V\. Vontimitta, V\. Ajayi, V\. Montanez, V\. Mohan, V\. S\. Kumar, V\. Mangla, V\. Ionescu, V\. Poenaru, V\. T\. Mihailescu, V\. Ivanov, W\. Li, W\. Wang, W\. Jiang, W\. Bouaziz, W\. Constable, X\. Tang, X\. Wu, X\. Wang, X\. Wu, X\. Gao, Y\. Kleinman, Y\. Chen, Y\. Hu, Y\. Jia, Y\. Qi, Y\. Li, Y\. Zhang, Y\. Zhang, Y\. Adi, Y\. Nam, Yu, Wang, Y\. Zhao, Y\. Hao, Y\. Qian, Y\. Li, Y\. He, Z\. Rait, Z\. DeVito, Z\. Rosnbrick, Z\. Wen, Z\. Yang, Z\. Zhao, and Z\. Ma \(2024\)The llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§4\.1\.3](https://arxiv.org/html/2606.24387#S4.SS1.SSS3.p1.1)\.
- R\. Guan, K\. L\. Man, F\. Chen, S\. Yao, R\. Hu, X\. Zhu, J\. Smith, E\. G\. Lim, and Y\. Yue \(2024\)FindVehicle and vehiclefinder: a ner dataset for natural language\-based vehicle retrieval and a keyword\-based cross\-modal vehicle retrieval system\.Multimedia Tools and ApplicationsExpert Syst\. Appl\.arXiv preprint arXiv:2304\.10428arXiv preprint arXiv:2305\.15444arXiv preprint arXiv:2505\.09388Journal of the American Medical Informatics Association83\(8\),pp\. 24841–24874\.External Links:[Document](https://dx.doi.org/10.1007/s11042-023-16373-y),[Link](https://doi.org/10.1007/s11042-023-16373-y),ISSN 1573\-7721Cited by:[§2\.2](https://arxiv.org/html/2606.24387#S2.SS2.p1.1)\.
- P\. He, J\. Gao, and W\. Chen \(2023\)DeBERTaV3: improving deberta using electra\-style pre\-training with gradient\-disentangled embedding sharing\.External Links:2111\.09543,[Link](https://arxiv.org/abs/2111.09543)Cited by:[§4\.1\.2](https://arxiv.org/html/2606.24387#S4.SS1.SSS2.p1.1)\.
- G\. Hripcsak and A\. S\. Rothschild \(2005\)Agreement, the f\-measure, and reliability in information retrieval\.12\(3\),pp\. 296–298\.Note:Epub 2005 Jan 31External Links:[Document](https://dx.doi.org/10.1197/jamia.M1733)Cited by:[footnote 6](https://arxiv.org/html/2606.24387#footnote6)\.
- S\. Hu and R\. Ma \(2024\)Named entity recognition of automotive parts based on roberta\-crf model\.In2024 4th International Conference on Neural Networks, Information and Communication Engineering \(NNICE\),Vol\.,pp\. 604–612\.External Links:[Document](https://dx.doi.org/10.1109/NNICE61279.2024.10499162)Cited by:[§2\.2](https://arxiv.org/html/2606.24387#S2.SS2.p1.1)\.
- A\. Q\. Jiang, A\. Sablayrolles, A\. Mensch, C\. Bamford, D\. S\. Chaplot, D\. de las Casas, F\. Bressand, G\. Lengyel, G\. Lample, L\. Saulnier, L\. R\. Lavaud, M\. Lachaux, P\. Stock, T\. L\. Scao, T\. Lavril, T\. Wang, T\. Lacroix, and W\. E\. Sayed \(2023\)Mistral 7b\.External Links:2310\.06825,[Link](https://arxiv.org/abs/2310.06825)Cited by:[§4\.1\.3](https://arxiv.org/html/2606.24387#S4.SS1.SSS3.p1.1)\.
- W\. Kwon, Z\. Li, S\. Zhuang, Y\. Sheng, L\. Zheng, C\. H\. Yu, J\. E\. Gonzalez, H\. Zhang, and I\. Stoica \(2023\)Efficient memory management for large language model serving with pagedattention\.InProceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles,Cited by:[Appendix E](https://arxiv.org/html/2606.24387#A5.p2.1),[§5\.3](https://arxiv.org/html/2606.24387#S5.SS3.p1.1)\.
- R\. Li, P\. Wang, L\. Wang, D\. Yang, and D\. Cai \(2024\)A corpus and method for Chinese named entity recognition in manufacturing\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\),pp\. 264–272\.External Links:[Link](https://aclanthology.org/2024.lrec-main.24/)Cited by:[§2\.3](https://arxiv.org/html/2606.24387#S2.SS3.p1.1)\.
- X\. Ling and D\. S\. Weld \(2012\)Fine\-grained entity recognition\.InProceedings of the Twenty\-Sixth AAAI Conference on Artificial Intelligence,pp\. 94–100\.External Links:[Link](https://www.aaai.org/ocs/index.php/AAAI/AAAI12/paper/view/5152)Cited by:[§2\.1](https://arxiv.org/html/2606.24387#S2.SS1.p1.1)\.
- Y\. Liu, M\. Ott, N\. Goyal, J\. Du, M\. Joshi, D\. Chen, O\. Levy, M\. Lewis, L\. Zettlemoyer, and V\. Stoyanov \(2019\)RoBERTa: a robustly optimized bert pretraining approach\.External Links:1907\.11692,[Link](https://arxiv.org/abs/1907.11692)Cited by:[§4\.1\.2](https://arxiv.org/html/2606.24387#S4.SS1.SSS2.p1.1)\.
- I\. Majid, V\. Mishra, R\. Ravindranath, and S\. Y\. Wang \(2024\)Evaluating the performance of large language models for named entity recognition in ophthalmology clinical free\-text notes\.Journal of the American Medical Informatics Association\.Note:Open AccessCited by:[§1](https://arxiv.org/html/2606.24387#S1.p2.1)\.
- M\. Naguib, X\. Tannier, and A\. Névéol \(2024\)Few\-shot clinical entity recognition in English, French and Spanish: masked language models outperform generative model prompting\.InFindings of the Association for Computational Linguistics: EMNLP 2024,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 6829–6852\.External Links:[Link](https://aclanthology.org/2024.findings-emnlp.400/),[Document](https://dx.doi.org/10.18653/v1/2024.findings-emnlp.400)Cited by:[§2\.3](https://arxiv.org/html/2606.24387#S2.SS3.p2.1)\.
- H\. Nakayama \(2018\)seqeval: a python framework for sequence labeling evaluation\.Note:Software available from https://github\.com/chakki\-works/seqevalExternal Links:[Link](https://github.com/chakki-works/seqeval)Cited by:[§4](https://arxiv.org/html/2606.24387#S4.p2.1)\.
- C\. Park, S\. Jeong, and J\. Kim \(2023\)ADMit: improving NER in automotive domain with domain adversarial training and multi\-task learning\.225\(C\)\.External Links:ISSN 0957\-4174,[Link](https://doi.org/10.1016/j.eswa.2023.120007),[Document](https://dx.doi.org/10.1016/j.eswa.2023.120007)Cited by:[§2\.2](https://arxiv.org/html/2606.24387#S2.SS2.p2.1)\.
- D\. Putthividhya and J\. Hu \(2011\)Bootstrapped named entity recognition for product attribute extraction\.InProceedings of the 2011 Conference on Empirical Methods in Natural Language Processing,pp\. 1557–1567\.External Links:[Link](https://aclanthology.org/D11-1144/)Cited by:[§2\.1](https://arxiv.org/html/2606.24387#S2.SS1.p2.1)\.
- R\. Richie, S\. Grover, and F\. \(\. Tsui \(2022\)Inter\-annotator agreement is not the ceiling of machine learning performance: evidence from a comprehensive set of simulations\.InProceedings of the 21st Workshop on Biomedical Language Processing,D\. Demner\-Fushman, K\. B\. Cohen, S\. Ananiadou, and J\. Tsujii \(Eds\.\),Dublin, Ireland,pp\. 275–284\.External Links:[Link](https://aclanthology.org/2022.bionlp-1.26/),[Document](https://dx.doi.org/10.18653/v1/2022.bionlp-1.26)Cited by:[footnote 6](https://arxiv.org/html/2606.24387#footnote6)\.
- A\. Singh, A\. Fry, A\. Perelman, A\. Tart, A\. Ganesh, A\. El\-Kishky, A\. McLaughlin, A\. Low, A\. Ostrow, A\. Ananthram, A\. Nathan, A\. Luo, A\. Helyar, A\. Madry, A\. Efremov, A\. Spyra, A\. Baker\-Whitcomb, A\. Beutel, A\. Karpenko, A\. Makelov, A\. Neitz, A\. Wei, A\. Barr, A\. Kirchmeyer, A\. Ivanov, A\. Christakis, A\. Gillespie, A\. Tam, A\. Bennett, A\. Wan, A\. Huang, A\. M\. Sandjideh, A\. Yang, A\. Kumar, A\. Saraiva, A\. Vallone, A\. Gheorghe, A\. G\. Garcia, A\. Braunstein, A\. Liu, A\. Schmidt, A\. Mereskin, A\. Mishchenko, A\. Applebaum, A\. Rogerson, A\. Rajan, A\. Wei, A\. Kotha, A\. Srivastava, A\. Agrawal, A\. Vijayvergiya, A\. Tyra, A\. Nair, A\. Nayak, B\. Eggers, B\. Ji, B\. Hoover, B\. Chen, B\. Chen, B\. Barak, B\. Minaiev, B\. Hao, B\. Baker, B\. Lightcap, B\. McKinzie, B\. Wang, B\. Quinn, B\. Fioca, B\. Hsu, B\. Yang, B\. Yu, B\. Zhang, B\. Brenner, C\. R\. Zetino, C\. Raymond, C\. Lugaresi, C\. Paz, C\. Hudson, C\. Whitney, C\. Li, C\. Chen, C\. Cole, C\. Voss, C\. Ding, C\. Shen, C\. Huang, C\. Colby, C\. Hallacy, C\. Koch, C\. Lu, C\. Kaplan, C\. Kim, C\. Minott\-Henriques, C\. Frey, C\. Yu, C\. Czarnecki, C\. Reid, C\. Wei, C\. Decareaux, C\. Scheau, C\. Zhang, C\. Forbes, D\. Tang, D\. Goldberg, D\. Roberts, D\. Palmie, D\. Kappler, D\. Levine, D\. Wright, D\. Leo, D\. Lin, D\. Robinson, D\. Grabb, D\. Chen, D\. Lim, D\. Salama, D\. Bhattacharjee, D\. Tsipras, D\. Li, D\. Yu, D\. Strouse, D\. Williams, D\. Hunn, E\. Bayes, E\. Arbus, E\. Akyurek, E\. Y\. Le, E\. Widmann, E\. Yani, E\. Proehl, E\. Sert, E\. Cheung, E\. Schwartz, E\. Han, E\. Jiang, E\. Mitchell, E\. Sigler, E\. Wallace, E\. Ritter, E\. Kavanaugh, E\. Mays, E\. Nikishin, F\. Li, F\. P\. Such, F\. de Avila Belbute Peres, F\. Raso, F\. Bekerman, F\. Tsimpourlas, F\. Chantzis, F\. Song, F\. Zhang, G\. Raila, G\. McGrath, G\. Briggs, G\. Yang, G\. Parascandolo, G\. Chabot, G\. Kim, G\. Zhao, G\. Valiant, G\. Leclerc, H\. Salman, H\. Wang, H\. Sheng, H\. Jiang, H\. Wang, H\. Jin, H\. Sikchi, H\. Schmidt, H\. Aspegren, H\. Chen, H\. Qiu, H\. Lightman, I\. Covert, I\. Kivlichan, I\. Silber, I\. Sohl, I\. Hammoud, I\. Clavera, I\. Lan, I\. Akkaya, I\. Kostrikov, I\. Kofman, I\. Etinger, I\. Singal, J\. Hehir, J\. Huh, J\. Pan, J\. Wilczynski, J\. Pachocki, J\. Lee, J\. Quinn, J\. Kiros, J\. Kalra, J\. Samaroo, J\. Wang, J\. Wolfe, J\. Chen, J\. Wang, J\. Harb, J\. Han, J\. Wang, J\. Zhao, J\. Chen, J\. Yang, J\. Tworek, J\. Chand, J\. Landon, J\. Liang, J\. Lin, J\. Liu, J\. Wang, J\. Tang, J\. Yin, J\. Jang, J\. Morris, J\. Flynn, J\. Ferstad, J\. Heidecke, J\. Fishbein, J\. Hallman, J\. Grant, J\. Chien, J\. Gordon, J\. Park, J\. Liss, J\. Kraaijeveld, J\. Guay, J\. Mo, J\. Lawson, J\. McGrath, J\. Vendrow, J\. Jiao, J\. Lee, J\. Steele, J\. Wang, J\. Mao, K\. Chen, K\. Hayashi, K\. Xiao, K\. Salahi, K\. Wu, K\. Sekhri, K\. Sharma, K\. Singhal, K\. Li, K\. Nguyen, K\. Gu\-Lemberg, K\. King, K\. Liu, K\. Stone, K\. Yu, K\. Ying, K\. Georgiev, K\. Lim, K\. Tirumala, K\. Miller, L\. Ahmad, L\. Lv, L\. Clare, L\. Fauconnet, L\. Itow, L\. Yang, L\. Romaniuk, L\. Anise, L\. Byron, L\. Pathak, L\. Maksin, L\. Lo, L\. Ho, L\. Jing, L\. Wu, L\. Xiong, L\. Mamitsuka, L\. Yang, L\. McCallum, L\. Held, L\. Bourgeois, L\. Engstrom, L\. Kuhn, L\. Feuvrier, L\. Zhang, L\. Switzer, L\. Kondraciuk, L\. Kaiser, M\. Joglekar, M\. Singh, M\. Shah, M\. Stratta, M\. Williams, M\. Chen, M\. Sun, M\. Cayton, M\. Li, M\. Zhang, M\. Aljubeh, M\. Nichols, M\. Haines, M\. Schwarzer, M\. Gupta, M\. Shah, M\. Y\. Guan, M\. Huang, M\. Dong, M\. Wang, M\. Glaese, M\. Carroll, M\. Lampe, M\. Malek, M\. Sharman, M\. Zhang, M\. Wang, M\. Pokrass, M\. Florian, M\. Pavlov, M\. Wang, M\. Chen, M\. Wang, M\. Feng, M\. Bavarian, M\. Lin, M\. Abdool, M\. Rohaninejad, N\. Soto, N\. Staudacher, N\. LaFontaine, N\. Marwell, N\. Liu, N\. Preston, N\. Turley, N\. Ansman, N\. Blades, N\. Pancha, N\. Mikhaylin, N\. Felix, N\. Handa, N\. Rai, N\. Keskar, N\. Brown, O\. Nachum, O\. Boiko, O\. Murk, O\. Watkins, O\. Gleeson, P\. Mishkin, P\. Lesiewicz, P\. Baltescu, P\. Belov, P\. Zhokhov, P\. Pronin, P\. Guo, P\. Thacker, Q\. Liu, Q\. Yuan, Q\. Liu, R\. Dias, R\. Puckett, R\. Arora, R\. T\. Mullapudi, R\. Gaon, R\. Miyara, R\. Song, R\. Aggarwal, R\. Marsan, R\. Yemiru, R\. Xiong, R\. Kshirsagar, R\. Nuttall, R\. Tsiupa, R\. Eldan, R\. Wang, R\. James, R\. Ziv, R\. Shu, R\. Nigmatullin, S\. Jain, S\. Talaie, S\. Altman, S\. Arnesen, S\. Toizer, S\. Toyer, S\. Miserendino, S\. Agarwal, S\. Yoo, S\. Heon, S\. Ethersmith, S\. Grove, S\. Taylor, S\. Bubeck, S\. Banesiu, S\. Amdo, S\. Zhao, S\. Wu, S\. Santurkar, S\. Zhao, S\. R\. Chaudhuri, S\. Krishnaswamy, Shuaiqi, Xia, S\. Cheng, S\. Anadkat, S\. P\. Fishman, S\. Tobin, S\. Fu, S\. Jain, S\. Mei, S\. Egoian, S\. Kim, S\. Golden, S\. Mah, S\. Lin, S\. Imm, S\. Sharpe, S\. Yadlowsky, S\. Choudhry, S\. Eum, S\. Sanjeev, T\. Khan, T\. Stramer, T\. Wang, T\. Xin, T\. Gogineni, T\. Christianson, T\. Sanders, T\. Patwardhan, T\. Degry, T\. Shadwell, T\. Fu, T\. Gao, T\. Garipov, T\. Sriskandarajah, T\. Sherbakov, T\. Korbak, T\. Kaftan, T\. Hiratsuka, T\. Wang, T\. Song, T\. Zhao, T\. Peterson, V\. Kharitonov, V\. Chernova, V\. Kosaraju, V\. Kuo, V\. Pong, V\. Verma, V\. Petrov, W\. Jiang, W\. Zhang, W\. Zhou, W\. Xie, W\. Zhan, W\. McCabe, W\. DePue, W\. Ellsworth, W\. Bain, W\. Thompson, X\. Chen, X\. Qi, X\. Xiang, X\. Shi, Y\. Dubois, Y\. Yu, Y\. Khakbaz, Y\. Wu, Y\. Qian, Y\. T\. Lee, Y\. Chen, Y\. Zhang, Y\. Xiong, Y\. Tian, Y\. Cha, Y\. Bai, Y\. Yang, Y\. Yuan, Y\. Li, Y\. Zhang, Y\. Yang, Y\. Jin, Y\. Jiang, Y\. Wang, Y\. Wang, Y\. Liu, Z\. Stubenvoll, Z\. Dou, Z\. Wu, and Z\. Wang \(2026\)OpenAI gpt\-5 system card\.External Links:2601\.03267,[Link](https://arxiv.org/abs/2601.03267)Cited by:[§4\.1\.3](https://arxiv.org/html/2606.24387#S4.SS1.SSS3.p1.1)\.
- G\. Team, R\. Anil, S\. Borgeaud, J\. Alayrac, J\. Yu, R\. Soricut, J\. Schalkwyk, A\. M\. Dai, A\. Hauth, K\. Millican, D\. Silver, M\. Johnson, I\. Antonoglou, J\. Schrittwieser, A\. Glaese, J\. Chen, E\. Pitler, T\. Lillicrap, A\. Lazaridou, O\. Firat, J\. Molloy, M\. Isard, P\. R\. Barham, T\. Hennigan, B\. Lee, F\. Viola, M\. Reynolds, Y\. Xu, R\. Doherty, E\. Collins, C\. Meyer, E\. Rutherford, E\. Moreira, K\. Ayoub, M\. Goel, J\. Krawczyk, C\. Du, E\. Chi, H\. Cheng, E\. Ni, P\. Shah, P\. Kane, B\. Chan, M\. Faruqui, A\. Severyn, H\. Lin, Y\. Li, Y\. Cheng, A\. Ittycheriah, M\. Mahdieh, M\. Chen, P\. Sun, D\. Tran, S\. Bagri, B\. Lakshminarayanan, J\. Liu, A\. Orban, F\. Güra, H\. Zhou, X\. Song, A\. Boffy, H\. Ganapathy, S\. Zheng, H\. Choe, Á\. Weisz, T\. Zhu, Y\. Lu, S\. Gopal, J\. Kahn, M\. Kula, J\. Pitman, R\. Shah, E\. Taropa, M\. A\. Merey, M\. Baeuml, Z\. Chen, L\. E\. Shafey, Y\. Zhang, O\. Sercinoglu, G\. Tucker, E\. Piqueras, M\. Krikun, I\. Barr, N\. Savinov, I\. Danihelka, B\. Roelofs, A\. White, A\. Andreassen, T\. von Glehn, L\. Yagati, M\. Kazemi, L\. Gonzalez, M\. Khalman, J\. Sygnowski, A\. Frechette, C\. Smith, L\. Culp, L\. Proleev, Y\. Luan, X\. Chen, J\. Lottes, N\. Schucher, F\. Lebron, A\. Rrustemi, N\. Clay, P\. Crone, T\. Kocisky, J\. Zhao, B\. Perz, D\. Yu, H\. Howard, A\. Bloniarz, J\. W\. Rae, H\. Lu, L\. Sifre, M\. Maggioni, F\. Alcober, D\. Garrette, M\. Barnes, S\. Thakoor, J\. Austin, G\. Barth\-Maron, W\. Wong, R\. Joshi, R\. Chaabouni, D\. Fatiha, A\. Ahuja, G\. S\. Tomar, E\. Senter, M\. Chadwick, I\. Kornakov, N\. Attaluri, I\. Iturrate, R\. Liu, Y\. Li, S\. Cogan, J\. Chen, C\. Jia, C\. Gu, Q\. Zhang, J\. Grimstad, A\. J\. Hartman, X\. Garcia, T\. S\. Pillai, J\. Devlin, M\. Laskin, D\. de Las Casas, D\. Valter, C\. Tao, L\. Blanco, A\. P\. Badia, D\. Reitter, M\. Chen, J\. Brennan, C\. Rivera, S\. Brin, S\. Iqbal, G\. Surita, J\. Labanowski, A\. Rao, S\. Winkler, E\. Parisotto, Y\. Gu, K\. Olszewska, R\. Addanki, A\. Miech, A\. Louis, D\. Teplyashin, G\. Brown, E\. Catt, J\. Balaguer, J\. Xiang, P\. Wang, Z\. Ashwood, A\. Briukhov, A\. Webson, S\. Ganapathy, S\. Sanghavi, A\. Kannan, M\. Chang, A\. Stjerngren, J\. Djolonga, Y\. Sun, A\. Bapna, M\. Aitchison, P\. Pejman, H\. Michalewski, T\. Yu, C\. Wang, J\. Love, J\. Ahn, D\. Bloxwich, K\. Han, P\. Humphreys, T\. Sellam, J\. Bradbury, V\. Godbole, S\. Samangooei, B\. Damoc, A\. Kaskasoli, S\. M\. R\. Arnold, V\. Vasudevan, S\. Agrawal, J\. Riesa, D\. Lepikhin, R\. Tanburn, S\. Srinivasan, H\. Lim, S\. Hodkinson, P\. Shyam, J\. Ferret, S\. Hand, A\. Garg, T\. L\. Paine, J\. Li, Y\. Li, M\. Giang, A\. Neitz, Z\. Abbas, S\. York, M\. Reid, E\. Cole, A\. Chowdhery, D\. Das, D\. Rogozińska, V\. Nikolaev, P\. Sprechmann, Z\. Nado, L\. Zilka, F\. Prost, L\. He, M\. Monteiro, G\. Mishra, C\. Welty, J\. Newlan, D\. Jia, M\. Allamanis, C\. H\. Hu, R\. de Liedekerke, J\. Gilmer, C\. Saroufim, S\. Rijhwani, S\. Hou, D\. Shrivastava, A\. Baddepudi, A\. Goldin, A\. Ozturel, A\. Cassirer, Y\. Xu, D\. Sohn, D\. Sachan, R\. K\. Amplayo, C\. Swanson, D\. Petrova, S\. Narayan, A\. Guez, S\. Brahma, J\. Landon, M\. Patel, R\. Zhao, K\. Villela, L\. Wang, W\. Jia, M\. Rahtz, M\. Giménez, L\. Yeung, J\. Keeling, P\. Georgiev, D\. Mincu, B\. Wu, S\. Haykal, R\. Saputro, K\. Vodrahalli, J\. Qin, Z\. Cankara, A\. Sharma, N\. Fernando, W\. Hawkins, B\. Neyshabur, S\. Kim, A\. Hutter, P\. Agrawal, A\. Castro\-Ros, G\. van den Driessche, T\. Wang, F\. Yang, S\. Chang, P\. Komarek, R\. McIlroy, M\. Lučić, G\. Zhang, W\. Farhan, M\. Sharman, P\. Natsev, P\. Michel, Y\. Bansal, S\. Qiao, K\. Cao, S\. Shakeri, C\. Butterfield, J\. Chung, P\. K\. Rubenstein, S\. Agrawal, A\. Mensch, K\. Soparkar, K\. Lenc, T\. Chung, A\. Pope, L\. Maggiore, J\. Kay, P\. Jhakra, S\. Wang, J\. Maynez, M\. Phuong, T\. Tobin, A\. Tacchetti, M\. Trebacz, K\. Robinson, Y\. Katariya, S\. Riedel, P\. Bailey, K\. Xiao, N\. Ghelani, L\. Aroyo, A\. Slone, N\. Houlsby, X\. Xiong, Z\. Yang, E\. Gribovskaya, J\. Adler, M\. Wirth, L\. Lee, M\. Li, T\. Kagohara, J\. Pavagadhi, S\. Bridgers, A\. Bortsova, S\. Ghemawat, Z\. Ahmed, T\. Liu, R\. Powell, V\. Bolina, M\. Iinuma, P\. Zablotskaia, J\. Besley, D\. Chung, T\. Dozat, R\. Comanescu, X\. Si, J\. Greer, G\. Su, M\. Polacek, R\. L\. Kaufman, S\. Tokumine, H\. Hu, E\. Buchatskaya, Y\. Miao, M\. Elhawaty, A\. Siddhant, N\. Tomasev, J\. Xing, C\. Greer, H\. Miller, S\. Ashraf, A\. Roy, Z\. Zhang, A\. Ma, A\. Filos, M\. Besta, R\. Blevins, T\. Klimenko, C\. Yeh, S\. Changpinyo, J\. Mu, O\. Chang, M\. Pajarskas, C\. Muir, V\. Cohen, C\. L\. Lan, K\. Haridasan, A\. Marathe, S\. Hansen, S\. Douglas, R\. Samuel, M\. Wang, S\. Austin, C\. Lan, J\. Jiang, J\. Chiu, J\. A\. Lorenzo, L\. L\. Sjösund, S\. Cevey, Z\. Gleicher, T\. Avrahami, A\. Boral, H\. Srinivasan, V\. Selo, R\. May, K\. Aisopos, L\. Hussenot, L\. B\. Soares, K\. Baumli, M\. B\. Chang, A\. Recasens, B\. Caine, A\. Pritzel, F\. Pavetic, F\. Pardo, A\. Gergely, J\. Frye, V\. Ramasesh, D\. Horgan, K\. Badola, N\. Kassner, S\. Roy, E\. Dyer, V\. C\. Campos, A\. Tomala, Y\. Tang, D\. E\. Badawy, E\. White, B\. Mustafa, O\. Lang, A\. Jindal, S\. Vikram, Z\. Gong, S\. Caelles, R\. Hemsley, G\. Thornton, F\. Feng, W\. Stokowiec, C\. Zheng, P\. Thacker, Ç\. Ünlü, Z\. Zhang, M\. Saleh, J\. Svensson, M\. Bileschi, P\. Patil, A\. Anand, R\. Ring, K\. Tsihlas, A\. Vezer, M\. Selvi, T\. Shevlane, M\. Rodriguez, T\. Kwiatkowski, S\. Daruki, K\. Rong, A\. Dafoe, N\. FitzGerald, K\. Gu\-Lemberg, M\. Khan, L\. A\. Hendricks, M\. Pellat, V\. Feinberg, J\. Cobon\-Kerr, T\. Sainath, M\. Rauh, S\. H\. Hashemi, R\. Ives, Y\. Hasson, E\. Noland, Y\. Cao, N\. Byrd, L\. Hou, Q\. Wang, T\. Sottiaux, M\. Paganini, J\. Lespiau, A\. Moufarek, S\. Hassan, K\. Shivakumar, J\. van Amersfoort, A\. Mandhane, P\. Joshi, A\. Goyal, M\. Tung, A\. Brock, H\. Sheahan, V\. Misra, C\. Li, N\. Rakićević, M\. Dehghani, F\. Liu, S\. Mittal, J\. Oh, S\. Noury, E\. Sezener, F\. Huot, M\. Lamm, N\. D\. Cao, C\. Chen, S\. Mudgal, R\. Stella, K\. Brooks, G\. Vasudevan, C\. Liu, M\. Chain, N\. Melinkeri, A\. Cohen, V\. Wang, K\. Seymore, S\. Zubkov, R\. Goel, S\. Yue, S\. Krishnakumaran, B\. Albert, N\. Hurley, M\. Sano, A\. Mohananey, J\. Joughin, E\. Filonov, T\. Kępa, Y\. Eldawy, J\. Lim, R\. Rishi, S\. Badiezadegan, T\. Bos, J\. Chang, S\. Jain, S\. G\. S\. Padmanabhan, S\. Puttagunta, K\. Krishna, L\. Baker, N\. Kalb, V\. Bedapudi, A\. Kurzrok, S\. Lei, A\. Yu, O\. Litvin, X\. Zhou, Z\. Wu, S\. Sobell, A\. Siciliano, A\. Papir, R\. Neale, J\. Bragagnolo, T\. Toor, T\. Chen, V\. Anklin, F\. Wang, R\. Feng, M\. Gholami, K\. Ling, L\. Liu, J\. Walter, H\. Moghaddam, A\. Kishore, J\. Adamek, T\. Mercado, J\. Mallinson, S\. Wandekar, S\. Cagle, E\. Ofek, G\. Garrido, C\. Lombriser, M\. Mukha, B\. Sun, H\. R\. Mohammad, J\. Matak, Y\. Qian, V\. Peswani, P\. Janus, Q\. Yuan, L\. Schelin, O\. David, A\. Garg, Y\. He, O\. Duzhyi, A\. Älgmyr, T\. Lottaz, Q\. Li, V\. Yadav, L\. Xu, A\. Chinien, R\. Shivanna, A\. Chuklin, J\. Li, C\. Spadine, T\. Wolfe, K\. Mohamed, S\. Das, Z\. Dai, K\. He, D\. von Dincklage, S\. Upadhyay, A\. Maurya, L\. Chi, S\. Krause, K\. Salama, P\. G\. Rabinovitch, P\. K\. R\. M, A\. Selvan, M\. Dektiarev, G\. Ghiasi, E\. Guven, H\. Gupta, B\. Liu, D\. Sharma, I\. H\. Shtacher, S\. Paul, O\. Akerlund, F\. Aubet, T\. Huang, C\. Zhu, E\. Zhu, E\. Teixeira, M\. Fritze, F\. Bertolini, L\. Marinescu, M\. Bölle, D\. Paulus, K\. Gupta, T\. Latkar, M\. Chang, J\. Sanders, R\. Wilson, X\. Wu, Y\. Tan, L\. N\. Thiet, T\. Doshi, S\. Lall, S\. Mishra, W\. Chen, T\. Luong, S\. Benjamin, J\. Lee, E\. Andrejczuk, D\. Rabiej, V\. Ranjan, K\. Styrc, P\. Yin, J\. Simon, M\. R\. Harriott, M\. Bansal, A\. Robsky, G\. Bacon, D\. Greene, D\. Mirylenka, C\. Zhou, O\. Sarvana, A\. Goyal, S\. Andermatt, P\. Siegler, B\. Horn, A\. Israel, F\. Pongetti, C\. "\. Chen, M\. Selvatici, P\. Silva, K\. Wang, J\. Tolins, K\. Guu, R\. Yogev, X\. Cai, A\. Agostini, M\. Shah, H\. Nguyen, N\. Ó\. Donnaile, S\. Pereira, L\. Friso, A\. Stambler, A\. Kurzrok, C\. Kuang, Y\. Romanikhin, M\. Geller, Z\. Yan, K\. Jang, C\. Lee, W\. Fica, E\. Malmi, Q\. Tan, D\. Banica, D\. Balle, R\. Pham, Y\. Huang, D\. Avram, H\. Shi, J\. Singh, C\. Hidey, N\. Ahuja, P\. Saxena, D\. Dooley, S\. P\. Potharaju, E\. O’Neill, A\. Gokulchandran, R\. Foley, K\. Zhao, M\. Dusenberry, Y\. Liu, P\. Mehta, R\. Kotikalapudi, C\. Safranek\-Shrader, A\. Goodman, J\. Kessinger, E\. Globen, P\. Kolhar, C\. Gorgolewski, A\. Ibrahim, Y\. Song, A\. Eichenbaum, T\. Brovelli, S\. Potluri, P\. Lahoti, C\. Baetu, A\. Ghorbani, C\. Chen, A\. Crawford, S\. Pal, M\. Sridhar, P\. Gurita, A\. Mujika, I\. Petrovski, P\. Cedoz, C\. Li, S\. Chen, N\. D\. Santo, S\. Goyal, J\. Punjabi, K\. Kappaganthu, C\. Kwak, P\. LV, S\. Velury, H\. Choudhury, J\. Hall, P\. Shah, R\. Figueira, M\. Thomas, M\. Lu, T\. Zhou, C\. Kumar, T\. Jurdi, S\. Chikkerur, Y\. Ma, A\. Yu, S\. Kwak, V\. Ähdel, S\. Rajayogam, T\. Choma, F\. Liu, A\. Barua, C\. Ji, J\. H\. Park, V\. Hellendoorn, A\. Bailey, T\. Bilal, H\. Zhou, M\. Khatir, C\. Sutton, W\. Rzadkowski, F\. Macintosh, R\. Vij, K\. Shagin, P\. Medina, C\. Liang, J\. Zhou, P\. Shah, Y\. Bi, A\. Dankovics, S\. Banga, S\. Lehmann, M\. Bredesen, Z\. Lin, J\. E\. Hoffmann, J\. Lai, R\. Chung, K\. Yang, N\. Balani, A\. Bražinskas, A\. Sozanschi, M\. Hayes, H\. F\. Alcalde, P\. Makarov, W\. Chen, A\. Stella, L\. Snijders, M\. Mandl, A\. Kärrman, P\. Nowak, X\. Wu, A\. Dyck, K\. Vaidyanathan, R\. R, J\. Mallet, M\. Rudominer, E\. Johnston, S\. Mittal, A\. Udathu, J\. Christensen, V\. Verma, Z\. Irving, A\. Santucci, G\. Elsayed, E\. Davoodi, M\. Georgiev, I\. Tenney, N\. Hua, G\. Cideron, E\. Leurent, M\. Alnahlawi, I\. Georgescu, N\. Wei, I\. Zheng, D\. Scandinaro, H\. Jiang, J\. Snoek, M\. Sundararajan, X\. Wang, Z\. Ontiveros, I\. Karo, J\. Cole, V\. Rajashekhar, L\. Tumeh, E\. Ben\-David, R\. Jain, J\. Uesato, R\. Datta, O\. Bunyan, S\. Wu, J\. Zhang, P\. Stanczyk, Y\. Zhang, D\. Steiner, S\. Naskar, M\. Azzam, M\. Johnson, A\. Paszke, C\. Chiu, J\. S\. Elias, A\. Mohiuddin, F\. Muhammad, J\. Miao, A\. Lee, N\. Vieillard, J\. Park, J\. Zhang, J\. Stanway, D\. Garmon, A\. Karmarkar, Z\. Dong, J\. Lee, A\. Kumar, L\. Zhou, J\. Evens, W\. Isaac, G\. Irving, E\. Loper, M\. Fink, I\. Arkatkar, N\. Chen, I\. Shafran, I\. Petrychenko, Z\. Chen, J\. Jia, A\. Levskaya, Z\. Zhu, P\. Grabowski, Y\. Mao, A\. Magni, K\. Yao, J\. Snaider, N\. Casagrande, E\. Palmer, P\. Suganthan, A\. Castaño, I\. Giannoumis, W\. Kim, M\. Rybiński, A\. Sreevatsa, J\. Prendki, D\. Soergel, A\. Goedeckemeyer, W\. Gierke, M\. Jafari, M\. Gaba, J\. Wiesner, D\. G\. Wright, Y\. Wei, H\. Vashisht, Y\. Kulizhskaya, J\. Hoover, M\. Le, L\. Li, C\. Iwuanyanwu, L\. Liu, K\. Ramirez, A\. Khorlin, A\. Cui, T\. LIN, M\. Wu, R\. Aguilar, K\. Pallo, A\. Chakladar, G\. Perng, E\. A\. Abellan, M\. Zhang, I\. Dasgupta, N\. Kushman, I\. Penchev, A\. Repina, X\. Wu, T\. van der Weide, P\. Ponnapalli, C\. Kaplan, J\. Simsa, S\. Li, O\. Dousse, F\. Yang, J\. Piper, N\. Ie, R\. Pasumarthi, N\. Lintz, A\. Vijayakumar, D\. Andor, P\. Valenzuela, M\. Lui, C\. Paduraru, D\. Peng, K\. Lee, S\. Zhang, S\. Greene, D\. D\. Nguyen, P\. Kurylowicz, C\. Hardin, L\. Dixon, L\. Janzer, K\. Choo, Z\. Feng, B\. Zhang, A\. Singhal, D\. Du, D\. McKinnon, N\. Antropova, T\. Bolukbasi, O\. Keller, D\. Reid, D\. Finchelstein, M\. A\. Raad, R\. Crocker, P\. Hawkins, R\. Dadashi, C\. Gaffney, K\. Franko, A\. Bulanova, R\. Leblond, S\. Chung, H\. Askham, L\. C\. Cobo, K\. Xu, F\. Fischer, J\. Xu, C\. Sorokin, C\. Alberti, C\. Lin, C\. Evans, A\. Dimitriev, H\. Forbes, D\. Banarse, Z\. Tung, M\. Omernick, C\. Bishop, R\. Sterneck, R\. Jain, J\. Xia, E\. Amid, F\. Piccinno, X\. Wang, P\. Banzal, D\. J\. Mankowitz, A\. Polozov, V\. Krakovna, S\. Brown, M\. Bateni, D\. Duan, V\. Firoiu, M\. Thotakuri, T\. Natan, M\. Geist, S\. tan Girgin, H\. Li, J\. Ye, O\. Roval, R\. Tojo, M\. Kwong, J\. Lee\-Thorp, C\. Yew, D\. Sinopalnikov, S\. Ramos, J\. Mellor, A\. Sharma, K\. Wu, D\. Miller, N\. Sonnerat, D\. Vnukov, R\. Greig, J\. Beattie, E\. Caveness, L\. Bai, J\. Eisenschlos, A\. Korchemniy, T\. Tsai, M\. Jasarevic, W\. Kong, P\. Dao, Z\. Zheng, F\. Liu, F\. Yang, R\. Zhu, T\. H\. Teh, J\. Sanmiya, E\. Gladchenko, N\. Trdin, D\. Toyama, E\. Rosen, S\. Tavakkol, L\. Xue, C\. Elkind, O\. Woodman, J\. Carpenter, G\. Papamakarios, R\. Kemp, S\. Kafle, T\. Grunina, R\. Sinha, A\. Talbert, D\. Wu, D\. Owusu\-Afriyie, C\. Du, C\. Thornton, J\. Pont\-Tuset, P\. Narayana, J\. Li, S\. Fatehi, J\. Wieting, O\. Ajmeri, B\. Uria, Y\. Ko, L\. Knight, A\. Héliou, N\. Niu, S\. Gu, C\. Pang, Y\. Li, N\. Levine, A\. Stolovich, R\. Santamaria\-Fernandez, S\. Goenka, W\. Yustalim, R\. Strudel, A\. Elqursh, C\. Deck, H\. Lee, Z\. Li, K\. Levin, R\. Hoffmann, D\. Holtmann\-Rice, O\. Bachem, S\. Arora, C\. Koh, S\. H\. Yeganeh, S\. Põder, M\. Tariq, Y\. Sun, L\. Ionita, M\. Seyedhosseini, P\. Tafti, Z\. Liu, A\. Gulati, J\. Liu, X\. Ye, B\. Chrzaszcz, L\. Wang, N\. Sethi, T\. Li, B\. Brown, S\. Singh, W\. Fan, A\. Parisi, J\. Stanton, V\. Koverkathu, C\. A\. Choquette\-Choo, Y\. Li, T\. Lu, A\. Ittycheriah, P\. Shroff, M\. Varadarajan, S\. Bahargam, R\. Willoughby, D\. Gaddy, G\. Desjardins, M\. Cornero, B\. Robenek, B\. Mittal, B\. Albrecht, A\. Shenoy, F\. Moiseev, H\. Jacobsson, A\. Ghaffarkhah, M\. Rivière, A\. Walton, C\. Crepy, A\. Parrish, Z\. Zhou, C\. Farabet, C\. Radebaugh, P\. Srinivasan, C\. van der Salm, A\. Fidjeland, S\. Scellato, E\. Latorre\-Chimoto, H\. Klimczak\-Plucińska, D\. Bridson, D\. de Cesare, T\. Hudson, P\. Mendolicchio, L\. Walker, A\. Morris, M\. Mauger, A\. Guseynov, A\. Reid, S\. Odoom, L\. Loher, V\. Cotruta, M\. Yenugula, D\. Grewe, A\. Petrushkina, T\. Duerig, A\. Sanchez, S\. Yadlowsky, A\. Shen, A\. Globerson, L\. Webb, S\. Dua, D\. Li, S\. Bhupatiraju, D\. Hurt, H\. Qureshi, A\. Agarwal, T\. Shani, M\. Eyal, A\. Khare, S\. R\. Belle, L\. Wang, C\. Tekur, M\. S\. Kale, J\. Wei, R\. Sang, B\. Saeta, T\. Liechty, Y\. Sun, Y\. Zhao, S\. Lee, P\. Nayak, D\. Fritz, M\. R\. Vuyyuru, J\. Aslanides, N\. Vyas, M\. Wicke, X\. Ma, E\. Eltyshev, N\. Martin, H\. Cate, J\. Manyika, K\. Amiri, Y\. Kim, X\. Xiong, K\. Kang, F\. Luisier, N\. Tripuraneni, D\. Madras, M\. Guo, A\. Waters, O\. Wang, J\. Ainslie, J\. Baldridge, H\. Zhang, G\. Pruthi, J\. Bauer, F\. Yang, R\. Mansour, J\. Gelman, Y\. Xu, G\. Polovets, J\. Liu, H\. Cai, W\. Chen, X\. Sheng, E\. Xue, S\. Ozair, C\. Angermueller, X\. Li, A\. Sinha, W\. Wang, J\. Wiesinger, E\. Koukoumidis, Y\. Tian, A\. Iyer, M\. Gurumurthy, M\. Goldenson, P\. Shah, M\. Blake, H\. Yu, A\. Urbanowicz, J\. Palomaki, C\. Fernando, K\. Durden, H\. Mehta, N\. Momchev, E\. Rahimtoroghi, M\. Georgaki, A\. Raul, S\. Ruder, M\. Redshaw, J\. Lee, D\. Zhou, K\. Jalan, D\. Li, B\. Hechtman, P\. Schuh, M\. Nasr, K\. Milan, V\. Mikulik, J\. Franco, T\. Green, N\. Nguyen, J\. Kelley, A\. Mahendru, A\. Hu, J\. Howland, B\. Vargas, J\. Hui, K\. Bansal, V\. Rao, R\. Ghiya, E\. Wang, K\. Ye, J\. M\. Sarr, M\. M\. Preston, M\. Elish, S\. Li, A\. Kaku, J\. Gupta, I\. Pasupat, D\. Juan, M\. Someswar, T\. M\., X\. Chen, A\. Amini, A\. Fabrikant, E\. Chu, X\. Dong, A\. Muthal, S\. Buthpitiya, S\. Jauhari, N\. Hua, U\. Khandelwal, A\. Hitron, J\. Ren, L\. Rinaldi, S\. Drath, A\. Dabush, N\. Jiang, H\. Godhia, U\. Sachs, A\. Chen, Y\. Fan, H\. Taitelbaum, H\. Noga, Z\. Dai, J\. Wang, C\. Liang, J\. Hamer, C\. Ferng, C\. Elkind, A\. Atias, P\. Lee, V\. Listík, M\. Carlen, J\. van de Kerkhof, M\. Pikus, K\. Zaher, P\. Müller, S\. Zykova, R\. Stefanec, V\. Gatsko, C\. Hirnschall, A\. Sethi, X\. F\. Xu, C\. Ahuja, B\. Tsai, A\. Stefanoiu, B\. Feng, K\. Dhandhania, M\. Katyal, A\. Gupta, A\. Parulekar, D\. Pitta, J\. Zhao, V\. Bhatia, Y\. Bhavnani, O\. Alhadlaq, X\. Li, P\. Danenberg, D\. Tu, A\. Pine, V\. Filippova, A\. Ghosh, B\. Limonchik, B\. Urala, C\. K\. Lanka, D\. Clive, Y\. Sun, E\. Li, H\. Wu, K\. Hongtongsak, I\. Li, K\. Thakkar, K\. Omarov, K\. Majmundar, M\. Alverson, M\. Kucharski, M\. Patel, M\. Jain, M\. Zabelin, P\. Pelagatti, R\. Kohli, S\. Kumar, J\. Kim, S\. Sankar, V\. Shah, L\. Ramachandruni, X\. Zeng, B\. Bariach, L\. Weidinger, T\. Vu, A\. Andreev, A\. He, K\. Hui, S\. Kashem, A\. Subramanya, S\. Hsiao, D\. Hassabis, K\. Kavukcuoglu, A\. Sadovsky, Q\. Le, T\. Strohman, Y\. Wu, S\. Petrov, J\. Dean, and O\. Vinyals \(2025a\)Gemini: a family of highly capable multimodal models\.External Links:2312\.11805,[Link](https://arxiv.org/abs/2312.11805)Cited by:[§4\.1\.3](https://arxiv.org/html/2606.24387#S4.SS1.SSS3.p1.1)\.
- G\. Team, A\. Kamath, J\. Ferret, S\. Pathak, N\. Vieillard, R\. Merhej, S\. Perrin, T\. Matejovicova, A\. Ramé, M\. Rivière, L\. Rouillard, T\. Mesnard, G\. Cideron, J\. Grill, S\. Ramos, E\. Yvinec, M\. Casbon, E\. Pot, I\. Penchev, G\. Liu, F\. Visin, K\. Kenealy, L\. Beyer, X\. Zhai, A\. Tsitsulin, R\. Busa\-Fekete, A\. Feng, N\. Sachdeva, B\. Coleman, Y\. Gao, B\. Mustafa, I\. Barr, E\. Parisotto, D\. Tian, M\. Eyal, C\. Cherry, J\. Peter, D\. Sinopalnikov, S\. Bhupatiraju, R\. Agarwal, M\. Kazemi, D\. Malkin, R\. Kumar, D\. Vilar, I\. Brusilovsky, J\. Luo, A\. Steiner, A\. Friesen, A\. Sharma, A\. Sharma, A\. M\. Gilady, A\. Goedeckemeyer, A\. Saade, A\. Feng, A\. Kolesnikov, A\. Bendebury, A\. Abdagic, A\. Vadi, A\. György, A\. S\. Pinto, A\. Das, A\. Bapna, A\. Miech, A\. Yang, A\. Paterson, A\. Shenoy, A\. Chakrabarti, B\. Piot, B\. Wu, B\. Shahriari, B\. Petrini, C\. Chen, C\. L\. Lan, C\. A\. Choquette\-Choo, C\. Carey, C\. Brick, D\. Deutsch, D\. Eisenbud, D\. Cattle, D\. Cheng, D\. Paparas, D\. S\. Sreepathihalli, D\. Reid, D\. Tran, D\. Zelle, E\. Noland, E\. Huizenga, E\. Kharitonov, F\. Liu, G\. Amirkhanyan, G\. Cameron, H\. Hashemi, H\. Klimczak\-Plucińska, H\. Singh, H\. Mehta, H\. T\. Lehri, H\. Hazimeh, I\. Ballantyne, I\. Szpektor, I\. Nardini, J\. Pouget\-Abadie, J\. Chan, J\. Stanton, J\. Wieting, J\. Lai, J\. Orbay, J\. Fernandez, J\. Newlan, J\. Ji, J\. Singh, K\. Black, K\. Yu, K\. Hui, K\. Vodrahalli, K\. Greff, L\. Qiu, M\. Valentine, M\. Coelho, M\. Ritter, M\. Hoffman, M\. Watson, M\. Chaturvedi, M\. Moynihan, M\. Ma, N\. Babar, N\. Noy, N\. Byrd, N\. Roy, N\. Momchev, N\. Chauhan, N\. Sachdeva, O\. Bunyan, P\. Botarda, P\. Caron, P\. K\. Rubenstein, P\. Culliton, P\. Schmid, P\. G\. Sessa, P\. Xu, P\. Stanczyk, P\. Tafti, R\. Shivanna, R\. Wu, R\. Pan, R\. Rokni, R\. Willoughby, R\. Vallu, R\. Mullins, S\. Jerome, S\. Smoot, S\. Girgin, S\. Iqbal, S\. Reddy, S\. Sheth, S\. Põder, S\. Bhatnagar, S\. R\. Panyam, S\. Eiger, S\. Zhang, T\. Liu, T\. Yacovone, T\. Liechty, U\. Kalra, U\. Evci, V\. Misra, V\. Roseberry, V\. Feinberg, V\. Kolesnikov, W\. Han, W\. Kwon, X\. Chen, Y\. Chow, Y\. Zhu, Z\. Wei, Z\. Egyed, V\. Cotruta, M\. Giang, P\. Kirk, A\. Rao, K\. Black, N\. Babar, J\. Lo, E\. Moreira, L\. G\. Martins, O\. Sanseviero, L\. Gonzalez, Z\. Gleicher, T\. Warkentin, V\. Mirrokni, E\. Senter, E\. Collins, J\. Barral, Z\. Ghahramani, R\. Hadsell, Y\. Matias, D\. Sculley, S\. Petrov, N\. Fiedel, N\. Shazeer, O\. Vinyals, J\. Dean, D\. Hassabis, K\. Kavukcuoglu, C\. Farabet, E\. Buchatskaya, J\. Alayrac, R\. Anil, Dmitry, Lepikhin, S\. Borgeaud, O\. Bachem, A\. Joulin, A\. Andreev, C\. Hardin, R\. Dadashi, and L\. Hussenot \(2025b\)Gemma 3 technical report\.External Links:2503\.19786,[Link](https://arxiv.org/abs/2503.19786)Cited by:[§4\.1\.3](https://arxiv.org/html/2606.24387#S4.SS1.SSS3.p1.1)\.
- E\. F\. Tjong Kim Sang and F\. De Meulder \(2003\)Introduction to the CoNLL\-2003 shared task: language\-independent named entity recognition\.InProceedings of the Seventh Conference on Natural Language Learning at HLT\-NAACL 2003,pp\. 142–147\.External Links:[Link](https://aclanthology.org/W03-0419/)Cited by:[§1](https://arxiv.org/html/2606.24387#S1.p2.1),[§2\.1](https://arxiv.org/html/2606.24387#S2.SS1.p1.1)\.
- F\. Ventirozos, I\. Nteka, T\. Nandy, J\. Baca, P\. Appleby, and M\. Shardlow \(2024\)Shifting NER into high gear: the Auto\-AdvER approach\.External Links:2412\.05655,[Link](https://arxiv.org/abs/2412.05655)Cited by:[§2\.2](https://arxiv.org/html/2606.24387#S2.SS2.p1.1)\.
- S\. Wang, X\. Sun, X\. Li, R\. Ouyang, F\. Wu, T\. Zhang, J\. Li, and G\. Wang \(2023\)GPT\-NER: named entity recognition via large language models\.External Links:[Link](https://arxiv.org/abs/2304.10428)Cited by:[Appendix B](https://arxiv.org/html/2606.24387#A2.p1.1),[§1](https://arxiv.org/html/2606.24387#S1.p5.1),[§2\.3](https://arxiv.org/html/2606.24387#S2.SS3.p2.1),[§4\.1\.3](https://arxiv.org/html/2606.24387#S4.SS1.SSS3.Px1.p1.1),[§4\.1\.3](https://arxiv.org/html/2606.24387#S4.SS1.SSS3.Px2.p1.1)\.
- B\. Warner, A\. Chaffin, B\. Clavié, O\. Weller, O\. Hallström, S\. Taghadouini, A\. Gallagher, R\. Biswas, F\. Ladhak, T\. Aarsen, G\. T\. Adams, J\. Howard, and I\. Poli \(2025\)Smarter, better, faster, longer: a modern bidirectional encoder for fast, memory efficient, and long context finetuning and inference\.InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),W\. Che, J\. Nabende, E\. Shutova, and M\. T\. Pilehvar \(Eds\.\),Vienna, Austria,pp\. 2526–2547\.External Links:[Link](https://aclanthology.org/2025.acl-long.127/),[Document](https://dx.doi.org/10.18653/v1/2025.acl-long.127),ISBN 979\-8\-89176\-251\-0Cited by:[§4\.1\.2](https://arxiv.org/html/2606.24387#S4.SS1.SSS2.p1.1)\.
- T\. Wolf, L\. Debut, V\. Sanh, J\. Chaumond, C\. Delangue, A\. Moi, P\. Cistac, T\. Rault, R\. Louf, M\. Funtowicz, J\. Davison, S\. Shleifer, P\. von Platen, C\. Ma, Y\. Jernite, J\. Plu, C\. Xu, T\. L\. Scao, S\. Gugger, M\. Drame, Q\. Lhoest, and A\. M\. Rush \(2020\)HuggingFace’s transformers: state\-of\-the\-art natural language processing\.External Links:1910\.03771,[Link](https://arxiv.org/abs/1910.03771)Cited by:[§4\.1\.2](https://arxiv.org/html/2606.24387#S4.SS1.SSS2.p1.1)\.
- A\. Yang, A\. Li, B\. Yang, B\. Zhang, B\. Hui, B\. Zheng, B\. Yu, C\. Gao, C\. Huang, C\. Lv, C\. Zheng, D\. Liu, F\. Zhou, F\. Huang, F\. Hu, H\. Ge, H\. Wei, H\. Lin, J\. Tang, J\. Yang, J\. Tu, J\. Zhang, J\. Yang, J\. Yang, J\. Zhou, J\. Zhou, J\. Lin, K\. Dang, K\. Bao, K\. Yang, L\. Yu, L\. Deng, M\. Li, M\. Xue, M\. Li, P\. Zhang, P\. Wang, Q\. Zhu, R\. Men, R\. Gao, S\. Liu, S\. Luo, T\. Li, T\. Tang, W\. Yin, X\. Ren, X\. Wang, X\. Zhang, X\. Ren, Y\. Fan, Y\. Su, Y\. Zhang, Y\. Zhang, Y\. Wan, Y\. Liu, Z\. Wang, Z\. Cui, Z\. Zhang, Z\. Zhou, and Z\. Qiu \(2025\)Qwen3 technical report\.Cited by:[§4\.1\.3](https://arxiv.org/html/2606.24387#S4.SS1.SSS3.p1.1)\.
- S\. Zhang, B\. Cao, and J\. Fan \(2024\)KCL: few\-shot named entity recognition with knowledge graph and contrastive learning\.InProceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation \(LREC\-COLING 2024\),pp\. 9681–9692\.External Links:[Link](https://aclanthology.org/2024.lrec-main.846/)Cited by:[§2\.3](https://arxiv.org/html/2606.24387#S2.SS3.p2.1)\.
## Appendix APreliminary Self\-Verification Experiments
We conducted preliminary experiments to assess the effect of the LLM self\-verification step\. Verification improved overall F1 for both Mistral\-7B and LLaMA3\-8B, mainly by increasing precision while slightly reducing recall\. The effect was strongest for Mistral, whose weighted F1 increased from 0\.375 to 0\.409\. For LLaMA3, the gain was smaller, with weighted F1 increasing from 0\.473 to 0\.481\. However, verification produced larger gains for some lower\-support labels\. For instance, with LLaMA3, the F1\-score forBATTERY\_CAPACITYincreased from 0\.291 to 0\.552 and forBOOT\_SIZEfrom 0\.112 to 0\.213\.
Table 6:Effect of self\-verification in preliminary LLM experiments\.
## Appendix BLLM Prompts
We use a per\-entity extraction strategy followingWanget al\.\([2023](https://arxiv.org/html/2606.24387#bib.bib21)\)\. Each advertisement is processed once per entity type using the prompts below\.
### Extraction Prompt
System:
> You are a NER Model\. Your task is to label \{entity\_type\} entities in the given sentence\.
User:
> Below are some examples\. \{few\_shot\_examples\}Input: \{text\} Output:
where each few\-shot example follows the template:
> Input: \{input\_text\} Output: \{output\_text\}
### Self\-Verification Prompt
System:
> You are a verification expert\. Your task is to verify if an extracted entity is correct\.
User:
> Original text: \{text\} Entity type: \{entity\_type\} Extracted entity: \{entity\} Is this extraction correct and accurate? Consider: 1\. Is the entity actually present in the text? 2\. Is the entity type classification correct? 3\. Are the boundaries accurate \(no extra or missing characters\)? Answer only with a ‘‘yes’’ if all statements are true, otherwise say ‘‘no’’\. Do not say anything more\. Verification result:
## Appendix CRules\-based Precision and Recall
Table[7](https://arxiv.org/html/2606.24387#A3.T7)show the per\-label precision and recall scores for the best\-performing rules\-based methods\.
Table 7:Precision and Recall Scores for Best Rules\-Based Methods\.
## Appendix DEncoder Training Set Size Analysis
Figure[2](https://arxiv.org/html/2606.24387#A4.F2)shows the evaluation loss of the encoder models relative to the number of training samples\. The plateauing of the loss curve suggests that the dataset size of 659 adverts was sufficient for the models to learn the primary patterns in the data\.
Figure 2:Evaluation loss vs\. number of training samples for encoder models\.
## Appendix EHardware and Runtime Details
All encoder\-based transformer models were fine\-tuned on an Apple M4 MacBook Pro equipped with 36GB unified memory\. Inference for the encoder\-based models required between a few seconds and approximately one minute to process the complete test set\.
For open\-source LLMs, models with fewer than 10 billion parameters were executed using vLLM\(Kwonet al\.,[2023](https://arxiv.org/html/2606.24387#bib.bib25)\)on an NVIDIA RTX PRO 6000 GPU with 96GB VRAM\. Larger models were executed on an NVIDIA H200 SXM GPU with 141GB VRAM\.
LLM inference included both the entity extraction and entity verification stages\. Processing the full evaluation set required approximately 1\.5 hours per model\. This corresponded to approximately109×15×2=3,270109\\times 15\\times 2=3,270model queries, where 109 is the number of test samples, 15 is the number of target entities, and 2 corresponds to the verification stage\.
Closed\-source models \(e\.g\., GPT and Gemini\) were accessed through their respective provider APIs\. All LLMs were interfaced using an OpenAI\-compatible API format\. To ensure deterministic outputs, we set the temperature to 0\.0 where supported by the API and used greedy decoding for all generations\.Similar Articles
DE-NER : Zero-shot Named Entity Recognition via Dialogue Elicitation of Large Language Models
Introduces DE-NER, a dialogue elicitation framework for zero-shot named entity recognition that uses self-play between questioner and roleplayer LLMs to clarify entity boundaries, achieving an average 3.75% F1 improvement over baselines.
Enhancing Scientific Named Entity Recognition via Large Language Models: A Type-driven Multi-task Learning Approach
This paper proposes TdSciNER, a type-driven multi-task learning approach that leverages LLMs to improve scientific named entity recognition by filtering entity types, adding an auxiliary typing task, and using a demonstration selection strategy. Experiments on three datasets show performance comparable to fully supervised models.
MiNER: Fine-Tuned Biomedical Natural Language Processing for Malaria Disease Entity Recognition in Clinical Texts
This paper proposes MiNER, a fine-tuned BioBERT model for extracting biomedical entities from malaria-related clinical texts, and releases a human-labeled dataset for future research.
DiZiNER: Disagreement-guided Instruction Refinement via Pilot Annotation Simulation for Zero-shot Named Entity Recognition
DiZiNER is a framework that uses disagreement between multiple LLMs to refine task instructions for zero-shot named entity recognition, achieving state-of-the-art results on 14 out of 18 benchmarks and significantly reducing the performance gap between zero-shot and supervised systems.
BERT-based Models vs. Large Language Models for Low-Resource Named Entity Recognition: A Comparative Study on Marathi
This paper compares fine-tuned MahaBERT-based models with large language models (Gemini, LLaMA-3.3-70B, Gemma) for Marathi named entity recognition, finding that the specialized BERT models significantly outperform the LLMs, achieving F1-scores of 0.88–0.91 versus 0.57–0.69.