Monitoring Transformative Technological Convergence Through LLM-Extracted Semantic Entity Triple Graphs

arXiv cs.CL Papers

Summary

This paper proposes a novel data-driven pipeline using LLMs to extract semantic triples from unstructured text and construct knowledge graphs for monitoring technological convergence. The method is validated on arXiv preprints and USPTO patent applications, demonstrating scalability for technology forecasting.

arXiv:2510.25370v2 Announce Type: replace Abstract: Forecasting transformative technologies remains a critical but challenging task, particularly in fast-evolving domains such as Information and Communication Technologies (ICTs). Traditional expert-based methods struggle to keep pace with short innovation cycles and ambiguous early-stage terminology. In this work, we propose a novel, data-driven pipeline to monitor the emergence of transformative technologies by identifying patterns of technological convergence. Our approach leverages advances in Large Language Models (LLMs) to extract semantic triples from unstructured text and construct a large-scale graph of technology-related entities and relations. We introduce a new method for grouping semantically similar technology terms (noun stapling) and develop graph-based metrics to detect convergence signals. The pipeline includes multi-stage filtering, domain-specific keyword clustering, and a temporal trend analysis of topic co-occurence. We validate our methodology on two complementary datasets: 278,625 arXiv preprints (2017--2024) to capture early scientific signals, and 9,793 USPTO patent applications (2018-2024) to track downstream commercial developments. Our results demonstrate that the proposed pipeline can identify both established and emerging convergence patterns, offering a scalable and generalizable framework for technology forecasting grounded in full-text analysis.
Original Article
View Cached Full Text

Cached at: 07/09/26, 07:54 AM

# Monitoring Transformative Technological Convergence Through LLM-Extracted Semantic Entity Triple Graphs
Source: [https://arxiv.org/html/2510.25370](https://arxiv.org/html/2510.25370)
\[ orcid=0009\-0008\-6801\-6160\]\\cormark\[1\]

1\]organization=Institute of Entrepeneurship and Management, HES\-SO, city=Sierre, country=Switzerland

\[orcid=0000\-0003\-0429\-8644\]

\[orcid=0000\-0003\-0429\-8644\]

\[orcid=0000\-0001\-6471\-772X\]

\[orcid=0000\-0002\-1002\-057X\]

\[\]

2\]organization=Cyber\-Defence Campus, armasuisse, Science and Technology, city=Thun, country=Switzerland

\\cortext

\[cor1\]Corresponding author

Andrei KucharavyDimitri Percia DavidAlain MermoudJulian Jang\-JaccardNathan Monnet\[

###### Abstract

Forecasting transformative technologies remains a critical but challenging task, particularly in fast\-evolving domains such as Information and Communication Technologies \(ICTs\)\. Traditional expert\-based methods struggle to keep pace with short innovation cycles and ambiguous early\-stage terminology\. In this work, we propose a novel, data\-driven pipeline to monitor the emergence of transformative technologies by identifying patterns of technological convergence\.

Our approach leverages advances in Large Language Models \(LLMs\) to extract semantic triples from unstructured text and construct a large\-scale graph of technology\-related entities and relations\. We introduce a new method for grouping semantically similar technology terms \(*noun stapling*\) and develop graph\-based metrics to detect convergence signals\. The pipeline includes multi\-stage filtering, domain\-specific keyword clustering, and a temporal trend analysis of topic co\-occurence\.

We validate our methodology on two complementary datasets: 278,625 arXiv preprints \(2017–2024\) to capture early scientific signals, and 9,793 USPTO patent applications \(2018–2024\) to track downstream commercial developments\. Our results demonstrate that the proposed pipeline can identify both established and emerging convergence patterns, offering a scalable and generalizable framework for technology forecasting grounded in full\-text analysis\.

###### keywords:

Bibliometrics\\sepEntity Extraction\\sepMachine Learning\\sepTechnological Forecasting\\sepLarge Language Models

## 1Introduction

Accurate forecasting of technological disruptions is essential for organizations seeking to remain competitive and resilient\. While incremental innovations can typically be anticipated and managed within existing planning frameworks,transformative technologiesintroduce paradigm shifts that reshape entire domains and often spill over into adjacent sectors\[Schumpeter1949,Ettlie1984OrganizationSA\]\. Because such technologies challenge established conceptual frameworks, they are inherently difficult to predict\[Christensen1997InnovatorsDilemma\]\.

Despite these challenges, the strategic importance of anticipating transformative change has long motivated the development of forecasting methods, particularly since the Cold War era\[TF1967Jantsch\]\. Seminal approaches such as the Delphi Method\[dalkey1969experimental\]and S\-curve Substitution Analysis\[fisher1971simple\]have inspired decades of methodological refinement\[national2010persistent,calleja\_sanz2020,technology2004technology\]\. However, these classical methods struggle to keep pace with the rapid innovation cycles characteristic of Information and Communication Technologies \(ICTs\)\. For example, the term “Large Language Models” \(LLMs\) emerged in 2019, and within just three years, ChatGPT had become a globally recognized application\[kucharavy2024deep\]\. In such contexts, the reliance of classical methods on expert panels and historical data renders them ill\-suited\[daim\_digital\_2022\]\.

As a result, ICT forecasting has increasingly shifted toward data\-driven methodologies\[daim\_anticipating\_2016\], with scientometrics emerging as a promising approach\[perciadavid2023\]\. Yet even data\-centric methods face difficulties in the early identification of transformative technologies, due in part to the absence of stable terminology and well\-defined application domains in the initial stages\. Continuing the example of LLMs, even a year after the release of ChatGPT, major bibliometric platforms such as OpenAlex still lacked ontology terms for many of the foundational sub\-technologies\[VelezEstevez2023NewTI,SinghChawla2022MassiveOI,würsch2023llms\]\. Whilewürsch2023llmsshowed that extracting technology\-related nouns from scientific texts is feasible, the process remains noisy and challenging to use effectively for large\-scale monitoring and forecasting\.

In this work, we present an alternative approach that leverages recent advances in Large Language Models \(LLMs\) to enable scalable extraction and analysis of semantic triples centered on technological concepts\[sternfeld\_eeke\]\. We focus ontechnological convergenceas an early indicator of transformative potential\. First introduced byrosenberg1963in his study of the machine tool industry, convergence describes how formerly separate fields of science and technology come to rely on a shared technical and knowledge base, blurring the boundaries between them\. Transformative technologies frequently arise at these points of overlap, so tracking where distinct fields begin to intersect offers an early signal of disruptive potential\[Dosi1982TechnologicalPA,Li2024ANI\]\.

Our main contributions are as follows:

1. \(i\)We introduce an unbiased pipeline for extracting semantic triples that involve technology\-designating nouns\.
2. \(ii\)We propose a novel syntactic similarity identification technique—termednoun stapling—to group related technology terms\.
3. \(iii\)We implement a coarse decomposition of technology components to improve interpretability\.
4. \(iv\)We develop a graph\-based method to detect and track technological convergence over time\.

These elements are integrated into a unified pipeline for identifying the emergence of transformative technologies through convergence patterns\. We demonstrate the utility of this pipeline through a case study on LLMs, applying it to a large corpus of 278,625 arXiv preprints \(2017–2024\) and 9,793 USPTO patent applications \(2018–2024\)\.

The remainder of this paper is structured as follows: Section[2](https://arxiv.org/html/2510.25370#S2)reviews the related literature\. Section[3](https://arxiv.org/html/2510.25370#S3)describes the data sources\. Section[4](https://arxiv.org/html/2510.25370#S4)outlines the methodology\. Section[5](https://arxiv.org/html/2510.25370#S5)presents the results\. Finally, Section[6](https://arxiv.org/html/2510.25370#S6)concludes the paper and discusses directions for future research\.

## 2Related work

In this section, three branches of previous research are discussed\. First, we discuss the main topics and trends in bibliometrics\. Then, we discuss related work on claim and triple extraction, where we highlight those methods that leverage machine learning techniques\. Last, we discuss the usage of topological analyses for technology forecasting\.

### 2\.1Bibliometrics

As a method for technological forecasting, bibliometrics leverages both qualitative and quantitative analysis of recorded information—primarily scientific books, articles, and patents\. The systematic study of written works has long been established, with the term*bibliometrics*coined in the early 20thcentury\. One of the most widely recognized classical bibliometric techniques is citation analysis, which examines citation patterns among academic publications\[smith1981citation\]\.

Early citation analysis focused primarily on citation counts and co\-citation networks\. However, the advent of advanced network analysis methods and increasing computational power has enabled more sophisticated analyses of citation networks\. Such methods are now employed across a range of fields for technology forecasting\. For instance,qiu2022conducted a citation network analysis of robotics patents to identify key pathways of innovation\. Similarly,LI2020analyzed both patents and academic papers in the domain of nanogenerator technologies, offering a more integrated view of scientific and technological progress\.

Traditionally, bibliometric studies were constrained to structured metadata\. Recent advances in Natural Language Processing \(NLP\), however, have opened new avenues for deeper analysis of full\-text documents\. Techniques such asnn\-gram analysis\[michel2011quantitative\], Latent Dirichlet Allocation \(LDA\)\[LDA2003\], semantic embeddings\[Mikolov2013DistributedRO\], deep neural language model encodings\[Context2Vec2016\], and, more recently, large language models \(LLMs\)\[Beltagy2019SciBERT\], have greatly expanded the scope of bibliometric inquiry\. These developments are further supported by the growing availability of full\-text documents through open\-access and preprint platforms, such as arXiv111[https://info\.arxiv\.org/](https://info.arxiv.org/), medXiv222[https://www\.medrxiv\.org/content/about\-medrxiv](https://www.medrxiv.org/content/about-medrxiv), bioRxiv333[https://www\.biorxiv\.org/content/about\-biorxiv](https://www.biorxiv.org/content/about-biorxiv), and PubMed Central444[https://pmc\.ncbi\.nlm\.nih\.gov/about/intro/](https://pmc.ncbi.nlm.nih.gov/about/intro/)\.

For example,perciadavid2023analyzed computer science preprints on arXiv using LLM\-derived sentence embeddings to extract author sentiment and relate it to technology security over time\. Likewise, Sandu et al\. \(2024\) constructed a collaboration network in the area of social media text mining and identified recurring research themes usingnn\-gram analysis\.

In parallel with developments in literature\-based bibliometrics, there has been increasing attention to patent analysis, given the rich technical detail and close ties to commercial applications that patents offer\. While academic papers often represent early\-stage scientific exploration, patents typically reflect efforts to translate these discoveries into deployable technologies\. Importantly, both European and U\.S\. patents follow the Cooperative Patent Classification \(CPC\) system, enabling systematic, cross\-jurisdictional analysis\. For example,KIM2017228cluster patents based on CPC codes and apply citation and claim analysis to forecast emerging technologies\. Similarly,you2017analyze light generator technology patents using Bass and ARIMA models to identify promising innovation trajectories\. By integrating bibliometric insights from both scholarly publications and patents, researchers gain a more comprehensive view of technological development—spanning foundational research to commercial realization\[LI2020,ADAMUTHE2019181\]\.

### 2\.2Claim and triple extraction

Semantic Triples, also known as RDF \(Resource Description Framework\) triples, are a fundamental representation for knowledge extraction and organization in semantic web technologies\. A semantic triple consists of three components: a subject, a predicate, and an object, and is used to represent factual statements in a machine\-readable format\[berners2001rdf\]\. This structure is foundational in knowledge graphs, enabling systems to store and reason about facts in a way that can be queried and expanded upon\[Hogan\_2021\]\. In the context of claim extraction from scientific literature, we aim to extract claims that can be represented as semantic triples, focusing on technological assertions and relationships\. By adhering to the semantic triple standard, these claims are organized into unrooted, unbiased knowledge graphs that facilitate downstream tasks such as pattern mining and knowledge discovery\.

Previous approaches to claim extraction can be broadly categorized into heuristic and machine learning methods\. Heuristic methods, such as the approach byjansen2017, do not require training data and have low computational cost\. However, they are limited in their ability to capture complex claims\.jansen2017extract core claims by scoring sentences based on pattern matching, term frequency, and sentence length\. However, they only succeed in rewriting claims to the required AIDA \(atomic, independent, declarative, and absolute\) standard in only 29% of cases\. In contrast, machine learning methods like the one proposed byli2021, using BiLSTMs for sentence classification, offer better performance, categorizing sentences with an F1 score over 0\.8\. However, these models require domain\-specific supervised training data, which must be updated regularly\.

Most triple extraction models rely on supervised training, such as RECON and sPERT, which require labeled training data\[bastos2021,eberts2020\]\. These methods are limited by their dependence on available training data and predefined relations\. A long\-standing research stream is Open Information Extraction \(OIE\), first framed byopenieas the schema\-free, domain\-independent extraction of relational tuples from text\. Early OIE systems relied on shallow syntactic patterns \(TextRunner\[textrunner\], ReVerb\[reverb\]\), later incorporating deeper linguistic features \(OLLIE\[ollie\], ClausIE\[clausie\]\) and clause\-based reasoning, of which Stanford OpenIE\[stanfordie\]is a widely used representative\. The recurring limitation across these generations is that schema\-free extraction favours coverage over precision, so systems tend to over\-generate surface\-level or non\-useful tuples\[liu2022\], which motivates approaches that constrain extraction without reverting to full supervision\.ottersen2024address this issue by using a language model to extract triples within a predefined set of relations, improving precision but requiring users to specify all possible relations\.

More recently, extraction has shifted toward*generative*formulations in which a language model emits triples directly as text\. REBEL\[rebel\]casts relation extraction as end\-to\-end sequence generation, while UIE\[uie\]unifies multiple extraction tasks through a single text\-to\-structure schema\. With instruction\-tuned LLMs, this line has moved further toward few\-shot and zero\-shot extraction driven by prompting rather than task\-specific training\. ChatIE\[chatie\]decomposes zero\-shot extraction into a multi\-turn question\-answering dialogue, and GPT\-RE\[gptre\]uses in\-context demonstrations to improve relation extraction\. A parallel concern is enforcing valid output structure: constrained or structured decoding restricts generation so that outputs conform to the target format\[uie,li2024\]\. For knowledge\-graph construction specifically, EDC\[edc\]combines open extraction with schema definition and canonicalisation, showing that LLMs can produce high\-quality triples without parameter tuning\. The suitability of LLM\-based extraction is domain\-dependent;würsch2023llmsreport that LLMs extract poorly matched knowledge entities from cybersecurity literature, which underscores that generative extraction must be paired with validation and filtering\. A comprehensive survey of this generative IE paradigm, including prompt design, zero\-shot learning, and constrained decoding, is provided byxu2023\.

Our work performs schema\-free triple extraction from full\-text scientific articles and patents using few\-shot prompting of LLMs, thereby retaining the domain independence of OIE while exploiting the contextual capacity of LLMs to curb its characteristic over\-generation\. Unlike generative IE methods evaluated on sentence\-level benchmarks\[rebel,uie,chatie,gptre,xu2023\], we apply open extraction at the scale of hundreds of thousands of documents to build large, unbiased technology graphs\.

### 2\.3Topological Analyses for Technology Forecasting

Traditional co\-word analysis techniques primarily relied on co\-occurrence frequencies to detect associations between terms\[callon1983coword\]\. More recent approaches, however, emphasize the connectivity and centrality of entities within complex graphs to identify zones of technological convergence and emerging innovation clusters\. At the core of these methods areknowledge graphs, which are constructed by extracting key entities such as technologies, methods, and applications from scientific publications and patents\. In these graphs, nodes represent concepts, and edges encode semantic, citation, or temporal relationships\. Keyword extraction tools, such as RAKE\[rake\]and KeyBERT, are often used to enrich these graphs\. The resulting structures can be analyzed as either static or dynamic representations of technological domains\.

The topological characteristics of knowledge graphs are crucial for forecasting applications\. For instance,DOTSIKA2017114demonstrated that keyword co\-occurrence networks derived from scientific and business literature can be effectively analyzed using centrality metrics to detect potentially disruptive or fast\-growing technological areas\. Building on this idea,LI2020developed dual\-layer networks that integrate scientific publications and patents to trace the evolution and commercialization of nanogenerator technologies\. More recently,WANG2025101606applied semantic term extraction in combination with Louvain\-based community detection to identify early\-stage convergence between disparate research fields\.

Community detection continues to be a foundational component of topological forecasting\. The Louvain algorithm, introduced byblondel2008LouvainMethod, is widely used to identify coherent topic clusters within large\-scale graphs\. Complementary measures, such as eigenvector centrality, highlight influential nodes that play a key role in the dissemination and integration of knowledge\. These techniques support dynamic analyses of how scientific topics emerge, grow, merge, or decline over time\. Dynamic knowledge graphs, such as those implemented in the Science4Cast project\[krenn2023science4cast\], enable the longitudinal monitoring of research landscapes and facilitate the detection of novel knowledge pathways as they form\.

In addition, predictive modeling on these graphs—using either neural networks or handcrafted topological features—can forecast whether previously unconnected concepts are likely to become linked in the future, and whether such connections are poised to attract significant attention\[gu2024forecasting,martinez2016survey\]\.

## 3Data

The methodology developed in this work can be used for any textual data source\. Several examples of such sources are scientific articles, patents, news articles, social media posts and web pages\. In this paper, we focus on two data sources: scientific articles and US patent applications\. In order to evaluate the capabilities of different methods of technology\-related semantic triples extraction, we created a golden dataset of manually labeled semantic triples in the LLM field\.

### 3\.1Golden Dataset for Triples Extraction

To fine\-tune the LLMs for technology\-related semantic triples extraction, a labeled dataset is required, both for evaluation of method performance, and for supervised training\. To the best of our knowledge, such a labeled dataset for semantic triples extraction from scientific papers does not exist\. Therefore, we manually construct a training dataset based on the paperA Survey of Large Language Models\[zhao2023survey\]\. We chose this paper given our focus on the LLMs transformative technology, since it is a comprehensive and general review of LLM component technologies\. In total, we manually annotated 547 triples in 100 text segments of 15 lines, given that amount of examples is generally considered to be sufficient for parameter\-efficient fine\-tuning \(PEFT\), while leaving room for a test holdout set\[falissard\-etal\-2023\-improving\]\. To evaluate the performance of the fine\-tuned LLMs, we use a holdout set of 20% of the annotated paragraphs\. The dataset is available for download from the data and code repository associated to this article:[https://github\.com/submissiontfscanonymized/tfsc2025](https://github.com/submissiontfscanonymized/tfsc2025)\.

### 3\.2arXiv data

In the first use case, preprints publicly available on arXiv are considered\. As arXiv is an open scholarly preprint archive, the documents are not peer\-reviewed\. We choose this platform for two key reasons\. First, as there is no peer review and only a short moderation process, submissions to arXiv are available rapidly, tracking the state of research as close to real time as possible\. The archive includes papers that may not have been accepted at conferences due to perceived lack of immediate interest, yet have gone on to receive significant citations — highlighting their eventual importance\. Second, arXiv ensures that all articles are classified by the authors into a category from a pre\-defined taxonomy\. To aid authors, arXiv has published an algorithm that can suggest the best fitting categories for a paper, based on the existing corpus of articles555https://info\.arxiv\.org/help/api/classify\.html\#cat\. Finally, all arXiv pdf’s and raw latex files can be downloaded through Google Cloud Storage Buckets, which are updated weekly666https://www\.kaggle\.com/datasets/Cornell\-University/arxiv, making text mining on arXiv preprings significantly easier\.

For an initial evaluation and refinement of our methodology, we only worked with papers from December 2023 \(4225 papers\), from the arXiv categories`cs\.AI`,`cs\.CL`and`cs\.LG`\. We choose to do so, as a small dataset reduces the computation time and required computational resources\. Therefore, such a dataset is suitable for early refinements of the pipeline\. After settling on the final methodology, all papers between 2018 and 2024 are considered, from the arXiv categories that are specified in Table[1](https://arxiv.org/html/2510.25370#S3.T1)for a total of 278,625 articles\. Both the categories and timeframe correspond to the timeline of convergence of technologies that led to the transformative LLM technology emergence\.

Table 1:The arXiv categories that are considered for this study, alongside their full names\.CategoryFull namecs\.CLComputation and Languagecs\.LGMachine Learningcs\.AIArtificial Intelligencecs\.IRInformation Retrievalstat\.MLMachine Learning
### 3\.3USPTO patent data

The second data source we consider is USPTO patent data\. We choose to consider patents as a secondary data source as it provides information on downstream applications of technologies\. While preprints on arXiv will show early signs of novel technological developments, the presence of technologies in patent applications shows that technologies are maturing and are being adopted in commercial products\.

Specifically, we consider patent applications to the USPTO, which are publicly available through the Bulk Data Directory777https://data\.uspto\.gov/bulkdata/datasets\. Similar to the arXiv papers, we consider the period 2018\-2024\. All patent applications are categorized according to the Cooperative Patent Classification \(CPC\) system, which is a classification system that is used both by te USPTO and the European Patent Office \(EPO\)\.

For our case study, we filter the data for those patents that contain at least one of the termslarge language model,large language models,llmorllmsin the abstract or title\. This results in a total of 9793 patent applications between 2018 and 2024\.

## 4Methodology

In this Section we describe the methodology of our triple extraction pipeline and downstream analyses\. We first describe the data preprocessing, after which two triple extraction methods are presented\. Then, the post\-processing and filtering of the triples is elaborated upon\. Last, we explain our proposed*noun stapling*methods to group similar triples, and our downstream analyses\. The full pipeline is displayed in Figure[1](https://arxiv.org/html/2510.25370#S4.F1)\.

### 4\.1Preprocessing

The`PyMuPDF`pdf processing library\[PyMuPDF\]is used to turn the pdfs into text\. The resulting text is then processed to remove citations, using a heuristic pattern matching\. We remove both square and round brackets with only numbers inside, as well as brackets with a year inside, matching most frequent citation formats found in the arXiv papers\. In addition to that, we remove the line breaks using a similar heuristic, although with dashes\.

In order to unify extracted entities, it is important to expand abbreviations in the text\. By doing so, we will be able to later on mapLLMandLarge Language Modelto the same entity\. To ensure that the abbreviation resolution is done properly, we benchmark different algorithms on theFundamentals of Generative Large Language Models and Perspectives in Cyber\-Defensereport, given its mixed usage of both ML, cybersecurity, and cyberdefense abbreviations\[kucharavy2023fundamentals\]\. Table[2](https://arxiv.org/html/2510.25370#S4.T2)shows the comparative performance of the Schwartz\-Hearst\[Schwartz2003\], scispaCy\[neumann\-etal\-2019\-scispacy\], NLPRe\[NLPRe\]and a fine\-tuned RoBERTa model\[zilio\-etal\-2022\-plod\]as abbreviation detection methods\. The results show that the Schwartz\-Hearst algorithm performs the best, with scispaCy coming a close second\. Given the speed of Schwartz\-Hearst, we select it for our pipeline\. Table[3](https://arxiv.org/html/2510.25370#S4.T3)shows an example of a sentence before and after preprocessing\.

Table 2:Performance of three abbreviation detection algorithms on theFundamentals of Generative Large Language Models and Perspectives in Cyber\-Defense\[kucharavy2023fundamentals\]\. Manual annotation resulted in 15 abbreviations being explicitly introduced in the report, on which the percentages are based\.Correctly detectedFalse positiveTimeSchwartz\-Hearst11 \(73%\)10\.04 sscispaCy11 \(73%\)48\.73 sNLPRe0 \(0%\)00\.01 sFine\-tuned RoBERTa10 \(67%\)3100 sTable 3:Illustration of the preprocessing effect, the parts in bold are altered during preprocessing\.UncleanedPreprocessedSociety has been affected byartificial intelligence \(AI\)and hasbecome morerel\- iantonAIproducts\.Society has been affected byartificial intelligenceand has becomemorereliantonartificial intelligenceproducts\.
### 4\.2Triple extraction

#### 4\.2\.1spaCy\-based triple extraction

##### Claim extraction\.

After preprocessing the raw text, the sentences that constitute the core claims of the text are identified\. To this end, theClaimDistillerframework is used, which is developed byWei2023\. In this work, both CNN and BiLSTM models have been employed for claim extraction tasks through training on the PubMED\-RCT and SciARK datasets\[fergadis\-etal\-2021\-argumentation,dernoncourt\-lee\-2017\-pubmed\]\. While incorporating supervised contrastive learning has been shown to enhance model performance, it also introduces additional computational cost\. To maintain a balance between effectiveness and efficiency, we opt for the BiLSTM model variant without supervised contrastive training\. This model is used to extract claims from academic papers, thereby reducing the number of sentences that need to be processed during the subsequent triple extraction phase\.

![Refer to caption](https://arxiv.org/html/2510.25370v2/images/pipelineV2.png)Figure 1:The complete pipeline for the triple extraction and downstream analysis, starting from either raw patent data or raw arXiv data\. The green components of the pipeline reflect the triple extraction procedure, including the pre\- and post\-processing steps\. The grey components of the pipeline illustrate the key\-term extraction, noun stapling and the downstream analyses\.
##### Triple extraction\.

Next, we aim to reduce the claims to \(subject, predicate, object\) semantic triples, analogous to the Resource Description Framework \(RDF\) format commonly used in the representation of knowledge in the Web Ontology Language \(OWL\)\. We choose this representation, as it will facilitate the comparison of claims across papers\. To achieve this, we use the Python library`textacy`, which is built on`spaCy`\[spaCy2020\]and has a built\-in extraction method that does not require the specification of relations in advance\.

#### 4\.2\.2LLM\-based triple extraction

As a second method for triple extraction, we consider an approach based on LLMs\. The main advantage of such an approach is that LLMs can take into account context when extracting triples from a sentence\. However, one has to be cautious, as LLMs can hallucinate and one can thus never be certain regarding the factuality of generations\. In fact, LLMs have been reported to struggle with niche concepts\[Wursch2023\]\. Therefore, prior to using LLMs it is critical to assess LLMs for hallucinations during triple extraction\. Three state\-of\-the\-art LLMs are considered:Mistral\-7B\-Instruct\-v0\.2,Meta\-Llama\-3\-8B\-Instructand Starling\-LM\-7B\-beta\[jiang2023mistral,starling2023,llama3modelcard\]\. Our selection of candidate models follows directly from the deployment constraints of the pipeline\. Because our aim is a triple\-extraction method that is scalable, replicable, and runnable on modest hardware \(16GB VRAM GPUs; Section[4\.8](https://arxiv.org/html/2510.25370#S4.SS8)\), we restrict attention to open\-weight models in the 7–8B parameter range\. Additionally, we exclude proprietary API\-only systems, as they compromise replicability and are costly at the scale of hundreds of thousands of documents\. We deliberately chose three instruction\-tuned models that cover three distinct pretraining lineages, and two alignment paradigms: supervised instruction tuning with RLHF for the Llama and Mistral models, and reinforcement learning from AI feedback for Starling\. Whereas the Starling and Llama models have a context length of 8192, the Mistral model has a context length of 32,000\.

We note that unlike the spaCy\-based triple extraction, for LLM\-based extraction we do not first extract claims from the text\. The LLM is prompted to extract triples from the full input text, as the strength of LLMs is to take into account the full context\. Therefore, we do not want to remove this context by isolating the claims\.

##### Few\-shot learning prompts and fine\-tuning\.

Generative LLMs per se are not specifically pretrained or instruction\-tuned for triple extraction\. In order to improve their performance at this task, the two common approaches are few\-shot learning prompts\[LLMFewShotLearners2020OpenAI\], or fine\-tuning\[OpenAI2018GPT1\]\.

Few\-shot learning is generally considered as simpler to implement, given that it is equivalent to providing examples of desired text transformation as part of the user prompt\. We formulated several of such example prompts based on the existing literature examples, and selected the best performing one\. The resulting prompt can be found in Appendix Section[A\.1](https://arxiv.org/html/2510.25370#A1.SS1)

Fine\-tuning LLMs is generally believed to be more effective, but requires significantly more computational power and a substantial amount of fine\-tuning examples that scales with the model size\. It is possible to mitigate both problems using parameter\-efficient fine\-tuning \(PEFT\), which reduces both the computational power requirements and offers better generalization with less data\[falissard\-etal\-2023\-improving\]\. Specifically, we use LoRA\[hu2021lora\], which freezes the majority of model parameters and fine\-tunes only the scaling matrix of a singular value decomposition of a random matrix added to the model weights\.

### 4\.3Post\-processing

Given that the end goal is to compare triples across papers, we normalize triple representation in order to facilitate the matching of nouns likely to designate technologies of interest\. Specifically, we:

1. 1\.Lowercase all words in the triple;
2. 2\.Remove triples where either the subject or object contains more than 6 words;
3. 3\.Remove stopwords from the triples based on the list included in`NLTK`;
4. 4\.Remove any non\-text characters;
5. 5\.Lemmatize verbs and nouns in the triple;
6. 6\.Remove words containing less than 3 characters;

Both cutoff values in 2\. and 6\. were chosen empirically, based on manual inspection of triple nouns\.

### 4\.4Filtering

As a final step, we filter the triples to identify the key domain terms on which the downstream analysis depends\. Specifically, subjects and objects that designate technologies or technology\-adjacent concepts specific to the field of interest, as opposed to rare, generic, or domain\-agnostic vocabulary\. We apply three complementary criteria, each removing a different class of non\-domain term: terms too rare to denote an established technology \(Frequency\), terms too generic to be field\-specific \(Bookcorpus\), and terms that are not discriminative of any particular field \(Entropy\)\.

#### Frequency

First, we filter out triples with subjects or objects that are too rare\. We do so by considering acontrol corpusof all papers from the last quarter of 2023, enriched by a sample of 400 papers per month from the period January 2015 \- September 2023\. For every termtt, we then definefc​c,tf\_\{cc,t\}, which is the number of papers in the control corpus that contains the term at least 3 times per page, on average\. If a triple has a subject or object of whichfc​c,t<5f\_\{cc,t\}<5for all termsttin the subject or object, we remove this triple\. This filtering step aims at removing nouns that are not used sufficiently frequently to correspond to a potential technology\.

#### Bookcorpus

Second, we use the general\-purpose Gutenberg book corpus to detect triples that are generic and thus carry little information for technological monitoring\. Letfb,tf\_\{b,t\}define the number of books in which termttappears at least 3 times per page\. Furthermore, we denote the number of papers in the control corpus asNpN\_\{p\}and the number of books in the book corpus asNbN\_\{b\}\. Then, each termttobtains the following scorests\_\{t\}:

si=\{l​o​g​\(fc​c,iNp\)−l​o​g​\(fb,iNb\)iffb,i\>0∞iffb,i=0s\_\{i\}=\\begin\{cases\}log\(\\frac\{f\_\{cc,i\}\}\{N\_\{p\}\}\)\-log\(\\frac\{f\_\{b,i\}\}\{N\_\{b\}\}\)&\\text\{if $f\_\{b,i\}\>0$\}\\\\ \\infty&\\text\{if $f\_\{b,i\}=0$\}\\end\{cases\}\(1\)
The score of a term thus increases when the frequency of the term in the book corpus is lower\. Simultaneously, the score increases when the frequency of the term in the control corpus is higher\. We keep the triples of which at least one word from the subject and object is in the top 10% of the term scores, with cutoff determined empirically\.

#### Entropy

Last, we filter terms according to how concentrated their usage is across arXiv categories\. A term that appears throughout the taxonomy, such as "methodology" or "lab", carries little information about the field of a paper\. In contrast, a term confined to a few categories is field\-discriminative and thus more likely to denote a technology or technology\-adjacent concept\. We quantify this concentration with thecategorical occurrence entropyof a term: the entropy of its occurrence pattern over arXiv categories\. Terms concentrated in few categories yield low entropy and are retained, whereas terms spread across many categories yield high entropy and are removed\.

Formally, let us denote the set of all arXiv categories as𝒞\\mathcal\{C\}, then the categorical occurrence entropy for termttcan be calculated as denoted below\.

Ht=∑c∈𝒞P​\(t∈p\|p∈c\)⋅l​o​g​\(1P​\(t∈p\|p∈c\)\)H\_\{t\}=\\sum\_\{c\\in\\mathcal\{C\}\}\{P\(t\\in p\|p\\in c\)\\cdot log\(\\frac\{1\}\{P\(t\\in p\|p\\in c\)\}\)\}\(2\)
Here, we denote a paper asppand the individual arXiv categories ascc\.

### 4\.5Noun stapling

At this point in the pipeline, we have triples involving nouns that designate with high probability technologies and technology\-adjacent terms, conformed to a somewhat standard representation\. However, in order to perform an emerging technology analysis through technological convergence, we need to connect technology\-designating nouns that indicate similar technologies despite naming variations\. We refer to this asnoun stapling\.

When stapling \(compound\) nouns, we want to group subjects or objects together that are semantically sufficiently similar\. As we are dealing with big datasets, we use a syntactic string similarity measure, which is fast and can combine detection of similarities at token and character level\. Appendix[A\.3](https://arxiv.org/html/2510.25370#A1.SS3)studies semantic embedding similarity as a measure for noun similarities, but we find that in niche domains embeddings are not able to capture similarities accurately, which is in line with the findings ofWursch2023\.

Specifically, we use a novel string similarity measure, introduced bySyntactic\_string\_measures, that utilizes a combination of character\- and token\-level similarity measures\. Specifically, the framework uses thesoft cardinality of a set, as introduced byvargas2008knowledge\. The aim of a soft cardinality measure is to capture the number of unique concepts in a string\. For example, the set\(dog, dogs\)should have a lower soft cardinality than the set\(dog, church\)\.

The results ofvargas2008knowledgeshowed that, in general, the combination of Dice \(token\-level\) and 3\-gram \(character\-level\) worked well in all settings\. Therefore, we choose to use this procedure, where we tokenize the compound nouns by splitting on spaces\. Let us first introduce the 3\-gram similarity measure between the stringss1s\_\{1\}ands2s\_\{2\}, which is defined as

Q​G​r​a​m​s​S​i​m​\(s1,s2\)\\displaystyle QGramsSim\(s\_\{1\},s\_\{2\}\)=1−\\displaystyle=1\-∑i=1n\|m​a​t​c​h​\(qi,Qs1\)−m​a​t​c​h​\(qi,Qs2\)\|\|Qs1\|\+\|Qs2\|,\\displaystyle\\hskip\-65\.00009pt\\frac\{\\sum\_\{i=1\}^\{n\}\\big\|match\(q\_\{i\},Q\_\{s\_\{1\}\}\)\-match\(q\_\{i\},Q\_\{s\_\{2\}\}\)\\big\|\}\{\|Q\_\{s\_\{1\}\}\|\+\|Q\_\{s\_\{2\}\}\|\},\(3\)
whereQs1Q\_\{s\_\{1\}\}andQs2Q\_\{s\_\{2\}\}are the sets of q\-grams froms1s\_\{1\}ands2s\_\{2\}respectively\. Furthermore,n=\|Qs1∪Qs2\|n=\|Q\_\{s\_\{1\}\}\\cup Q\_\{s\_\{2\}\}\|, andm​a​t​c​h​\(qi,Qs1\)match\(q\_\{i\},Q\_\{s\_\{1\}\}\)is the number of times the q\-gramqiq\_\{i\}appears inQs1Q\_\{s\_\{1\}\}\. This measure can then be used to calculate the soft cardinality of a setT=\(T1,T2,…,Tn\)T=\(T^\{1\},T^\{2\},\.\.\.,T^\{n\}\), which is defined as

\|T\|s​o​f​t=∑i=1n1∑j=1nQ​G​r​a​m​s​S​i​m​\(Ti,Tj\)\\displaystyle\|T\|\_\{soft\}=\\sum\_\{i=1\}^\{n\}\{\\frac\{1\}\{\\sum\_\{j=1\}^\{n\}\{QGramsSim\(T^\{i\},T^\{j\}\)\}\}\}\(4\)
Using the soft cardinalities, one can then compute the similarity between the two compound nounst1t\_\{1\}andt2t\_\{2\}through an adjustment of the token\-level measure\. For the DICE measure, this then becomes:

S​o​f​t​D​I​C​E​\(t1,t2\)=2×\|t1∩t2\|s​o​f​t\|t1\|s​o​f​t\+\|t2\|s​o​f​t\\displaystyle SoftDICE\(t\_\{1\},t\_\{2\}\)=\\frac\{2\\times\|t\_\{1\}\\cap t\_\{2\}\|\_\{soft\}\}\{\|t\_\{1\}\|\_\{soft\}\+\|t\_\{2\}\|\_\{soft\}\}\(5\)
Here, the soft cardinality of a set intersection can be computed as

\|t1∩t2\|s​o​f​t=\|t1\|s​o​f​t\+\|t2\|s​o​f​t−\|t1∪t2\|s​o​f​t\\displaystyle\|t\_\{1\}\\cap t\_\{2\}\|\_\{soft\}=\|t\_\{1\}\|\_\{soft\}\+\|t\_\{2\}\|\_\{soft\}\-\|t\_\{1\}\\cup t\_\{2\}\|\_\{soft\}\(6\)Last, one must choose a threshold above which the two compound nounst1t\_\{1\}andt2t\_\{2\}are considered to be similar\. As we want to be conservative, we choose a threshold of 0\.85\. Above this threshold, one can be relatively certain that two compound nouns are sufficiently similar to be marked as such\.

Table 4:Pseudo algorithm for keyword clustering, based on the soft dice similarity as introduced in equation[5](https://arxiv.org/html/2510.25370#S4.E5)\.Pseudo algorithm:keyword clustering based on similarity1:Inputs:2:keywordsList:A list of keywords ordered by extraction count3:similarKeywords:A dictionary which gives, for every keyword, a list of similar keywords4:Output:5:Grouped keywords:A dictionary of groups of keywords6:assignedKeywords←\\leftarrowempty set7:groups←\\leftarrowempty dictionary8:foreach keyword1 in keywordsListdo9:ifkeyword1 in groupsthen10:continue11:end if12:currentGroup←\\leftarrowempty list13:add keyword1 to currentGroup14:foreach keyword2 in similarKeywords\[keyword1\]15:ifkeyword2 is in assignedKeywordsthen16:continue17:end if18:add keyword2 to currentGroup19:add keyword2 to assignedKeywords20:end for21:groups\[keyword1\]←\\leftarrowcurrentGroup22:add keyword1 to assignedKeywords23:end for24:returngroups
### 4\.6Key term extraction and clustering

From the triple extraction pipeline, we want to create a large graph of technology\-related topics \(nodes\)\. To derive insights from this graph, it is essential to focus on a specific domain of interest\. To this end, one requires a list of relevant topics\. The manual curation of such a list requires a strong domain expertise, which is regularly not readily available\.

Therefore, we develop a separate technology\-designating keyword extraction for the first step of our approach\. After manually specifying a few terms that broadly specify the field of interest, we select papers from arXiv that contain at least one of these terms in the abstract or title\. Then, we extract a maximum of 10 key terms for each abstract using KeyBERT, withallenai/specteras the embedding model\[grootendorst2020keybert,cohan2020specterdocumentlevelrepresentationlearning\]\.

Next, we use the noun stapling method as described in Section[4\.5](https://arxiv.org/html/2510.25370#S4.SS5)to calculate similarities between the extracted key terms\. We then use the pseudo\-algorithm as given in Table[4](https://arxiv.org/html/2510.25370#S4.T4)to cluster the key terms intotopics\. The algorithm takes a simple approach, leveraging the extraction count for each term\. Starting from the most popular term, we iterate over the terms and check whether they are already assigned to a group\. If not, we make a new group to which we add the term and all similar terms that are not yet assigned to a group\. Last, the groups are named based on the most common substring across the key terms\. If two substrings are equally common, we prefer to choose the longest one as it is likely more informative\.

Finally, we map the extracted triples onto these topic clusters to obtain the technology network used in the downstream analyses\. Each subject or object is assigned to a topic if at least one key term in that cluster, where similarity is again measured with the SoftDICE similarity of Equation[5](https://arxiv.org/html/2510.25370#S4.E5), using the threshold of 0\.85 as discussed in Section[4\.5](https://arxiv.org/html/2510.25370#S4.SS5)\. We note that it is possible for subjects and objects to be assigned to multiple topics simultaneously, if they are matched to keywords from multiple topics\. This results in a graph where each node represents a topic, and the nodes are connected through the associated triples\. Specifically, the predicates of the triples form undirected edges between the topics\.

Table 5:Comparative benchmarking of triple extraction methods on the golden dataset\. A \(↑\\uparrow\) in the column header means that higher values are preferred, whereas a \(↓\\downarrow\) means lower values are better\.Avg\. numberof triples \(↑\\uparrow\)Perc\. withcorrect format \(↑\\uparrow\)Perc\. inconsistentwith Levenshtein dist\. 2 \(↓\\downarrow\)Perc\. inconsistentwith Levenshtein dist\. 3 \(↓\\downarrow\)Avg\. time / line\(s\) \(↓\\downarrow\)Annotated triples8\.1100%7\.2%2\.6%\-spaCy extraction3\.9100%0%0%0\.013Starling \- base\-0%\-\-0\.206Mistral \- base\-0%\-\-0\.221Llama \- base8\.010%0%0%0\.097Starling \- base \+ few\-shot13\.0100%13\.2%4\.6%0\.217Mistral \- base \+ few\-shot5\.360%28%16%0\.243Llama \- base \+ few\-shot15\.3100%7\.9%0\.6%0\.104Starling \- fine\-tuned \+ few\-shot10\.8100%40\.8%24\.8%0\.220Mistral \- fine\-tuned \+ few\-shot17\.075%50\.4%41\.1%0\.244Llama \- fine\-tuned \+ few\-shot13\.430%29\.8%21\.3%0\.102

### 4\.7Graph analyses

Through the extraction and clustering of key terms in the domain of interest, and assigning the processed and filtered triples to the corresponding key topics, we have built a network of technological topics\. We then perform several analyses on the network to derive useful insights from it\.

First, we identify established related technology clusters by partitioning the network into densely connected sub\-networks through the Louvain community detection method\[blondel2008LouvainMethod\]\. Specifically, we use the implementation in Gephi with a resolution of 0\.85\[gephi\]\. We scale the nodes based on the number of triples corresponding to the topics\. The labels are scaled through the eigenvector centrality, which measures the influence a node has in a connected network\.

To assess technology convergences over time, we consider the topic co\-publication frequency\. Specifically, we use the Jaccard similarity between topics\. LetA​\(t\)A\(t\)andB​\(t\)B\(t\)be the sets of triplets that are related to topicsAAandBBat timett, respectively\. Then, the Jaccard similarity between the topicsAAandBBat timettis defined as

JA,B​\(t\)=\|A​\(t\)∩B​\(t\)\|\|A​\(t\)∪B​\(t\)\|\\displaystyle J\_\{A,B\}\(t\)=\\frac\{\|A\(t\)\\cap B\(t\)\|\}\{\|A\(t\)\\cup B\(t\)\|\}\(7\), The Jaccard similarity will always be between 0 and 1, with a higher similarity indicating a high co\-occurrence between the topics\. A technology convergence will be identifiable by an increase in the Jaccard similarity\.

### 4\.8Implementation

To download the data, preprocess the text and extract the triples, we used 4 NVIDIA A100 Tensor Core GPU’s with 40GB RAM\. For the deployment of the pipeline, at least one GPU with 16GB of VRAM is required to run the LLMs\. Depending on the size of the data, more GPUs may be required for swift processing\. Furthermore, it is recommended to use at least 10 cores for swift downloading and preprocessing of arXiv papers and USPTO patent applications\.

Table 6:Eight triples selected from the paperSimple Recurrence Improves Masked Language Models\[lei2022\], to illustrate the extraction, normalization and filtering of triples\.Raw triple \(LM output\)Normalized tripleOutcomeTransformer; rely; attention mechanismstransformer; rely; attention mechanismsRetainedTransformer architecture; interleave; multiheaded attention blocktransformer architecture; interleave; multiheaded attention blockRetainedTransformer architecture; add; residual connectiontransformer architecture; add; residual connectionRetainedscalar vectors; optimize; during trainingscalar vectors; optimize; trainingRetained\(preposition*during*dropped\)works; demonstrate; superior performanceworks; demonstrate; superior performanceRemoved\(entropy filtering of*works*\)our model; perform; consistentlymodel; perform; consistentlyRemoved\(entropy filtering of*model*\)example; improve; translationexample; improve; translationRemoved\(bookcorpus filtering of*example*model; achieve; average improvementmodel; achieve; average improvementRemoved\(entropy filtering of*model*and*average improvement*\)

## 5Results

Our results are organized into two distinct parts\. In the first part \(Section[5\.1](https://arxiv.org/html/2510.25370#S5.SS1)\) we compare the candidate triple\-extraction methods and select the model that is used throughout the rest of the paper\. The second part \(Sections[5\.2](https://arxiv.org/html/2510.25370#S5.SS2)and[5\.3](https://arxiv.org/html/2510.25370#S5.SS3)\) is applied: we deploy the selected pipeline end\-to\-end on two independent case studies: arXiv preprints \(Section[5\.2](https://arxiv.org/html/2510.25370#S5.SS2)\) and USPTO patent applications \(Section[5\.3](https://arxiv.org/html/2510.25370#S5.SS3)\)\. Here we show that it surfaces both established and emerging convergence patterns and that it generalizes across data sources\.

### 5\.1Selecting the Triple Extraction Method

In this first section, we will discuss the difference in performance between the spaCy and different LLM triple extraction methods, justifying our use of a specific LLM model and adaptation approach down the line\. We are interested in four dimension of triple extraction performance: the number of triples extracted, the format of the generated text, signs of hallucination and the computation time\.

Although more extracted triples potentially mean a more granular technology identification, it is crucial to preserve their quality\. Part of such assessment is evaluation of generative LLMs for hallucination instead of extraction, which we do by mapping the triples back to the reference text\. Specifically, we consider whether the subjects and objects are present in the original text\. We use the Levenshtein distance, which is defined as the number of single\-character edits to change one word into the other\. We note that the evaluation is performed on the raw triples produced by each extraction method, prior to the processing and filtering discussed in Sections[4\.3](https://arxiv.org/html/2510.25370#S4.SS3)and[4\.4](https://arxiv.org/html/2510.25370#S4.SS4)\. This is essential, as the formatting and hallucination verifications require the raw output\. The results, presented in Table[5](https://arxiv.org/html/2510.25370#S4.T5)shows that there are also inconsistencies in the human\-annotated triples\. Through manual inspection, we identify that these inconsistencies in the human\-annotated triples are caused by the stemming of verbs or nouns\. This shows that the Levenshtein distance can cause false positives in the hallucination detection\. Given that this is limited to 7% for the human annotated triples, we accept this ground level of noise\.

Table[5](https://arxiv.org/html/2510.25370#S4.T5)also shows that the spaCy\-based extraction method retrieves, on average, 3\.9 triples per 15 lines\. In contrast, on average, the LLMs extract up to 17 triples per text segment\.

A second aspect of extracted triple quality evaluation is formatting\. In order to automatically parse the LLM output to use the triples for a downstream analysis, we need the formatting to be respected, which we observe to be only occurring for few\-shot learning prompted models\. Additionally, we observe that fine\-tuning the LLM degrades the percentage of correctly formatted generations for Mistral and Llama\. We expect that this is caused by the models being already extensively fine\-tuned and being positioned on a Pareto frontier, with degeneration kicking in with further fine\-tuning\[bai2022constitutional\]\.

Overall, theMeta\-Llama\-3\-8B\-Instructmodel with few\-shot prompting alone performs best, as it produces a large number of triples that are correctly formatted, while showing few to no signs of hallucinations\. The main disadvantage of using a LLM for triple extraction is the computational overhead, as it is 10 times slower than the spaCy\-based triple extraction\. For this reason, we do not consider even larger LLMs\. However, as a sufficiently large number of triples of high quality is essential for downstream performance, we choose to proceed with the Llama model\.

### 5\.2Case study part 1: technology forecasting on LLMs through arXiv e\-prints

#### 5\.2\.1Key term extraction and clustering

From the 278,625 arXiv e\-prints, we extract a total of 20,822,710 triples after post\-processing and filtering, creating to our knowledge the largest unbiased graph of technology\-related semantic triples\. Table[6](https://arxiv.org/html/2510.25370#S4.T6)shows 8 examples of extracted triples, highlighting which ones were retained and which ones were removed in the filtering stage\. Next, following the approach described in Section[4\.6](https://arxiv.org/html/2510.25370#S4.SS6), we extract key terms from those papers related to LLMs\. Here, we consider the same 278,625 papers from between 2018\-2024, from the arXiv categories as specified in Table[1](https://arxiv.org/html/2510.25370#S3.T1)\. Specifically, we choose those papers with an abstract that contains at least one of the termslarge language model,large language modelsorllm\. The extraction results in a total of 53,337 key terms from 4,182 papers\. Next, the key terms are clustered into topics, following the approach that is described in Section[4\.6](https://arxiv.org/html/2510.25370#S4.SS6)\. The following analyses focus on the subjects and objects of the extracted triples, as we found them to be most informative\. Additionally, Appendix[A\.5](https://arxiv.org/html/2510.25370#A1.SS5)provides an analysis of the triple predicates\.

#### 5\.2\.2Domain overview

We start by leveraging the triples to provide a comprehensive overview of current research on large language models \(LLMs\)\. For that, we analyze the 20 most prominent topics in the field, presented in Table[7](https://arxiv.org/html/2510.25370#S5.T7), along with the five most frequently occurring keywords associated with each topic\. According to that table, the most prevalent topics areLLM trainingandConversational AI, highlighting the significant focus on optimizing model performance and enhancing human\-computer interactions\.

Beyond these, several other key areas emerge, includinglanguage understanding,question answering, andretrieval\-augmented generation \(RAG\)\. Research inlanguage understandingfocuses on refining LLMs’ ability to process, interpret, and generate natural language with greater accuracy and contextual awareness\. Question answering remains a critical application of LLMs, driving advancements in knowledge extraction, reasoning, and response generation across various domains\. Additionally, RAG has gained traction as a method for improving factual consistency by integrating external knowledge retrieval with generative models, mitigating issues related to hallucination and enhancing the reliability of generated content\.

To analyse the landscape in more detail, let us consider the size of the subtopics and their corresponding arXiv categories\. Figure[3](https://arxiv.org/html/2510.25370#S5.F3)shows the 15 largest topics and the corresponding arXiv categories from the papers to which the extracted triples belong\. One can see that the most prominent arXiv categories are`cs\.LG`\(machine learning\) and`cs\.CL`\(computation and language\), but that there are also applications in`eess\.AS`\(audio and speech processing\) and`cs\.CR`\(cryptography and security\)\.

Table 7:For each of the 20 largest topics in the field, the 5 most frequent key terms are displayed\. The key terms are determined through the method described in Section[4\.6](https://arxiv.org/html/2510.25370#S4.SS6)\.CategoryTop 5 KeywordsLLM trainingtrained language, training strat, training large, training large language, model pretrainingConversational AIgeneration, generative adversarial, conversation, generating, conversationalDeep Learningdeep learning, deep learning models, art deep, deep learning architectures, art deep learningLanguage Understandingunderstanding, planning, language understand, language understanding, language understandinQuestion Answeringquestion, answering, question answer, question answering, question answerinSemanticsemantic, semantic communication, automatic summarization, image semantic, semantic evaluationNatural Languagenatural language, natural language pro, natural language process, natural language processing, natural language inferenceRetrieval Augmentedretrieval, retrieval augmented, retrieval augmented generation, generative retrieval, retrieval augmentationInstruction Tuninginstruction, instructions, instruction following, instruction gen, based instructionLanguage Processinglanguage processing, prompting large language, grounding language robotic, language modeling pretraining, prompted languageMachine Translationtranslation, machine transla, machine translation, neural machine translation, translation modelText Generationgenerated text, al language generation, text generation, conversational agent, generation modelsVideovideos, video generation, based video, video models, generated videosSummarizationsummarization, narrative, summarization model, summarization models, abstractive textChatGPTchatgpt, based chatbot, chatgpt model, openai chatgpt, chatgpt1Multilingualmultilingual, multilingual bert, multilingual lan, multimodal large, multilingual largeDialoguedialogue, dialogue evaluation, dialogue task, dialogue tasks, dialogue modelingSpeechtext model, speech recognition models, text speech, speech text, based speechN\-Shot Learningshot learning, shot learner, shot learners, zero shot, shot learning methodsSource Codesource code, code data, source code data, code test, source code sequenceTime Seriestime series, time series forecasting, time series prediction, multivariate time series, time series dataset![Refer to caption](https://arxiv.org/html/2510.25370v2/images/tech_trends.png)Figure 2:Multiple line plot showing for each of the 10 most common topics the number of papers in which at least one triple appears for that topic\. The gray dashed line shows the aggregate seasonal component, whereas the colored lines show the trend for each topic\.![Refer to caption](https://arxiv.org/html/2510.25370v2/images/stacked_bar_chart.png)Figure 3:The number of papers for each of the 15 most common topics in the field of LLMs, based on the extracted and grouped key terms\. Each of the bars is divided into the arXiv categories from which the papers originate\.
#### 5\.2\.3Emerging technologies

Next, we consider emerging technologies by performing a time series analysis of the most popular topics\. As publications follow yearly conference cycles, it is necessary to decompose the paper counts into a trend and a seasonal component\. We therefore employ a simple seasonal decomposition based on moving averages\. That is, we model the number of papersNp​\(t\)N\_\{p\}\(t\)at timettfor topicppas

Np​\(t\)=Tp​\(t\)\+Sp​\(t\)\+ep​\(t\),N\_\{p\}\(t\)=T\_\{p\}\(t\)\+S\_\{p\}\(t\)\+e\_\{p\}\(t\),\(8\)
whereTp​\(t\)T\_\{p\}\(t\)is the trend,Sp​\(t\)S\_\{p\}\(t\)is the seasonal component andep​\(t\)e\_\{p\}\(t\)is the residual\. Figure[2](https://arxiv.org/html/2510.25370#S5.F2)shows the trends for each of the ten most common topics, alongside the common seasonal component\. We note that the trends are computed using separate seasonal components for each of the topics, but we display a common seasonal component for ease of interpretation\.

When considering the seasonal component, one can see two spikes a year, one around May\-June and one in October\. These spikes correspond to the biggest conferences for NLP related work: the Association for Computational Linguistics \(ACL\) and Empirical Methods in Natural Language Processing \(EMNLP\)\. The ACL conference is commonly around July, with submissions generally uploaded in the months before\. The EMNLP conference is in November, with submissions uploaded to arXiv beforehand\.

When considering the trends of the various topics around LLMs on Figure[2](https://arxiv.org/html/2510.25370#S5.F2), one can see that, in general, there is a surge in interest for topics in LLMs in early 2023\. This aligns with the release of ChatGPT in November 2022, which caused a surge in public interest in LLMs\. Furthermore, this event seems to have had little effect in the popularity of older technologies such asdeep learningandreinforcement learning\. In contrast, the biggest emerging technologies are related toconversational aiandlanguage understanding\. This shows that the current focal point in the field is the simplified interaction between humans and LLMs\.

For a more fine\-grained analysis, we investigate the trends of individual key terms within the topics\. Again, we distinguish between a seasonal component and the trend\. To identify those key terms with the most interesting trends, we select the 5 key terms for each category that show the largest increase in prevalence between 2018 and 2024\. Figure[4](https://arxiv.org/html/2510.25370#S5.F4)displays the trends for these key terms for each topic, alongside the aggregate seasonal components\. The results show that there are specific research directions that surged over the past years, most notablyevidential deep learning,semantic communication,reinforcement learning with human feedbackandquery rewriting, corresponding to emergent technologies\. Such topics are all driven by advances in the capabilities of LLMs and are likely to continue growing in the near future\.

![Refer to caption](https://arxiv.org/html/2510.25370v2/images/grid_temporal.png)Figure 4:The trends for key sub\-technology terms for the 10 most common emerging technology topics in the LLM space, based on the extracted and grouped key terms\. The gray dashed line shows the aggregate seasonal component that the key terms share, whereas the colored lines show the trend for each key term\.![Refer to caption](https://arxiv.org/html/2510.25370v2/x1.png)\(a\)Summarization and factuality
![Refer to caption](https://arxiv.org/html/2510.25370v2/x2.png)\(b\)Language and image understanding
![Refer to caption](https://arxiv.org/html/2510.25370v2/x3.png)\(c\)Core NLP concepts
![Refer to caption](https://arxiv.org/html/2510.25370v2/x4.png)\(d\)Instruction\-tuning and prompting
![Refer to caption](https://arxiv.org/html/2510.25370v2/x5.png)\(e\)Retrieval augmentation
![Refer to caption](https://arxiv.org/html/2510.25370v2/x6.png)\(f\)Question answering

Figure 5:Clusters of topics composed using the Louvain method for community detection\. Relations in red occur over 70% of the time in 2022 or later, whereas relations in blue occur over 70% of the time in 2021 or earlier\. All other relations are in black\. Node/edge sizes correspond to absolute frequency, and text size to eigenvector network centrality\.![Refer to caption](https://arxiv.org/html/2510.25370v2/x7.png)\(a\)Emerging connections between key topics between 2018 and 2024\.
![Refer to caption](https://arxiv.org/html/2510.25370v2/x8.png)\(b\)Connections between key topics that are losing prevalence between 2018 and 2024\.
![Refer to caption](https://arxiv.org/html/2510.25370v2/x9.png)\(c\)Persistent connections between key topics between 2018 and 2024\.

Figure 6:Emerging \(red\), disappearing \(blue\), and persistent \(black\) relations between technology\-related term clusters based on the arXiv paper data\. Relations in red occur over 70% of the time in 2022 or later, whereas relations in blue occur over 70% of the time in 2021 or earlier\. All other relations are displayed as black\. The node color reflects the term cluster they belong to, the sizes of the nodes and edges reflect the frequency of their appearances, and the size of the text labels reflects the eigenvector network centrality\.
#### 5\.2\.4Technology Convergence analysis

The technology convergence analysis is the heart of our approach, and we perform it through topological analysis of semantic triple graph linking together technology\-designating terms\. We perform a network analysis, where each node represents a key topics and the edges represent the number of triplets connecting the topics, with edges mapping both to the connecting verb and the paper from which the triple was extracted\. The edges are then aggregated, leading to a single weighted undirected edge, with weight proportional to triple occurrence\. We also assign a "color" to the edge, based on whether the occurrence of triples involving the two technologies has been increasing \(red\), decreasing \(blue\) or remained consistent \(black\)\.

##### Technology clusters

We identified established related technology clusters by partitioning the network into densely connected sub\-networks, by using the Louvain community detection method\[blondel2008LouvainMethod\]888For legibility reasons we are unable to provide the entire network\. Figure[5](https://arxiv.org/html/2510.25370#S5.F5)shows the six largest detected clusters, with node color indicating the cluster\. The size of nodes corresponds to the frequency of technology\-related term occurrence, while the size of the term itself \- to the node eigenvector centrality on the entire network\. The eigenvector centrality corresponds how important the node is to connecting the entire network\.

![Refer to caption](https://arxiv.org/html/2510.25370v2/images/jaccard_similarity.png)Figure 7:The Jaccard similarities for the 10 pairs of topics for which the Jaccard similarity had the largest increase between in the Jaccard similarity between 2018 and 2024\. We consider all topics, as identified by the extracted and grouped key terms in for the topic of LLMs\.
##### Technology convergences

In order to identify emerging technological convergences, we then focus on the technology\-related terms connecting different technology clusters\. Figure[6](https://arxiv.org/html/2510.25370#S5.F6)shows three networks, the first presenting emerging technological connections, the second showing waning ones, and the last one showing persisting ones\. We observe an emerging technological convergence betweeninstruction tuningandnatural language\. This highlights the surge in conversational AI that we discussed previously, which was unlocked by instruction\-tuned LLMs, that allowed general public access to conversational LLMs and made them useful as a human assistant emulators\[ouyang2022training\]\. Furthermore, we see pivotal roles forprompt engineeringandretrieval augmentation\. These relatively new branches of research aim at improving the performance of LLMs\. For instance, we see that retrieval augmentation is linked to information retrieval, illustrating its goal of improving the output of LLMs by injecting data from an existing database\.

When we consider the relations that declined in prevalence between 2018 and 2024, we most notably see a link betweenlanguage understandingandfact checking\. This seems to suggest that fact checking is increasingly less related to language understanding\. One hypothesis is that fact checking is now more often performed through stylistic patterns, leveraging ML methods\.

Finally, we consider persistent relations between 2018 and 2024, which often involve core concepts in LLMs\. For instance,question answeringlinks withlanguage understandingandreasoning capabilities\. Similarly,retrieval augmentation, introduced early, has remained stable in its connection toquestion answering\.

##### Temporal trends of convergence

Last, we analyze the temporal aspect of convergence by considering the topic co\-publication frequency\. To analyse which technologies are converging, we study the pairs of topics that have the largest increase in the Jaccard similarity between 2018 and 2024\. The Jaccard similarity measure is defined in equation[7](https://arxiv.org/html/2510.25370#S4.E7)\. Figure[7](https://arxiv.org/html/2510.25370#S5.F7)shows the evolution of the Jaccard similarities for these pairs of topics\. One can see that from the release of ChatGPT in November 2022, several technology convergences emerge rapidly\. However, it can also be observed that in 2024 several of such convergences start to decline again, indicating that the effect was only temporary\. In contrast, one can see that the convergence betweenretrieval augmentedandconversational agentis stronger and is not yet stagnating\. In combination with the network analysis that highlighted these two topics as technological convergences as well, we can designateretrieval augmentationandconversational agentsas emerging transformative technologies in the field of LLMs, as of end 2024, based on research preprints published on arXiv\.

### 5\.3Case study part 2: technology forecasting on LLMs through USPTO patent applications

![Refer to caption](https://arxiv.org/html/2510.25370v2/images/patents_line_plot.png)Figure 8:Multiple line plot showing for each of the 10 most common topics, as identified by the extracted and grouped key terms, the number of patent applications in which at least one triple appears for that topic\.In order to demonstrate the generalization of our approach and its usefulness in the fields with less emphasis on early preprint publication, we use USPTO patent data to provide a different perspective on the technological developments related to LLMs\. From the 9,793 patent applications, our pipeline extracts 3,027,121 triples after post\-processing and filtering\. As before, to perform an informative downstream analysis, we require a set of key terms that are most relevant to the field of LLMs\. For this purpose, we extract key terms from the abstracts of the 9,793 patents, following the same methodology as described in Section[4\.6](https://arxiv.org/html/2510.25370#S4.SS6)\.

We focus on two analyses to showcase the unique perspective that patents give on the technological developments surrounding LLMs compared to arXiv preprints\. As the primary goal of this section is to highlight the generalizability of our method across different data sources, we do not cover all of the perspectives that were taken using the arXiv papers, but merely highlight two approaches\.

#### 5\.3\.1Emerging technologies

We start by performing a technology\-related term usage analysis similar to Section[5\.2\.3](https://arxiv.org/html/2510.25370#S5.SS2.SSS3), where we are looking for terms with recently increasing usage\. Whereas papers uploaded on arXiv follow a strong seasonal pattern due to yearly conference cycles, this pattern is not present for patent applications, due to the lack of yearly conference deadlines\.

Figure[8](https://arxiv.org/html/2510.25370#S5.F8)shows the most popular emerging terms associated with LLMs in patent applications\. Across the topics, there is a surge in LLM\-related technology terms halfway through 2024\. This rise in frequency of term usage tracks the rise of similar terms usage on arXiv at the end of 2022, suggesting the analysis of scientific articles preprints allows an early trend emergence identification\.

We see that LLMs are connected to a variety of fields, for instancecomputer storage,cloud computing,source code, andspeech recognition\. Furthermore, we note that the rise in patent applications has only started recently, and it will be interesting to see if future trends follow the patterns identified in the temporal analysis of arXiv preprints, with some rapidly growing technolgy topics pleateauing shortly\.

![Refer to caption](https://arxiv.org/html/2510.25370v2/x10.png)Figure 9:Network representation of the key topics based on the USPTO patent application data\. Relations in red are occurring over 70% of the time in 2022 or later, whereas relations in blue occur over 70% of the time in 2021 or earlier\. All other relations are displayed as black\. The node color corresponds to the technology term cluster they belong to, the sizes of the nodes and edges correspond to the absolute frequency of their appearances, and the size of the text labels corresponds to the eigenvector network centrality\.
#### 5\.3\.2Network analysis

Just as for arXiv articles, we analyze technological convergence through technology\-related terms semantic graph analysis\. Figure[9](https://arxiv.org/html/2510.25370#S5.F9)displays the network, with node and edge color and size corresponding to the same properties as for arXiv semantinc graph analysis in section[5\.2\.4](https://arxiv.org/html/2510.25370#S5.SS2.SSS4)\.

From this visualization, we can make several observations regarding the developments in industry applications of LLMs\. First, we see that there is a significant overlap between LLMs and self\-attention models, covering multimodal models, visual attention models and speech attention models\. For waning topic connections, we see that there is a decline in image processing connection to blockchain and payment processing, likely corresponding to the deflation of the non\-fungible tokens \(NFT\) bubble; as well as to self\-driving vehicle and unmanned aerial vehicles \(UAVs\), likely corresponding to a switch to other types of sensors and manual controls\. The rupture of the link between speech recognition and speech recognition server seems to indicate the move of speech recognition onto devices\.

When considering emerging topics, one can observe that there is a rise in interest for knowledge graphs and knowledge bases\. Specifically, this seems to be closely related with the instruction\-tuning of LLMs\. Furthermore, one can see a rise in attention to source code manipulation, bridging the classical abstract syntax tree analysis with LLM\-based machine learning methods, corresponding to the trend of LLM\-based code generation and analysis integration into programming workflows\.

Based on the USPTO patent applications, our analysis suggests thatcoding LLMsandon\-device speech recognitionwith self\-attention models are emerging transformative topics in the field of LLMs and self\-attention models, as of end 2024\.

## 6Conclusion

The monitoring and forecasting of emerging transformative technologies is a notoriously important and yet difficult topic in the field of technological forecasting\. The accelerated accumulation and availability of potentially relevant data presents an opportunity for more grounded forecasts, but also poses a challenge due to the scale and presence of noise\. We developed and presented in this paper a pipeline that leverages advances in NLP and big data analysis to perform a technology convergence analysis while addressing the challenges that previously made such analysis at scale impossible\. Specifically, we \(I\) leveraged the recent advances in NLP and recently published benchmarking datasets for to extract information relevant to technological forecasting in an unbiased structured format, and at a scale that has been previously unattainable; \(II\) developed a new approach to account for variation in technologies designation through varying terms allowing their systematic analysis; \(III\) constructed a novel graph\-based technique for identification of emerging technologies and likely transformative emerging technologies for specific fields\.

We then use the pipeline we developed in order to analyze the trends in the LLM\-related technology field based on two different unstructured data sources on a previously unseen scale and granularity\. Specifically, we perform a full\-text analysis of 278,625 arXiv scientific preprints and 9,793 USPTO patent applications, resulting in 53,337 LLM technology\-related terms and 23,849,831 triples\. By leveraging big data techniques on the resulting semantic graph analysis, we were able to identify two emerging transformative technologies as of 2024:retrieval\-augmented generationandconversational agents, which seem to be supported by the data\-grounded agentic AI interest in 2025; as well as two emerging transformative applications:source\-code generation and processingwith LLMs andon\-device speech processing with multimodal models\.

We leave to future research the task of validating the identified transformative technologies, assessing the pipeline’s use with other data sources — such as social media, blogs, and grey literature like reports and white papers — and examining its longer\-term predictive capabilities\.

## References

## Appendix AAppendix

### A\.1Few\-shot prompting

We used few\-shot prompting for the triplet extraction\. We attempted various different prompts and achieved the best results with the setting that is illustrated below\. Theinput textis the text for which we will extract the triples\.

User:You will extract the subject\-predicate\-object triples from the text and return them in the form \[\(subject\_1; predicate\_1; object\_1\), \(subject\_2; predicate\_2; object\_2\), \.\.\., \(subject\_n; predicate\_n; object\_n\)\]\. I want the subjects and objects in the triples to be specific, they cannot be pronouns or generic nouns\. The first text is: Ever since the Turing Test was proposed in the 1950s, humans have explored the mastering of language intelligence by machine\. Language is essentially a complex, intricate system of human expressions governed by grammatical rules\.Example output:\[\(turing test; proposed in; 1950s\), \(human; explored; language intelligence\), \(language; is; system of human expressions\), \(grammatic rule; govern; language\)\]User:The next text is: Currently, LLMs are mainly built upon the Transformer architecture, where multi\-head attention layers are stacked in a very deep neural network\. Existing LLMs adopt similar Transformer architectures and pre\-training objectives \(e\.g\., lan\-guage modeling\) as small language models\.Example output:\[\(large language model; built upon; transformer architecture\), \(multi\-head attention layer; stacked in; deep neural network\), \(large language model; adopt; transformer architecture\)\]User:The next text is: After collecting a large amount of text data, it is essential to preprocess the data for constructing the pre\-training corpus, especially removing noisy, redundant, irrelevant, and potentially toxic data, which may largely affect the capacity and performance of LLMs\. Example output:\[\(preprocessing; remove; noisy data\), \(preprocessing; remove; redundant data\), \(preprocessing; remove; irrelevant data\), \(preprocessing; remove; toxic data\), \(noisy data; affect; performance large language model\), \(redundant data; affect; performance large language model\), \(irrelevant data; affect; performance large language model\), \(toxic data; affect; performance large language model\)\]User:The next text is:input text Assistant:Output

### A\.2Threshold Selection

Here, we provide justification for several threshold selections for filtering and noun stapling, as described in Sections[4\.4](https://arxiv.org/html/2510.25370#S4.SS4)and[4\.5](https://arxiv.org/html/2510.25370#S4.SS5)\. Specifically, we discuss the length threshold, frequency threshold, and the SoftDICE threshold\.

#### A\.2\.1Post\-processing of Triples

##### Subject and Object length\.

We first discuss the choice for removing subjects and objects with more than 6 terms\. We opted for this threshold through manual inspection of the samples\. Specifically, we considered 50 randomly selected entities of lengths 1 through 10\. Here, in Table[8](https://arxiv.org/html/2510.25370#A1.T8)we show a representative sample of 10 entries from length 4 to 8\. This shows that above length 6, entries are never clear entities, and are safe to remove without losing information\. We note that for length 5 and 4, there are also entries that are no clear entities\. However, we choose to be conservative with the threshold choice to not lose information, and also rely on subsequent filtering steps\.

Additionally, Figure[10](https://arxiv.org/html/2510.25370#A1.F10)shows the percentage of entities with a specific number of terms\. We see that most entities have 1 or 2 terms, and we are removing only a small fraction of subjects and objects in this filtering step\. This is also expected, as the few\-shot examples should induce the LLM to not often extract subjects and objects with multiple terms\.

Table 8:Representative subjects/objects by length\. At lengths 4\-6 the spans are still mostly \(compound\) technology entities alongside some noun phrases; at lengths 7\-8 they tend to be clauses that no longer name a single entity\.LengthExamples4deep feedforward neural network; neural machine translation systems; multi\-agent markov decision process; bayesian experimental design algorithm; instance\-wise adaptive adversarial training; user interest\-aware capsule network; machine learning pipeline management; renewable energy sources generation; concentration of particulate matter; number of oracle calls5fully connected deep neural network; semi\-markov conditional random fields model; exponential family variational kalman filter; hierarchical attention multi\-task learning networks; improved phishing spam detection model; other graph contrastive learning methods; balance between efficiency and quality; discrepancy between training and inference; finite amount of training data; model performance and training time6security information and event management system; gumbel tree\-based long short\-term memory network; bounded information rate variational autoencoder algorithm; high resolution image synthesis and generation; computer vision and natural language processing; signal processing and machine learning techniques; competitiveness and effectiveness of proposed method; powerful representation capability of neural networks; detecting misinformation at an early stage; analysis on impact of network parameters7models may be sensitive to few epochs; weights do not move far from initialization; probability of document belonging to each class; disparities extend far beyond our evaluated domains; reason the intended purpose of the code; rewards linked to amount of pixels acquired; distance from data points to decision boundary; effect of number of sampled image patches; quick adaptation to patterns in input prompts; model more sensitive to relative positional information8more training data helps to improve certified accuracy; repairs do not impact large language model performance; approach performs well on synthetic and real\-world datasets; preserving direction of momentum leads to better performance; using convolutional neural networks to solve poisson equation; whether one object is completely inside the other; how many countries are diplomatically related to italy; required for emergency vehicle to arrive at destination; power of each negotiator building on its history; importance of aligning storytelling approach with customer journey![Refer to caption](https://arxiv.org/html/2510.25370v2/images/wordcount.png)Figure 10:Distribution of the subjects and objects versus the number of words they contain\.
##### Word length\.

Second we discuss the choice to remove words with 2 characters or less\. Table[9](https://arxiv.org/html/2510.25370#A1.T9)shows 10 randomly sampled 1\-character, 2\-character and 3\-character terms from the raw triples from the arXiv papers\. We see that the 1\-character and 2\-character terms tend to be noise or stopwords, such asofandin\. For the 3\-character terms, there are also filler terms \(such asandandthe\)\. However, there are also meaningful terms included, such asfit,sumandlow\. Therefore, we choose to be conservative and choose to remove terms with 2 characters or less\. We note that due to incomplete abbreviation resolution, there may still be two character abbreviations that we remove, this illustrates a trade\-off between noise reduction and loss of signal\.

![Refer to caption](https://arxiv.org/html/2510.25370v2/images/charactercount.png)Figure 11:Distribution of the character counts of each of the words in the subjects and objects in the processed but unfiltered triples\.Table 9:Examples of 1\-character, 2\-character and 3\-character subjects and objects from raw triples\. The chosen threshold to remove tokens was set at less than 3 characters, hence 1\-character and 2\-character tokens are removed\.Examples1\-character tokensg, e, k, x, r, \., a, p, %, v2\-character tokensof, in, al, we, et, ci, as, is, on, ar3\-character tokenssum, dsc, low, set, the, and, liu, al\., fit, all
##### Frequency Threshold\.

We now discuss the chosen frequency threshold, where we remove a triple if for either the subject or object all terms appear less than 5 times in the control corpus, as defined in Section[4\.4](https://arxiv.org/html/2510.25370#S4.SS4)\. We take this step to remove entities that are too rare\. After all, we are interested in temporal trends and relations, hence entities need to have a minimum presence\. This filtering step removes 5\.5% of all triples, indicating that most triples contain entities that are relatively present\. To verify that our threshold is not too low, we randomly sample 20 of the removed entities, which are displayed in Table[10](https://arxiv.org/html/2510.25370#A1.T10)\. We see that most removed entities are names, or niche terms such assocratic inquiry\.

Table 10:Twenty randomly sampled subject/object spans removed by the frequency filter \(every term hasfc​c,t<5f\_\{cc,t\}<5in the control corpus\)\. They are predominantly author names and other tokens absent from the general research vocabulary\.Removed subject/object spansadam arvidssonsangwoo chiheon sungwoonggervet romoff batra parikh rabbat pineausicheng wang sabato marco siniscalchi chin\-huisennrich alham fikriyining hongsocratic inquirymichael whalenalex fiannacamontral newtechagnes luhtaruderek xiuyu felix hohne ser\-namchristian hempelmannaalbers haarenbenjamin piwowarskibertossi salimilaying fireplacehuaping zhongiyengar singalmustafa bozdag

#### A\.2\.2Noun Stapling: SoftDICE threshold

The noun\-stapling step discussed in Section[4\.5](https://arxiv.org/html/2510.25370#S4.SS5)merges two compound nouns when their Soft\-DICE similarity \(Eq\.[5](https://arxiv.org/html/2510.25370#S4.E5)\) is at least0\.850\.85\. We chose this threshold to be conservative, so that only surface variants of the same entity are merged and distinct technologies are kept apart\. This subsection shows more background on this choice\.

Table[11](https://arxiv.org/html/2510.25370#A1.T11)lists example pairs at each value\. We find that at values of0\.70\.7and0\.750\.75, the entities can be significantly different, such asdeep neural network watermarkanddeep neural network modularization\. At values of0\.850\.85and above, the compound nouns are highly similar, with only minor variations\. In the end, we opted to put the threshold at0\.850\.85, as for0\.850\.85there are still meaningful differences, for exampleproposed attacksandproposed poisoning attacks\. We note that with this matching, there are still minor variations allowed between paired compound nouns\. This is the balance we aim to strike between grouping identical signal together, while maintaining a high precision and reducing noise\.

Table 11:Example noun pairs at each achievable Soft\-DICE value around the0\.850\.85threshold, we allow for an allowed deviation of0\.020\.02around the target value for each row\. There were no samples with SoftDice values around0\.950\.95\. they are variants of the same entity\.SoftDICE≈0\.7\\approx 0\.7learnable background tokenlearnable class\-specific tokendeferral loss functionvariance\-based loss functioncurrent deep neural networkshard dual deep neural networksstate training datapredictions training datareplaced causal modelcausal imitative modelSoftDICE≈0\.75\\approx 0\.75texture convolutional neural networkboosted convolutional neural networkdeep neural network watermarkdeep neural network modularizationanalogy deep neural networkspoisoned deep neural networksspatiotemporal gaussian process modelgaussian process density modelunknown deep neural networkiterative deep neural networkSoftDICE≈0\.80\\approx 0\.80noise estimatenoise contrastive estimatetext\-generation modelretrieval\-augmented text\-generation modelproposed attacksproposed poisoning attacksunique propertiesunique chemical propertiesmodel accuracy costaccuracy costSoft\-DICE≈0\.85\\approx 0\.85framework multi\-objective bayesian optimizationbayesian multi\-objective optimizationradar waveform selection processradar waveform selectionloss function employedwidely employed loss functionstepwise feature selection lassostepwise feature selectiondense neural networkthree\-layer dense neural networkSoftDICE≈0\.9\\approx 0\.9using standard binary cross\-entropy lossusing binary cross\-entropy lossgated recurrent unit relation\-awarerelation\-aware gated recurrent unit networkadaptive multi\-hierarchical attention modulenovel adaptive multi\-hierarchical attention moduleopen\-source proprietary large languageopen\-source proprietary large language modelsautonomous robot bimanual manipulation behaviorrobot bimanual manipulation behaviorSoftDICE=1\.0=1\.0classification accuracy rateaccuracy classification rateshape pose parameterspose shape parametersvisual linguistic modalitieslinguistic visual modalitiesstudent ensemble modelensemble student modelcharge predictionprediction charge

### A\.3Semantic noun\-stapling

Our work considers syntactic similarity for the noun stapling, as discussed in Section[4\.5](https://arxiv.org/html/2510.25370#S4.SS5)\. However, this misses cases where two terms are semantically similar, but lexically different\. For example, a missed abbreviation, such asllmandlarge language model\. Or, alternatively, two different names for the same concept, such aspromptandinstruction\.

To test whether an embedding\-based semantic normalization layer would help, we embedded the graph vocabulary with four sentence\-embedding models spanning size and training regime,all\-MiniLM\-L6\-v2\(22M\)999https://huggingface\.co/sentence\-transformers/all\-MiniLM\-L6\-v2,all\-mpnet\-base\-v2\(110M\)101010https://huggingface\.co/sentence\-transformers/all\-mpnet\-base\-v2,allenai\-specter\(110M\)\[specter\], andgoogle/embeddinggemma\-300m\(300M\)\[embeddinggemma\]\.

The idea is that through semantic similarity, we would be able to link entities that are semantically identical but syntactically different\. However, there are two concerns\. First, one needs to verify that embedding models can accurately encode the information and identify whether two scientific entities are semantically equal or not, which research has shown to not always be the case\[Wursch2023,alexandria\]\. Second, computing embedding similarities is computationally more costly than syntactic similarities\.

To assess the performance of embedding models, we manually curated three sets of entity pairs\. The first category consists of semantically equal, but syntactically different entities \(e\.g\.data annotationanddata labelling\)\. The second category consists of semantically different, but syntactically close entities \(e\.g\.logistic regressionandlinear regression\)\. The last category consists of entities that are both syntactically and semantically distinct \(e\.g\.convolutional neural networkandsentiment analysis\)\. We refer to these categories asSynonym,DistinctandUnrelated\. For semantic similarity to work, we need to be able to find a threshold where we both attain a low level of false negatives and false positives\. The curated pairs are available in the repository\.

Figure[12](https://arxiv.org/html/2510.25370#A1.F12)shows the cosine similarities for the three categories of pairs, for each embedding model\. We find that none of the embedding model cleanly manages to pair synonyms, while not falsely pairing distinct entities\. We find thatMiniLM\-L6\(22M\) on average separates synonyms and distinct pairs better thanSpecter\(110M\)\. However, it also misses more synonym pairs, whereasSpecter\(110M\) scores almost all synonym pairs above 0\.80\. Due to the computational cost of embedding similarity computation, and the inability of embedding models to cleanly pair entities, our work chooses to use syntactic similarity measures\.

![Refer to caption](https://arxiv.org/html/2510.25370v2/images/semnorm-separability.png)Figure 12:Cosine similarity of synonym, distinct and unrelated pairs for each embedding model\.
### A\.4Topic Modeling Methods

In Section[4\.6](https://arxiv.org/html/2510.25370#S4.SS6), we describe how our work defines topics, to which the extracted triples are assigned later\. To this end, we use KeyBERT, withallenai/specteras the embedding model\. A natural question is how this approach compares to classical topic models such as Latent Dirichlet Allocation \(LDA\)\[lda\]and Non\-negative Matrix Factorization \(NMF\)\[nmf\]\. This appendix provides a comparison across these three methods, on the4,1824\{,\}182LLM\-related abstracts\. We run LDA and NMF twice, once to fit with2020topics and once to fit with5050topics, on unigram, bigram and trigram features, so that they can also surface multi\-word terms\.

Figure[13](https://arxiv.org/html/2510.25370#A1.F13)shows the terms surfaced by each method, with for LDA and NMF 20 topics at the top row, and 50 topics on the bottom row\. We find that LDA and NMF mostly surface generic single terms, such astasksanddata\. In contrast, KeyBERT surfaces more specific entities, such aslarge language modelandquestion answering\. We validate this by randomly selecting 5 papers, and comparing the extracted entities from KeyBERT, LDA and NMF side by side\. The results are displayed in Table[12](https://arxiv.org/html/2510.25370#A1.T12), where we again see that LDA and NMF lean towards single terms, that are relatively generic\. In contrast, KeyBERT manages to extract more specific and fine\-grained entities\. For these reasons, we opt to use KeyBERT in our pipeline for topic identification\.

![Refer to caption](https://arxiv.org/html/2510.25370v2/images/wc_keybert.png)\(a\)KeyBERT
![Refer to caption](https://arxiv.org/html/2510.25370v2/images/wc_lda.png)\(b\)LDA \(20 topics\)
![Refer to caption](https://arxiv.org/html/2510.25370v2/images/wc_nmf.png)\(c\)NMF \(20 topics\)
![Refer to caption](https://arxiv.org/html/2510.25370v2/images/wc_keybert.png)\(d\)KeyBERT
![Refer to caption](https://arxiv.org/html/2510.25370v2/images/wc_lda50.png)\(e\)LDA \(50 topics\)
![Refer to caption](https://arxiv.org/html/2510.25370v2/images/wc_nmf50.png)\(f\)NMF \(50 topics\)

Figure 13:Word clouds of the terms surfaced by each method on the LLM\-related abstracts\. In the top row LDA and NMF at2020topics, in the bottom row at5050topics, with KeyBERT repeated for reference\. KeyBERT surfaces specific, multi\-word technologies \(e\.g\.*visual question answering*,*abstractive summarization*,*instruction tuning*\); LDA and NMF are dominated by broad, generic single words \(*tasks*,*data*,*learning*,*knowledge*\), and remain so even at5050topics\.Table 12:Extracted terms by KeyBERT, LDA and NMF, for five random papers\. For LDA/NMF we use2020topics and unigram–trigram extraction\.UQE: A Query Engine for Unstructured Databases\[uqe\]KeyBERTunstructured data analytics, structured data, query execution, query language uql, query engineLDAdata, retrieval, methods, natural, propose, newNMFdata, retrieval, queries, query, natural language, methodsLLM Agents Improve Semantic Code Search\[llmagents\]KeyBERTsemantic code search, code retrieval, code retrieval systems, enhance code retrievalLDAretrieval, information, context, rag, code, generationNMFretrieval, rag, generation rag, code, informationEvaluating the Logical Reasoning Ability of ChatGPT and GPT\-4\[liu2023\]KeyBERTchatgpt performs significantly, generative pretrained transformer, comprehension natural languageLDAreasoning, tasks, gpt, natural language, benchmarkNMFreasoning, reasoning tasks, logical, gpt, abilityDEE: Dual\-stage Explainable Evaluation Method for Text Generation\[dee\]KeyBERTtext generation, machine generated texts, generated texts, quality text generationLDAtext, generation, evaluation, generated, dataset, systemsNMFtext, evaluation, human, generation, qualityEnd\-to\-End Ontology Learning with Large Language Models\[lo2024\]KeyBERTontology learning, ontology finetuning, target ontology finetuning, partial ontology learningLDAcode, datasets, methods, github, https, learningNMFcode, knowledge, github, https, github com
### A\.5Predicate Analysis

In the case studies discussed in Sections[5\.2](https://arxiv.org/html/2510.25370#S5.SS2)and[5\.3](https://arxiv.org/html/2510.25370#S5.SS3), we focus on the subjects and objects of the triples\. In this appendix, we perform a high\-level analysis on the predicates of the triples, to investigate whether they provide us with additional insights\.

##### Most Frequent Predicates\.

Figure[14](https://arxiv.org/html/2510.25370#A1.F14)shows the most frequent predicate verbs across the2\.12\.1M final graph edges\. The distribution is dominated by generic scientific\-discourse and usage verbs \(*use*,*be*,*write*,*have*,*propose*,*provide*\): the ten most frequent head verbs already account for22\.7%22\.7\\%of all edges and describe how authors write about their work rather than a specific relation between two technologies\.

![Refer to caption](https://arxiv.org/html/2510.25370v2/images/predicate_head_verbs.png)Figure 14:The most frequent predicate head verbs, as a share of all graph edges\. The distribution is dominated by generic discourse and usage verbs\.
##### Predicate Categorization\.

To characterize the predicate space, we assign each head verb to its lexicographic class in WordNet111111https://wordnet\.princeton\.edu/documentation/lexnames5wn\[wordnet\]\. Because many common verbs are highly polysemous, and their single most frequent WordNet sense can be misleading \(for example the dominant sense of*play*would place “plays a role” in*competition*\), we use a sense\-agreement filter: a verb is assigned to a class only when a majority of its verb senses share that lexname, and is otherwise left*ambiguous*\. This labels59\.9%59\.9\\%of edges \(Table[13](https://arxiv.org/html/2510.25370#A1.T13)\)\. The predicates that express a directed technical relation, such as*change*,*creation*, and*competition*, are a minority\. The larger classes \(*stative*,*cognition*,*communication*\) are reporting and description verbs\. This aligns with Figure[14](https://arxiv.org/html/2510.25370#A1.F14), which shows that the most common verbs are generic\.

However, for a qualitative evaluation, one can still consider the smaller informative predicate classes, such as*competition*\. Here, for example, we see that it is dominated by the predicate*outperform*\(21,20621\{,\}206edges\)\. These predicates encode*technology succession*: a triple\(X,outperform,Y\)\(\\textrm\{X\},\\textrm\{outperform\},\\textrm\{Y\}\)states that technology X supersedes technology Y on some task\. Aggregating the objects of*outperform*therefore traces which established technologies are most frequently reported as surpassed \(Table[14](https://arxiv.org/html/2510.25370#A1.T14)\)\. The result is intuitive: recent, strong baselines such as*ChatGPT*,*multilingual BERT*,*in\-context learning*,*fine\-tuning*and*federated learning*are the technologies most often outperformed\. Individual edges are equally interpretable, for example*deep learning*→\\rightarrow*traditional machine learning\-based methods*, and*structured prompt tuning*→\\rightarrow*standard prompt tuning*\.

Table 13:Predicate head verbs of the graph edges, categorized by WordNet lexicographic class with a sense\-agreement filter\. Ambiguous covers highly polysemous verbs with no majority class\.WordNet classEdges%ambiguous847,08840\.1creation215,40310\.2stative211,61610\.0cognition182,4638\.6change165,0677\.8consumption149,7857\.1communication141,0276\.7social77,2743\.7possession37,8011\.8competition31,3791\.5contact26,4521\.3perception18,0230\.9other \(motion, body, emotion, weather\)9,0800\.4Table 14:Top\-5 technologies on each side of the predicate*outperform*, by year\. The lists track the shift from deep learning, to BERT and prompt/fine\-tuning, to ChatGPT and retrieval\-augmented generation\.Winner \(X outperforms⋅\\cdot\)Surpassed \(⋅\\cdotoutperformed by X\)2018deep learningsingle\-task learninggated path planning networkvalue iteration networkbiomedical translation modelmulti\-task learningself\-paced partial\-label learningonline learning algorithmssemantically conditioned variational autoencoderdeep learning methods2020deep learningmultilingual bertmultilingual bertdeep learning modelsfederated learningmultilingual modelmulti\-task learningmodel\-agnostic meta\-learninglearning to retrievemultilingual models2022prompt tuningfine\-tuningmultilingual bertmultilingual bertfederated learningdeep learning modelsdeep learning modelsmultilingual modelsdeep learningstrong baselines2024chatgptchatgptfederated learningin\-context learningin\-context learningzero\-shotretrieval\-augmented generationdeep learning modelsdeep learningzero\-shot prompting

Similar Articles