Multi-Functional Embedding Models for Funder Name Disambiguation in Scientific Publication Records

arXiv cs.CL Papers

Summary

This paper presents a framework for funder name disambiguation in scientific publications using multi-functional embedding models trained with multi-task learning, achieving over 0.90 accuracy and outperforming general-purpose LLMs like GPT-5.2.

arXiv:2609.09984v1 Announce Type: new Abstract: Understanding the historical allocation and distribution of research funding advances our knowledge of how scientific research is supported across fields, institutions, and regions. However, large-scale analyses are hindered by the lack of comprehensive funder name disambiguation solutions, as funder names often exhibit spelling variations, translations, abbreviations, and inconsistent levels of granularity. In this paper, we present a framework for developing multilingual, multi-functional funder name disambiguation models and demonstrate its application to research publications in biodiversity conservation. To construct a training dataset, we integrated the Research Organization Registry (ROR), which provides unique identifiers for research organizations, with two publication datasets: the Web of Science (WoS) and the Crossref Open Funder Registry (OFR). We used multi-task learning with Contrastive Loss and Multiple Negatives Ranking Loss to fine-tune three open-weight embedding models from the Sentence Transformer, Gemma, and Qwen3 families. The best-performing models achieved accuracy above 0.90 when matching WoS funder names to ROR identifiers, outperforming general-purpose LLMs, including GPT-5.2, Claude-Sonnet-4.6, and Gemini-2.5-Flash, by more than 0.1. For funder names not indexed in ROR, we constructed a similarity network among funder names and identified clusters within it. Finally, we analyzed the disambiguation results and highlighted challenges arising from limited knowledge of smaller funders and funders from non-English-speaking countries. This work provides a reusable framework for funder name disambiguation with potential applicability across different model architectures and datasets, featuring cost-effective training data creation and multi-task learning and disambiguation.
Original Article
View Cached Full Text

Cached at: 09/10/26, 08:17 AM

# Abstract
Source: [https://arxiv.org/html/2609.09984](https://arxiv.org/html/2609.09984)
Multi\-Functional Embedding Models for Funder Name Disambiguation in Scientific Publication Records111Journal of Information ScienceFirst submitted: 7 September 2025Accepted: 7 September 2026[View journal version](https://doi.org/10.1177/01655515261490074)Corresponding author:Kanyao HanEmail:[kanyaoh2@gmail\.com](mailto:[email protected])

Kanyao Han1, Zhiwen You1, Jinseok Kim2and Jana Diesner1,3,4,5

1University of Illinois at Urbana\-Champaign, Champaign, USA

2University of Michigan, Ann Arbor, USA

3Technical University of Munich, Munich, Germany

4Munich Center for Machine Learning, Munich, Germany

5Munich Data Science Institute, Munich, Germany

Understanding the historical allocation and distribution of research funding advances our knowledge of how scientific research is supported across fields, institutions, and regions\. However, large\-scale analyses are hindered by the lack of comprehensive funder name disambiguation solutions, as funder names often exhibit spelling variations, translations, abbreviations, and inconsistent levels of granularity\. In this paper, we present a framework for developing multilingual, multi\-functional funder name disambiguation models and demonstrate its application to research publications in biodiversity conservation\. To construct a training dataset, we integrated the Research Organization Registry \(ROR\), which provides unique identifiers for research organizations, with two publication datasets: the Web of Science \(WoS\) and the Crossref Open Funder Registry \(OFR\)\. We used multi\-task learning with Contrastive Loss and Multiple Negatives Ranking Loss to fine\-tune three open\-weight embedding models from the Sentence Transformer, Gemma, and Qwen3 families\. The best\-performing models achieved accuracy above 0\.90 when matching WoS funder names to ROR identifiers, outperforming general\-purpose LLMs, including GPT\-5\.2, Claude\-Sonnet\-4\.6, and Gemini\-2\.5\-Flash, by more than 0\.1\. For funder names not indexed in ROR, we constructed a similarity network among funder names and identified clusters within it\. Finally, we analyzed the disambiguation results and highlighted challenges arising from limited knowledge of smaller funders and funders from non\-English\-speaking countries\. This work provides a reusable framework for funder name disambiguation with potential applicability across different model architectures and datasets, featuring cost\-effective training data creation and multi\-task learning and disambiguation\.

Keywords

funder name disambiguation; publication records; biodiversity conservation; multi\-task learning; embedding models

## 1Introduction

Research funding is essential to advance science and scholarship\[[46](https://arxiv.org/html/2609.09984#bib.bib46),[54](https://arxiv.org/html/2609.09984#bib.bib54)\]\. Analyses of funding allocation to individuals, sociodemographic groups, fields, and organizations can improve our understanding of issues in funding acquisition or provision\. A considerable body of research has examined funding allocation practices\[[8](https://arxiv.org/html/2609.09984#bib.bib8),[48](https://arxiv.org/html/2609.09984#bib.bib48),[45](https://arxiv.org/html/2609.09984#bib.bib45),[53](https://arxiv.org/html/2609.09984#bib.bib53)\]\. Much of this research examines the funding allocation practices of one or a few prestigious funders, such as the National Science Foundation \(NSF\)\[[36](https://arxiv.org/html/2609.09984#bib.bib36)\]and the National Institutes of Health \(NIH\)\[[19](https://arxiv.org/html/2609.09984#bib.bib19),[3](https://arxiv.org/html/2609.09984#bib.bib3)\]in the United States, the National Natural Science Foundation of China \(NSFC\)\[[53](https://arxiv.org/html/2609.09984#bib.bib53)\], or the Natural Sciences and Engineering Research Council of Canada \(NSERC\)\[[12](https://arxiv.org/html/2609.09984#bib.bib12)\]\. For instance, Zhi and Meng\[[53](https://arxiv.org/html/2609.09984#bib.bib53)\]analyzed how the NSFC distributed its funds across institutions, cities, and life science fields, uncovering significant disparities\. Despite the large body of literature, large\-scale comparative analyses across multiple funders and countries remain scarce, primarily due to the lack of accessible, comprehensive, and high\-quality funding data\[[4](https://arxiv.org/html/2609.09984#bib.bib4)\]\.

One major obstacle to funding data analysis is the lack of comprehensive mappings that link funder name occurrences, including variations in spelling, translations, and abbreviations, to standardized records of unique funder identifiers\[[27](https://arxiv.org/html/2609.09984#bib.bib27),[44](https://arxiv.org/html/2609.09984#bib.bib44)\]\. Many bibliometric data providers or sources, such as the Web of Science \(WoS\), Scopus, and PubMed, record funder names for some of the papers they index\. However, the coverage and quality of these funder name records vary across data sources\[[26](https://arxiv.org/html/2609.09984#bib.bib26),[29](https://arxiv.org/html/2609.09984#bib.bib29),[34](https://arxiv.org/html/2609.09984#bib.bib34)\]\. For example, prior studies\[[26](https://arxiv.org/html/2609.09984#bib.bib26),[29](https://arxiv.org/html/2609.09984#bib.bib29)\]found that WoS has the highest coverage of papers with funder name records \(i\.e\., a list of funder names per paper\)\. Funder names indexed in WoS are mainly extracted from the acknowledgment sections of papers\[[35](https://arxiv.org/html/2609.09984#bib.bib35)\], such that the same funder might be referred to by different names\. For instance, the National Science Foundation, a major science funder based in the United States, may be recorded as “NSF”, “National Science Foundation”, “the U\.S\. National Science Foundation”, among other variations\. This issue is further complicated by translations, name changes, misspellings, and the level of resolution for funders \(e\.g\., NSF Graduate Research Fellowship\)\. Recently, academic communities have been curating cleaner versions of funder names by integrating publishing industry data and using crowdsourcing approaches, such as through the Crossref Open Funder Registry \(OFR\)\. The OFR, originally contributed by Elsevier and now maintained by Crossref, allows authors to report funder names of their papers and select a unique OFR funder ID for each reported funder name\. Existing entries are also reviewed to ensure accuracy\[[17](https://arxiv.org/html/2609.09984#bib.bib17)\]\. However, OFR has lower coverage of papers with funder records, for example, in the domain of biodiversity conservation\. Specifically, we found that 122,508 papers categorized as biodiversity conservation have been published since 1900 in the WoS corpus\. Of those, 52,760 papers have funder name records in the WoS data\. In the OFR data, the number of papers with funder records is 20,793\.

To address the gap between the lack of high\-quality, high\-coverage funder records and the data needed for funding\-allocation studies, we leverage multiple funder data sources to develop a funder disambiguation model for biodiversity conservation\. Because research in this domain is supported by organizations and programs worldwide across both the natural and social sciences, it captures much of the complexity of funder naming and provides a useful testbed for a broadly applicable model\.

Name disambiguation is an established NLP task, with a large body of literature dedicated to designing and discussing methods for resolving named entities\[[1](https://arxiv.org/html/2609.09984#bib.bib1),[6](https://arxiv.org/html/2609.09984#bib.bib6)\]to advance studies in the science of science\[[11](https://arxiv.org/html/2609.09984#bib.bib11),[23](https://arxiv.org/html/2609.09984#bib.bib25),[25](https://arxiv.org/html/2609.09984#bib.bib24),[24](https://arxiv.org/html/2609.09984#bib.bib23),[30](https://arxiv.org/html/2609.09984#bib.bib30)\], among other fields\. Although much is known about the impact of name ambiguity\[[25](https://arxiv.org/html/2609.09984#bib.bib24)\]and how to disambiguate author names\[[39](https://arxiv.org/html/2609.09984#bib.bib39),[51](https://arxiv.org/html/2609.09984#bib.bib51)\], little attention has been paid to the disambiguation of funder names\. In addition to studies related to author name disambiguation, there are a few prior studies on institution name disambiguation\[[2](https://arxiv.org/html/2609.09984#bib.bib2),[21](https://arxiv.org/html/2609.09984#bib.bib21),[40](https://arxiv.org/html/2609.09984#bib.bib40),[41](https://arxiv.org/html/2609.09984#bib.bib41)\]\. However, these studies on institution name disambiguation focus on author affiliation names rather than funder names\. Specifically, although there is a large overlap between author affiliations and funders as the main task for both is to disambiguate certain entity names222This paper uses three related terms:organization,institution, andfunder\.Organizationrefers broadly to an organized entity, such as a research organization, government agency, or company, whileinstitution, especially in bibliometric research, is commonly used to refer to an academic or research organization with which an author of a paper is affiliated\.Funderrefers to an entity reported as providing research funding and does not necessarily correspond to an organization or institution, as researchers may also report funding programs or even individuals as funders\., funder names include more non\-research organizations \(such as governmental departments, non\-governmental organizations, and small businesses\) because authors of scholarly papers are mainly from research organizations\. Furthermore, most of these prior studies\[[2](https://arxiv.org/html/2609.09984#bib.bib2),[21](https://arxiv.org/html/2609.09984#bib.bib21)\]tried to identify shared components in institution names, including but not limited to characters, words, and n\-grams, as well as the sequence of these components\. These approaches often fail when a name disambiguation task requires external knowledge\. For example, the abbreviated name “MacArthur Foundation” is often used to refer to the “John D\. and Catherine T\. MacArthur Foundation” in the U\.S\., rather than the “Ellen MacArthur Foundation”, a distinct organization based in the U\.K\. We need such external and contextualizing knowledge to identify which funder an abbreviated name such as “MacArthur Foundation” refers to rather than relying solely on shared components\. Therefore, fine\-tuning pre\-trained models using funder name data offers an alternative approach that leverages knowledge acquired through large\-scale pre\-training while incorporating domain\-specific knowledge during fine\-tuning\[[38](https://arxiv.org/html/2609.09984#bib.bib38)\]\.

Figure 1:Overall funder name disambiguation workflowFigure[1](https://arxiv.org/html/2609.09984#S1.F1)presents our workflow for model fine\-tuning and the overall funder name disambiguation process\. The workflow begins with the creation of a training dataset to address the limited availability of labeled funder\-name pairs\. Fine\-tuning state\-of\-the\-art embedding models typically requires training data in the format of \[text1, text2, label\]\[[37](https://arxiv.org/html/2609.09984#bib.bib37)\], where each pair of funder names is accompanied by a label indicating whether the two name strings refer to the same funder\. To construct such a dataset, we identified papers that are indexed by both WoS and OFR and contain funder name records\. We constructed pairs of funder names per paper \(i\.e\., text1 = funder name in WoS, text2 = funder name in OFR\) that refer to the same funder but are recorded differently in the two sources, based on similarity scores computed by a pre\-trained Sentence Transformer \(ST\) model\[[37](https://arxiv.org/html/2609.09984#bib.bib37)\]\. We then replaced the names derived from OFR in each pair with the funder name from the Research Organization Registry \(ROR\), a dataset that maps name variants of funders to shared unique funder IDs with OFR, for data augmentation\. This data augmentation step ensures that the model learns as many name variants of the same funder as possible\. We used the resulting dataset to fine\-tune the pre\-trained ST\[[37](https://arxiv.org/html/2609.09984#bib.bib37)\], Gemma\[[43](https://arxiv.org/html/2609.09984#bib.bib43)\], and Qwen3\[[52](https://arxiv.org/html/2609.09984#bib.bib52)\]embedding models using two loss functions: Contrastive Loss, which classifies whether two funder names refer to the same funder or not, and Multiple Negatives Ranking Loss, which ranks candidate names based on their likelihood of referring to the same funder\. Among the funders with unique funder IDs documented in the ROR data, the fine\-tuned ST and Gemma embedding models achieved accuracies of 0\.98 for names occurring in ROR and over 0\.91 for names occurring in WoS, outperforming GPT\-5\.2333[https://openai\.com/index/introducing\-gpt\-5\-2/](https://openai.com/index/introducing-gpt-5-2/), Claude\-Sonnet\-4\.6444[https://www\.anthropic\.com/news/claude\-sonnet\-4\-6](https://www.anthropic.com/news/claude-sonnet-4-6), and Gemini\-2\.5\-Flash555[https://ai\.google\.dev/gemini\-api/docs/models/gemini\-2\.5\-flash](https://ai.google.dev/gemini-api/docs/models/gemini-2.5-flash)by more than 0\.1 \(p<\.001p<\.001for all pairwise comparisons based on two\-sided McNemar tests\)\.

Using the fine\-tuned ST model \(with an F1 score of 0\.92\) for classification, we identified approximately 37\.4% of funder names in WoS that were not indexed in ROR data and thus cannot be assigned an existing unique funder ID\. To address this gap, we constructed a similarity network by linking highly similar funder names using the similarity scores computed by our disambiguation model\. We then applied the Louvain algorithm\[[5](https://arxiv.org/html/2609.09984#bib.bib5)\]to cluster these funder names to identify and disambiguate major funders in biodiversity conservation research, particularly those from non\-English\-speaking countries, such as the “National Basic Research Program \(973 Program\)\.” This example also highlights linguistic differences in how the term “funder” is defined and referenced\. For instance, in China, the term “program” is often used to refer to funders, even when the official name of the organization should be, for example, the “Office of the 973 Program\.”

With this work, we make four contributions:

- •First, we introduce a comprehensive and reusable framework for funder name disambiguation, covering training data creation, model fine\-tuning, evaluation, and application\. The framework supports multiple disambiguation tasks, including determining whether two names refer to the same funder or organization, matching funder names against a list of candidates, and clustering funder names\. It is also compatible with diverse model architectures, including Sentence Transformer, Gemma, and Qwen\.
- •Second, through extensive evaluation, we demonstrate that fine\-tuned embedding models using our framework substantially improve funder name disambiguation compared with pre\-trained embedding models and generative LLMs\. Although the models are trained using data from biodiversity conservation publications, the training data include funders operating across research domains, suggesting the potential applicability of the approach beyond biodiversity conservation\.
- •Third, we construct a large\-scale disambiguated funder dataset for biodiversity conservation publications by integrating WoS funder records with publicly available organizational information\. Due to licensing restrictions on the underlying WoS data, the disambiguated dataset and fine\-tuned embedding models cannot be publicly released666The disambiguated dataset and fine\-tuned embedding models may be shared with researchers who have appropriate access to the underlying licensed WoS data\. Researchers with access to the underlying WoS data may contact the authors and Clarivate to inquire about access to the disambiguated dataset and fine\-tuned models\.\. To facilitate reproducibility, we provide the publication identifiers \(WoS IDs, DOIs, and paper titles\), code, and documentation needed to reconstruct the dataset and reproduce the model training \(see Section[8](https://arxiv.org/html/2609.09984#S8)\)\.
- •Fourth, we characterize and detail the challenges in funder name disambiguation, such as the lack of comprehensive funder records from non\-English\-speaking countries and smaller funders\. These findings highlight limitations in existing funder registries and bibliometric data that should be considered in large\-scale funding analyses\.

## 2Related Work

Funder name disambiguation is a special case of name disambiguation and is most similar to the task of institution \(i\.e\., author affiliation\) disambiguation in scholarly publications\. We refer to the latter task as institution name disambiguation in this paper\. An example institution name is “Department of Statistics, Virginia Tech, Blacksburg, USA, address@vt\.edu\.”

There are generally two approaches to institution name disambiguation\[[18](https://arxiv.org/html/2609.09984#bib.bib18),[40](https://arxiv.org/html/2609.09984#bib.bib40)\]: 1\) matching names from a given dataset to a dictionary in which each institution is assigned a unique identifier, and 2\) grouping names from a given dataset into multiple clusters, so that each cluster represents a single institution\. A common challenge in both approaches is determining whether two names refer to the same institution\. Similar names may refer to different institutions, resulting in false\-positive matches, whereas dissimilar names may refer to the same institution, resulting in false\-negative matches\. Therefore, effective disambiguation requires similarity measures that capture institutional identity beyond superficial name similarity\.

A popular method for similarity measurement is to identify shared components \(e\.g\., n\-grams or characters in names, locations such as cities, states, and countries, as well as components in email addresses\) within institution names\[[2](https://arxiv.org/html/2609.09984#bib.bib2),[21](https://arxiv.org/html/2609.09984#bib.bib21),[40](https://arxiv.org/html/2609.09984#bib.bib40)\]\. For instance, Ancona et al\.\[[2](https://arxiv.org/html/2609.09984#bib.bib2)\]computed the number of common words between name pairs as well as consecutive common characters, which they then used as features for disambiguation\. Various auxiliary lexical resources, such as dictionaries and knowledge bases, have also been leveraged to normalize and identify components in institution names\. Examples include dictionaries in the Geoworldmap database for country mapping777[http://www\.geobytes\.com/freeservices\.htm](http://www.geobytes.com/freeservices.htm)\[[21](https://arxiv.org/html/2609.09984#bib.bib21)\]and the Authority File for Affiliations, an affiliation knowledge base containing 113,700 affiliation concepts and approximately 583,700 affiliation names\[[42](https://arxiv.org/html/2609.09984#bib.bib42)\]\. The key innovation of these studies typically lies in the rules or algorithms used for similarity measurement between names based on components\. Weighting components based on human\-crafted rules has been used to measure the similarity of names\[[31](https://arxiv.org/html/2609.09984#bib.bib31),[33](https://arxiv.org/html/2609.09984#bib.bib33)\]\. Specifically, the frequency of each component appearing in institution names is calculated, and each component is weighted based on a set of human\-crafted rules\. The final weighted scores of institution names are used to measure the similarity between names: Institution names with similar weighted scores are considered similar\. In addition to simply weighting occurrences of components, algorithms have also been used to compute edit distance, which represents the minimum number of operations required to transform one name into another based on component insertion, deletion, and substitution\[[18](https://arxiv.org/html/2609.09984#bib.bib18)\]\. Normalized compression distance \(NCD\) has also been introduced, based on the idea that compressing two names together should result in a smaller size than compressing them separately and adding the results\[[20](https://arxiv.org/html/2609.09984#bib.bib20)\]\. A shorter distance indicates higher similarity\.

Based on the assumption that an author often publishes multiple papers under the same institution, Huang et al\.\[[18](https://arxiv.org/html/2609.09984#bib.bib18)\]leveraged author information to improve institution name disambiguation\. They created an author–institution table where each entry lists an author’s name alongside the corresponding institution names from their publications\. They then applied a series of human\-crafted rules to this table, based on shared words and edit distances as well as author names, to group institution names into clusters\. They tested this multi\-method approach, which integrates most of the methods described above, on data for institution names collected from the Web of Science for the domains of mathematics, computer science, psychology, and economics\. They achieved precision scores ranging from 0\.84 to 0\.94 and recall scores from 0\.50 to 0\.87 across domains, with lower performance in social sciences and higher performance in natural sciences\. We speculate that this discrepancy may arise from a broader range of funders in social science research, making funder disambiguation more challenging\.

More recently, Dalsgaard et al\.\[[10](https://arxiv.org/html/2609.09984#bib.bib10)\]developed a large\-scale pipeline for linking funder names in WoS to standardized organizations in OpenAlex and ROR\. Their approach combines lexical normalization, similarity\-based clustering, rule\-based matching, named entity recognition, and manual validation\. Among 7\.4 million unique funder names, 1\.9 million were assigned at least one potential match, covering 72% of all funder mentions in WoS \(i\.e\., a unique funder name may be mentioned repeatedly across publications\), while approximately 74% of unique funder names remained unmatched\. These results demonstrate the effectiveness of large\-scale linkage for frequently occurring funders while also highlighting the continuing challenge of disambiguating long\-tail funder names\.

Shared component identification, human\-crafted rules, and curated knowledge bases can be effective for certain tasks, but the rules used for institution name disambiguation may not generalize to funder name disambiguation for several reasons\. First, institution names are usually followed by well\-documented or highly structured informative components, such as locations, zip codes, and domain names from email addresses\[[14](https://arxiv.org/html/2609.09984#bib.bib14)\]\. Second, the level of resolution in funder names can also be more diverse than that in institution names\. For example, the name of a research institute typically includes the name of the institution and its sub\-component \(such as a school or department\)\. In contrast, a funder name might also include the name of a funding program or project\. Third, funder names tend to have broader coverage, including more non\-research organizations \(such as a variety of small businesses and companies\)\. These differences suggest that funder name disambiguation may require more complex rules and extensive knowledge, especially when handling large\-scale data\.

With the advancement of large pre\-trained models and the growing prominence of fine\-tuning techniques in NLP, the potential for using these approaches for funder name disambiguation remains underexplored\. Accordingly, we focus on task\-specific fine\-tuning of pre\-trained models for biodiversity conservation publications\.

## 3Data

This study uses three datasets produced by prior efforts to curate funder information\.

- •The Web of Science Core Collection \(WoS\)888[https://clarivate\.com/products/scientific\-and\-academic\-research/research\-discovery\-and\-workflow\-solutions/webofscience\-platform/web\-of\-science\-core\-collection/](https://clarivate.com/products/scientific-and-academic-research/research-discovery-and-workflow-solutions/webofscience-platform/web-of-science-core-collection/)consists of publication records with bibliometric meta\-information and funder information\. It has recorded funding information since 2008\. Funding\-related data from 2008 onward are generally considered more complete compared to other data such as Scopus and PubMed\[[26](https://arxiv.org/html/2609.09984#bib.bib26),[29](https://arxiv.org/html/2609.09984#bib.bib29)\]\. As of February 3, 2022, 122,508 conservation publications were categorized as biodiversity\-related papers by the WoS corpus, all of which we downloaded\. Biodiversity conservation publications offer three advantages for studying funder name disambiguation, particularly regarding the diversity of funders\. First, this domain spans both natural and social sciences as work in this domain is supported by funders from a wide range of fields\. Moreover, as noted above, previous research has shown that disambiguating institution names is more challenging in the social sciences than in the natural sciences\[[18](https://arxiv.org/html/2609.09984#bib.bib18)\], such that using biodiversity conservation as a domain enhances the generalizability of our work\. Second, conservation research involves geographically diverse funding flows\. For example, between 2015 and 2022, Africa, Asia, and Latin America and the Caribbean received 30%, 21%, and 17% of biodiversity\-related official development finance, respectively\[[32](https://arxiv.org/html/2609.09984#bib.bib32)\]\. Biodiversity conservation also draws on diverse domestic and international funding sources\[[9](https://arxiv.org/html/2609.09984#bib.bib9)\], extending the funder landscape beyond major funders in the United States, China, and the European Union\. Third, the geographic diversity of funders results in funder names appearing across multiple languages, requiring disambiguation methods to account for multilingual name variations\. Specifically, we identified 112 languages in our dataset using fast\-langdetect999[https://github\.com/LlmKira/fast\-langdetect](https://github.com/LlmKira/fast-langdetect)\[[22](https://arxiv.org/html/2609.09984#bib.bib22)\]\. Approximately 85\.6% of the funder names are in English, followed by Portuguese \(5\.0%\), Spanish \(3\.8%\), French \(1\.7%\), and German \(1\.2%\), while each of the remaining languages accounts for less than 1% of the funder names\. This multilingual distribution introduces additional challenges, including language\-specific naming conventions, spelling variations, and transliteration differences\. Taken together, the disciplinary diversity, broad global distribution of funders, and multilingual nature of the corpus create a particularly challenging benchmark for funder name disambiguation\. A model that performs well on this dataset is therefore expected to generalize more effectively to other heterogeneous, real\-world bibliographic data\.
- •The Crossref Open Funder Registry \(OFR\)101010[https://www\.crossref\.org/services/funder\-registry/](https://www.crossref.org/services/funder-registry/)\[[17](https://arxiv.org/html/2609.09984#bib.bib17)\]is a recent effort to curate funding information\. Authors of scholarly publications who use the OFR system can find the unique IDs for the funders they would like to acknowledge, standardize the metadata of their publications, and deposit the standardized metadata in OFR\. Compared with WoS, OFR has more standardized funder name records\. However, we found that the OFR dataset does not include a significant portion of funder records that are indexed in WoS as mentioned above\. In short, WoS provides more comprehensive coverage of funder records than OFR, but its funder names are less standardized and contain more variations: typos \(e\.g\., U\.S National Science Foundation\), name variants \(e\.g\., National Council for Science and Technology Consejo Nacional de Ciencia y Tecnologia\), and different levels of resolution for one funder \(e\.g\., Stanford Medicine, which belongs to Stanford University\)\. Thus, a disambiguation effort is needed for more accurate funding analysis\. ![Refer to caption](https://arxiv.org/html/2609.09984v1/ROR.png)Figure 2:Example of ROR data
- •The Research Organization Registry \(ROR\)111111[https://ror\.org/](https://ror.org/)\[[13](https://arxiv.org/html/2609.09984#bib.bib13),[27](https://arxiv.org/html/2609.09984#bib.bib27)\]was launched in 2019 and currently indexes more than 102,000 organizations\. Each organization in ROR has a unique ID, a primary name, an associated set of name variants, and meta information such as organization type\. Moreover, the inclusion of funder identifiers in other databases \(see OTHER IDENTIFIERS in Fig\.[2](https://arxiv.org/html/2609.09984#S3.F2)\) helps to establish connections between them\. This dataset is a dictionary for organizational information and does not contain bibliometric metadata\.

## 4Methodology

### 4\.1Detailed Workflow

To obtain cleaner and more comprehensive funding information for analysis, we use two complementary approaches\. For funders indexed in well\-curated organization registries with unique identifiers \(e\.g\., ROR\), we map WoS funder names to their corresponding records based on name similarity\. For funders not indexed in ROR, we cluster funder names based on their similarity, with each cluster representing a distinct funder\. One strategy for mapping and clustering funder names is to directly use Sentence Transformer \(ST\) models, which are a widely used for semantic search due to its stable performance for tasks based on textual similarity calculation\[[37](https://arxiv.org/html/2609.09984#bib.bib37)\]\. ST models generate text embeddings that capture semantic similarity between texts\. The original ST model was fine\-tuned on Natural Language Inference \(NLI\) data by combining the SNLI\[[7](https://arxiv.org/html/2609.09984#bib.bib7)\]and MultiNLI\[[47](https://arxiv.org/html/2609.09984#bib.bib47)\]datasets, totaling approximately one million manually labeled sentence pairs \(e\.g\., “A smiling costumed woman is holding an umbrella” vs\. “A happy woman in a fairy costume holds an umbrella”\)\. A variety of ST models have been subsequently trained or fine\-tuned on broader and more diverse datasets, commonly with pairs of texts that capture different types of semantic relationships\. These models compute the cosine similarity between the embeddings of two candidate strings\. Other embedding models with similar functionality but different architectures have also emerged in recent years, including those built upon the widely adopted Qwen3\[[52](https://arxiv.org/html/2609.09984#bib.bib52)\]and Gemma\[[43](https://arxiv.org/html/2609.09984#bib.bib43)\]families\. Although these foundation models demonstrate strong general\-purpose matching capabilities, they are not tailored to the domain\-specific nuances of funder name disambiguation, necessitating fine\-tuning on our target task\. Given that our dataset encompasses funder names across multiple languages, we selected paraphrase\-multilingual\-mpnet\-base\-v2121212[https://huggingface\.co/sentence\-transformers/paraphrase\-multilingual\-mpnet\-base\-v2](https://huggingface.co/sentence-transformers/paraphrase-multilingual-mpnet-base-v2)\(a multilingual ST model\), EmbeddingGemma131313[https://huggingface\.co/google/embeddinggemma\-300m](https://huggingface.co/google/embeddinggemma-300m), and Qwen3\-Embedding141414[https://huggingface\.co/Qwen/Qwen3\-Embedding\-0\.6B](https://huggingface.co/Qwen/Qwen3-Embedding-0.6B)models as our starting models\.

Figure 3:Steps for training data creation, model fine\-tuning, and funder name disambiguation\.Paperrefers to a published research article\.Funder namerefers to a single textual representation of a funding organization as it appears in a paper\.Funder IDrefers to the unique identifier assigned to a funder\. Multiple funder names across different papers may refer to a single funder\. Therefore, we use the funder ID to link these name variants to the same underlying entity\.Figure[3](https://arxiv.org/html/2609.09984#S4.F3)illustrates the detailed workflow of the disambiguation process, which consists of three main parts: training data creation, model fine\-tuning, and application of the fine\-tuned model for funder name disambiguation\.

### 4\.2Training Data Restructuring

A major challenge in building a funder name disambiguation model is the lack of training data with funder information for model training or fine\-tuning\. Manual curation of a large\-scale training set is time\-consuming and costly, because WoS contains a very large number of funders and name variants\. We thus created a training set by restructuring existing data, namely WoS, OFR, and ROR\. In the training set, each instance represents a pair consisting of an anchor name \(input name\) and a positive/negative name as well as a label indicating whether or not these two names refer to the same funder \(positive or negative pair\): \[name 1, name 2, 1 \(positive\) or 0 \(negative\)\]\.

To create positive pairs, we used OFR as a bridge between WoS and ROR\. ROR and OFR can be linked directly because both sources include Crossref Funder IDs for individual funder names in their records\. However, WoS and OFR can be linked only at the paper level through Digital Object Identifiers \(DOIs\), not at the level of individual funders\. Funder records for the same paper may still differ across the two datasets\. For example, a paper in the OFR data might list “University of Illinois at Urbana\-Champaign” and “National Science Foundation” as funders, while the same paper in the WoS data might list “the Graduate College of the University of Illinois at Urbana\-Champaign” and “U\.S\. National Science Foundation” as funders\. Fortunately, matching names within a single paper is substantially easier than matching them across papers, primarily because each paper contains relatively few funders, with a median of two\. The pre\-trained ST model can effectively match funder names within a single paper\. Within\-paper matches with cosine similarity above 0\.70 achieved accuracy above 0\.98, we retained all such pairs\. This process assigned a unique ID to funder names in 20,793 of the 52,760 papers\. We then obtained all possible positive pairs among names sharing the same ID, yielding 54,517 positive pairs\.

After identifying the positive pairs, we treated the remaining funder name combinations, other than positive pairs, as negative pairs in the 20,793 papers\. For each funder name, we retrieved its 20 most similar names and removed any known positive pairs\. The remaining high\-similarity pairs served as hard negatives\. We excluded lower\-similarity negative pairs, or soft negatives, because prior work has found that hard negatives provide stronger learning signals\[[49](https://arxiv.org/html/2609.09984#bib.bib49)\]\. This procedure yielded 215,163 negative pairs\.

Given that the positive and negative pairs we created might be limited in range and scope as they are from 20,793 out of 52,760 papers, we applied the same methods solely to ROR to expand the funder list and identified 140,684 positive pairs and 3,061,200 negative pairs\. Of these ROR\-derived pairs, we used all positive pairs and 500,000 randomly selected negative pairs\. This is because we found that the pre\-trained model performs well \(98% accuracy\) for funder names in ROR, as these names are cleaner than those in WoS\. Including substantially more such pairs would increase the training cost while providing little additional improvement\.

Finally, we created and selected 195,201 \(54,517 from WoS \+ 140,684 from ROR\) positive pairs and 715,163 \(215,163 from WoS \+ 500,000 from ROR\) negative pairs\. Among these, 5,000 positive pairs were randomly sampled for out\-of\-sample testing, and the remaining 905,364 pairs were used for model fine\-tuning \(40,000 for validation and 865,364 for training\)\.

### 4\.3Multi\-task Learning

Instead of using a single loss function for fine\-tuning, we adopted a multi\-task learning approach that involves two loss functions for two key reasons\. First, our training set contains two types of funder name pairs \(positive and negative\), with an imbalance of approximately 3\.7 negative pairs for every positive pair\. Employing one loss function to learn primarily from positive pairs and another from both positive and negative pairs allows us to better manage this imbalance\. Second, given that WoS includes many small or lesser\-known funders that cannot be mapped to ROR, our model must both map funder names to ROR and determine whether a valid ROR match exists\. Using one loss function to update the parameters for mapping and another to update those for classification enables the model to handle both tasks effectively\. Therefore, we utilized both Contrastive Loss \(CL\) and Multiple Negatives Ranking Loss \(MNRL\)\.

- •Contrastive Loss\[[15](https://arxiv.org/html/2609.09984#bib.bib15)\]: This loss function uses both positive and negative pairs and can be written as: ℒCL=12​\[yi⋅D​\(ui,vi\)2\+\(1−yi\)⋅max⁡\(0,m−D⁡\(ui,vi\)\)2\]\\mathcal\{L\}\_\{\\text\{CL\}\}=\\frac\{1\}\{2\}\\left\[y\_\{i\}\\cdot D\(u\_\{i\},v\_\{i\}\)^\{2\}\+\(1\-y\_\{i\}\)\\cdot\\max\\bigl\(0,m\-D\(u\_\{i\},v\_\{i\}\)\\bigr\)^\{2\}\\right\]\(1\) In the equation,D⁡\(ui,vi\)=1−S⁡\(ui,vi\)D\(u\_\{i\},v\_\{i\}\)=1\-S\(u\_\{i\},v\_\{i\}\)represents the cosine distance derived from the cosine similarityS⁡\(ui,vi\)S\(u\_\{i\},v\_\{i\}\)between the two normalized embeddings \(uiu\_\{i\}andviv\_\{i\}\)\. The binary labelyi∈\{0,1\}y\_\{i\}\\in\\\{0,1\\\}indicates whether the two names refer to the same funder entity \(yi=1y\_\{i\}=1\) or distinct entities \(yi=0y\_\{i\}=0\)\. The marginmmacts as a minimum separation threshold, penalizing dissimilar pairs whenever their distance falls belowmm\(m=0\.5m=0\.5in our implementation, corresponding to an upper cosine similarity threshold of1−m=0\.51\-m=0\.5\)\. This loss function is particularly effective at separating negative funder pairs in the embedding space151515Technically, contrastive loss reduces the distance between matching funder names as well\. In practice, however, its primary strength lies in separating dissimilar pairs, actively penalizing hard negative pairs that fall within the margin radius rather than merely pulling positive pairs closer together\.by ensuring their distance remains above the required distinguishing threshold\.
- •Multiple Negatives Ranking Loss\[[16](https://arxiv.org/html/2609.09984#bib.bib16)\]: It exclusively uses positive pairs in our training set as input\. In each batch, there areKKpositive pairs, ranging from \(u1u\_\{1\},v1v\_\{1\}\) to \(uku\_\{k\},vkv\_\{k\}\), whereuiu\_\{i\}is theit​hi^\{th\}anchor name, andviv\_\{i\}is theit​hi^\{th\}positive name\. MNRL assumes thatui=viu\_\{i\}=v\_\{i\}andui≠vju\_\{i\}\\neq v\_\{j\}wheni≠ji\\neq j\. Therefore, for each anchor nameuiu\_\{i\}, there is one positive name \(viv\_\{i\}\) and K\-1 negative names \(vjv\_\{j\}\)\. For a single batch, it can be expressed as: ℒMNRL=−1K∑i=1KlogPapprox\(vi∣ui\)=−1K∑i=1K\[S\(ui,vi\)−log∑j=1KeS⁡\(ui,vj\)\]\\mathcal\{L\}\_\{\\text\{MNRL\}\}=\-\\frac\{1\}\{K\}\\sum\_\{i=1\}^\{K\}\\log P\_\{\\text\{approx\}\}\(v\_\{i\}\\mid u\_\{i\}\)=\-\\frac\{1\}\{K\}\\sum\_\{i=1\}^\{K\}\\left\[S\(u\_\{i\},v\_\{i\}\)\-\\log\\sum\_\{j=1\}^\{K\}e^\{S\(u\_\{i\},v\_\{j\}\)\}\\right\]\(2\) In the equation,S⁡\(ui,vi\)S\(u\_\{i\},v\_\{i\}\)represents the cosine similarity between the two embeddings \(uiu\_\{i\}andviv\_\{i\}\)\.∑j=1KeS⁡\(ui,vj\)\\sum^\{K\}\_\{j=1\}e^\{S\(u\_\{i\},v\_\{j\}\)\}is the sum of the exponentiated cosine similarity scores for allKKembeddings\. It is the key part of the softmax normalization, where it normalizes raw scores into a probability distribution\. Therefore, this loss function aims to minimize the approximate mean negative log\-likelihood of the data\. This loss function is particularly effective at minimizing the distance between positive names because its task is to identify the positive name of each anchor name from a list ofKKcandidates \(220 in our case161616Technically, a largerKKcould lead to higher matching accuracy\. However, it will require larger GPU memory\.\)\.

### 4\.4Baselines

- •Pre\-trained embedding models:Widely\-used pre\-trained embedding models are evaluated as embedding\-based retrieval baselines\. Given a funder name, each model generates a dense embedding that is compared against candidate funder names using cosine similarity\. The evaluation is restricted to models specifically optimized for semantic matching tasks because our preliminary experiments with general\-purpose encoder models such as BERT and RoBERTa achieved an accuracy below 0\.3\. The evaluated models include both traditional sentence embedding models and recent large language model \(LLM\)\-based embedding models: - – - –EmbeddingGemma: A lightweight embedding model developed by Google DeepMind, recognized as one of the most capable models under 0\.5 billion parameters\[[43](https://arxiv.org/html/2609.09984#bib.bib43)\]\. For brevity, we refer to this model asGemmathroughout the remainder of this paper\. - –Qwen3\-Embedding: A family of recent embedding models based on the Qwen3 architecture\[[52](https://arxiv.org/html/2609.09984#bib.bib52)\]\. We evaluated the 0\.6B, 4B, and 8B variants to investigate how model size influences matching performance\. Throughout this paper,Qwen3refers specifically to these embedding models rather than the generative LLMs\.
- •Generative LLMs:We also evaluated state\-of\-the\-art generative LLMs \(GPT\-5\.2, Gemini\-2\.5\-Flash, and Claude\-Sonnet\-4\.6\)\. We used the following prompt to identify the official name for each funder in our test set: Please let me know the official name of the following funder or organization\. Only return one name without any acronym\. Funder name: <Name from WoS\> Official name: Since generative LLMs may return different official names for the same funder \(e\.g\., “Alexander von Humboldt\-Stiftung” and “Alexander von Humboldt Foundation” or “Natural Sciences and Engineering Research Council of Canada” and “Natural Sciences and Engineering Research Council”\), the generated names still require post\-processing to map them to entities in ROR\. We experimented with two post\-processing methods: - –Normalized exact matching, which normalizes both the generated name and ROR names by lowercasing and removing special characters before performing exact string matching\. - –Embedding\-based matching, which combines a generative LLM with a pre\-trained multilingual embedding model to encode the generated official name and retrieve the most semantically similar ROR entity based on cosine similarity\.

## 5Matching Accuracy on Test Data

Using the fine\-tuned models, the cosine similarity is computed between each anchor funder name in the test data and every funder name in the ROR corpus\. The funder name with the highest cosine similarity score in the ROR corpus is considered a match for the anchor name\. Since each matched name in the ROR corpus has a unique funder ID, the anchor name inherits this ID\. Finally, the anchor funder name is considered correctly disambiguated if its inherited ID matches the ID of the corresponding positive name in the test data\.

Table 1:Model performance in terms of accuracy\. Bold values denote the highest accuracy in each column across all approaches, and underlined values denote the highest accuracy within each approach\.Table[1](https://arxiv.org/html/2609.09984#S5.T1)presents the matching accuracy on the test data\. The test dataset is divided into two subsets according to whether the anchor funder name originated from ROR or WoS\. The ROR names are generally clean and standardized because they are curated organization names, whereas the WoS names are substantially noisier, as many are extracted directly from funding acknowledgment texts\. Additionally, because a funder may appear multiple times in the test data under different name variants, we constructed an alternative dataset \(denoted as “Unique”\) by retaining only one occurrence of each funder\.

Frequently occurring funders are easier to disambiguate\.Across nearly all models, accuracy is consistently higher on the complete WoS dataset \(denoted as “All”\) than on the unique subset, where duplicate funders have been removed\. This observation suggests that frequently occurring funder names are easier to identify than rare funders\. One possible explanation is that common funders contribute more training examples during fine\-tuning\. In addition, for pre\-trained models, these funders might be more likely to appear in the large\-scale corpora used to pre\-train modern language models, enabling the models to learn richer semantic representations for well\-known funders than for rarely occurring funders\.

Pre\-trained embedding models perform well on standardized funder names but struggle with noisy funder names\.For the ROR data, all embedding models, including the Multilingual ST model, Gemma, and the Qwen3 family, achieve an accuracy of approximately 0\.98, comparable to our fine\-tuned models\. This result suggests that recent embedding models are already highly effective for matching clean, standardized funder names curated in ROR\. However, their performance drops substantially on the WoS data\. For example, the best\-performing pre\-trained model \(Gemma\) achieves only an accuracy of 0\.76 on the WoS data\. These results indicate that the primary challenge in funder name disambiguation lies in handling noisy, real\-world funder names rather than standardized ones\.

Generative LLMs are not suitable as standalone models for funder name disambiguation\.Using generative LLMs followed by exact matching yields substantially lower accuracy than embedding\-based approaches, achieving only around 0\.7 on the ROR data and 0\.6–0\.67 on the WoS data\. Unlike embedding models specifically designed for semantic retrieval and similarity matching, generative LLMs do not consistently produce the canonical funder names required for reliable exact matching against the ROR data\. These results suggest that generative LLMs alone are not well suited for large\-scale funder name disambiguation\.

Generative LLMs can improve the standardization of noisy funder names, but may also propagate errors to downstream matching\.Combining generative LLMs with pre\-trained embedding models consistently improves performance on the WoS data, increasing the accuracy from 0\.76 to 0\.81–0\.82\. This improvement suggests that generative LLMs can effectively standardize noisy funder names before semantic matching\. However, the same approach performs substantially worse on the standardized ROR data, where the best pre\-trained embedding models achieve approximately an accuracy of 0\.98, compared with only 0\.84–0\.93 for the hybrid approaches\. The same pattern is observed when generative LLMs are combined with fine\-tuned embedding models\. These results suggest that generative LLMs can only improve funder name matching when they can successfully standardize noisy funder names\. However, when the standardization is incorrect, the error is inevitably propagated to the downstream embedding matching stage\. Therefore, generative LLMs could reduce the overall matching accuracy\. In other words, generative LLMs represent a double\-edged sword for funder name disambiguation: They can improve performance by standardizing noisy funder names, but LLM errors can also propagate to the matching step and degrade the final results\.

Our fine\-tuning approach substantially improves the funder name matching task\.On the WoS data, the fine\-tuned ST and Gemma models achieve an accuracy of approximately 0\.91 for all funder names and 0\.89 for unique funder names, outperforming their corresponding pre\-trained models by 0\.15–0\.21 and 0\.15–0\.23, respectively\. The fine\-tuned Qwen3\-0\.6B model also substantially improves the accuracy, although its performance is lower than the other two fine\-tuned models\. These consistent improvements across multiple embedding architectures demonstrate that our fine\-tuning approach generalizes well and is not limited to a particular model\. The fine\-tuned ST and Gemma models also outperform the best hybrid approach \(generative LLM \+ pre\-trained embedding model\) by 0\.09 for all funder names and 0\.14 for unique funder names, respectively\. Overall, these results demonstrate that domain\-specific fine\-tuning is an effective strategy for funder name matching\.

Larger embedding models do not necessarily achieve better performance\.

Larger pre\-trained models do not consistently outperform smaller ones\. Before fine\-tuning, Gemma achieves better performance than the larger Qwen3 models on the WoS dataset\. Even within the Qwen3 family, increasing model size does not consistently improve performance\. After fine\-tuning, the ST and Gemma models also outperform the fine\-tuned Qwen3\-0\.6B model despite their smaller sizes\. These results suggest that simply increasing model size does not necessarily lead to better funder name matching performance\. Possible explanations include differences in model architectures and pre\-training objectives, as well as the possibility that larger models require more training data or different fine\-tuning strategies to fully realize their potential\[[28](https://arxiv.org/html/2609.09984#bib.bib28),[50](https://arxiv.org/html/2609.09984#bib.bib50)\]\.

## 6Disambiguation on the WoS Corpus

### 6\.1Distinguishing Funder Names within ROR’s Coverage

We applied the fine\-tuned ST model to the WoS corpus because of its best performance, retrieving the most similar funder name and its corresponding ID from the ROR data\. Unlike the test data, the WoS corpus includes funder names that extend beyond ROR’s coverage\. This means that in addition to matching funder names, we must also identify those that are outside the ROR’s coverage, and thus cannot be disambiguated by mapping them to the ROR data\.

Table 2:Cosine similarity, matching accuracy, and proportion of instancesGiven that our model is partially fine\-tuned using Contrastive Loss, we can establish a threshold similarity score, below which a WoS funder name is unlikely to be matched with any names in the ROR data\. We categorized the data into four groups based on the cosine similarity score between each funder name in WoS and its matched name in ROR:<<0\.8, 0\.8\-0\.85, 0\.85\-0\.9, and≥\\geq0\.9\. We then randomly selected around 200 instances from each group and manually annotated whether the matching result was correct\. As shown in Table[2](https://arxiv.org/html/2609.09984#S6.T2), 52% of the WoS funder names have a match in ROR with a similarity score of≥\\geq0\.9, achieving a matching accuracy of 0\.98\. The accuracy drops to 0\.75, 0\.38, and 0\.0625 for similarity scores of 0\.85\-0\.9, 0\.8\-0\.85, and<<0\.8, respectively\.

We evaluated cosine similarity scores of 0\.80, 0\.85, and 0\.90 as classification thresholds\. For each threshold, funder names with similarity scores below the threshold were classified as unmatched, whereas those with scores equal to or above the threshold were classified as matched\. For illustration, the equations below show the calculation for a threshold of 0\.85\. We applied the same procedure to thresholds of 0\.80 and 0\.90 by regrouping the similarity intervals according to whether they fell below or at or above each threshold\.

T​r​u​e​p​o​s​i​t​i​v​e​s=\(M≥0\.9⋅I≥0\.9\+M0\.85−0\.9⋅I0\.85−0\.9\)⋅nTrue\\penalty\\ positives=\(M\_\{\\geq 0\.9\}\\cdot I\_\{\\geq 0\.9\}\+M\_\{0\.85\-0\.9\}\\cdot I\_\{0\.85\-0\.9\}\)\\cdot n\(3\)F​a​l​s​e​p​o​s​i​t​i​v​e​s=\(\(1−M≥0\.9\)⋅I≥0\.9\+\(1−M0\.85−0\.9\)⋅I0\.85−0\.9\)⋅nFalse\\penalty\\ positives=\(\(1\-M\_\{\\geq 0\.9\}\)\\cdot I\_\{\\geq 0\.9\}\+\(1\-M\_\{0\.85\-0\.9\}\)\\cdot I\_\{0\.85\-0\.9\}\)\\cdot n\(4\)T​r​u​e​n​e​g​a​t​i​v​e​s=\(\(1−M0\.8−0\.85\)⋅I0\.8−0\.85\+\(1−M<0\.8\)⋅I<0\.8\)⋅nTrue\\penalty\\ negatives=\(\(1\-M\_\{0\.8\-0\.85\}\)\\cdot I\_\{0\.8\-0\.85\}\+\(1\-M\_\{<0\.8\}\)\\cdot I\_\{<0\.8\}\)\\cdot n\(5\)F​a​l​s​e​n​e​g​a​t​i​v​e​s=\(\(M0\.8−0\.85⋅I0\.8−0\.85\+M<0\.8⋅I<0\.8\)\)⋅nFalse\\penalty\\ negatives=\(\(M\_\{0\.8\-0\.85\}\\cdot I\_\{0\.8\-0\.85\}\+M\_\{<0\.8\}\\cdot I\_\{<0\.8\}\)\)\\cdot n\(6\)
whereMiM\_\{i\}andIiI\_\{i\}denote the correct match rate and the proportion of instances for groupii, andnnis the number of instances in the corpus\.

We then calculated the precision, recall, and F1 scores for our classification strategy using different thresholds\. As shown in Table[3](https://arxiv.org/html/2609.09984#S6.T3), a cosine similarity threshold of 0\.85 yields the highest F1 score \(0\.92\) and provides the best balance between precision \(0\.94\) and recall \(0\.90\)\. Applying this threshold yielded a disambiguated dataset covering 93,155 \(62\.6% of\) funder names in the WoS corpus with an F1 score of 0\.92\. These names correspond to 7,953 unique funders indexed in ROR\. The remaining set contained 55,635 funder names, representing 37\.4% of the funder corpus, including an estimated 9,392 false negatives that should have been matched to ROR and 46,243 true negatives that were not indexed\. In the following sections, we will refer to the disambiguated dataset \(62\.6%\) as the main dataset and the remaining one as the remaining dataset \(37\.4%\)\.

Table 3:Estimated precision, recall, and F1 scores under different cosine similarity acceptance thresholds\.
### 6\.2Clustering Funder Names Beyond ROR’s Coverage

Although the previous steps could not confidently match WoS funder names in the remaining dataset to the ROR data, we can still cluster the unmatched funder names using our model to group funder names for a second round of disambiguation and also to understand why the prior disambiguation failed \(i\.e\., what types of funders are not indexed by ROR, and why mismatches and misclassifications occurred\)\. We computed the cosine similarity between every pair of names in the remaining dataset\. If the similarity score between two funder names was 0\.9 or higher, we established a link between them, ultimately forming a funder name network\. In this network, funder names are represented as nodes, links with a similarity score of 0\.9 or higher as edges, and similarity scores as edge weights\. We then applied the Louvain algorithm\[[5](https://arxiv.org/html/2609.09984#bib.bib5)\]for network community detection, identifying 28,357 unique funders\. Each community or cluster was considered to represent either a single funder or a group of similar funders\.

The occurrences of funders in prior conservation publications are highly right\-skewed\. Only 1\.5%, 2\.9%, and 14\.1% of funders in the main \(matching\) dataset appear in at least 100, 50, and 10 publications, respectively, while 46\.8% of funders occur in only one publication\. For the remaining \(clustering\) dataset, according to our clustering results, these percentages are 0\.067%, 0\.19%, 1\.66%, and 76\.2%, respectively\. The more right\-skewed distribution in the remaining dataset compared to the main dataset indicates that a greater number of funders in the remaining dataset occur infrequently\. Examples of these funders that appear once in our data include “Instituto Biotropicos” \(a small NGO from Minas Gerais, Brazil\), “Pride of Maui” \(a small travel company from Wailuku, Hawaii\), and “Guizhou R&D Program for Social Development” \(a local government fund from Guizhou, China\)\. The high proportion of infrequently occurring funders reflects the long\-tail nature of the funding landscape, where a diverse range of small, local, and specialized organizations contribute to research funding but may have limited representation in existing crowdsourced resources such as ROR, making their names more challenging to disambiguate\.

Table 4:Top 10 most frequently occurring funders in the main and remaining datasets, identified through matching and clustering, respectively\.Table[4](https://arxiv.org/html/2609.09984#S6.T4)presents the most frequently occurring funders identified through matching in the main dataset and clustering in the remaining dataset, respectively\. Seven of the ten most frequently occurring funders in the clustering dataset are outside the scope of ROR\. For instance, we identified various funders affiliated with the European Commission/Union in the first two clusters, including two major funders—“European Framework Programme” and “European Regional Development Fund”—neither of which was indexed by ROR\. Similarly, some major Chinese funders are also missing from ROR, such as the “National Basic Research Program of China”, “National Key R&D Program of China”, and the “Program for New Century Excellent Talents in University\.” Whether a “program” should be regarded as an organization is debatable, but in some countries or regions like EU and China, researchers frequently report large funding programs as funders, further complicating funder name disambiguation\. This also contributes to their absence from ROR, as ROR typically excludes programs that fall outside its organizational scope\.

Among the three funders that are within the scope of ROR, only one, “Darwin Initiative”, is directly indexed by ROR\. Our fine\-tuned model matched this funder to ROR in most instances; however, the similarity score was only around 0\.8 \(below the threshold of 0\.85\), leading to the omission of these correct matches\. This lower similarity score is primarily due to variations in the funder’s name, such as “Darwin Initiative for the Survival of Species”, which includes additional long postfixes that reduce the similarity score\. Two cases, “DST\-NRF Centre of Excellence for Invasion Biology \(India\)” and “Spanish Ministry of Science and Innovation”, require further discussion\. “DST\-NRF Centre of Excellence for Invasion Biology” typically refers to a funder in South Africa, yet the WoS dataset frequently associates this name with an organization in India\. It is unclear whether these refer to the same funder, as there is little information about the Indian organization available online\. “Spanish Ministry of Science and Innovation” is frequently mentioned in the WoS corpus, yet its official name should be “Ministry of Science, Innovation and Universities”\. However, this name is often matched to “Ministry of Science and Innovation” in other countries by our model with a low similarity score\.

These examples demonstrate that clustering helps identify major funders not indexed by ROR and correct some mismatches\. Furthermore, a comparison of the most frequently occurring funders in the well\-disambiguated main dataset versus the remaining dataset reveals that funders from non\-English\-speaking countries present additional challenges\. Many of these funders are not indexed in ROR\. Additionally, funder names from non\-English\-speaking countries are often written in other languages \(though English names are provided in the table\), and the definition of “funder” may vary, with some regions using terms like “programs” to refer to funders\.

## 7Discussion and Conclusion

In this paper, we introduced \(1\) a comprehensive and reusable framework for training\-data creation, model fine\-tuning, and funder\-name disambiguation across multiple model architectures, including Sentence Transformer, Qwen3, and Gemma models; \(2\) empirical evidence that fine\-tuned embedding models substantially outperform pre\-trained embedding models and generative LLMs on this task; \(3\) a large\-scale dataset of biodiversity conservation publications with cleaned and standardized funder names; and \(4\) an analysis of the opportunities and challenges in funder\-name disambiguation, including limitations in existing funder records and organizational information\.

The added value of combining multiple funder\-related data sources:Previous studies have typically concentrated on a single source of publication records, occasionally supplemented with dictionaries, ontologies, or registries related to organizational information to support rule\-based disambiguation\[[2](https://arxiv.org/html/2609.09984#bib.bib2),[21](https://arxiv.org/html/2609.09984#bib.bib21),[40](https://arxiv.org/html/2609.09984#bib.bib40)\]\. Our research demonstrates that linking and matching multiple sources of publication records \(such as WoS and OFR\) with data or tools related to organizational information \(such as ROR\) can facilitate the creation of training datasets for model fine\-tuning or training\. The lack of supervised learning efforts in previous organization name disambiguation work could be attributed to the absence of large\-scale training data\. Our approach introduces a method for training data creation without requiring costly manual annotation\.

Multi\-task learning and multi\-functional model:We experimented with multi\-task learning for model fine\-tuning, resulting in a multi\-functional model capable of handling several steps in funder name disambiguation\. These functions include mapping funder names in WoS to those in ROR, classifying whether funder names are mappable in the first phase, and clustering WoS funder names when they are unlikely to be indexed in ROR\. Since no single step can resolve all disambiguation challenges, training or fine\-tuning a multi\-functional model can be highly beneficial\.

Challenge of disambiguating rarely occurring funders and those from non\-English\-speaking countries:The greatest challenge is to disambiguate funders that occur infrequently, particularly small, local, or specialized organizations\. Pre\-trained large models may not capture sufficient information about these funders, even with large\-scale training data\. Also, creating our own training data that includes funder information is challenging because these funders appear too infrequently in publication records and are not indexed by curated data related to organization or funder information\. One possible approach might be to label them as unknown or uncertain when analyzing them and use statistical methods to mitigate their negative impact on analysis\. Funders from non\-English\-speaking countries present a related challenge\. Although some of these funders may occur frequently in scholarly publication data, they may still have limited representation in existing curated resources like ROR because of differences in language, naming conventions, and funding systems\. For example, some regions frequently use the term “program” to refer to funding entities, which can complicate their representation and disambiguation in organization\-focused resources\. Our work suggests that clustering can help identify some frequently occurring funders that are not represented in curated resources\.

Limitations and future work:This work has a significant limitation, in addition to the challenges previously discussed: Some correct matches \(e\.g\., “Darwin Initiative” vs\. “Darwin Initiative for the Survival of Species”\) identified in the first step of the matching task may be filtered out in the second step of the classification task due to low similarity scores caused by lengthy prefixes or suffixes\. Generative LLMs could potentially assist in splitting lengthy names and extracting the core components before applying our model for matching\. In other words, while our experiments demonstrate that generative LLMs do not perform well when used directly for name disambiguation \(i\.e\., official name prediction\), they may still be useful for supporting intermediate steps in the process, such as name cleaning and cluster verification\. Future work can explore integrating generative LLMs into these intermediate stages to improve the robustness and accuracy of the overall disambiguation process\.

Cautious Use of the Model and Data:Given the challenges and limitations discussed above, it is important to acknowledge that disambiguation models are not infallible and will still make errors\. We highly recommend that users of the model and data follow these steps before conducting funder analysis:

- •The ROR dataset contains over 102,000 organization names across various domains and countries, some of which may not be relevant to specific studies\. Users should consider removing organizations that are not of interest before using the model for matching\. Narrowing down the pool of matching candidates can improve matching accuracy\. Additionally, we strongly recommend that users remove acronyms from the ROR data when matching WoS names to ROR names, as many funders share the same acronym\. For example, NSF can represent both the National Science Foundation and the National Sleep Foundation\.
- •We set a threshold of 0\.85 for the similarity score to classify whether a funder is accurately matched to ROR, aiming for a high F1 score\. This allowed us to split the data into the main dataset, which is accurately matched, and the remaining dataset, which likely contains funders not indexed by ROR\. Since a higher threshold increases precision but decreases recall \(and vice versa\), users should adjust the threshold based on their research priorities—whether they prioritize precision, recall, or F1 score\.
- •When constructing similarity networks for clustering funder names in the remaining dataset, we used a threshold of 0\.9\. Users can also modify this threshold according to their needs\. A higher threshold will result in smaller, more specific clusters, while a lower threshold will produce larger clusters\.
- •Finally, we strongly recommend that users review the data or the model results and make necessary manual adjustments based on their needs before conducting any funder analysis\.

## 8Data and Code

The datasets utilized in this study are derived from the Web of Science \(WoS\) and are therefore subject to Clarivate’s data licensing agreements\. These terms prohibit the direct redistribution of raw publication records alongside our disambiguated funder names\. To maintain compliance while supporting reproducibility, we provide a mapping file containing only publicly accessible identifiers: article titles, WoS IDs and DOIs\. This allows researchers with authorized access to WoS to reconstruct the complete dataset\. Furthermore, we have made our implementation code available to the community\. All source code and the aforementioned identifier data can be accessed in our GitHub repository181818[https://github\.com/khan1792/Funder\_name\_disambiguation](https://github.com/khan1792/Funder_name_disambiguation)\. Researchers with appropriate access to the underlying WoS data may contact the authors to inquire about access to the disambiguated dataset and fine\-tuned models\.

## 9Funding Statement

This research was supported by the John D\. and Catherine T\. MacArthur Foundation \(Grant No\. 18\-1802\-152800\-CSD\)\.

## References

- \[1\]\(2022\)Named entity recognition using deep learning: a review\.In2022 international conference on business analytics for technology and security \(ICBATS\),Dubai, UAE,pp\. 1–7\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p4.1)\.
- \[2\]A\. Ancona, R\. Cerqueti, and G\. Vagnani\(2023\)A novel methodology to disambiguate organization names: an application to EU Framework Programmes data\.Scientometrics128\(8\),pp\. 4447–4474\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p4.1),[§2](https://arxiv.org/html/2609.09984#S2.p3.1),[§7](https://arxiv.org/html/2609.09984#S7.p2.1)\.
- \[3\]J\. M\. Ballreich, C\. P\. Gross, N\. R\. Powe, and G\. F\. Anderson\(2021\)Allocation of National Institutes of Health funding by disease category in 2008 and 2019\.JAMA Network Open4\(1\),pp\. e2034890\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p1.1)\.
- \[4\]C\. Bloch and M\. P\. Sørensen\(2015\)The size of research funding: trends and implications\.Science and Public Policy42\(1\),pp\. 30–43\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p1.1)\.
- \[5\]V\. D\. Blondel, J\. Guillaume, R\. Lambiotte, and E\. Lefebvre\(2008\)Fast unfolding of communities in large networks\.Journal of Statistical Mechanics: Theory and Experiment2008\(10\),pp\. P10008\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p6.1),[§6\.2](https://arxiv.org/html/2609.09984#S6.SS2.p1.1)\.
- \[6\]W\. Bouarroudj, Z\. Boufaida, and L\. Bellatreche\(2022\)Named entity disambiguation in short texts over knowledge graphs\.Knowledge and Information Systems64\(2\),pp\. 325–351\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p4.1)\.
- \[7\]S\. R\. Bowman, G\. Angeli, C\. Potts, and C\. D\. Manning\(2015\)A large annotated corpus for learning natural language inference\.InProceedings of the 2015 conference on empirical methods in natural language processing,Lisbon, Portugal,pp\. 632–642\.Cited by:[§4\.1](https://arxiv.org/html/2609.09984#S4.SS1.p1.1)\.
- \[8\]L\. Bromham, R\. Dinnage, and X\. Hua\(2016\)Interdisciplinary research has consistently lower funding success\.Nature534\(7609\),pp\. 684–687\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p1.1)\.
- \[9\]Convention on Biological Diversity\(2022\)Kunming–Montreal Global Biodiversity Framework\.External Links:[Link](https://www.cbd.int/gbf)Cited by:[1st item](https://arxiv.org/html/2609.09984#S3.I1.i1.p1.1)\.
- \[10\]J\. A\. Dalsgaard, F\. N\. Silva, and J\. Ai\(2026\)Linking global science funding to research publications\.arXiv preprint arXiv:2603\.24147\.Cited by:[§2](https://arxiv.org/html/2609.09984#S2.p5.1)\.
- \[11\]J\. Diesner, C\. Evans, and J\. Kim\(2015\)Impact of entity disambiguation errors on social network properties\.InProceedings of the international AAAI conference on web and social media,Vol\.9,Oxford, UK,pp\. 81–90\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p4.1)\.
- \[12\]J\. Fortin and D\. J\. Currie\(2013\)Big science vs\. little science: how scientific impact scales with funding\.PLoS One8\(6\),pp\. e65263\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p1.1)\.
- \[13\]A\. French\(2022\)Emerging uses of the research organization registry\.InSeptentrio Conference Series,Cited by:[3rd item](https://arxiv.org/html/2609.09984#S3.I1.i3.p1.1)\.
- \[14\]Y\. Guan\(2025\)Disambiguating academic institution names: a comprehensive study of authority files, linguistic variations, and computational evaluation in PubMed affiliations\.Ph\.D\. Dissertation,University of Illinois Urbana\-Champaign,Urbana, Illinois\.External Links:[Link](https://hdl.handle.net/2142/130147)Cited by:[§2](https://arxiv.org/html/2609.09984#S2.p6.1)\.
- \[15\]R\. Hadsell, S\. Chopra, and Y\. LeCun\(2006\)Dimensionality reduction by learning an invariant mapping\.In2006 IEEE computer society conference on computer vision and pattern recognition \(CVPR’06\),Vol\.2,New York, USA,pp\. 1735–1742\.Cited by:[1st item](https://arxiv.org/html/2609.09984#S4.I1.i1.p1.1)\.
- \[16\]M\. Henderson, R\. Al\-Rfou, B\. Strope, Y\. Sung, L\. Lukács, R\. Guo, S\. Kumar, B\. Miklos, and R\. Kurzweil\(2017\)Efficient natural language response suggestion for smart reply\.arXiv preprint arXiv:1705\.00652\.Cited by:[2nd item](https://arxiv.org/html/2609.09984#S4.I1.i2.p1.1)\.
- \[17\]G\. Hendricks, D\. Tkaczyk, J\. Lin, and P\. Feeney\(2020\)Crossref: the sustainable source of community\-owned scholarly metadata\.Quantitative Science Studies1\(1\),pp\. 414–427\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p2.1),[2nd item](https://arxiv.org/html/2609.09984#S3.I1.i2.p1.1)\.
- \[18\]S\. Huang, B\. Yang, S\. Yan, and R\. Rousseau\(2014\)Institution name disambiguation for research assessment\.Scientometrics99\(3\),pp\. 823–838\.Cited by:[§2](https://arxiv.org/html/2609.09984#S2.p2.1),[§2](https://arxiv.org/html/2609.09984#S2.p3.1),[§2](https://arxiv.org/html/2609.09984#S2.p4.1),[1st item](https://arxiv.org/html/2609.09984#S3.I1.i1.p1.1)\.
- \[19\]B\. A\. Jacob and L\. Lefgren\(2011\)The impact of research grant funding on scientific productivity\.Journal of Public Economics95\(9\-10\),pp\. 1168–1177\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p1.1)\.
- \[20\]Y\. Jiang, H\. Zheng, X\. Wang, B\. Lu, and K\. Wu\(2011\)Affiliation disambiguation for constructing semantic digital libraries\.Journal of the American Society for Information Science and Technology62\(6\),pp\. 1029–1041\.Cited by:[§2](https://arxiv.org/html/2609.09984#S2.p3.1)\.
- \[21\]S\. Jonnalagadda and P\. Topham\(2010\)NEMO: extraction and normalization of organization names from PubMed affiliation strings\.Journal of Biomedical Discovery and Collaboration5,pp\. 50\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p4.1),[§2](https://arxiv.org/html/2609.09984#S2.p3.1),[§7](https://arxiv.org/html/2609.09984#S7.p2.1)\.
- \[22\]A\. Joulin, E\. Grave, P\. Bojanowski, and T\. Mikolov\(2017\)Bag of tricks for efficient text classification\.InProceedings of the 15th conference of the European chapter of the association for computational linguistics,Valencia, Spain,pp\. 427–431\.Cited by:[1st item](https://arxiv.org/html/2609.09984#S3.I1.i1.p1.1)\.
- \[23\]J\. Kim, J\. Diesner, H\. Kim, A\. Aleyasen, and H\. Kim\(2014\)Why name ambiguity resolution matters for scholarly big data research\.In2014 IEEE international conference on big data,Washington, D\.C\., USA,pp\. 1–6\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p4.1)\.
- \[24\]J\. Kim and J\. Diesner\(2015\)The effect of data pre\-processing on understanding the evolution of collaboration networks\.Journal of Informetrics9\(1\),pp\. 226–236\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p4.1)\.
- \[25\]J\. Kim and J\. Diesner\(2016\)Distortive effects of initial\-based name disambiguation on measurements of large\-scale coauthorship networks\.Journal of the Association for Information Science and Technology67\(6\),pp\. 1446–1461\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p4.1)\.
- \[26\]P\. Kokol and H\. B\. Vošner\(2018\)Discrepancies among Scopus, Web of Science, and PubMed coverage of funding information in medical journal articles\.Journal of the Medical Library Association106\(1\),pp\. 81\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p2.1),[1st item](https://arxiv.org/html/2609.09984#S3.I1.i1.p1.1)\.
- \[27\]R\. Lammey\(2020\)Solutions for identification problems: a look at the research organization registry\.Science Editing7\(1\),pp\. 65–69\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p2.1),[3rd item](https://arxiv.org/html/2609.09984#S3.I1.i3.p1.1)\.
- \[28\]E\. Liu, A\. Bertsch, L\. Sutawika, L\. Tjuatja, P\. Fernandes, L\. Marinov, M\. Chen, S\. Singhal, C\. Lawrence, A\. Raghunathan,et al\.\(2025\)Not\-just\-scaling laws: towards a better understanding of the downstream impact of language model design decisions\.InProceedings of the 2025 conference on empirical methods in natural language processing,Suzhou, China,pp\. 16407–16438\.Cited by:[§5](https://arxiv.org/html/2609.09984#S5.p9.1)\.
- \[29\]W\. Liu\(2020\)Accuracy of funding information in Scopus: a comparative case study\.Scientometrics124\(1\),pp\. 803–811\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p2.1),[1st item](https://arxiv.org/html/2609.09984#S3.I1.i1.p1.1)\.
- \[30\]S\. Mishra, B\. D\. Fegley, J\. Diesner, and V\. I\. Torvik\(2018\)Self\-citation is the hallmark of productive authors, of any gender\.PLoS One13\(9\),pp\. e0195773\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p4.1)\.
- \[31\]F\. Morillo, J\. Aparicio, B\. González\-Albo, and L\. Moreno\(2013\)Towards the automation of address identification\.Scientometrics94\(1\),pp\. 207–224\.Cited by:[§2](https://arxiv.org/html/2609.09984#S2.p3.1)\.
- \[32\]OECD\(2024\)Biodiversity and development finance 2015–2022: contributing to target 19 of the kunming–montreal global biodiversity framework\.OECD Publishing,Paris\.Cited by:[1st item](https://arxiv.org/html/2609.09984#S3.I1.i1.p1.1)\.
- \[33\]N\. Onodera, M\. Iwasawa, N\. Midorikawa, F\. Yoshikane, K\. Amano, Y\. Ootani, T\. Kodama, Y\. Kiyama, H\. Tsunoda, and S\. Yamazaki\(2011\)A method for eliminating articles by homonymous authors from the large number of articles retrieved by author search\.Journal of the American Society for Information Science and Technology62\(4\),pp\. 677–690\.Cited by:[§2](https://arxiv.org/html/2609.09984#S2.p3.1)\.
- \[34\]A\. Paul\-Hus, N\. Desrochers, and R\. Costas\(2016\)Characterization, description, and considerations for the use of funding acknowledgement data in Web of Science\.Scientometrics108\(1\),pp\. 167–182\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p2.1)\.
- \[35\]R\. Pranckutė\(2021\)Web of Science \(WoS\) and Scopus: the titans of bibliographic information in today’s academic world\.Publications9\(1\),pp\. 12\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p2.1)\.
- \[36\]T\. M\. Rabovsky and W\. C\. Ellis\(2014\)Higher education and congressional influence on administrative decisions: an examination of NSF and NIH research grant funding to four\-year universities\.Social Science Quarterly95\(3\),pp\. 740–759\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p1.1)\.
- \[37\]N\. Reimers and I\. Gurevych\(2019\)Sentence\-BERT: sentence embeddings using Siamese BERT\-networks\.InProceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing \(EMNLP\-IJCNLP\),Hong Kong, China,pp\. 3982–3992\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p5.1),[§4\.1](https://arxiv.org/html/2609.09984#S4.SS1.p1.1)\.
- \[38\]A\. Roberts, C\. Raffel, and N\. Shazeer\(2020\)How much knowledge can you pack into the parameters of a language model?\.InProceedings of the 2020 conference on empirical methods in natural language processing \(EMNLP\),pp\. 5418–5426\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p4.1)\.
- \[39\]D\. K\. Sanyal, P\. K\. Bhowmick, and P\. P\. Das\(2021\)A review of author name disambiguation techniques for the PubMed bibliographic database\.Journal of Information Science47\(2\),pp\. 227–254\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p4.1)\.
- \[40\]Z\. Shao, X\. Cao, S\. Yuan, and Y\. Wang\(2020\)ELAD: an entity linking based affiliation disambiguation framework\.IEEE Access8,pp\. 70519–70526\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p4.1),[§2](https://arxiv.org/html/2609.09984#S2.p2.1),[§2](https://arxiv.org/html/2609.09984#S2.p3.1),[§7](https://arxiv.org/html/2609.09984#S7.p2.1)\.
- \[41\]K\. Song, Y\. Li, L\. Yao, and Y\. Wang\(2020\)Research on organization name matching based on word vector\.InJournal of Physics: Conference Series,Vol\.1684,pp\. 012085\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p4.1)\.
- \[42\]H\. Sun, J\. Li, Y\. Wu, L\. Wang, and K\. W\. Fung\(2017\)Using an ontology\-based approach to handle author affiliations in a large biomedical citation database\.Studies in Health Technology and Informatics245,pp\. 1338\.Cited by:[§2](https://arxiv.org/html/2609.09984#S2.p3.1)\.
- \[43\]H\. S\. Vera, S\. Dua, B\. Zhang, D\. Salz, R\. Mullins, S\. R\. Panyam, S\. Smoot, I\. Naim, J\. Zou, F\. Chen,et al\.\(2025\)EmbeddingGemma: powerful and lightweight text representations\.arXiv preprint arXiv:2509\.20354\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p5.1),[2nd item](https://arxiv.org/html/2609.09984#S4.I2.i1.I1.i2.p1.1),[§4\.1](https://arxiv.org/html/2609.09984#S4.SS1.p1.1)\.
- \[44\]A\. Waldron, D\. C\. Miller, D\. Redding, A\. Mooers, T\. S\. Kuhn, N\. Nibbelink, J\. T\. Roberts, J\. A\. Tobias, and J\. L\. Gittleman\(2017\)Reductions in global biodiversity loss predicted from conservation spending\.Nature551\(7680\),pp\. 364–367\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p2.1)\.
- \[45\]A\. Waldron, A\. O\. Mooers, D\. C\. Miller, N\. Nibbelink, D\. Redding, T\. S\. Kuhn, J\. T\. Roberts, and J\. L\. Gittleman\(2013\)Targeting global conservation funding to limit immediate biodiversity declines\.Proceedings of the National Academy of Sciences110\(29\),pp\. 12144–12148\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p1.1)\.
- \[46\]R\. Whitley, J\. Gläser, and G\. Laudel\(2018\)The impact of changing funding and authority relationships on scientific innovations\.Minerva56,pp\. 109–134\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p1.1)\.
- \[47\]A\. Williams, N\. Nangia, and S\. R\. Bowman\(2018\)A broad\-coverage challenge corpus for sentence understanding through inference\.InProceedings of the 2018 conference of the North American chapter of the association for computational linguistics: human language technologies,New Orleans, Louisiana,pp\. 1112–1122\.Cited by:[§4\.1](https://arxiv.org/html/2609.09984#S4.SS1.p1.1)\.
- \[48\]Y\. Xie\(2014\)“Undemocracy”: inequalities in science\.Science344\(6186\),pp\. 809–810\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p1.1)\.
- \[49\]Z\. Yang, M\. Ding, T\. Huang, Y\. Cen, J\. Song, B\. Xu, Y\. Dong, and J\. Tang\(2024\)Does negative sampling matter? a review with insights into its theory and applications\.IEEE Transactions on Pattern Analysis and Machine Intelligence46\(8\),pp\. 5692–5711\.Cited by:[§4\.2](https://arxiv.org/html/2609.09984#S4.SS2.p3.1)\.
- \[50\]B\. Zhang, Z\. Liu, C\. Cherry, and O\. Firat\(2024\)When scaling meets LLM finetuning: the effect of data, model and finetuning method\.InInternational conference on learning representations,Vol\.2024,Vienna, Austria,pp\. 44694–44713\.Cited by:[§5](https://arxiv.org/html/2609.09984#S5.p9.1)\.
- \[51\]L\. Zhang, W\. Lu, and J\. Yang\(2023\)LAGOS\-AND: a large gold standard dataset for scholarly author name disambiguation\.Journal of the Association for Information Science and Technology74\(2\),pp\. 168–185\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p4.1)\.
- \[52\]Y\. Zhang, M\. Li, D\. Long, X\. Zhang, H\. Lin, B\. Yang, P\. Xie, A\. Yang, D\. Liu, J\. Lin,et al\.\(2025\)Qwen3 embedding: advancing text embedding and reranking through foundation models\.arXiv preprint arXiv:2506\.05176\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p5.1),[3rd item](https://arxiv.org/html/2609.09984#S4.I2.i1.I1.i3.p1.1),[§4\.1](https://arxiv.org/html/2609.09984#S4.SS1.p1.1)\.
- \[53\]Q\. Zhi and T\. Meng\(2016\)Funding allocation, inequality, and scientific research output: an empirical study based on the life science sector of Natural Science Foundation of China\.Scientometrics106\(2\),pp\. 603–628\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p1.1)\.
- \[54\]P\. Zhou, X\. Cai, and X\. Lyu\(2020\)An in\-depth analysis of government funding and international collaboration in scientific research\.Scientometrics125,pp\. 1331–1347\.Cited by:[§1](https://arxiv.org/html/2609.09984#S1.p1.1)\.

Similar Articles

Robust Text Watermarking for Large Language Models via Dual Semantic Embeddings

arXiv cs.CL

This paper presents Dual-Embedding Watermarking (DEW), a semantic watermarking scheme for LLMs that improves robustness against paraphrasing and translation by leveraging contextual and token-level embeddings. Experimental results show improved detection after paraphrasing and translation compared to prior methods.