Observatorio Lazaro: A self-populating database of anglicism usage in the Spanish press

arXiv cs.CL Papers

Summary

This paper presents Observatorio Lázaro, a continuously updated database and public web/API resource that monitors anglicism usage in the Spanish digital press using a neural sequence-labeling model, recording over two million borrowings from 2020 to 2026.

arXiv:2608.00713v1 Announce Type: new Abstract: This paper describes Observatorio L\'azaro, a language resource that monitors unassimilated lexical borrowings (predominantly English lexical borrowings or anglicisms) in the Spanish digital press. Since April 2020 the system has automatically processed the daily output of a collection of news outlets, detected borrowings with a neural sequence-labeling model, and made the results available through a public web interface and API. The result is a continuously updated diachronic database which, at the time of writing, records more than two million borrowings across 1.88 million articles and 993 million running tokens of text (2020-2026). The paper documents the resource: we describe the end-to-end pipeline (acquisition, detection, post-processing, storage and access), the data model and the terms of availability; we evaluate the resource through the detector's held-out performance (span-level F1=0.86 for the borrowing class), inter-annotator agreement on the training corpus (Cohen's kappa=0.91) and a manual precision audit of 1,000 spans from the deployed data; and we situate it with respect to Spanish borrowing lexicography, annotated borrowing corpora and neology-monitoring observatories. The data shows that unassimilated anglicisms are used in the Spanish press at a frequency of approximately two anglicisms per thousand tokens, and that this rate remains stable. Our statistical analysis over six years reveals that the anglicism vocabulary in Spanish behaves as an open and growing class, with 58.7% of its types attested only once (53.6% after correcting for detection precision), and that its density is highest in the fashion, technology and lifestyle sections and lowest in political and institutional news. The resource is intended to complement static borrowing dictionaries and one-off annotated corpora by providing a continuously updated record of borrowing in the Spanish press.
Original Article
View Cached Full Text

Cached at: 08/04/26, 07:44 AM

# Observatorio Lázaro: A self-populating database of anglicism usage in the Spanish press
Source: [https://arxiv.org/html/2608.00713](https://arxiv.org/html/2608.00713)
\[1\]\\fnmElena\\surÁlvarez\-Mellado

1\]\\orgdivDepartment of Linguistics,\\orgnameUniversidad Autónoma de Madrid,\\orgaddress\\countrySpain

###### Abstract

This paper describes Observatorio Lázaro, a language resource that monitors unassimilated lexical borrowings \(predominantly English lexical borrowings or*anglicisms*\) in the Spanish digital press\. Since April 2020 the system has automatically processed the daily output of a collection of news outlets, detected borrowings with a neural sequence\-labeling model, and made the results available through a public web interface and API\. The result is a continuously updated diachronic database which, at the time of writing, records more than two million borrowings across 1\.88 million articles and 993 million running tokens of text \(2020–2026\)\. The paper documents the resource: we describe the end\-to\-end pipeline \(acquisition, detection, post\-processing, storage and access\), the data model and the terms of availability; we evaluate the resource through the detector’s held\-out performance \(span\-level F1 = 0\.86 for the borrowing class\), inter\-annotator agreement on the training corpus \(Cohen’sκ\\kappa= 0\.91\) and a manual precision audit of 1,000 spans from the deployed data; and we situate it with respect to Spanish borrowing lexicography, annotated borrowing corpora and neology\-monitoring observatories in other languages\. The data shows that unassimilated anglicisms are used in the Spanish press at a frequency of approximately two anglicisms per thousand tokens, and that this rate remains stable\. Our statistical analysis over six years reveals that the anglicism vocabulary in Spanish behaves as an open and growing class, with 58\.7% of its types attested only once \(53\.6% after correcting for detection precision\), and that its density is highest in the fashion, technology and lifestyle sections and lowest in political and institutional news\. The resource is intended to complement static borrowing dictionaries and one\-off annotated corpora by providing a continuously updated record of borrowing in the Spanish press\.

###### keywords:

anglicisms, lexical borrowing, language resource, diachronic corpus, neology monitoring, Spanish

## 1Introduction

Multilingualism and contact between languages have been the norm throughout history rather than the exception\. Asthomason\_sarah\_g\_social\_2003puts it, “all languages are mixed in a weak sense: there is no natural human language in which foreign material is wholly lacking”\. Contact is in fact one of the major inducers of language change: speakers take patterns belonging to one language and incorporate them into another, a process known as linguistic borrowing\[haugen\_analysis\_1950,weinreich\_languages\_1963\]\. When what is transferred belongs to the lexicon, we speak of*lexical borrowing*\.

Lexical borrowing is a particularly informative object of study because it sits at the intersection of the social and the systemic\. A word may be imported together with an artefact or a technique, documenting a cultural exchange between communities; or it may be imported because the foreign form is perceived as more prestigious than the native one, reflecting a social hierarchy that shapes language use\. Borrowings are thus evidence of contact between languages and of the social dynamics between the communities that speak them\[zenner\_cognitive\_2012\]\. At the same time, which words get borrowed, what status they acquire, and how far they are adapted to the recipient language reveal the grammatical patterns and speaker expectations that operate in that language, and the general constraints on language change\[van1994modeling,haspelmath\_ii\_2009\]\.

The borrowing of English vocabulary into other languages \(words such asinfluencer,streaming,podcastandlook\) is the most salient instance of this process in contemporary European languages\[furiassi\_anglicization\_2012\]\. Readers of the French press were estimated to encounter a new lexical borrowing roughly every thousand words, with English borrowings outnumbering all other donor languages combined\[chesley\_predicting\_2010,chesley\_lexical\_2010\]\. Spanish is no exception to this\[nunez\_nogueroles\_up\-\-date\_2017\]: in the Chilean press, borrowings were found to account for some 30% of neologisms, 80% of them anglicisms\[gerding\_anglicism\_2014\]; and for European Spanish, anglicisms were estimated to make up around 2% of the vocabulary used in the newspaperEl Paísin 1991\[rodriguez\_gonzalez\_spanish\_2002\], a figure that is likely to be considerably higher today, but that no continuously updated measurement has been able to confirm\.

Existing estimates of how pervasive lexical borrowings are are scarce, dated and hard to compare, as studying the phenomenon empirically has traditionally been difficult\. The lexicographic tradition documents borrowings authoritatively but statically and with a lag of years\[pratt1980,lorenzo1996,rodriguezlillo1997,rodriguez2017\]\. Corpus\-based studies capture a synchronic slice and, crucially, have depended on the manual inspection of general reference corpora, which forces researchers either to annotate an entire corpus by hand or to restrict the enquiry to a predefined list of query terms \(Section[2\.3](https://arxiv.org/html/2608.00713#S2.SS3)\)\. Neither strategy scales to an on\-going process: borrowings are notoriously volatile\[poplack2017borrowing\], and because they are also a sparse phenomenon, exhaustive repertoires require very large corpora, which is not feasible without automation\.

There has been a recent wave of computational work on borrowing detection, which has produced annotated datasets and models\[alvarezmellado2020headlines,adobo2021,alvarezmellado2022,ahmadi2025conloan\], but not a standing observation of the phenomenon as it unfolds\. What has been lacking is a resource that combines four properties: large scale, diachronic depth, open availability and continuous updating, that is, a standing record of which English forms enter the Spanish press, when, where and how often\.

In this paper we describe Observatorio Lázaro111[https://observatoriolazaro\.es](https://observatoriolazaro.es/)\[alvarezmellado2020obs\], an automatic monitoring system that, every day since April 2020, retrieves the articles published by a collection of Spanish news outlets, detects the unassimilated borrowings they contain with a neural sequence\-labeling model, and publishes the resulting data through a public website and API\. Its by\-product is a cumulative diachronic database that, as of July 2026, records 2,007,647 borrowing tokens in 1,880,377 articles spanning 993\.4 million running tokens\.

The aim of the resource is to make the analysis of anglicism usage possible in real context, at large scale, and in a systematic and data\-driven fashion, so as to support both lexicographic work and corpus\-linguistic research\. The detection model and the annotated corpus on which it was trained were introduced in prior work\[alvarezmellado2022\]; the contribution of the present paper is the operational monitoring system and the diachronic resource built on top of that model: its pipeline, its data model, its terms of availability and reuse, its evaluation as a resource, and a first characterization of what six years of continuous monitoring reveal, including a measurement of anglicism density that the estimates cited above could not provide\.

It may be useful to state plainly what this adds to the work it builds on\.alvarezmellado2022introduced thecoalascorpus and a set of models trained on it, and evaluated those models on held\-out annotated text; that paper is about an annotated dataset and a modeling task\. The present paper is about neither\. Its object is the system that has been running the resulting model in production since 2020 and the database that six years of doing so have produced: the acquisition and preprocessing pipeline, the data model and the access layer \(Section[3](https://arxiv.org/html/2608.00713#S3)\); the conditions under which the data may be reused \(Section[4\.3](https://arxiv.org/html/2608.00713#S4.SS3)\); an evaluation of what the detector does when applied to unfiltered newswire rather than to curated test text, which turns out to differ substantially from the benchmark figure \(Section[5\.3](https://arxiv.org/html/2608.00713#S5.SS3)\); and a characterization of the accumulated output \(Section[6](https://arxiv.org/html/2608.00713#S6)\)\. None of these is present in the earlier work, and the deployed evaluation in particular could not have been carried out until the system had run long enough to produce a database worth auditing\. A model and an annotated corpus are inputs to a resource of this kind; they are not themselves the resource\.

The paper is organized as follows\. Section[2](https://arxiv.org/html/2608.00713#S2)defines the phenomenon the resource tracks, reviews the descriptive, corpus\-based and computational literature on anglicisms in Spanish, and positions Lázaro among comparable resources\. Section[3](https://arxiv.org/html/2608.00713#S3)documents the end\-to\-end pipeline\. Section[4](https://arxiv.org/html/2608.00713#S4)describes the resulting database: its data model, size, coverage and its availability and licensing\. Section[5](https://arxiv.org/html/2608.00713#S5)evaluates the resource along three axes: model performance, annotation reliability and a manual precision audit of the deployed data\. Section[6](https://arxiv.org/html/2608.00713#S6)characterizes the resource statistically over 2020–2026\. Section[7](https://arxiv.org/html/2608.00713#S7)discusses applications and reuse; Section[8](https://arxiv.org/html/2608.00713#S8)states limitations\.

## 2Background and related resources

### 2\.1Lexical borrowing: scope of the phenomenon

Linguistic borrowing is the process of reproducing in one language elements and patterns that come from another\[haugen\_analysis\_1950\], and has been studied extensively within contact linguistics\[weinreich\_languages\_1963\]\. Several typologies have been proposed to characterize borrowings according to the levels of language involved and the degree of integration of the borrowed element\[thomason\_language\_1992,matras\_grammatical\_2007,haspelmath\_loanwords\_2009\]\. Lexical borrowing in particular involves the incorporation of single lexical units, usually accompanied by morphological and phonological modification to conform to the patterns of the recipient language\[poplack\_social\_1988,onysko\_anglicisms\_2007\]\. In Spanish, for instance, borrowed verbs must take the suffix\-aror\-earto enter the verbal paradigm \(tuitear, fromto tweet\), and phonological adaptation may surface in spelling, so that the unadaptedspoilercoexists with the adaptedespóiler\.

All these phonological, orthographic and morphological transformations can lead to a borrowing eventually becoming fully assimilated into the recipient language lexicon, which may lead to native speakers losing the perception of the term being “foreign”\[lipski\_code\-switching\_2005\]\. Some authors establish the need of a borrowing being recognized as foreign by native speakers in order to be considered as such\[zenner\_cognitive\_2012\]\. For instance, a word likebarwas originally borrowed from English into Spanish, but it has been so assimilated that it is now perceived as a native word by monolingual speakers of Spanish, and its English nature is only seen as etymological\. On the other hand, a word likewhisky\(that has also been used in Spanish for some time\) has never been fully assimilated and is perceived as a foreign word\.

In this paper we deal with unassimilated lexical borrowings in Spanish, that is, words from other languages that have been incorporated into Spanish but that have not undergone the process of assimilation\. We focus on lexical borrowings from English \(also known as anglicisms\), as they represent the vast majority of current borrowings in Spanish\.

A further boundary is the one between borrowing and codeswitching, the alternation between two or more languages in discourse that is typical of bilingual communities\[poplack\_sometimes\_1980\]\. Unlike borrowings, codeswitches are by definition not integrated into the recipient language and do not produce a change in its lexicon; ashaspelmath\_ii\_2009puts it, codeswitching “is not a kind of contact\-induced language change, but rather a kind of contact\-induced speech behavior”\. A clear\-cut distinction nonetheless remains elusive, particularly for non\-integrated foreign items appearing in otherwise monolingual contexts \(so\-called “lone other\-language items”\)\. Some authors treat borrowing and codeswitching as a continuum with a fuzzy frontier\[clyne\_dynamics\_2003\], whereaspoplack\_myths\_2012argue that they are distinct phenomena and that integration is abrupt rather than gradual\[poplack\_borrowing\_1984,poplack\_social\_1988\]\. Criteria proposed to separate them include frequency\[stammers\_testing\_2012\], level of integration\[poplack\_borrowing\_1984\], length\[calude\_modelling\_2020\], speaker bilingual competence\[pfaff\_constraints\_1979\]and listedness\[treffers\-daller\_simple\_2023\]; the debate is unresolved\[lipski\_code\-switching\_2005\], and some authors accordingly prefer agnostic terminology such as “donor\-language items”\[deuchar\_english\-origin\_2016\]or code\-mixing\[muysken\_bilingual\_2000\]\. Aspoplack\_myths\_2012put it, “distinguishing codeswitching and borrowing is the thorniest issue in the field of contact linguistics today”\.

For the purpose of this work we follow the approach taken bypoplack\_myths\_2012and consider codeswitching and borrowing as two separate phenomena\. We consider that codeswitches are fluent multiword interferences that normally comply with grammatical restrictions in both languages and that are produced by bilingual speakers in bilingual discourses \(usually in spoken language\), while we define lexical borrowings as words of foreign origin that are used by monolingual individuals without knowledge of the donor language\.

### 2\.2Anglicisms in Spanish

The influence of English on Spanish has been a sustained topic of linguistic research for decades\[pratt1980,furiassi\_anglicization\_2012,nunez\_nogueroles\_up\-\-date\_2017\]\. Prior work has proposed different analysis and classifications of anglicism usage in Spanish\[lorenzo1996,medina\_lopez\_anglicismo\_1998,rodriguez\_gonzalez\_anglicisms\_1999,nunez\_nogueroles\_comprehensive\_2018\]and examined orthographic integration\[nunez\_nogueroles\_typographical\_2017\], diachronic shifts\[gimeno\_menendez\_desplazamiento\_2003\], typological characteristics\[gomez\_capuz\_towards\_1997\], syntactic anglicisms\[rodriguez\_medina\_anglicismos\_2002\], sociocultural dimensions\[gomez\_capuz\_prestamos\_2004\]and lexicographic coverage\[balteiro\_reassessment\_2011\]\. This tradition is crystallised in reference works such as theNuevo diccionario de anglicismos\[rodriguezlillo1997\]and theGran diccionario de anglicismos\[rodriguez2017\]\. Authoritative as they are, such works are static: they describe a curated inventory at a point in time, and cannot track frequency, diffusion, or the constant churn of ephemeral borrowings that a living press produces\.

### 2\.3Limitations of corpus\-based study of anglicisms

Empirical research on anglicism usage in Spanish has been informed chiefly by corpus linguistics\[rodriguez\_medina\_anglicismos\_2002,gimeno\_menendez\_desplazamiento\_2003,balteiro\_reassessment\_2011,nunez\_nogueroles\_corpus\-based\_2018\], relying either on general reference corpora such as CREA222[https://corpus\.rae\.es/creanet\.html](https://corpus.rae.es/creanet.html)and CORPES333[https://www\.rae\.es/corpes/](https://www.rae.es/corpes/)or on tailor\-made corpora built for a specific genre or variety\. The main limitation of these corpus\-assisted projects is that the retrieval of anglicisms and their concordances relies on the manual lookup of the corpus\. That implies that either the whole corpus is processed and annotated by a human \(in order to identify all anglicisms it contains\), or the research is limited to a list of predefined terms of interest to be queried\. This approach seems insufficient to account for an on\-going phenomenon like anglicism incorporation, because borrowings are notorious for being volatile\[poplack2017borrowing\]\. In addition, as lexical borrowings are a sparse phenomenon, large corpora are required in order to collect exhaustive repertoires of anglicisms in use, something that is simply not feasible without efficient automatization\.

Some corpus interfaces do allow searching for borrowings in general, such as CORPES\. Their results, however, leave a lot to be desired\. At the time of writing, the concordances obtained by CORPES when querying for borrowings include true borrowings, along with foreign proper nouns \(such asSpice Girls,Modern FamilyorMadama Butterfly\), unusual Spanish words \(such aschundachundaorfu, in the idiomni fu ni fa\), ill\-tokenized words or complete Latin or English phrases \(which would correspond more to the phenomenon of codeswitching than borrowing\)\.

Ultimately, both the dissatisfaction with the CORPES anglicism results and the obvious limitations of having to manually process the corpus point in the same direction: the need for robust computer\-assisted methods in corpus linguistics in general and in borrowing studies in particular that can facilitate data\-driven research on language contact, language change and the lexicon\. The resource presented here is a direct response to that need\.

### 2\.4Automatic borrowing detection

The task of extracting unassimilated anglicisms from Spanish text is a more challenging undertaking than it might appear to be at first\. To begin with, lexical borrowings can be single or multiword expressions \(e\.g\.,prime time,tie breakormachine learning\)\. Second, linguistic assimilation is a diachronic process and, as a result, what constitutes an unassimilated borrowing is not clear\-cut\. For example, words likebarorclubwere unassimilated lexical borrowings in Spanish at some point in the past, but have become so widespread and frequent that the process of phonological and morphological adaptation is now complete and they cannot be considered unassimilated borrowings anymore\. In addition, not all English words that appear in Spanish texts will be anglicisms: proper names, literal quotations and other language\-mixing phenomena can naturally occur in Spanish text, without any of them being true examples of lexical borrowing\. To make things worse, there are non\-related words that exist both as English and Spanish words \(quince,come,primer,pie\)\. These words may or may not be a borrowing depending on the context they appear in\. For instance, the wordpiewill be a native word in Spanish when it means “foot”, but it will be a borrowing inpie de limón\(“lemon pie”\)\. Similarly, the wordcashwill be a proper noun when referred to singerJohnny Cash, but it will be a borrowing inEstoy sin cash\(“I have no cash”\)\. All these subtleties make the automatic retrieval of lexical borrowings non\-trivial\. Consequently, in prior work on anglicism extraction from Spanish text, plain dictionary lookup produced very limited results, with F1 scores of 0\.47\[serigos\_applying\_2017\]and 0\.26\[alvarezmellado2020headlines\]\.

Beyond Spanish, computational approaches to borrowing and foreign\-word detection have been explored for a range of languages and purposes, from lexicographic and corpus work\[hofland\_self\-expanding\_2000,furiassi\_retrieval\_2007,losnegaard\_data\-driven\_2012,chesley\_lexical\_2010,serigos\_applying\_2017\]to speech and cross\-lingual applications\[mansikkaniemi\_unsupervised\_2012,leidig\_automatic\_2014\], and more recently with neural and multilingual models\[miller\_using\_2020,nath\_generalized\_2022,pugh\_itml\_2023,dinu\_robocop\_2023\]; related work in historical linguistics addresses loanword detection in etymological databases\[list\_automated\_2019,haspelmath\_2014\_11137\]\.

For Spanish specifically, the line of work on which this resource rests reframes the problem as sequence labeling, in which relevant spans \(either single\-word or multiword\) are extracted from sentences, much as in named\-entity or multiword\-expression recognition works\.alvarezmellado2020headlinesreleased an annotated corpus of emerging anglicisms in Spanish newspaper headlines; the ADoBo shared task\[adobo2021\]benchmarked systems for the automatic detection of unassimilated borrowings in the Spanish press; andalvarezmellado2022introducedcoalas, a manually annotated corpus of 372,701 tokens of Spanish newswire containing 3,161 unassimilated\-borrowing annotations \(3,038 English and 123 other\-language; 1,683 unique types\), together with a suite of models\. The best of these, a BiLSTM\-CRF combining bilingual and sub\-word embeddings, is the detector deployed in Lázaro \(Section[3\.4](https://arxiv.org/html/2608.00713#S3.SS4)\)\. These are the immediate predecessors of the present resource: they provide the annotation scheme, the training data and the model, whereas Lázaro turns them into a deployed, longitudinal resource\.

Automatic borrowing detection is also of interest beyond linguistics: borrowings and neologisms contribute to out\-of\-vocabulary words, which can degrade model performance\[zheng\_neo\-bench\_2024\], and their identification has proved relevant for parsing\[alex\_automatic\_2008\], text\-to\-speech synthesis\[leidig\_automatic\_2014\]and as a bootstrapping technique in machine translation for low\-resource languages that borrow heavily\[tsvetkov\_constraint\-based\_2015,tsvetkov\_cross\-lingual\_2016\]\.

### 2\.5Neology and borrowing observatories

Internationally, the closest analogues are neology and borrowing monitoring systems built on “monitor corpora” \(corpora that grow over time so that new usage can be observed as it appears\)\. Notable examples include the Observatori de Neologia \(OBNEO/BOBNEO\) at the Universitat Pompeu Fabra, a long\-running project on Catalan and Spanish neology\[obneo2004\]; large news\-monitoring corpora such as the News on the Web \(NOW\) corpus\[davies2013\], which is diachronic and continuously extended but not specialised for borrowing; and Néoveille, a multilingual web platform for tracking and documenting neologisms whose working languages include Spanish\[cartier2017\], although the project seems to be discontinued and its site was not reachable at the time of writing\. Table[1](https://arxiv.org/html/2608.00713#S2.T1)situates Lázaro with respect to these resources\. It differs from them in combining four properties: it is specialised for unassimilated borrowing detection, fully automatic and updated daily, openly available together with a public API, and now six years deep\. We are not aware of another Spanish resource that combines all four, although each is present individually in one or more of the resources listed\.

Table 1:Observatorio Lázaro among related resources \(schematic comparison\)\.

## 3The Observatorio Lázaro pipeline

### 3\.1Overview

The main component of Observatorio Lázaro is an automatic pipeline that extracts anglicisms from Spanish journalistic texts\. This pipeline consists of three steps: article retrieval, anglicism extraction and anglicism storage\. The output of the pipeline is then aggregated and published through the project website and API \(Section[3\.6](https://arxiv.org/html/2608.00713#S3.SS6)\)\.

The observatory monitors the type of borrowings that the model can extract: unassimilated lexical borrowings, with a special focus on unassimilated anglicisms\. Consequently, other borrowings such as adapted borrowings \(words whose spelling has been modified to comply with the morphophonological and orthographic patterns of the Spanish language, likefútbolfromfootball\) or assimilated borrowings \(words that have already been registered in reference dictionaries without any italics or quotations, such asbar\) are not expected to be extracted by the model and are therefore not tracked by the observatory\. Similarly, other phenomena like semantic calques, syntactic anglicisms, acronyms and proper names are also considered beyond the scope of the project and are therefore not considered by the observatory\. Residual codeswitches and partially assimilated borrowings nevertheless surface among the detection errors \(Section[5\.3](https://arxiv.org/html/2608.00713#S5.SS3)\)\.

The observatory started monitoring 8 Spanish newspaper sources in 2020\. The collection of sources tracked was expanded in September 2022\. Table[2](https://arxiv.org/html/2608.00713#S3.T2)gives the daily throughput of the pipeline in its two configurations\. In its initial period the system processed some 640 articles and 195,000 tokens a day from 8 sources, extracting around 440 borrowing occurrences \(180 distinct\) and roughly 19 previously unseen borrowings\. After the 2022 expansion it processes some 890 articles and 566,000 tokens a day, extracting around 1,110 occurrences \(430 distinct\) and some 35 new borrowings daily\.

Table 2:Average daily throughput of the pipeline, before and after the mid\-2022 expansion\. The boundary is the outlet expansion of September 2022, and averages are taken over days on which at least one article was retrieved \(964 and 1,423 days respectively\)\. “Distinct” is the mean number of distinct lemmas detected in a day; “new” counts lemmas attested for the first time\.
### 3\.2Source acquisition

Every day the system retrieves the articles published by a collection of Spanish news outlets \(Appendix[A](https://arxiv.org/html/2608.00713#A1)\), including the most popular general national dailies \(El País,El Mundo,La Vanguardia,elDiario\.es,ABC,20minutos,El Confidencial\), a news agency \(Agencia EFE\), the economic press \(El Economista,Cinco Días\), sports \(Marca\), and specialist magazines \(Elle,Fotogramas,Rolling Stone,Men’s Health,Muy Interesante\)\.

The starting point of the pipeline is an automatic system that connects daily to the RSS feeds of several major news sources from Spain\. Every day, the pipeline connects to these sources and extracts the texts of the articles published within the last 24 hours\. For each article the system stores the URL, headline, publication date, outlet, section and a running\-token count; the article text is passed to the detector\. Duplicate articles are avoided by keeping a persistent record of the URLs already scraped, so that each article is processed exactly once\. The eight outlets covering general national news have been monitored since 2020 ; a further group was added in September 2022 \(Section[8](https://arxiv.org/html/2608.00713#S8)\), so composition\-controlled longitudinal analyses restrict attention to the stable “core” set\.

### 3\.3Text preprocessing

These articles are then preprocessed \(for HTML tag removal, social media embeds, cookie notices, navigation menus and similar elements\) and made ready for the next step of the system\. The resulting text is then segmented into sentences and tokenized prior to detection usingspaCy\[ines\_montani\_2023\_10009823\]\. For every detected borrowing the system stores the surrounding sentence as*context*, the borrowing’s start and end token offsets, and two Boolean flags indicating whether the borrowing occurs in a headline and whether it falls within quotation marks or italics\.

### 3\.4Borrowing detection

Detection is framed as BIO sequence labeling over the tokenized text\. The model deployed since August 2022 is the best system ofalvarezmellado2022: a BiLSTM\-CRF that combines bilingual Spanish–English word embeddings with sub\-word representations \(byte\-pair\-encoding and character\-level embeddings\), trained on the manually annotatedcoalascorpus\. This architecture was adopted because it outperformed other alternatives, such as fine\-tuning multilingual and Spanish Transformer encoders \(mBERT, BETO and XLM\-RoBERTa\), which did not yield higher results\[alvarezmellado2022,delarosa2021\]\. More recently, large language models have likewise produced only modest results on anglicism and loanword detection, failing to surpass fine\-tuned Transformers\[sousaahmadi2026,iberbench2025,alvarezmellado2025\]\.

The current configuration achieves a span\-level F1 of 0\.86 for borrowing detection on thecoalastest set \(Section[5\.1](https://arxiv.org/html/2608.00713#S5.SS1)\)\. The model also predicts a per\-span language label \(English vs\. other\); in practice the other\-language class is rarely predicted \(Section[5\.1](https://arxiv.org/html/2608.00713#S5.SS1)\), so nearly all detections are labeled English and the resource functions as a monitor of anglicisms\. From April 2020 to August 2022 an earlier Conditional Random Field model was used\[alvarezmellado2020headlines\]; the model change is one component of the mid\-2022 discontinuity discussed in Section[8](https://arxiv.org/html/2608.00713#S8)\.

### 3\.5Post\-processing and storage

Each detection is \(i\) assigned a*lemma*using thePattern666[https://github\.com/clips/pattern](https://github.com/clips/pattern)library\[de2012pattern\], run in English mode, since the forms to be normalized are of English or English\-like origin and inflect on English rather than Spanish patterns, \(ii\) tagged with the language label predicted by the detector \(English vs\. other; reliable only for English, Section[5\.1](https://arxiv.org/html/2608.00713#S5.SS1)\), and \(iii\) collected and stored in a MySQL database\. For every anglicism, the date, context, newspaper and link to the article where the anglicism was found are stored\. In parallel, the system maintains a per\-form*index*that aggregates, for each borrowing type, its total number of occurrences, its first\-attestation date and a hapax flag\.

### 3\.6Access layer: website, API and library

The output of the pipeline is aggregated and published daily at[https://observatoriolazaro\.es](https://observatoriolazaro.es/)\. The site is the primary point of access to the resource, and since a substantial part of what the observatory offers is available only through it, we describe it here in some detail\.

Five tools sit alongside the front page\. The*search*interface queries occurrences and their context sentences, with filters on outlet, section, date range and language label\. The*lexicon*browses the inventory of recorded borrowings\. The*trends*view ranks borrowings over a configurable window \(the last week, month, quarter, year or three years\) under three headings: the most frequent, those whose frequency in the period most exceeds their own historical average, and those detected for the first time\. The*comparison*tool plots the frequency series of up to five lemmas against one another, which is the form in which competition between rival borrowings is most legible \(runningagainstfooting, orstreamingagainstpodcast\) \(see Figure[1](https://arxiv.org/html/2608.00713#S3.F1)\)\. The*outlets*view ranks every monitored publication by borrowing density per million published words, reporting for each the number of articles, the running\-token count and the raw number of detections, with a parallel ranking by section\.

![Refer to caption](https://arxiv.org/html/2608.00713v1/img/startup.png)Figure 1:Comparison between competing anglicisms:*startup*,*start up*,*start\-up*\.The lexical entry page is the part of the site closest to a dictionary entry, and it aggregates everything the observatory holds on one borrowing \(Fig\.[2](https://arxiv.org/html/2608.00713#S3.F2)\)\. For each lemma it reports the total number of occurrences and the normalized rate per million words over several periods; a frequency series, weekly or smoothed over four weeks; the surface forms grouped under the lemma with their individual counts; distributions by outlet and by section; the words that most often occur next to it; other borrowings related to it or containing it; and the full set of concordances, sortable by form, outlet, section and date\. Listing the surface forms individually is what allows a user to inspect the composition of a lemma group rather than take the normalization on trust \(Section[3\.5](https://arxiv.org/html/2608.00713#S3.SS5)\), and it is also what makes competing spellings visible \(as withcoming of agebesidecoming\-of\-age, or the English and Spanish plurals ofcelebrity\)\. The page additionally displays an automatically derived usage profile, giving the proportion of occurrences inside quotation marks and an assignment of part of speech and grammatical gender\.

Programmatic access is provided by a public JSON API, documented on the site’s data page and summarized in Table[3](https://arxiv.org/html/2608.00713#S3.T3)\. Responses are UTF\-8 encoded and rate\-limited to between twenty and sixty requests per minute depending on the endpoint, with excess requests returning HTTP 429\. For bulk use the same page offers the database as monthly CSV files covering April 2020 to the present, so that a user can reconstruct the full record without paging through the API; a snapshot of the database is also available via Zenodo\[alvarez\_mellado\_2026\_21721949\]: searches run in the browser export up to 5,000 rows\. Because the site is updated daily while the analyses reported in this paper are computed on a snapshot taken on 24 July 2026, figures retrieved from the site will exceed those reported here by whatever has accumulated since\. Finally, thepylazarolibrary allows the detector itself to be run over new text, and all code is released as open source in the project repository\.777[https://github\.com/lirondos/lazaro](https://github.com/lirondos/lazaro)

Table 3:Public API endpoints\. All return JSON; paths are relative tohttps://observatoriolazaro\.es/api/\.![Refer to caption](https://arxiv.org/html/2608.00713v1/img/lawfare.png)Figure 2:Upper part of the lexical entry page forlawfare, in the English interface\.

## 4The resource: database description

### 4\.1Data model

The database is organized around three related tables \(Table[4](https://arxiv.org/html/2608.00713#S4.T4)\)\. The occurrences table stores one row per detected borrowing token; the*index*table stores one row per borrowing type; and the articles table stores one row per processed article, providing the denominators required to normalize frequencies\. A snapshot of the database has been publicly released in Zenodo\[alvarez\_mellado\_2026\_21721949\]\.

Table 4:Principal fields of the Lázaro database \(abbreviated\)\.
### 4\.2Size and coverage

Table[5](https://arxiv.org/html/2608.00713#S4.T5)summarizes the resource by year\. It grows steeply: the annual volume of processed text rises from 40 million running tokens in 2020 to over 220 million in 2025, reflecting both continuous accumulation and the mid\-2022 expansion of the outlet set\. In total the resource records 2,007,647 borrowing tokens \(68,424 distinct case\-folded lemmas\) in 1,880,377 articles over 993\.4 million running tokens\.888Throughout, ‘tokens’ \(or ‘running tokens’\) meansspaCytokens, the unit in which the text is processed and in which the corpus size is recorded\. These include punctuation marks, which are tokenized separately:spin\-off, for instance, is three tokens \(spin,\-,off\)\. Densities expressed per token are therefore slightly lower than the equivalent rate per orthographic word, by the share of punctuation in the text \(of the order of 10–15% in news prose\)\.The overall density is approximately 2,020 borrowing tokens per million tokens, or \(in more intuitive terms\) about two unassimilated anglicisms per thousand tokens of Spanish press prose, roughly one every 500 tokens\. Section[5\.3](https://arxiv.org/html/2608.00713#S5.SS3)shows that this estimate is stable under correction for detection error\.

Table 5:Coverage of the resource by year \(2020–2026; 2026 partial, to 24 July\)\.
### 4\.3Availability, licensing and reuse

All components of Lázaro are open\. The live data are browsable and queryable through the website and its data/API endpoint, and the complete 2020–2026 database is downloadable in full from[https://observatoriolazaro\.es](https://observatoriolazaro.es/); the detection model and thepylazarolibrary are released in the project repository; and thecoalastraining corpus is distributed withalvarezmellado2022\. To respect the copyright of the source outlets, neither the website nor the released database reproduces full articles: only the isolated sentence in which each detected borrowing occurs is stored and displayed, which complies with reproduction law for the purposes of scientific research\. Articles are retrieved from the outlets’ public RSS feeds, and no full article text is redistributed\. A citable, versioned snapshot of the exact 2020–2026 database described here is available with a persistent identifier via Zenodo999[https://doi\.org/10\.5281/zenodo\.21721949](https://doi.org/10.5281/zenodo.21721949)\.

Following FAIR principles, the resource is findable \(persistent website and repository\), accessible \(open web, API and library, plus a full download\), interoperable \(tabular records with documented fields\) and reusable \(open license and companion tooling\)\.

## 5Resource evaluation

Because a monitoring resource is only as trustworthy as the detector that populates it, we evaluate it along three complementary axes: the detector’s benchmark performance, the reliability of the annotation it learned from, and the precision of the data it actually produces in deployment\.

### 5\.1Model performance

On thecoalastest set \(58,997 tokens containing 1,239 English and 46 other\-language borrowing spans\), the deployed BiLSTM\-CRF attains a span\-level F1 of 0\.86 for the borrowing class \(precision 0\.90, recall 0\.82\), the best result reported for the task\[alvarezmellado2022\]\. Table[6](https://arxiv.org/html/2608.00713#S5.T6)breaks the result down by predicted language label\. English borrowings, the overwhelming majority of the data, are detected well \(F1 0\.87\), but theotherclass is detected poorly: precision is high \(0\.86\) yet recall is only 0\.13 \(F1 0\.23\)\. The cause is a severe class imbalance in the training data, whose training split contains just 28 other\-language borrowings against 1,493 English ones \(the corpus\-level counts of 123 and 3,038 given in Section[2\.4](https://arxiv.org/html/2608.00713#S2.SS4)are distributed across the training, development and test splits\), so the model rarely predictsotherand tends to label any borrowing it detects as English\. Two consequences follow for users: \(i\) the resource is best understood as a monitor of*anglicisms*specifically; and \(ii\) the language tag stored with each occurrence is trustworthy for English but not for non\-English borrowing, whose counts are substantially under\-estimated \(Section[8](https://arxiv.org/html/2608.00713#S8)\)\.

Table 6:Span\-level detection performance on thecoalastest set, overall and by language label\[alvarezmellado2022\]\.
### 5\.2Annotation reliability

Thecoalascorpus was manually annotated under a documented guideline\. To assess its reliability, a sample of 9,110 tokens \(450 sentences: 60% from the test split, 20% from training, 20% from development\) was doubly annotated by nine linguists; the mean inter\-annotator pair\-wise agreement, computed with Cohen’sκ\\kappa, was 0\.91, well above the 0\.8 threshold conventionally taken to indicate reliable annotation\[artstein2008,alvarezmellado2022\]\. This high agreement bounds the quality of the signal the detector learned from and provides a human ceiling against which its F1 \(Section[5\.1](https://arxiv.org/html/2608.00713#S5.SS1)\) can be read\.

### 5\.3Precision of the deployed data

The audit reported in this section was carried out on output of the current detector, and therefore characterizes the portion of the database produced from August 2022 onwards, which accounts for 79\.5% of all occurrences\. The deployed precision of the superseded CRF model, and hence of the remaining 20\.5% of the database, has not been measured\.

The F1 score of 0\.86 reported inalvarezmellado2022was measured on a curated test set in lab conditions\. In order to assess the performance of the model when deployed on the observatory to identify borrowings on noisier real world data, we manually reviewed a subset of 1,000 spans from the observatory database \(along with the context sentence they appeared in\)101010Due to copyright and storage limitations, the observatory only stores sentences where a borrowing was identified, which prevents us from evaluating recall retrospectively on the data already collected\. The limitation is one of what has been stored rather than of the design: annotating a sample of complete articles as they are retrieved would measure recall directly without requiring the full text to be retained, and such a probe is planned \(Section[8](https://arxiv.org/html/2608.00713#S8)\)\.\. The spans were randomly selected but, given the Zipfian distribution of borrowings, we imposed the following two conditions when selecting the sample:

1. 1\.The 1,000 spans belonged to 1,000 distinct lemmas, so that high\-frequency borrowings easily spotted by the model \(such aslookoronline\) did not dominate the sample\.
2. 2\.The 1,000 distinct spans were stratified along five frequency levels, to ensure that different types of borrowings were represented\. Table[7](https://arxiv.org/html/2608.00713#S5.T7)displays the number of spans per frequency tier\.

Table 7:Number of spans selected for reannotation per frequency level\. Total occurrences reports the total number of occurrences in the observatoryOut of the 1,000 manually annotated spans, 683 were true positives, which yields a raw precision of 0\.68 on the reannotation sample\. Because the sample is stratified by type across frequency tiers \(Table[7](https://arxiv.org/html/2608.00713#S5.T7)\), it deliberately over\-represents the rare tail, so this raw figure is not itself a population estimate; we derive interpretable token\- and type\-weighted precisions from the per\-tier results below \(Section[5\.3](https://arxiv.org/html/2608.00713#S5.SS3), Table[9](https://arxiv.org/html/2608.00713#S5.T9)\)\. Even so, it already points to a gap between processing clean, curated text under lab conditions and detecting borrowings on real\-world data in the wild\.

Table[8](https://arxiv.org/html/2608.00713#S5.T8)displays number of errors per type in the 1,000 spans\. The most frequent type of error was caused by labeling as anglicism a span that was a non\-English borrowing \(such asbratwurst,cocottesortofu\), French being the most frequent non\-English donor language\. This result is not surprising: results fromalvarezmellado2022already showed that the model was not reliable when identifying non\-English borrowings\. Boundary errors are the second most prevalent cause of error\. Boundary errors account for partially retrieved multitoken spans \(33; for instance, retrieving onlydividendindividend recap\), spans that overlap with a true borrowing but also include non\-borrowing tokens \(9;brunch modernete, when it should have retrieved onlybrunch\) and adjacent spans that were incorrectly fused into one single span \(19;shopping bag beige, instead ofshopping bagandbeige\)\.

In both cases \(wrong language error and boundary error\), the model correctly identified that there was a relevant span in the sentence, but it incorrectly labeled it as English origin or misidentified its boundaries\. If we perform a more lenient evaluation in which a span is considered correct regardless of the assigned label and as long as there is some overlap between the predicted span and the goldstandard span \(which is usually known as relaxed evaluation\), precision goes up to 0\.82, which is still below but closer to the precision scores reported on the official test set\. Put differently, the two error types just described \(wrong language and wrong boundaries\) together account for 44% of all errors: in almost half of the cases in which the model is wrong, a genuine unassimilated borrowing is nonetheless present in the sentence, and what fails is the language label or the delimitation of the span\. This distinction matters for downstream use: a study of which anglicisms are used must apply the strict precision, whereas a study of how often borrowing occurs at all can rely on the relaxed precision\.

The rest of the errors are due to foreign proper nouns, Spanish words \(either odd\-looking, already assimilated borrowings or words whose shape is equal to English words, such ashorror\), English words not used as a true borrowing but as in metalinguistic discourse, codeswitches and other interferences \(such as when quoting what someone said in English\)\. Interestingly, 20 of the errors were caused by ill\-tokenized URLs, hashtags, usernames and other Internet symbols, which should have been handled during preprocessing by simple heuristics\. Other phenomena that confused the model and were a source of errors include typos, scientific names, acronyms and demonyms\.

Table 8:Number of true positives, false positives and error types on the selected 1,000 spansError typeNumberExampleTrue positives683WRONG LANGUAGE79boutique,tofuBOUNDARY ERROR61\[pay to\] playPROPER NOUN50DSportSPANISH WORDS40tiki\-taka,post\-covid,horrorMETALINGUISTIC USAGE30yeahURLs, EMAIL, HASHTAGS20change\.org,googlemailCODESWITCHES15starringOTHER22PVP,vol\.,dB,LOGSE\\botruleBecause the sample was stratified, we can assess precision per frequency level\. Table[9](https://arxiv.org/html/2608.00713#S5.T9)shows that the model’s worst result is obtained on nonce borrowings \(those attested only once in the observatory\), with a precision of 0\.54: of all borrowings detected only once, roughly half are errors\.

These errors are, however, of the same kinds as those observed in the sample as a whole\. Of the 115 errors in the nonce tier, 47 \(40%\) are language or boundary errors, that is, cases in which a borrowing is in fact present in the sentence but has been given the wrong language label or the wrong span\. Only the remaining 60% are cases in which no borrowing is present at all\. Applied to a strict precision of 0\.54, this means that about 25% of one\-off detections are not borrowings, while the other 75% do mark a real borrowing, even if it is mislabeled or misdelimited: the relaxed precision for the nonce tier is therefore approximately 0\.72\. The distinction matters for reuse, since a study of which anglicisms are used requires the strict figure, whereas a study of how often borrowing occurs at all can rely on the relaxed one\.

Precision rises sharply with frequency, reaching 0\.94 for the very\-high tier: borrowings detected more than a thousand times are correct 94% of the time\. The real challenge, then, lies not in the frequent core but in the long tail of rare borrowings\.

These per\-tier figures also let us correct the raw sample precision of 0\.68, which the stratified design inflates towards the tail\. Weighting each tier’s precision by the tier’s actual share of the database \(Table[9](https://arxiv.org/html/2608.00713#S5.T9)\) yields two interpretable population estimates\. The token\-weighted precision \(the probability that a randomly retrieved occurrence is a genuine borrowing\) is 0\.87, close to the 0\.90 obtained under lab conditions, because the token stream is dominated by high\-frequency borrowings that the model detects reliably\. The type\-weighted precision \(the probability that a randomly retrieved distinct borrowing is genuine\) is much lower, 0\.59, because the type inventory is dominated by the error\-prone rare tail\. This contrast is the crux of the resource’s reliability profile: frequency\- and trend\-based analyses rest on data nearly as clean as the benchmark, whereas analyses that reach into the rare tail \(e\.g\. neology or hapax studies\) must contend with a substantially higher error rate and should apply the per\-tier precisions reported here\.

These figures also allow the resource’s headline density estimate to be error\-corrected\. The raw detection rate is 2\.02 borrowings per thousand tokens \(Section[4\.2](https://arxiv.org/html/2608.00713#S4.SS2)\)\. Discounting false positives by the token\-weighted precision \(0\.87\) gives 1\.76 genuine anglicisms per thousand tokens actually captured; further scaling by the detector’s recall \(0\.82 on the test set\) to account for borrowings that were missed gives an estimated true rate of 2\.14 per thousand\. The two error sources thus largely offset each other, and the estimate is stable at approximately two unassimilated anglicisms per thousand tokens of Spanish press prose \(roughly one in every 500 tokens\) across all three variants\. This should be read as an order\-of\-magnitude estimate rather than a point value, since the recall correction assumes that test\-set recall transfers to deployed data, which we cannot verify directly \(Section[8](https://arxiv.org/html/2608.00713#S8)\)\.

Table 9:Precision audit per frequency tier, with each tier’s share of the type inventory and of the token stream\. Weighting the per\-tier precisions by these shares yields a token\-weighted precision of 0\.87 \(a randomly retrieved occurrence\) and a type\-weighted precision of 0\.59 \(a randomly retrieved distinct borrowing\)\.

## 6The resource in use: a statistical overview \(2020–2026\)

We close with a first characterisation of the resource, both to demonstrate its research value and to give prospective users a sense of its contents\.

### 6\.1Scale, skew and the most frequent anglicisms

The anglicism vocabulary is dominated by a small high\-frequency core over a very long tail\. Counted over occurrences, the ten most frequent lemmas account for 21% of all tokens and the top hundred for 51%\. Counted over the vocabulary, those hundred lemmas are 0\.15% of the 68,424 distinct types, while 58\.7% of the types occur exactly once and together contribute only 2% of the tokens \(53\.6% once each tier is discounted by its measured precision, Section[5\.3](https://arxiv.org/html/2608.00713#S5.SS3)\)\. The gap between the two is the skew itself: a handful of borrowings carries half the running text, and the majority of the recorded vocabulary is barely attested at all\. Table[10](https://arxiv.org/html/2608.00713#S6.T10)lists the twenty most frequent lemmas; they are overwhelmingly the vocabulary of digital media, technology and everyday evaluative register \(look,online,influencer,app,streaming\)\. The rank–frequency distribution \(Fig\.[3](https://arxiv.org/html/2608.00713#S6.F3)\) is steep and heavy\-tailed, with a log–log slope ofa=1\.365a=1\.365over ranks 10–5,000\. The fitted range excludes the first nine ranks, where the head is flattened by a number of high\-frequency borrowings of comparable rank, and stops at 5,000 to stay clear of the region in which counts fall to single figures and rank ties become pervasive\. The distribution is thus reliably heavy\-tailed, which is the usual finding for word frequencies\.

Table 10:The twenty most frequent anglicism lemmas \(case\-folded\), with total occurrences\.![Refer to caption](https://arxiv.org/html/2608.00713v1/img/fig_zipf.png)Figure 3:Rank–frequency distribution of anglicism lemmas \(log–log\); dashed line, fitted Zipfian slope−1\.365\-1\.365\.
### 6\.2A productive open class

The vocabulary is not a closed inventory but a continuously renewed open class\. Baayen’s potential productivityP=V​\(1\)/NP=V\(1\)/N\(the share of the token mass contributed by hapax legomena\) is0\.0200\.020\(roughly one new type per fifty borrowing tokens\), Herdan’sC=log⁡V/log⁡NC=\\log V/\\log Nis0\.7670\.767, and vocabulary growth follows the Heaps–Herdan lawV∝NβV\\propto N^\{\\beta\}withβ=0\.655\\beta=0\.655\(R2=0\.999R^\{2\}=0\.999; Fig\.[4](https://arxiv.org/html/2608.00713#S6.F4)\)\.

These are raw detection counts\. Discounting every lemma by the measured precision of its frequency tier \(Table[9](https://arxiv.org/html/2608.00713#S5.T9)\), and correcting the token stream on the same basis so that numerator and denominator are treated alike, givesP=0\.012P=0\.012,C=0\.738C=0\.738andβ=0\.611\\beta=0\.611\(R2=0\.9999R^\{2\}=0\.9999\)\. The correction is not a rescaling: it removes 46% of the nonce types but only 6% of the most frequent ones, so it changes the shape of the accumulation curve and not merely its intercept\. All four constants fall, but the exponent remains well below one and the hapax share remains a majority of the inventory, so the vocabulary is appreciably less productive than the raw counts suggest while still behaving as an open and growing class\.

![Refer to caption](https://arxiv.org/html/2608.00713v1/img/fig_heaps.png)Figure 4:Vocabulary growth \(Heaps–Herdan law\) from January 2023, the first full year after both the detector change and the outlet expansion: cumulative types against cumulative borrowing tokens \(log–log\), with fitted curveV=5\.06​N0\.655V=5\.06\\,N^\{0\.655\}\.The same curve can be read from the point of view of a reader rather than of the corpus, which gives a sense of how much text one has to go through before meeting a borrowing that is new to them\. Over the period as a whole the observatory records 0\.069 first attestations per thousand tokens, and the marginal rate at the current corpus size is 0\.048\. Discounted by the relaxed precision of the nonce tier \(0\.72, Section[5\.3](https://arxiv.org/html/2608.00713#S5.SS3)\), this gives between 0\.034 and 0\.050 new borrowings per thousand tokens, that is, one previously unseen borrowing every 20,000 to 29,000 tokens\. In other words, this amounts to one new borrowing every 38 to 56 articles of average length, or roughly one every four to six days for a reader who goes through ten articles a day; only about one in 51 borrowing encounters involves a word that is new to the reader\. These figures should be read as a lower bound, since recall is unmeasured on the deployed data and novel forms are precisely where a static detector is most likely to fail \(Section[8](https://arxiv.org/html/2608.00713#S8)\)\.

### 6\.3Borrowing across newspaper sections

Because every occurrence records the newspaper*section*in which it appears, the resource supports a data\-driven view of where borrowing concentrates\. Table[11](https://arxiv.org/html/2608.00713#S6.T11)reports, for the main sections, the borrowing density: occurrences per million running tokens, using each section’s own token count as the denominator\. The section labels are the ones the outlets themselves assign, as retrieved from the feed, and they are stored unmodified\.

Borrowing is stratified by domain\. Consumer, technology and entertainment sections lie above the corpus average of≈\\approx2,020 per million: fashion shows the highest density at≈\\approx10,500 per million, roughly five times the average, followed by technology and the women’s section \(≈\\approx5,000\), music, the men’s section, television, cinema and lifestyle \(≈\\approx2,500–4,300\)\. Hard\-news and knowledge sections fall below the average, with science, international news, politics and national news between≈\\approx590 and≈\\approx880 per million\. The spread across sections is therefore of more than an order of magnitude, from≈\\approx10,500 down to≈\\approx592 per million\. The most frequent borrowings nonetheless include domain\-general register words \(look,top,boom,shock\) that recur across all sections, so borrowing is partly stylistic and not only terminological\.

Table 11:Borrowing density by newspaper section, 2020–2026 \(sections with≥\\geq20,000 occurrences, sorted by density\)\. Density is normalized by each section’s own token count\.
### 6\.4Temporal dynamics

We now analyze borrowing density diachronically\. We restrict our analysis to 2023 onwards, as it was the period of time when the same model was used and the collection of outlets remained stable111111The detection model changed in 2022 from the CRF to the current BiLSTM\-CRF model, and the collection of outlets monitored was expanded\.\.

![Refer to caption](https://arxiv.org/html/2608.00713v1/img/fig_density.png)Figure 5:Normalized anglicism density on the composition\-stable core outlets, per million running tokens, by year\. The series begins in 2023, the first full year after both the August 2022 detector change and the September 2022 outlet expansion; 2026 is hatched because it runs only to 24 July, and is excluded from the pooled rate quoted in the text\.Our data suggests that over the three years for which a single model has been in operation, the normalized rate of anglicism use in the Spanish press has remained stable \(Fig\.[5](https://arxiv.org/html/2608.00713#S6.F5)\), with a rate of 1,817 borrowings per million tokens\. This finding should however be taken cautiously: the detector is static, so its recall on borrowings that entered Spanish after training may decline over time\. Because the observatory stores only sentences containing a detection, recall cannot be measured directly on the deployed data \(Section[8](https://arxiv.org/html/2608.00713#S8)\)\. The flatness reported above is therefore best read as an upper bound on any decline and a lower bound on any rise\.

## 7Applications and reuse

Observatorio Lázaro is designed for reuse across communities\. For corpus and theoretical linguistics, its diachronic, frequency\-annotated record supports the study of borrowing diffusion, productivity and establishment\. For computational linguistics, the deployed data and thepylazarolibrary provide training and evaluation material for borrowing and code\-switching detection\. For lexicography and language planning, the continuously updated lexicon is a candidate\-detection feed for dictionaries of neologism\. For teaching, the searchable interface offers attested, dated examples of contemporary borrowing\. What these uses have in common is that they depend on observing the phenomenon as it occurs, at scale and with open access to the underlying data\.

The content of the website and social media accounts provide a social dimension to the project geared towards a non\-specialized public\. In this regard, Observatorio Lázaro seeks to contribute to the public conversation about language change and anglicism usage from a descriptivist perspective\. We hope that the development of Observatorio Lázaro can contribute to raise awareness about the possibilities that machine learning has to offer to the study of language contact and corpus linguistics in general, and the process of lexical borrowing in particular\.

Finally, the purpose of Observatorio Lázaro is to analyze the usage of borrowings in the Spanish press\. This project does not seek to promote or stigmatize the usage of borrowings, or those who use them\. The motivation behind our research is not to defend an alleged linguistic purity, but to study the phenomenon of lexical borrowing from a descriptive and data\-driven point of view\.

### 7\.1Documented uptake

The resource is already in third\-party use\. To date, at least eleven independent publications have drawn on Observatorio Lázaro’s data, mostly within Hispanic linguistics and applied English studies, covering anglicisms in the language of economics\[de2023anglicismos\], Spanish toponymy\[lillo2022anglicismos\], the leisure and tourism domains\[10553\_77321,10553\_127747\], food and drink\[garcia2021anglicisms,lujan2022we,lujan2023anglicisms,lujan2023drink\], information technology\[10553\_114908,NúñezNogueroles\_Luján\-García\_2022\], and sports anglicisms and their metaphorical extensions in the digital press\[lujan2024political\]\. This uptake preceded any formal description of the resource, which is part of the motivation for documenting it here\.

## 8Limitations

Five limitations should guide use\.

#### Detection is automatic and therefore imperfect\.

The precision audit \(Section[5\.3](https://arxiv.org/html/2608.00713#S5.SS3)\) quantifies the error and shows it to be modest and concentrated in the low\-frequency tail, but studies of rare types should apply the reported rates\. Relatedly, the English/other language tag is reliable only for English \(Section[5\.1](https://arxiv.org/html/2608.00713#S5.SS1)\): non\-English borrowings are heavily under\-detected and frequently mislabeled as English, so the resource should be used to study anglicisms rather than borrowing from other languages\.

#### The resource contains a mid\-2022 discontinuity\.

The detection model was upgraded \(CRF→\\rightarrowBiLSTM\-CRF\) in August 2022 and the outlet set was expanded in September 2022, so measurements taken on either side of that boundary are not comparable\. We therefore exclude cross\-changeover comparisons altogether rather than qualify them, and restrict all diachronic statements to the single\-model window 2023–2025 on the stable core outlets \(Section[6\.4](https://arxiv.org/html/2608.00713#S6.SS4)\)\.

#### Recall is unmeasured on the deployed data\.

Because only sentences containing a detection are stored, borrowings the model misses leave no trace, so recall can be estimated only on the curated test set \(0\.82\) and not in the wild\. This matters especially for diachronic claims: the model has been unchanged since August 2022, and its recall on borrowings that entered Spanish after training may well be lower, so measured frequencies \(and above all the apparent influx of new types\) may increasingly under\-represent genuine usage\. Users making claims about emerging or rare borrowings over time should accordingly treat any measured trend as a lower bound on a rise and an upper bound on a decline \(Section[6\.4](https://arxiv.org/html/2608.00713#S6.SS4)\)\. Periodic retraining on freshly annotated data, and a recall probe on full \(undetected\) sentences, are the natural remedies and are planned\.

#### Lemmatization has not been validated against human judgment\.

Lemmatization affects around a tenth of lemma groups, which bounds its influence, but every type\-level figure in Section[6](https://arxiv.org/html/2608.00713#S6)inherits this uncertainty\.

#### The resource is a monitor of the Spanish written press\.

It reflects the register and topical priorities of news media in Spain and should not be treated as representative of the Spanish language as a whole\.

## 9Conclusion

In this paper we have presented Observatorio Lázaro, an observatory of anglicism usage in the Spanish press\. The observatory automatically monitors the presence of English lexical borrowings in a collection of Spanish media sites in real time, and occupies a position between static borrowing dictionaries and one\-off annotated corpora\. The result is a database of over two million borrowings extracted from real context from Spanish press, the largest self\-populating database of its kind\. We are not aware of another continuously growing resource devoted specifically to monitoring borrowing usage in the wild\. It is built on a previously published detector, and what the present paper contributes is the operational system, the diachronic database and the documentation of its evaluation and reuse\.

The pipeline facilitates tracking anglicism frequency over time, and documents the incorporation of novel anglicisms along with their context, which can assist lexicographic work and corpus linguistics research\. The observatory shows that computer\-assisted methods for lexical borrowing detection can successfully be used to build real\-world applications that can inform linguistic work in a systematic and data\-driven fashion\.

## Appendix AOutlet inventory

Table[12](https://arxiv.org/html/2608.00713#A1.T12)lists every outlet monitored by the observatory, with the date on which it entered the record and its contribution in running tokens and borrowing occurrences over the period 2020–2026\.

Table 12:Outlets monitored by Observatorio Lázaro, ordered by token contribution\.
## Declarations

- •Competing interests\.The authors declare no competing interests\.
- •
- •

## References

Similar Articles