A Data-Driven Approach to Idiomaticity Based on Experts' Criteria in Theoretical Linguistics

arXiv cs.CL Papers

Summary

This paper presents a data-driven analysis of multi-word expressions (MWEs) based on 16 theoretical criteria, annotated by linguistics experts, finding that no expressions are absolutely idiomatic and that lexical criteria are most influential.

arXiv:2605.19575v1 Announce Type: new Abstract: The article observes data analysis of 286 multi-word expressions (MWEs) based on 16 lexical, grammatical and other criteria described in theoretical books and papers on the notion of idiomaticity. MWEs were collected from the same theoretical sources, and a set of experts in linguistics annotated them with these categories. The distribution of categories shows that there are no absolutely idiomatic expressions. Lexical criteria seem to be the most influential; grammatical criteria are bound to certain conditions; presence of obsolete words and grammar influence ability of an MWE to be replaced with one word.
Original Article
View Cached Full Text

Cached at: 05/20/26, 08:26 AM

# A Data-Driven Approach to Idiomaticity Based on Experts’ Criteria in Theoretical Linguistics
Source: [https://arxiv.org/html/2605.19575](https://arxiv.org/html/2605.19575)
###### Abstract

The article observes data analysis of 286 multi\-word expressions \(MWEs\) based on 16 lexical, grammatical and other criteria described in theoretical books and papers on the notion of idiomaticity\. MWEs were collected from the same theoretical sources, and a set of experts in linguistics annotated them with these categories\. The distribution of categories shows that there are no absolutely idiomatic expressions\. Lexical criteria seem to be the most influential; grammatical criteria are bound to certain conditions; presence of obsolete words and grammar influence ability of an MWE to be replaced with one word\.

\\NAT@set@cites

A Data\-Driven Approach to Idiomaticity Based on Experts’ Criteria in Theoretical Linguistics

Elena Mikhalkova, Anastasiya Vishnyakova, Anastasiya Drozdova,Polina Gavin, Aleksander Zhmykhov, Timofey ProtasovTyumen State University, Volodarskogo, 6, Tyumen, RussiaAston University, Aston St, Birmingham B4 7ET, Great Britainevrog2009@gmail\.com, a\.y\.vishnyakova@utmn\.ru, a\.o\.drozdova@utmn\.ru160232923@aston\.ac\.uk, linkyspie@gmail\.com, t\.a\.protasov@utmn\.ruAbstract content

## 1\. Introduction

Multi\-word expressions \(MWEs\), groups of lexemes occurring in a text and more complex linguistically, than just a free word\-group, can confuse automatic text processing in many ways\. Even the terminology surrounding them is quite extensive and lacks commonly accepted definitions, hence, often relying on an approach to their treatment\. What unites these approaches is understanding that lexemes in MWEs cannot be treated fully as their equals in free word\-groups\.

Works, summing up recent approaches to MWE processing, are published every now and then\(Pearce,[2002](https://arxiv.org/html/2605.19575#bib.bib1); Piaoet al\.,[2005](https://arxiv.org/html/2605.19575#bib.bib8); Constantet al\.,[2017](https://arxiv.org/html/2605.19575#bib.bib26); Ashoket al\.,[2019](https://arxiv.org/html/2605.19575#bib.bib25)\)\. And, probably, the main world event in this topic is the annual ACL\-affiliated workshop on MWEs\. In recent years, we observe a trend to evaluate methods for particular languages, e\.g\. LREC Workshop Towards a Shared Task for Multiword Expressions \(MWE 2008\) addressed the issue in English, German, Czech, Estonian and other languages\(Grégoireet al\.,[2008](https://arxiv.org/html/2605.19575#bib.bib10)\)\. Later appeared Arabic\(Attiaet al\.,[2010](https://arxiv.org/html/2605.19575#bib.bib12)\), Russian\(Tutubalina,[2015](https://arxiv.org/html/2605.19575#bib.bib34)\), Polish\(Chrząszcz,[2016](https://arxiv.org/html/2605.19575#bib.bib19)\), Spanish and Basque\(Iñurrietaet al\.,[2017](https://arxiv.org/html/2605.19575#bib.bib18)\), Irish\(Walshet al\.,[2019](https://arxiv.org/html/2605.19575#bib.bib17)\), Bulgarian and Romanian\(Barbu Mititeluet al\.,[2019](https://arxiv.org/html/2605.19575#bib.bib16)\), Serbian\(Stankovićet al\.,[2020](https://arxiv.org/html/2605.19575#bib.bib15)\)and many others\. Currently, there is also a trend to discuss translation and alignment\(Lamet al\.,[2015](https://arxiv.org/html/2605.19575#bib.bib20); Fisaset al\.,[2020](https://arxiv.org/html/2605.19575#bib.bib13); Hanet al\.,[2020](https://arxiv.org/html/2605.19575#bib.bib14)\)\. The last LREC workshop on MWEs\(Bhatiaet al\.,[2022](https://arxiv.org/html/2605.19575#bib.bib11)\)addressed heuristic and machine\-learning approaches to their detection, tool\-kits for annotation, low\-resource corpora, figurative language, etc\. However, discussions about the nature of MWEs remain\. In our study, we join these discussions and try to oversee one side of it that has not been granted due attention\.

In our study, we describe theoretical and modern applied approaches to classification of collocations paying more attention to what can be called “data\-driven” approaches\. Second, we propose a model of 16 linguistic criteria that were derived from theoretical works by linguists\. The model encompasses lexical, semantic, grammatical and pragmatic criteria\. Third, we take MWEs from the same works and label them with these criteria \(whether a related feature is manifested in an MWE or not\)\. Why we use these particular method of collecting MWEs is that we are of the opinion that these examples, suggested by theorists, are a gold standard demonstrating typical features of idiomaticity\. Fourth, we group the criteria into four sets and vectorize our corpus so that it can be modelled as a group of points in a 3D cube\. Finally, we observe clusters of these points and conclude about how they are grouped in the multidimensional vector space and what it shows about the nature of idiomaticity\.

## 2\. Approaches to Classification of Collocations

Statistical criteria of “MWEhood”\.Baldwin and Kim \([2010](https://arxiv.org/html/2605.19575#bib.bib2)\)underline that MWEs allow to use a comparatively brief lexicon to create nuances of meaning\. The lack of freedom in MWEs, orvice versastrength of connection, is usually referred to asidiomaticity– “markedness or deviation from the basic properties of the component lexemes” \(Ibid\.\)\. Idiomaticity shows in “lexical, syntactic, semantic, pragmatic, and/or statistical levels” \(Ibid\.\)\.

Statistical methods focus on inferring idiomaticity from co\-occurrence of lexemes inside a certain word\-group and in free contexts\. Among the most common statistical measures used in this task aremutual information\(Church and Hanks,[1990](https://arxiv.org/html/2605.19575#bib.bib5)\),likelihood ratio tests\(Dunning,[1993](https://arxiv.org/html/2605.19575#bib.bib4)\),cost criteria\(Kitaet al\.,[1994](https://arxiv.org/html/2605.19575#bib.bib6)\)\.Pecina \([2008](https://arxiv.org/html/2605.19575#bib.bib9)\)enlists 55 “lexical association measures used for ranking MWE candidates”\. The output of these methods is a number that evaluates the strength of idiomaticity\. Based on it, MWEs are arranged \(ranked\), but are hardly classified\.Vice versathe process usually leads to understanding what type of MWE is better derived with the method\. It would not be wrong to state that there is no universal criterion to MWE extraction, and statistical methods are applied to existing classifications\.

The termcollocation, often used in papers describing statistical methods, can be considered a synonym toMWE, althoughBaldwin and Kim \([2010](https://arxiv.org/html/2605.19575#bib.bib2)\)put it that collocations are statistically idiomatic MWEs\. In fact, we tend to observe that statistical methods extracting this of that type of MWE are designed without limitations\. Hence, any MWE can look statistically significant with a proper measure\. Another way to disambiguate between an MWE and collocation is that some collocations lacknon\-compositionality– they arecompositional, their meaning is easily extracted from lexemes composing them\. E\.g\.Many thanks\!is a statistically significant proper way of being thankful, but its meaning is clear from the words composing it\. The \(non\-\)compositionality cannot be statistically inferred from word frequencies\.

In some literature, MWEs are synonymous tomulti\-word units\(MWUs\), “lexical items that go beyond single word items”\(Shin and Chon,[2019](https://arxiv.org/html/2605.19575#bib.bib21)\), andwords\-with\-spaces, “idiosyncratic interpretations that cross word boundaries \(or spaces\)”\(Saget al\.,[2002](https://arxiv.org/html/2605.19575#bib.bib3)\)\. In our paper, we intentionally do not make any particular distinction between MWEs, collocations, MWUs and words\-with\-spaces, considering them to be manifestations of the same phenomenon – idiomaticity\.

Expert classifications\.Baldwin and Kim \([2010](https://arxiv.org/html/2605.19575#bib.bib2)\)suggest a classification that first splits MWEs into two main classes \(see fig\.[1](https://arxiv.org/html/2605.19575#S2.F1)\): institutionalized phrases are collocations proper \(statistically common phrases such as “Many thanks\!”\) and lexicalized phrases show idiomaticity to a certain degree and are marked by different features \(e\.g\. decomposable / non\-decomposable\)\. We believe that such an approach to classification, although it is supported by examples, cannot be called data\-driven; rather, it is an expert view of a complex phenomenon\. Also, the bottom level of the obtained hierarchy \(VNIC, nominal, VPC, LVC\) is based on parts\-of\-speech analysis of English collocations, hence, binding idiomaticity to one particular linguistic feature\. However, in Table 12\.2 from the same chapterBaldwin and Kim \([2010](https://arxiv.org/html/2605.19575#bib.bib2)\)approach classification of MWEs from another perspective: they enlist properties of MWEs and annotate several examples with these properties, acquiring a matrix of feature distribution from which they conclude about the probability of “MWEhood”, see fig\.[2](https://arxiv.org/html/2605.19575#S2.F2)\. In our opinion, this approach could be called data\-driven, as classes are inferred from annotations\. But the examples were few and were chosen so as to demonstrate several cases referring to pre\-designed classes\.

![Refer to caption](https://arxiv.org/html/2605.19575v1/x1.png)Figure 1:Classification of MWEs by \(Baldwin and Kim, 2010\)\. VNIC \- verb\-noun idiomatic combination; VPC \- verb\-particle construction; LVC \- light\-verb construction\.![Refer to caption](https://arxiv.org/html/2605.19575v1/x2.png)Figure 2:Classification of MWEs in terms of their idiomaticity, by \(Baldwin and Kim, 2010\)\.- •Nominal MWEs - –Multiword named entities - –NN compounds - –Other nominal MWEs
- •Verbal MWEs - –Phrasal verbs - –Light verb constructions - –VP idioms - –Other verbal MWEs
- •Prepositional MWEs
- •Adjectival MWEs
- •MWEs of other categories
- •Proverbs

Resembling classification in figure[1](https://arxiv.org/html/2605.19575#S2.F1), this project lays weight on part\-of\-speech properties of the headword in a phrase and splits the set into distinct subgroups\. We tend to believe that such an approach was organizational: it is important to split jobs in large projects\. Some other possible drawbacks in it are that it has to group non\-standard examples into “other” and that elements in the resulting hierarchy are not of the same level, theoretically\. E\.g\., from the point of view of theoretical linguistics, named entities are represented by proper nouns \(if we disregard anaphora\) as opposed to all common nouns – at the same time, Noun\+Noun compounds are a subgroup of common nouns\.

Another approach, yet leading to abandoning all classifications, is found in\(Schneideret al\.,[2014](https://arxiv.org/html/2605.19575#bib.bib23)\)\. The authors aimed to annotate a corpus for DiMSUM \(Ibid\.\), a SemEval task for detecting minimal semantic units and their meanings\. They collected a set of classes of idioms in English, that totaled 15 as illustrated in their article\. Beside some of the classes, mentioned byBaldwin and Kim \([2010](https://arxiv.org/html/2605.19575#bib.bib2)\), it included named entities, compound words \(motion picture\), support and phrasal verbs \(make decision, cry foul\), coordinated phrases \(cut and dry\), phatic phrases \(You’re welcome\!\), proverbs \(To each his own\), etc\. Annotators in this project only marked what they considered to just be an MWE\. It is of interest that the authors did not find any particular POS pattern that could be associated with a collocation type: “Categorizing MWEs by their coarse POS tag sequence, we find only 8 of these patterns that occur more than 100 times”\(Schneideret al\.,[2014](https://arxiv.org/html/2605.19575#bib.bib23)\)\.

Expert criteria\.Another way to build a classification theoretically is to enlist various linguistic criteria according to which something can be defined as an MWE and terminologically label gold examples demonstrating these criteria\. Such is an approach byVinogradov \([1977](https://arxiv.org/html/2605.19575#bib.bib28)\), who adapted a classification by the Swiss scholar Charles Bally and singled out two types of MWEs – a less and a more idiomatic:

- •combinations \(to conclude an agreement\)
- •fusions - –with archaic words –to eke out - –with archaic grammatical forms –hither and thither - –that were changed so that they do not resemble lexemes from which they were composed –lo and behold\(lofromlook\) - –complete loss of initial meaning –caught red\-handed\(initially meant catching someone who hunted an animal they were not allowed to hunt\)

Outside the scope of his classification,Vinogradov \([1977](https://arxiv.org/html/2605.19575#bib.bib28)\)placed terminological groups and named entities\.

Without building a classification,Manning and Schutze \([1999](https://arxiv.org/html/2605.19575#bib.bib27)\)describe three criteria that characterize collocations: non\-compositionality \(discussed above\), non\-substitutability \(lexemes cannot be substituted with synonyms\), non\-modifiability \(lexemes cannot change grammatically\)\. And again several types are mentioned separately: light verbs, verb\-particle constructions, proper nouns, terminological expressions\.

Cowie and Howarth \([1996](https://arxiv.org/html/2605.19575#bib.bib39)\)suggest the following criteria:

- •familiarity to speakers
- •ability to be stored in memory as ready\-made units
- •limited and arbitrary variability
- •opaque semantics

Tarasevitch \([1991](https://arxiv.org/html/2605.19575#bib.bib29)\)makes use of her own list:

- •stability of use
- •structural separateness
- •complexity of meaning
- •being not built on the generative pattern of free word\-groups

Mel’čuk \([1960](https://arxiv.org/html/2605.19575#bib.bib30)\)considers that idiomaticity influences translation of a phrase or its parts\. In a more idiomatic and stable expression it is hard to find an exact match to every lexeme and the whole phrase is easier to translate with a singe word\.Baldwin and Kim \([2010](https://arxiv.org/html/2605.19575#bib.bib2)\)also mention pragmatic idiomaticity \(being associated with a certain situation\), proverbiality \(describing a situation of social interest\), prosody, but we will leave these criteria outside the scope of our research as they require to go beyond the study of a written text\.

In this paragraph, we might have overlooked some criteria, but as far as we know in other works approximately the same criteria repeat\.

Which approach can be called data\-driven?Aminet al\.\([2021](https://arxiv.org/html/2605.19575#bib.bib31)\), who call their approach data\-driven in the title of their paper, design metrics that help to infer some of the mentioned above criteria for n\-grams in a text\. A similar scheme is traced in\(Rossyaykin and Loukachevitch,[2019](https://arxiv.org/html/2605.19575#bib.bib32)\)who use a set of statistical, context and distributional measures to infer MWEs from a corpus\. Another clustering method, based only on association measure, is found in\(Tutubalina,[2015](https://arxiv.org/html/2605.19575#bib.bib34)\)\.Nissim and Zaninello \([2013](https://arxiv.org/html/2605.19575#bib.bib35)\)introduce variation patterns as an alternative to association measure\.Wahl and Gries \([2018](https://arxiv.org/html/2605.19575#bib.bib36)\), who call their approach “bottom\-up”, introduce the MERGE algorithm, again as an alternative to association measure \(the project is developed byGries \([2022](https://arxiv.org/html/2605.19575#bib.bib38)\)\)\. Summing up, data\-driven are projects in which MWEs are represented as vectors in multi\-dimensional space and analysed statistically\. Often these vectors are visualised in diagrams to see whether there are clusters that attract more MWEs\.

## 3\. Experiment Setup

The intuition lying behind our approach is that the theoretical linguistic expertise about the phenomenon of idiomaticity makes it look like a single unity that can be split into sectors – classes, or types of MWEs\. However, seeing it as an umbrella term that unites phenomena of different linguistic nature333E\.g\., as mentioned, lexical: named entities versus noun compounds; grammatical: diversity of POS\-patterns\.leads us to a hypothesis that we can describe this heterogeneity and outline these phenomena as clusters in a multi\-dimensional space, based on vectorization of annotated examples\. To visualize these clusters, we propose a data\-driven approach\. We believe that we do not need a large collection for such a task if we have a gold standard that manifests the main features of idiomaticity\. Further, in our experiment we propose: a\. a set of criteria \(features of idiomaticity\), b\. the gold standard dataset of MWEs, c\. expert annotation, and d\. a method for clustering visualisation\.

Criteria\.The linguistic criteria that we took from theoretical literature, described earlier, total 15\. We do not claim that this list is full\. The chosen criteria were picked up so as to be able to give a precise instruction on how to annotate an MWE\. E\.g\.subordinate word cannot be replaced with a synonymis checked with an attempt to change a subordinate word in an MWE in several possible contexts, e\.g\.‘‘Я набрал белых грибов\.’’\(I’ve picked some penny buns\.\) does not allow such a change\. An example criterion that we did not include is “Lexemes \(partially\) lose independent meaning” as it is the essence of the idiomaticity and, hence, we view it as the dependent \(target\) variable\. Another left out criterion is POS pattern: as mentioned, its importance has not been determined statistically\.

The criteria were re\-formulated so as to demonstrate idiomaticity, which means that 0 denotes its lack and 1 – its presence\. An annotated example is a vector of zeros and ones, e\.g\. the vector for‘‘белый гриб’’\(En\.penny bun\) is \(1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 1, 0, 1, 1, 1\), see Table[1](https://arxiv.org/html/2605.19575#S3.T1); its vector sum is 9\. The higher is the vector sum of an annotated MWE, the more idiomatic an MWE is\.

We will later describe that we visualize the scores as a 3D cube with resulting vectors\. For this purpose, we grouped the mentioned 16 criteria into four categories: each category is an axis of the cube and the fourth category would be demonstrated with the color\. Grouping was arbitrary, based on the levels of language \(lexicalandgrammatical\) and our analysis of which criteria depend on each other\. E\.g\. if a subordinate word cannot be replaced with a synonym, then it would probably be hard to translate an MWU word by word\. The analysis resulted in singling out two groups:obsolescence– features, related to a long use, e\.g\. obsolete and unique words444To be called archaic, obsolete, a word needs other contexts\. But it is unique and remains only in this particular MWE\.andreplacement– ability of an MWE to “shrink” into a shorter phrase or even into one word\.

The resulting list of groups resembles the list byTarasevitch \([1991](https://arxiv.org/html/2605.19575#bib.bib29)\)and looks as follows:

- •Lexical change
- •Grammatical change
- •Obsolescence
- •Replacement

Table 1:An example of annotation of 16 criteria\. \*By ellipsis, we mean that a word or more can be omitted without change of meaning\. \*\*When words are “sewn” into one word\. In this example, we disregard a very rare adjectiveбелогрибный\. \*\*\*We checked translation into \(sic\!\) English\.Note that inGrammatical changetwo criteria are mutually exclusive \(vi\. Never changes grammatical formandvii\. Only headword changes grammatical form\)\. Hence, an expression cannot score more than 15\. Also, criteria iii\., vii\., xiv\. are not applicable to MWEs with a sentence\-like structure \(containing a subject and/or predicate\) and expressions with coordination due to absence of a headword\. Criterion ii\. is applied to all subordinate words in MWEs: if at least one of them can be replaced with a synonym, the score is 1\. In cases when a criterion is not applicable it was annotated with 0\.

To our annotation, we also added 4 linguistic features that are not directly bound to idiomaticity, but are often found as arbitrary criteria: POS\-pattern, “Is a sentence?” \(whether the MWE contains a subject and/or predicate\), headword \(if it is not a sentence\), phrase structure \(e\.g\. government, agreement, etc\.\)\. An example is given in Table[2](https://arxiv.org/html/2605.19575#S3.T2)\. These features can be used in further research\. All annotated data can be found at[REDACTEDFORANONYMITY](https://arxiv.org/html/2605.19575v1/REDACTEDFORANONYMITY)\.

Table 2:An example of annotation of 4 linguistic features\.Gold standard of MWEs\.For annotation we took 286 examples of MWEs from the works by Russian scholars mentioned in[2](https://arxiv.org/html/2605.19575#S2)\. The choice of the language was conditioned by our team being experts in the Russian language\. However, we later describe how our results can be applied to other languages\. The MWEs were collected without any pre\-selection – all examples that we could find in the papers excluding duplicates\.

Expert annotation\.Our annotators, experts with higher education in linguistics, were instructed about the criteria described above\. One expert annotated a column, then another expert looked it through and checked if they agree with the result\. Questionable cases were discussed at seminars and led to a final decision about the annotation\. However, we expect that the annotation can slightly change due to new arguments about this or that case555It would be impossible to manually test some of the properties in all possible contexts, even using NLP\-tools for support\.\. Also, although the experts used grammar and other reference books, dictionaries, corpora and web to check their expertise, it is possible that they could overlook something, or still disagree even after the final decision about an MWE was made666Many such cases were about the question whether a word is already obsolete\.\. Hence, our annotation should be considered as an expertapproximationof the real world\.

To double\-check some of the experts’ annotation with NLP tools, we used Russian textual corpora777https://ruscorpora\.ru/\. The criterion “v\. Does not allow insertions of lexemes” was checked with an expression “WORD1 \* WORD2” where \* shows that any word can be placed inside an MWE\. The group of criteriaGrammatical changewas checked similarly with \* replacing the grammatical morphemes\. In cases when only one corresponding example was found, we treated it as an occasional creative use of the language and, hence, equal to zero\.

Clustering visualisation\.The aim of our visualisation is to find patterns that can provide us with a idea of how to cluster MWEs in the multi\-dimensional vector space\. To demonstrate it, we use a 3D scatter plot\. Points are calculated as the vector sum in eachgroup\. E\.g\. the vector for‘‘белый гриб’’from Table[1](https://arxiv.org/html/2605.19575#S3.T1)is \(5, 0, 0, 4\)\. Limited by the three dimensions, to visualize the distribution of criteria, we take one group at a time and demonstrate with a color the average score of all MWEs that fall into this point\. For example, only 1 MWE has the same score for the first three groups as‘‘белый гриб’’: \(5, 0, 0\)\. Its fourth score is 1\. The sum of the fourth scores is 4\+1=5\. Divided by the number of MWEs, it is now 2\.5, and this number determines the color of the point\.

## 4\. Data Analysis

We will now describe our resulting dataset and then attempt to find patterns in our 3D visualization\. Due to our criteria being demonstrative of idiomaticity, we expect that the more of them score 1 for a given expression, the more idiomatic it is\. Figure[3](https://arxiv.org/html/2605.19575#S4.F3), left, shows distribution of scores among the vectors\. The distribution is right\-skewed: 191 MWEs \(67%\) score below the median – on average, they do not tend to score much in all the groups of criteria immediately\. And there are no MWEs that score from 13 to 15\. Hence, the perfect MWE does not exist: some of the criteria either exclude others or make it very difficult to combine them in one expression\.

The lowest idiomaticity score \(vector sum\) achieved is 3 – for 14 MWEs\. It is hard to notice any common features in them:богу ведомо, всякой твари по паре, заклятый враг, игральные карты\.\.\(God knows, motley crew, fast foe, playing cards\.\.\)\. The highest score is 12:так и быть\(So be it\!\)\. And close to it are four expressions with the score of 11:была не была; должно быть; ему и горюшка мало\!; при пиковом интересе\(Come what may\!; It must be that\.\.; He doesn’t give a damn\!; \(leave someone\) holding the bag\)\. We note, that this set is more “sentence\-like”, tends to have a subject and a predicate\.

The sum of scores \(ranged from smallest to highest\) for every criterion in Figure[3](https://arxiv.org/html/2605.19575#S4.F3), right, shows that several of them have very low values for the curve to be smooth\. These are:x\. Contains unique lexemes; xi\. Archaic syntax and/or morphology; ix\. Contains lexical archaisms; xiv\. Can be replaced with headword; xvi\. Can be translated with one word\. This can be due to a difficulty in annotating with these criteria\. Translation requires a very good knowledge of foreign language \(English, in our case\); obsolete and unique words need vast expertise in both old and modern Russian; also, it can be hard to distinguish between an archaism and a bookish, formal modern word\.

Some of the criteria score very high \(have positive values for many MWEs – Figure[3](https://arxiv.org/html/2605.19575#S4.F3), right, the right part of the curve\), and that means for our task that they do not contribute to a stricter division between probable classes \(they fill the vector space densely\)\. The top\-scoring \(above 200\) criteria are:i\. Lexemes \(partially\) lose independent meaning \(207 MWEs\), iv\. Cannot be translated word by word \(212 MWEs\), iii\. Headword cannot be replaced with a synonym \(233 MWEs\), xv\. Can be replaced with one word \(234 MWEs\)\. Intuitively, Criterion i\. appears like it is what theorists look for in all idiomatic expressions, regardless of their type\. Also, these criteria seem to be based more on lexical rather than grammatical properties\.

![Refer to caption](https://arxiv.org/html/2605.19575v1/x3.png)Figure 3:Number of MWEs scoring from 3 to 13 – left\. Sum of scores for each of the 16 categories \(ranged\) – right\.### 4\.1\. Patterns in Distribution of Vectors

Out of 286 MWEs 216 are unique vectors \(76%\)\. It is hard to say whether our annotation exhibits strong general patterns in distribution of all the criteria across the dataset\. We will try to generalize about what happens in each of the defined groups\.

Lexical change\.This group is the champion in earned scores: 954 \(50% of all scores\)\. Which, probably, means that it embodies idiomaticity and should not be considered as a basis of “MWEhood”\. I\.e\. any candidate for an MWE should score high in it\. An example of a low\-scorer in this group isбеспокойный человекa restless person\. This is what we earlier called an MWU, an expression that learners of a language memorize to sound more like native speakers, and it scores only 4 in all categories\. We believe thatLexical changeshould be used for classifying between MWEs and non\-MWEs and, probably, for defining MWUs\.

Grammatical change\.This category has a very high score \(481, 25%\) and can be called influential as well\. Although there can be the mentioned bias in selection of examples: those selected by scholars tend to be more fixed structurally, to better demonstrate fixedness\. Expressions that score 0 in it are, for example, \(беспробудное пьянство, потрясающее впечатлениеdeep drinking, stunning impression\) – also typical MWUs\. Also, presence of a verb usually makes expressions in Russian \(as well as in English\) less fixed\. Compare:пуститься во все тяжкие – во все тяжкиеdo whatever it takes – whatever it takes; the verbdocan change its grammatical form and can be substituted with a synonym\. This leads us to one of the main conclusions that what is singled out as an MWE can be be several expressions of a different degree of idiomaticity\.

Obsolescence\.As we mentioned earlier, this category seldom earned a positive decision from our annotators\. There are only 9 collocations that score 3 in it \(e\.g\.и вся недолга\!; не до жиру, быть бы живу; ничтоже сумняшесяend of story, survive before thrive; nothing doubting\)\. The collocations from our example score 8 and more in all criteria \(with 7\.5 being the median value, and 7\.4 – mean\)\. Although it is probably clear without the data, but our experiment supports it that this category is a strong marker of idiomaticity\. To add, obsolete words make it harder to modify an expression; archaic grammar hinders morphological and other changes in different contexts\.

Replacement\.Counter\-intuitively, this category seems to be in a conflict withGrammatical changeandObsolescence\. Often, if an MWE scores high in it, it scores low \(or medium\) in one or both\. The example in Table[1](https://arxiv.org/html/2605.19575#S3.T1)demonstrates it\. There are 231 expressions that score less than 3 both inObsolescenceandReplacementwhich might mean that these two categories are not crucially important in formation of an WME, but they probably point at two distinct sub\-classes\.

3D\-model of vector distributions\.We now want to demonstrate how the vectors are distributed in a multi\-dimensional space with a 3D scatter plot\. We described in Chapter[3](https://arxiv.org/html/2605.19575#S3)how the cubes were built\. In each of the four cubes in Diagram[4](https://arxiv.org/html/2605.19575#S4.F4), color demonstrates the average value of the fourth category\. The angle is chosen so as to illustrate our hypotheses\.

![Refer to caption](https://arxiv.org/html/2605.19575v1/x4.png)Figure 4:3D scatter plot of the sums of scores of MWEs in each category\.The top\-left cube where the color points the average value ofLexical changeshows that there is little correlation between it andReplacement: with any value ofReplacementLexical changecan be high \(yellow and green points\)\. It positively correlates withGrammatical changeandObsolescence\.

It looks likeGrammatical change\(top\-right\) requires certain conditions to score high: \[2:4\] forLexical changeand \[1:2\] forReplacement\.

Obsolescence\(bottom\-left\) positively correlates withLexical change, but it is unclear how it correlates with the two other groups\.

And the groupReplacement, bottom\-right, also positively correlates withLexical changeand negatively correlates withObsolescence\. It does not seem to correlate withGrammatical change\.

## 5\. Conclusion

The paper describes an attempt to search for theoretical grounds in the notion of idiomaticity with the help of linguistic annotation and data analysis of the gold standard MWEs\. We have proposed a model of 16 criteria that were grouped into four categories based on linguistic analysis\. We annotated a corpus of 286 Russian MWEs found in the same theoretical books from which we took the criteria\. Our analysis revealed several trends:

- •the group of criteria that we calledLexical changepositively correlates with other criteria which makes it look rather like a target category and an idiomaticity “test”\. A better proof requires comparison to non\-MWEs;
- •either scholars tend to choose MWEs that show mediumGrammatical changeor it is a stable property of all such expressions;
- •some MWE criteria are, probably, harder to annotate \(e\.g\. identifying obsolete words and contexts allowing to replace an MWE\);
- •there is no or negative correlation between presence of archaic words and grammar and the property of an expression to be shortened or replaced by a single word;
- •what is determined as an MWE can be several expressions with a different degree of idiomaticity\.

Relation to other languages\.It is hard to say whether the same conclusions will stand for other languages\. As we mentioned, some of the described features can be already observed in English, e\.g\. verbs in MWEs make expressions less idiomatic\. Our annotated dataset contains some translation into English, and also examples can be taken from the studied English literature to create a similar corpus\. A further extension can lie in application of NLP\-tools to automatic annotation of such corpora as PARSEME collections with the criteria that we described\.

Future work\. Future work requires a larger annotated collection\. The main obstacle here is that manual annotation that we performed is very time\-consuming\. We can see several more ways, beside the mentioned ones, of making it partially automatic, e\.g\. impossibility of translating an MWE word by word as well as translation with one word can be checked in parallel corpora\. Also, the criteria should be checked in free word groups\. And finally the stated correlations require quantitative analysis\.

## References

- Data\-driven identification of idioms in song lyrics\.InProceedings of the 17th Workshop on Multiword Expressions \(MWE 2021\),pp\. 13–22\.Cited by:[§2](https://arxiv.org/html/2605.19575#S2.p20.1)\.
- A\. Ashok, R\. Elmasri, and G\. Natarajan \(2019\)Comparing different word embeddings for multiword expression identification\.InNatural Language Processing and Information Systems: 24th International Conference on Applications of Natural Language to Information Systems, NLDB 2019, Salford, UK, June 26–28, 2019, Proceedings 24,pp\. 295–302\.Cited by:[§1](https://arxiv.org/html/2605.19575#S1.p2.1)\.
- M\. Attia, A\. Toral, L\. Tounsi, P\. Pecina, and J\. Van Genabith \(2010\)Automatic extraction of arabic multiword expressions\.InProceedings of the 2010 Workshop on Multiword Expressions: from Theory to Applications,pp\. 19–27\.Cited by:[§1](https://arxiv.org/html/2605.19575#S1.p2.1)\.
- T\. Baldwin and S\. N\. Kim \(2010\)Multiword expressions\.InHandbook of natural language processing,N\. Indurkhya and F\. J\. Damerau \(Eds\.\),pp\. 267–292\.Cited by:[§2](https://arxiv.org/html/2605.19575#S2.p1.1),[§2](https://arxiv.org/html/2605.19575#S2.p18.1),[§2](https://arxiv.org/html/2605.19575#S2.p3.1),[§2](https://arxiv.org/html/2605.19575#S2.p5.1),[§2](https://arxiv.org/html/2605.19575#S2.p9.1)\.
- V\. Barbu Mititelu, I\. Stoyanova, S\. Leseva, M\. Mitrofan, T\. Dimitrova, and M\. Todorova \(2019\)Hear about verbal multiword expressions in the Bulgarian and the Romanian wordnets straight from the horse’s mouth\.InProceedings of the Joint Workshop on Multiword Expressions and WordNet \(MWE\-WN 2019\),Florence, Italy,pp\. 2–12\.External Links:[Link](https://aclanthology.org/W19-5102),[Document](https://dx.doi.org/10.18653/v1/W19-5102)Cited by:[§1](https://arxiv.org/html/2605.19575#S1.p2.1)\.
- A\. Bhatia, P\. Cook, S\. Taslimipoor, M\. Garcia, and C\. Ramisch \(Eds\.\) \(2022\)Proceedings of the 18th workshop on multiword expressions @lrec2022\.European Language Resources Association,Marseille, France\.External Links:[Link](https://aclanthology.org/2022.mwe-1.0)Cited by:[§1](https://arxiv.org/html/2605.19575#S1.p2.1)\.
- P\. Chrząszcz \(2016\)Extraction and recognition of Polish multiword expressions using Wikipedia and finite\-state automata\.InProceedings of the 12th Workshop on Multiword Expressions,Berlin, Germany,pp\. 96–106\.External Links:[Link](https://aclanthology.org/W16-1815),[Document](https://dx.doi.org/10.18653/v1/W16-1815)Cited by:[§1](https://arxiv.org/html/2605.19575#S1.p2.1)\.
- K\. Church and P\. Hanks \(1990\)Word association norms, mutual information, and lexicography\.Computational linguistics16\(1\),pp\. 22–29\.Cited by:[§2](https://arxiv.org/html/2605.19575#S2.p2.1)\.
- M\. Constant, G\. Eryiğit, J\. Monti, L\. Van Der Plas, C\. Ramisch, M\. Rosner, and A\. Todirascu \(2017\)Multiword expression processing: a survey\.Computational Linguistics43\(4\),pp\. 837–892\.Cited by:[§1](https://arxiv.org/html/2605.19575#S1.p2.1)\.
- A\. P\. Cowie and P\. Howarth \(1996\)Phraseological competence and written proficiency\.British studies in applied linguistics11,pp\. 80–93\.Cited by:[§2](https://arxiv.org/html/2605.19575#S2.p14.1)\.
- T\. Dunning \(1993\)Accurate methods for the statistics of surprise and coincidence\.Computational linguistics19\(1\),pp\. 61–74\.Cited by:[§2](https://arxiv.org/html/2605.19575#S2.p2.1)\.
- B\. Fisas, L\. Espinosa Anke, J\. Codina\-Filbá, and L\. Wanner \(2020\)CollFrEn: rich bilingual English–French collocation resource\.InProceedings of the Joint Workshop on Multiword Expressions and Electronic Lexicons,online,pp\. 1–12\.External Links:[Link](https://aclanthology.org/2020.mwe-1.1)Cited by:[§1](https://arxiv.org/html/2605.19575#S1.p2.1)\.
- N\. Grégoire, S\. Evert, and B\. Krenn \(2008\)Proceedings of the lrec workshop: towards a shared task for multiword expressions \(mwe 2008\)\.European Language Resources Association \(ELRA\) Marrakech, Morocco\.Cited by:[§1](https://arxiv.org/html/2605.19575#S1.p2.1)\.
- S\. T\. Gries \(2022\)Multi\-word units \(and tokenization more generally\): a multi\-dimensional and largely information\-theoretic approach\.Lexis\. Journal in English Lexicology\(19\)\.Cited by:[§2](https://arxiv.org/html/2605.19575#S2.p20.1)\.
- L\. Han, G\. Jones, and A\. Smeaton \(2020\)AlphaMWE: construction of multilingual parallel corpora with MWE annotations\.InProceedings of the Joint Workshop on Multiword Expressions and Electronic Lexicons,online,pp\. 44–57\.External Links:[Link](https://aclanthology.org/2020.mwe-1.6)Cited by:[§1](https://arxiv.org/html/2605.19575#S1.p2.1)\.
- U\. Iñurrieta, I\. Aduriz, A\. Díaz de Ilarraza, G\. Labaka, and K\. Sarasola \(2017\)Rule\-based translation of Spanish verb\-noun combinations into Basque\.InProceedings of the 13th Workshop on Multiword Expressions \(MWE 2017\),Valencia, Spain,pp\. 149–154\.External Links:[Link](https://aclanthology.org/W17-1720),[Document](https://dx.doi.org/10.18653/v1/W17-1720)Cited by:[§1](https://arxiv.org/html/2605.19575#S1.p2.1)\.
- K\. Kita, Y\. Kato, T\. Omoto, and Y\. Yano \(1994\)A comparative study of automatic extraction of collocations from corpora: mutual information vs\. cost criteria\.Journal of Natural Language Processing1\(1\),pp\. 21–33\.Cited by:[§2](https://arxiv.org/html/2605.19575#S2.p2.1)\.
- K\. N\. Lam, F\. Al Tarouti, and J\. Kalita \(2015\)Phrase translation using a bilingual dictionary and n\-gram data: a case study from Vietnamese to English\.InProceedings of the 11th Workshop on Multiword Expressions,Denver, Colorado,pp\. 65–69\.External Links:[Link](https://aclanthology.org/W15-0911),[Document](https://dx.doi.org/10.3115/v1/W15-0911)Cited by:[§1](https://arxiv.org/html/2605.19575#S1.p2.1)\.
- C\. Manning and H\. Schutze \(1999\)Foundations of statistical natural language processing\.MIT press\.Cited by:[§2](https://arxiv.org/html/2605.19575#S2.p13.1)\.
- I\.A\. Mel’čuk \(1960\)On the terms "stability" and "idiomaticity" \(o terminakh "ustoichivost" i "idiomatichnost"\)\.Voprosy Yazykoznaniya4,pp\. 73–80\.Cited by:[§2](https://arxiv.org/html/2605.19575#S2.p18.1)\.
- M\. Nissim and A\. Zaninello \(2013\)Modeling the internal variability of multiword expressions through a pattern\-based method\.ACM Transactions on Speech and Language Processing \(TSLP\)10\(2\),pp\. 1–26\.Cited by:[§2](https://arxiv.org/html/2605.19575#S2.p20.1)\.
- D\. Pearce \(2002\)A comparative evaluation of collocation extraction techniques\.InProceedings of the Third International Conference on Language Resources and Evaluation \(LREC’02\),Las Palmas, Canary Islands \- Spain\.External Links:[Link](http://www.lrec-conf.org/proceedings/lrec2002/pdf/169.pdf)Cited by:[§1](https://arxiv.org/html/2605.19575#S1.p2.1)\.
- P\. Pecina \(2008\)A machine learning approach to multiword expression extraction\.InProceedings of the LREC Workshop Towards a Shared Task for Multiword Expressions \(MWE 2008\),Vol\.2008,pp\. 54–61\.Cited by:[§2](https://arxiv.org/html/2605.19575#S2.p2.1)\.
- S\. S\. Piao, P\. Rayson, D\. Archer, and T\. McEnery \(2005\)Comparing and combining a semantic tagger and a statistical tool for mwe extraction\.Computer Speech & Language19\(4\),pp\. 378–397\.Cited by:[§1](https://arxiv.org/html/2605.19575#S1.p2.1)\.
- P\. Rossyaykin and N\. Loukachevitch \(2019\)Measure clustering approach to mwe extraction\.InKomp’juternaja Lingvistika i Intellektual’nye Tehnologii,pp\. 562–575\.Cited by:[§2](https://arxiv.org/html/2605.19575#S2.p20.1)\.
- I\. A\. Sag, T\. Baldwin, F\. Bond, A\. Copestake, and D\. Flickinger \(2002\)Multiword expressions: a pain in the neck for nlp\.InComputational Linguistics and Intelligent Text Processing: Third International Conference, CICLing 2002 Mexico City, Mexico, February 17–23, 2002 Proceedings 3,pp\. 1–15\.Cited by:[§2](https://arxiv.org/html/2605.19575#S2.p4.1)\.
- N\. Schneider, S\. Onuffer, N\. Kazour, E\. Danchik, M\. T\. Mordowanec, H\. Conrad, and N\. A Smith \(2014\)Comprehensive annotation of multiword expressions in a social web corpus\.Cited by:[§2](https://arxiv.org/html/2605.19575#S2.p9.1)\.
- D\. Shin and Y\. Chon \(2019\)A multiword unit analysis coca multiword unit list 20 and collogram\.Journal of Asia TEFL16,pp\. 608–623\.External Links:[Document](https://dx.doi.org/10.18823/asiatefl.2019.16.2.11.608)Cited by:[§2](https://arxiv.org/html/2605.19575#S2.p4.1)\.
- R\. Stanković, J\. Mitrović, D\. Jokić, and C\. Krstev \(2020\)Multi\-word expressions for abusive speech detection in Serbian\.InProceedings of the Joint Workshop on Multiword Expressions and Electronic Lexicons,online,pp\. 74–84\.External Links:[Link](https://aclanthology.org/2020.mwe-1.10)Cited by:[§1](https://arxiv.org/html/2605.19575#S1.p2.1)\.
- M\. Tarasevitch \(1991\)Soviet phraseology: problems in the analysis and teaching of idioms\.Linguistics and Language Pedagogy: The State of the Art,pp\. 484–488\.Cited by:[§2](https://arxiv.org/html/2605.19575#S2.p16.1),[§3](https://arxiv.org/html/2605.19575#S3.p5.1)\.
- E\. Tutubalina \(2015\)Clustering\-based approach to multiword expression extraction and ranking\.InProceedings of the 11th Workshop on Multiword Expressions,pp\. 39–43\.Cited by:[§1](https://arxiv.org/html/2605.19575#S1.p2.1),[§2](https://arxiv.org/html/2605.19575#S2.p20.1)\.
- V\.V\. Vinogradov \(1977\)Ob osnovnikh tipakh phraseologicheskikh edinits v russkom yazyke \(on the main types of phraseological units in the russian language\)\.InIzbranniye Trudy: Lexicologia i Lexikographia \(Selected Works on Lexicology and Lexicography,pp\. 140–161\.Cited by:[§2](https://arxiv.org/html/2605.19575#S2.p11.1),[§2](https://arxiv.org/html/2605.19575#S2.p12.1)\.
- A\. Wahl and S\. T\. Gries \(2018\)Multi\-word expressions: a novel computational approach to their bottom\-up statistical extraction\.Lexical collocation analysis: advances and applications,pp\. 85–109\.Cited by:[§2](https://arxiv.org/html/2605.19575#S2.p20.1)\.
- A\. Walsh, T\. Lynn, and J\. Foster \(2019\)Ilfhocail: a lexicon of Irish MWEs\.InProceedings of the Joint Workshop on Multiword Expressions and WordNet \(MWE\-WN 2019\),Florence, Italy,pp\. 162–168\.External Links:[Link](https://aclanthology.org/W19-5120),[Document](https://dx.doi.org/10.18653/v1/W19-5120)Cited by:[§1](https://arxiv.org/html/2605.19575#S1.p2.1)\.

Similar Articles

IdioLink: Retrieving Meaning Beyond Words Across Idiomatic and Literal Expressions

arXiv cs.CL

Introduces IdioLink, a retrieval benchmark of 10,700 documents and 2,140 queries across 107 idioms that tests whether models can link idiomatic expressions to conceptually equivalent literal or paraphrased meanings. Evaluations show current embedding models struggle with this task, highlighting gaps in idiom-aware semantic retrieval.

Conceptual Networks for Cross-Linguistic Idiomatic Expressions:A Feature-Based Graph Approach

arXiv cs.CL

This paper presents an interpretable network-based framework for representing idiomatic expressions across eight languages using binary conceptual features. Community detection reveals that idioms cluster by conceptual schema rather than language, and the framework improves downstream idiom detection and cross-lingual transfer over embedding-based baselines.