Normalisation-Based Likelihood Ratio Estimation for Forensic Authorship Verification
Summary
This paper proposes two novel normalisation techniques (Square Root Correction and Hapax Correction) for deriving likelihood ratios from the LambdaG authorship verification method without needing a separate calibration model, reducing data requirements and complexity while achieving comparable or superior performance in forensic text comparison.
View Cached Full Text
Cached at: 07/13/26, 07:58 AM
# Normalisation-Based Likelihood Ratio Estimation for Forensic Authorship Verification
Source: [https://arxiv.org/html/2607.09501](https://arxiv.org/html/2607.09501)
###### Abstract
Authorship verification \(AV\) is the task of determining whether two texts were written by the same author\. In a forensic context, the strength of AV evidence can be quantified using likelihood ratios\. Most AV methods are score\-based and deriving well\-calibrated likelihood ratios from these scores requires a separate calibration model\. This, in turn, requires additional amounts of case\-relevant data, which is often time\-consuming to obtain and prepare\. This study proposes two novel normalisation techniques, the Square Root Correction and the Hapax Correction, for deriving likelihood ratios from the AV methodLambdaGwithout the need of a calibration model\(Niniet al\.,[2026](https://arxiv.org/html/2607.09501#bib.bib112)\)\. These corrections are designed to mitigate the overestimation of evidential strength that may result from long or highly repetitive texts\. Performance is evaluated against logistic regression calibration across fifteen corpora and a range of text lengths \(100\-9,500 tokens\), using the log\-likelihood ratio cost \(CllrC\_\{llr\}\)\. The proposed methods achieve performance comparable to logistic regression calibration, with the Hapax Correction outperforming it in approximately 45% of tests \(weighted by corpora\)\. Furthermore, performance was more frequently close \(within 5%\) when the Hapax Correction was outperformed by logistic regression calibration, compared with the reverse comparison\. Eliminating the need to train a calibration model reduces data\-requirements, time and complexity, thereby increasing the accessibility and transparency of forensic text comparison\. This combination of empirical performance and practical advantages supports the adoption of the proposed methods in forensic settings\.
###### keywords:
Authorship Verification , Calibration , Likelihood Ratios
††journal:Forensic Science International\\affiliation
\[1\]organization=The University of Manchester, Department of Linguistics and English Language,addressline=Oxford Road,city=Manchester,postcode=M13 9PL,postcodesep=\\affiliation\[2\]organization=The University of Manchester, Department of Computer Science,addressline=Oxford Road,city=Manchester,postcode=M13 9PL,postcodesep=
t1t1footnotetext:This work was supported by the North West Social Science Doctoral Training Partnership \(NWSSDTP\) \[ES/P000665/1\]\.## 1Introduction
Across forensic science, the Likelihood Ratio Framework has emerged as the ’logically and legally’ endorsed standard for evaluation evidenceIshiharaet al\.,[2022](https://arxiv.org/html/2607.09501#bib.bib46), p\.183;Forensic Science Regulator,[2021](https://arxiv.org/html/2607.09501#bib.bib26), p\.26\. The framework requires the comparison of the probability of evidence under two competing hypotheses\. While the Likelihood Ratio Framework is firmly established in DNA analysis\(Balding and Nichols,[1994](https://arxiv.org/html/2607.09501#bib.bib5); Tayloret al\.,[2013](https://arxiv.org/html/2607.09501#bib.bib99)\)and well\-validated in forensic voice comparison\(Rose,[2006](https://arxiv.org/html/2607.09501#bib.bib87); Morrison,[2011](https://arxiv.org/html/2607.09501#bib.bib68)\), its application in forensic text comparison remains comparatively nascent\(Ishihara,[2017](https://arxiv.org/html/2607.09501#bib.bib49); Grant,[2022](https://arxiv.org/html/2607.09501#bib.bib35); Niniet al\.,[2026](https://arxiv.org/html/2607.09501#bib.bib112)\)\.
For score\-based systems, Log\-Likelihood Ratios \(LLRs\) can be derived through a process calledcalibrationMorrison,[2013](https://arxiv.org/html/2607.09501#bib.bib75), p\.174;van der Vloed,[2024](https://arxiv.org/html/2607.09501#bib.bib101)\. This involves training a separate calibration model, which learns the necessary transformation to align predicted probabilities with empirical outcome frequenciesGuoet al\.,[2017](https://arxiv.org/html/2607.09501#bib.bib38);Silva Filhoet al\.,[2023](https://arxiv.org/html/2607.09501#bib.bib90), p\.3215;Pull and Hurlin,[2025](https://arxiv.org/html/2607.09501#bib.bib85), p\.2\.
In practice, calibration is most often achieved using logistic regression\(Brümmer and du Preez,[2006](https://arxiv.org/html/2607.09501#bib.bib14); Gonzalez\-Rodriguezet al\.,[2007](https://arxiv.org/html/2607.09501#bib.bib33); Morrison,[2013](https://arxiv.org/html/2607.09501#bib.bib75)\)\(detailed in Section[3](https://arxiv.org/html/2607.09501#S3)\), which requires comprehensive and case\-specific training data\. Compiling such datasets can often be complex and computationally demanding\. Data should ideally reflect the case conditions\(Morrison,[2024](https://arxiv.org/html/2607.09501#bib.bib65)\), however, what this means is not obvious\.
To demonstrate whether the calibration data was sufficient to produce properly calibrated LLRs, additional tests must be performed\. If a LLR is not properly calibrated, it is not considered to be valid and reliable forensic evidenceRamos and Gonzalez\-Rodriguez,[2013](https://arxiv.org/html/2607.09501#bib.bib86), p\.164;Vergeeret al\.,[2021](https://arxiv.org/html/2607.09501#bib.bib106), p\.14\. In many forensic disciplines, including forensic text comparison, calibration is therefore an essential step for producing meaningful LLRs\.
Many cases of forensic text comparison concern authorship verification \(AV\)\. This is the task of determining whether a text of disputed authorship was authored by the same individual as a text of known authorshipKoppelet al\.,[2012](https://arxiv.org/html/2607.09501#bib.bib59), p\.284;Potha and Stamatatos,[2014](https://arxiv.org/html/2607.09501#bib.bib84);Juola,[2021](https://arxiv.org/html/2607.09501#bib.bib52)\. AV is possible because individuals consistently reuse linguistic patterns\(Coulthard,[2004](https://arxiv.org/html/2607.09501#bib.bib21); Mollin,[2009](https://arxiv.org/html/2607.09501#bib.bib64); Wright,[2017](https://arxiv.org/html/2607.09501#bib.bib109)\), and the unique combination of these patterns can be used to identify an author\.
A recently proposed AV method, LambdaG, is the focus of this study\(Niniet al\.,[2026](https://arxiv.org/html/2607.09501#bib.bib112)\)\. LambdaG is outlined in Section[4](https://arxiv.org/html/2607.09501#S4)\. This method returns a score in the form of an uncalibrated LLR\. To ensure that the score is interpretable, a separate post\-hoc calibration is still required\.
This paper investigates the nature of AV scores and examines whether normalisation, the scaling of values to a common scale, offers a functionally equivalent alternative to calibration in this context\. In Section[5](https://arxiv.org/html/2607.09501#S5), we introduce two novel normalisation techniques \- the Square Root Correction, which addresses the effect of text length on AV scores, and the Hapax Correction, which adjusts scores based on a measure of uniqueness relative to size\.
These techniques are evaluated using 15 diverse datasets, spanning academic papers, newspaper articles, blog posts, four corpora of reviews, Wikipedia talk pages, two forum corpora, two email corpora, chat logs, text messages , and tweets\. For a subset of the corpora, the length of the text was systematically varied, ranging from100100to9,5009,500tokens\. This allowed us to test whether text length introduces a bias into the normalisation of scores\.
The results of this evaluation are summarised in Section[7](https://arxiv.org/html/2607.09501#S7)\. They indicate that, on average across the corpora, the best performing normalisation technique \(the Hapax correction\) outperforms the baseline of logistic regression calibration in 45\.4% of tests\. In 56\.91% of losses, this correction performed comparably to the baseline, remaining within 5% of logistic regression calibration\.
By eliminating the need for a separate calibration model, whilst maintaining comparable performance, the limitations that calibration models entail are also removed\. In particular, data requirements are reduced and in turn, the time cost and complexity of the method are lowered\. In sum, this research contributes to the development of more accessible, efficient and interpretable methods for forensic AV\.
## 2Likelihood Ratios
A LLR is a measure of evidential value, quantifying the relative probability of observing given evidence under two opposing hypotheses\(Biedermannet al\.,[2016](https://arxiv.org/html/2607.09501#bib.bib10)\)\. It is defined as the logarithm of the ratio of the probability of the observed evidence \(EE\) if the prosecution hypothesis \(HpH\_\{p\}\) is true, to the probability of the evidence if the defence hypothesis \(HdH\_\{d\}\) is true\(Ishihara,[2021](https://arxiv.org/html/2607.09501#bib.bib48), eq\. 1\):
LLR=logP\(E\|Hp\)P\(E\|Hd\)\{LLR=\\log\\frac\{P\(E\|H\_\{p\}\)\}\{P\(E\|H\_\{d\}\)\}\}\(1\)
In the context of forensic AV, the two opposing hypotheses may be:
> 𝐇p\\mathbf\{H\}\_\{p\}: the candidate author wrote the disputed document
> 𝐇d\\mathbf\{H\}\_\{d\}: the candidate author did not write the disputed document
The magnitude of the LLR represents the strength of the evidence in support of one hypothesis over the other\.
In forensic text comparison, however, leading methods typically output similarity scores rather than calibrated LLRs\(Evertet al\.,[2017](https://arxiv.org/html/2607.09501#bib.bib25); Grieveet al\.,[2019a](https://arxiv.org/html/2607.09501#bib.bib36)\)\. Currently, no AV method produces a well\-calibrated LLR without a separate calibration model\.
For example, the Impostors Method, which is one of the most widely known AV methods, produces a score between 0 and 1 that needs to be calibrated into a LLR\(Koppel and Winter,[2014](https://arxiv.org/html/2607.09501#bib.bib58)\)\. Similarly, LUAR and STAR, state\-of\-the\-art AV systems based on neural networks, produce cosine similarity scores, which, without calibration into a LLR, are suitable only for comparative evaluation\(Rivera\-Sotoet al\.,[2021](https://arxiv.org/html/2607.09501#bib.bib72); Huertas\-Tatoet al\.,[2024](https://arxiv.org/html/2607.09501#bib.bib45)\)\. Furthermore, a recently proposed AV method, LambdaG, produces a score in the form of an uncalibrated LLR, a value that is constructed as a LLR but cannot be interpreted probabilistically without further processing\(Niniet al\.,[2026](https://arxiv.org/html/2607.09501#bib.bib112)\)\.
Therefore, all existing AV methods require calibration, as we explain in the following section\.
## 3Calibration
Calibration involves the adjustment of raw model scores through transformations such as scaling and shifting\(Morrisonet al\.,[2011](https://arxiv.org/html/2607.09501#bib.bib67), p\.61\), so that the resulting values allow meaningful LLR interpretation\. For a perfectly calibrated system, calculating the LLR of the systems output returns the same value as the output itself\(Morrison,[2024](https://arxiv.org/html/2607.09501#bib.bib65)\)\.
### 3\.1Applied Calibration
To elucidate the concept of calibration and its role in AV, consider the analogy of weather forecasting\(Dawid,[1982](https://arxiv.org/html/2607.09501#bib.bib22); DeGroot and Fienberg,[1983](https://arxiv.org/html/2607.09501#bib.bib23)\)\. Intuitively, a well\-calibrated weather forecasting system is one where 70% of the time that there is a 70% chance of rain, it rains\(DeGroot and Fienberg,[1983](https://arxiv.org/html/2607.09501#bib.bib23), p\.13\)\. The calibrated forecast is therefore producing a prediction that corresponds to observed relative frequencies\(Silva Filhoet al\.,[2023](https://arxiv.org/html/2607.09501#bib.bib90), p\.3215\), rendering it suitable for probabilistic interpretation\.
In contrast, if rain follows a 70% forecast only 60% of the time, the model is systematically overestimating the likelihood of rainfall\. In this case, the reported probabilities cannot be interpreted as reliable estimates of empirical frequency, indicating the need for calibration\.
For LLR systems that are not inherently well\-calibrated, calibration involves training a model to transform raw outputs into interpretable values on an appropriate scale\. In weather forecasting, the calibration model is trained on historical forecast data paired with observed weather outcomes\(Wilks,[2009](https://arxiv.org/html/2607.09501#bib.bib107)\)\. The model learns parameters that transform raw predictions to better align with real\-world results\. This enables the forecast to be interpreted in probabilistic terms that can then be reliably used for decision making \- e\.g\. “how likely is it to rain, and thus, should I carry an umbrella?”
### 3\.2Calibration in AV
As in weather forecasting, the outputs of AV systems are not always well\-calibrated\. An AV system may produce an output that resembles a Likelihood Ratio \(LR\), in that it is comprised of similarity and typicality measures, but the probabilities \(P\(E\|Hp\)P\(E\|H\_\{p\}\)andP\(E\|Hd\)P\(E\|H\_\{d\}\)\) may be inaccurate\.
This could reflect that even when an AV system is grounded in cognitive linguistic theory\(Nini,[2023](https://arxiv.org/html/2607.09501#bib.bib80)\), the complexity and incomplete understanding of individual language production constrain the AV model to simplify the underlying generative mechanisms\(Argamon,[2008](https://arxiv.org/html/2607.09501#bib.bib2); Niniet al\.,[2026](https://arxiv.org/html/2607.09501#bib.bib112)\)\. If the assumptions of the model are not accurate, it is unlikely that the estimated probabilities will reflect the true probability of the observed evidence\.
While uncalibrated LLRs may correctly order the strength of evidence, their numerical magnitudes are not interpretable\. This means that the values do not reliably represent the likelihood of the evidence given each hypothesis\.
To transform the system’s outputs into a calibrated LLR, a further process is therefore required\. Most often, this is performed using logistic regression calibrationIshihara,[2011](https://arxiv.org/html/2607.09501#bib.bib47), p\.51;Morrison,[2013](https://arxiv.org/html/2607.09501#bib.bib75)\.
### 3\.3Logistic Regression Calibration
Logistic regression calibration was first established for speaker recognition\(Gonzalez\-Rodriguezet al\.,[2007](https://arxiv.org/html/2607.09501#bib.bib33); Brümmer and Doddington,[2013](https://arxiv.org/html/2607.09501#bib.bib15); Morrison,[2013](https://arxiv.org/html/2607.09501#bib.bib75)\)and is now applied across a variety of forensic science disciplines\(Biosaet al\.,[2020](https://arxiv.org/html/2607.09501#bib.bib11); Macarulla Rodriguezet al\.,[2024](https://arxiv.org/html/2607.09501#bib.bib61); Ishiharaet al\.,[2022](https://arxiv.org/html/2607.09501#bib.bib46)\)\. In practice, one of the main challenges of logistic regression calibration is acquiring and preparing the data needed to train a logistic regression model\.
#### 3\.3\.1Data for Calibration
The training data required for logistic regression calibration consists of outputs from the given forensic analysis system with accompanying ground\-truth labels\. In the context of AV, this requires a set of text pairs labelled as same\-source \(written by the same author\) or different\-source \(written by different authors\)\. It is necessary to have pairs of texts belonging to both source categories\.
As language use varies across contexts\(Biber,[2012](https://arxiv.org/html/2607.09501#bib.bib9)\), different conditions can lead to systematic variation in linguistic features, and consequently, AV scores\. To ensure the logistic regression transformation generalises to the case data, the texts must therefore reflect the conditions of the case\(Morrison,[2024](https://arxiv.org/html/2607.09501#bib.bib65)\)\. However, the threshold of sufficient similarity to the case data is not clearly defined\(Morrison,[2021](https://arxiv.org/html/2607.09501#bib.bib66)\)and neither are the precise parameters that should be more closely aligned\. The choices are therefore left to the expert, which, in turn, opens these choices to debate with opposing experts\.
In practice, perfectly matched data can be difficult to obtain\. For example, if the AV task concerns a malicious forensic text, like a threat, comparable texts may be rare or inaccessible\(Nini,[2017](https://arxiv.org/html/2607.09501#bib.bib79)\)\. In such cases, controlling for additional variables, such as matching the dialect to the relevant population\(Morrisonet al\.,[2012](https://arxiv.org/html/2607.09501#bib.bib114)\), may not be possible, whilst maintaining a sufficiently large dataset\.
The forensic linguist is therefore required to make decisions regarding how much and which data to use for calibration\. To ensure that this subjective judgement is transparent and defensible\(Morrison,[2021](https://arxiv.org/html/2607.09501#bib.bib66)\), these decisions must be validated through additional tests that verify the validity and reliability of the system\(Morrisonet al\.,[2012](https://arxiv.org/html/2607.09501#bib.bib114)\)\. Validation tests add further complexity, as well as time and computational cost\.
Further time cost is incurred as the texts in the calibration dataset require preprocessing, whereby features that may introduce noise or bias into the analysis are removed\. The specific preprocessing steps depend on the dataset and the AV system used and typically requires both automation and manual validation\. As calibration data is often obtained from different sources to the case data, the steps required to achieve a standardised output may differ between the two\.
Moreover, when pairing texts to create simulated authorship verification problems for the calibration dataset, experts must make a series of decisions\. These decisions, while not arbitrary, may not be grounded in the existing literature due to limited research on the issue\.
One example of this is the choice of anchoring method\(Reinderset al\.,[2022](https://arxiv.org/html/2607.09501#bib.bib69)\)\. Anchoring pertains to the construction of the different\-source comparisons in the calibration dataset\. In AV, trace\-anchoring is where the disputed text \(*Q*\) is paired with texts from randomly selected alternative authors\. In contrast, in source\-anchoring, it is the known\-author text \(*K*\) that is paired with texts from randomly selected authors\. Alternatively, the different\-source problem set may be constructed only from authors that do not feature in the case data\.
Importantly, each of these anchoring methods can result in systematically different LLR estimates\(Hepleret al\.,[2012](https://arxiv.org/html/2607.09501#bib.bib43)\)\. However, there is limited empirical validation regarding which anchoring strategy is optimal, particularly in the context of AV\. Once again, the forensic linguistic expert must therefore validate such decisions for every case, through additional tests\.
Once the pairs of texts are prepared, they are analysed using the selected AV model\. For each pair of texts, the model may return a score that indicates both similarity and typicality but is not directly interpretable as a LLR\. For the calibration dataset, each score is associated with a known ground\-truth label \(a probability of 1 if the pair of text is same\-source and 0 for different\-source pairs\)\.
Each step of this process must be explained and justified in the forensic report\. This requires the forensic expert to demonstrate the validity of the procedure in a clear and accessible outline, that can be understood by juries or legal practitioners without specialist knowledge\. Communicating all decisions clearly and transparently can be challenging, but is essential for maintaining the credibility of the analysis, and ensuring that the court can interpret the findings in an informed and appropriate way\.
#### 3\.3\.2Model Estimation
Once the calibration data is prepared it can be used to train a logistic regression model, which can subsequently be used to calculate a calibrated LLR\. This section provides an overview of logistic regression model estimation\. For a more in\-depth tutorial see Morrison\([2013](https://arxiv.org/html/2607.09501#bib.bib75)\)\.
Logistic regression performs calculations in log\-odds space, so the relationship between the uncalibrated score and the LLR is linear\. The transformation is performed using Equation[2](https://arxiv.org/html/2607.09501#S3.E2)\. whereα\\alphais the intercept andβ\\betais the slope\.
y=α\+βx\{y=\\alpha\+\\beta x\}\(2\)
These parameters are estimated from the training dataset of same\-source and different\-source scores, and the resulting transformation can then be applied to the case data, where the source category is unknown\. These scores are shifted byα\\alphaand scaled byβ\\betato obtain calibrated LLRs\.
The model is fitted using maximum likelihood estimation\. To ensure that the calibration model outputs correspond to the LLR, equal class priors must be assumed\.
Unlike generative models, which explicitly model the score distributions for same\-source and different\-source comparisons, logistic regression models the decision boundary between the two classes\. This boundary follows a sigmoidal function, the position and steepness of which can be adjusted, but does not directly represent the underlying score distributions\. Consequently, logistic regression calibration may be considered a relatively opaque approach\.
### 3\.4Calibration Evaluation
To determine whether a logistic regression model has produced well\-calibrated LLRs, it must undergo a validation test\. The metric used to evaluate the validity of LLR systems is the cost of log\-likelihood ratio \(CllrC\_\{llr\}\)\(Brümmer and du Preez,[2006](https://arxiv.org/html/2607.09501#bib.bib14); Morrison,[2011](https://arxiv.org/html/2607.09501#bib.bib68)\)\. This metric is the established standard in the forensic analysis of behavioural biometrics, with van Lierop et al\.\([2024](https://arxiv.org/html/2607.09501#bib.bib104), figure 1\)reporting its use in all reviewed publications on stylometric analysis, and the majority of those on speaker recognition\.
To calculate theCllrC\_\{llr\}, a validation dataset is required, consisting of LRs for same\- and different\-source pairs of texts, with ground\-truth labels specifying whether each pair was written by the same or different authors\. Ideally, this validation dataset should resemble the case conditions as closely as possible to ensure meaningful performance evaluation\. Consequently, the validation data is subject to the same data\-dependency limitations as the calibration data\.
In the validation dataset, the ground\-truth labels enable the calculation of whether the AV model’s prediction aligns with the true observation\. TheCllrC\_\{llr\}is calculated using the following equation\(Morrison,[2011](https://arxiv.org/html/2607.09501#bib.bib68), eq\. 3\):
Cllr=12\(1Ns∑i=1Nslog2\(1\+1LRsi\)\+1Nd∑j=1Ndlog2\(1\+LRdj\)\)\{C\_\{llr\}=\\frac\{1\}\{2\}\\left\(\\frac\{1\}\{N\_\{s\}\}\\sum\_\{i=1\}^\{N\_\{s\}\}\\log\_\{2\}\\Big\(1\+\\frac\{1\}\{LR\_\{s\_\{i\}\}\}\\Big\)\+\\frac\{1\}\{N\_\{d\}\}\\sum\_\{j=1\}^\{N\_\{d\}\}\\log\_\{2\}\\Big\(1\+LR\_\{d\_\{j\}\}\\Big\)\\right\)\}\(3\)
In this calculation, for same\-author problems, a cost is incurred when the LR \(LRsLR\_\{s\}\) fails to strongly support the prosecution hypothesis \(1\+1/LRs1\+1/LR\_\{s\}\)\. This cost decreases as evidence in favour of the correct hypothesis increases\. Conversely, the term\(1\+LRd\)\(1\+LR\_\{d\}\)penalises LRs that incorrectly favour the prosecution hypothesis\. The logarithmic transformation converts LRs into interpretable, additive quantities\. Averaging over both trial types \(1/Ns1/N\_\{s\}and1/Nd1/N\_\{d\}\) and weighting them equally yields a single scalar measure that jointly reflects discrimination and calibration quality\.
The lower theCllrC\_\{llr\}, the better the performance of the AV system\. However, whilst it is established that aCllrC\_\{llr\}of 1 indicates a system is uninformative, it is not well\-established what a goodCllrC\_\{llr\}is\(van Lieropet al\.,[2024](https://arxiv.org/html/2607.09501#bib.bib104)\)\.
## 4Authorship Verification
### 4\.1Properties of Authorship Verification Methods
AV methods can be categorised into three distinct groups: unary, binary intrinsic and binary extrinsic methods\(Halvaniet al\.,[2019](https://arxiv.org/html/2607.09501#bib.bib117)\)\. The categories are defined by the data required and how this data is utilised to establish a decision criteria for AV\.
A unary AV method\(e\.g\., Noecker Jr and Ryan,[2012](https://arxiv.org/html/2607.09501#bib.bib119); Halvaniet al\.,[2018](https://arxiv.org/html/2607.09501#bib.bib118)\)only requires the case data in order to establish its decision criteria\. The case data refers to the questioned text \(QQ\) and sample text\(s\) known to be written by the candidate author \(KK\)\.
As a result of this characteristic data requirement, it is an intrinsic property of unary AV methods that only one class is considered\. As such, a LLR cannot be obtained from a unary AV method, as, by definition \(Equation[1](https://arxiv.org/html/2607.09501#S2.E1)\), the probability of the evidence if the candidate author did not write the questioned text, must also be considered\.
Comparatively, a binary method utilises data that is external to the case data, allowing for LLR estimation\. Binary intrinsic methods\(e\.g\., Halvaniet al\.,[2017](https://arxiv.org/html/2607.09501#bib.bib41); Boenninghoffet al\.,[2019](https://arxiv.org/html/2607.09501#bib.bib12); Huertas\-Tatoet al\.,[2024](https://arxiv.org/html/2607.09501#bib.bib45)\)use training data to learn how to distinguish between same\-author and different\-author classes\. Comparatively, binary extrinsic methods\(e\.g\., Koppel and Winter,[2014](https://arxiv.org/html/2607.09501#bib.bib58); Potha and Stamatatos,[2017](https://arxiv.org/html/2607.09501#bib.bib83); Niniet al\.,[2026](https://arxiv.org/html/2607.09501#bib.bib112)\)compare the language in external reference documentsℝ\\mathbb\{R\}, to the language used inQQ\. This provides a point of comparison for the similarities found betweenQQandKK\.
TableLABEL:tbl\-compdemonstrates that across these three categories, the existing methods share a particular property: that meaningful likelihood ratios cannot be calculated without a separate calibration model, if at all\. The present research proposes two novel approaches that do not share this property, creating a new class of authorship verification method\.
Table 1:Comparative overview of AV categories, highlighting data requirementsMethodCalibration Data RequiredReference Corpus RequiredProposed MethodsNoYesUnaryN/ANoBinary IntrinsicYesNoBinary ExtrinsicYesYesThe proposed methods constitute an adjustment of an existing binary extrinsic method\. By eliminating the need for a separate calibration model, this, in turn, reduces both the data requirement and the complexity of the method\.
### 4\.2LambdaG
The binary\-extrinsic method employed in this paper is LambdaG\(Niniet al\.,[2026](https://arxiv.org/html/2607.09501#bib.bib112)\), which relies on n\-gram language modelling\. This method has a unique property that establishes it as a suitable focus for this study\. That is, it produces scores intrinsically formulated as LLRs, albeit uncalibrated prior to the application of logistic regression calibration\. Therefore, whilst LambdaG currently requires post\-hoc calibration, its formulation may allow for alternative, simpler transformations to achieve a calibrated LLR output, which would not be possible with any other existing AV method\.
To apply the method, first the input data \(QQ,KK, and the set of reference textsℝ\\mathbb\{R\}\) must becontent masked\. This is a computational procedure which accounts for the influence of topic on authorship analysis by obscuring words that carry lexical meaning, as opposed to serving a grammatical function\.Niniet al\.\([2026](https://arxiv.org/html/2607.09501#bib.bib112)\)employ thePOS\-Noisemethod of content masking\(Halvani and Graner,[2021](https://arxiv.org/html/2607.09501#bib.bib39)\), which substitutes tokens that contain topical information with their part\-of\-speech\. For example, the sentence*the cow jumped over the moon*becomes*the N V over the N*\. The tokens deemed topic\-agnostic are identified using a pre\-defined list111The list can be found here: https://bit\.ly/ARES\-2021\(Halvani and Graner,[2021](https://arxiv.org/html/2607.09501#bib.bib39), p\.3\)\.
The method then consists in calculatingλG\\lambda\_\{G\}, the ratio of the the likelihood of each token in context given agrammar modelof the author ofKK,GKG\_\{K\}, and the likelihood of each token in context given a set of grammar models from the reference population,GR=\{G1,G2,…Gr\}G\_\{R\}=\\\{G\_\{1\},G\_\{2\},\.\.\.G\_\{r\}\\\}\.
To estimate these grammar models, each of the content masked texts across the questioned, candidate and reference data are tokenised into sentences\. The set of sentences in the content maskedKK\(𝕊K\\mathbb\{S\}\_\{K\}\) are then used to train an n\-gram grammar model \(GKG\_\{K\}\)\.
An n\-gram language model estimates the conditional probability of a word by applying the Markov assumption of order\(Jurafsky and Martin,[2026](https://arxiv.org/html/2607.09501#bib.bib53), p\.40\)\. In the context of text generation, a first\-order Markov assumption states that the probability of the token at positionjjin a sentence depends only on the token itself\. Higher order Markov assumptions, e\.g\.N=10N=10, allow more context into the n\-gram language model\. The probability of a word sequence of lengthnn\(t1…tnt\_\{1\}\.\.\.t\_\{n\}\) using a 10\-gram language model would be calculated using Equation[4](https://arxiv.org/html/2607.09501#S4.E4)\. For the first nine tokens, beginning of sentence tokens \(\[BOS\]\[BOS\]\) are prepended\.
P\(t1…tn\)≈∏j=1nP\(tj∣tj−9,tj−8,…,tj−1\)P\(t\_\{1\}\.\.\.t\_\{n\}\)\\approx\\prod\_\{j=1\}^\{n\}P\(t\_\{j\}\\mid t\_\{j\-9\},t\_\{j\-8\},\\dots,t\_\{j\-1\}\)\(4\)
N\-gram languages models estimate probabilities from corpus counts using maximum likelihood estimation\. For LambdaG, these probabilities are then adjusted using Kneser\-Ney smoothing\(Kneser and Ney,[1995](https://arxiv.org/html/2607.09501#bib.bib56); Chen and Goodman,[1999](https://arxiv.org/html/2607.09501#bib.bib17)\)\. Kneser\-Ney smoothing redistributes probability mass by backing off to lower\-order distributions when higher\-order n\-grams are sparse or unseen\. The probabilities of tokens are weighted according to the number of distinct contexts in which they appear\.
To train the reference grammar models, sentences are randomly sampled fromℝ\\mathbb\{R\}and concatenated\. This produces a set of sentences written by multiple different authors\. The number of sentences sampled is equal to the number of sentences inKK\. On iterationiithis set of sentences is used to trainGiG\_\{i\}\. This process of sampling and training is repeatedrrnumber of times\.
For each token \(tjt\_\{j\}\), in each sentence, theλG\\lambda\_\{G\}is calculated by computing the probability thattjt\_\{j\}is the next token, conditioned on all preceding tokens in that sentence \(t<jt\_\{<j\}\) as Equation[5](https://arxiv.org/html/2607.09501#S4.E5)\.
λG\(tj\|t<j\)=1r∑i=1rlogP\(tj\|t<j;Gk\)P\(tj\|t<j;Gi\)\\lambda\_\{G\}\(t\_\{j\}\|t\_\{<j\}\)=\\frac\{1\}\{r\}\\sum\_\{i=1\}^\{r\}\\log\\frac\{P\(t\_\{j\}\|t\_\{<j\};G\_\{k\}\)\}\{P\(t\_\{j\}\|t\_\{<j\};G\_\{i\}\)\}\(5\)
The probability ofttgivenGKG\_\{K\}constitutes a similarity measure betweenQQand the texts known to be written by the candidate author \(KK\)\. The mean probability ofttgiven the reference grammar models is a typicality measure capturing distributional norms across language users\. By calculating the ratio of these probabilities to produce aλG\\lambda\_\{G\}value, that value is inherently formulated as a LLR\.
The final score for the questioned text is the sum of theλG\\lambda\_\{G\}scores for each token in theQQtext \(Equation[6](https://arxiv.org/html/2607.09501#S4.E6)\)\.
λG\(Q\)=∑λG\(tj\|t<j\)\\lambda\_\{G\}\(Q\)=\\sum\{\\lambda\_\{G\}\(t\_\{j\}\|t\_\{<j\}\)\}\(6\)
The full algorithm is provided in[\\thechapter\.A](https://arxiv.org/html/2607.09501#X.A1)\.
To illustrate this process, consider a content maskedQQthat contains the sentence*The N V over the N*\. TheλG\\lambda\_\{G\}value of the token*over*is the logarithm of the probability that*over*is the next token in the sequence*The N V*givenGKG\_\{K\}, divided by the mean probability that*over*is the next token in the sequence according to therrgrammars inGRG\_\{R\}\. The process is repeated for every token in every sentence inQQ\. TheλG\\lambda\_\{G\}for each token is added to produce an final overallλG\\lambda\_\{G\}for the text\.
Beyond its LLR formulation and state\-of\-the\-art performance, LambdaG also presents a number of other advantages\. These include that it is designed to exhibit minimal sensitivity to topic variation, performs consistently across registers, including very short texts, and is grounded in cognitive linguistic theories\.
The following section will outline that despiteλG\\lambda\_\{G\}being formulated as a LLR, it does not reliably reflect the true strength of the linguistic evidence\. It is therefore not a meaningful LLR, it is an uncalibrated score\. This may indicate that, despite being grounded in valid linguistic assumptions, LambdaG does not model language productivity with complete accuracy\. As a result, like other AV methods, LambdaG requires further calibration to obtain a calibrated LLR \(ΛG\\Lambda\_\{G\}\)\.Niniet al\.\([2026](https://arxiv.org/html/2607.09501#bib.bib112)\)achieve this using logistic regression calibration\.
## 5Score Inflation and Corrections
We propose that, in the case of LambdaG, calibration is necessary because, as LambdaG sums independent token\-level contributions, it adopts the naïve Bayes assumption of conditional independence between long\-distance input features\. This simplifying assumption introduces a systematic weakness affecting all naïve Bayes algorithms, whereby the reliability of the resulting probability may be compromised\.
We hypothesise that the issue is the inherent repetition in natural language production\. As token\-level contributions are aggregated, repeated evidence compounds\. However, it is plausible to assume that subsequent repetitions of the same pattern contribute less evidential value than initial occurrences\. This may lead to the strength of authorship evidence being overstated in longer or highly repetitive texts, with repeated stylistic features contributing disproportionately to the overall score\. The following section outlines the supporting evidence that grounds this hypothesis\.
### 5\.1Distribution of Raw Scores
Examining theλG\\lambda\_\{G\}distributions for same\-source and different\-source problems demonstrates substantial overestimation of evidential strength\. Figure[1](https://arxiv.org/html/2607.09501#S5.F1)presents these distributions across the corpora used to evaluate LambdaG inNiniet al\.\([2026](https://arxiv.org/html/2607.09501#bib.bib112)\)\.λG\\lambda\_\{G\}values greater than 600 were omitted from Figure[1](https://arxiv.org/html/2607.09501#S5.F1)for visualisation purposes\. All excluded values, the most extreme of which was 3,321\.84, correspond to theAll\-the\-newscorpus\. These values were retained for all later analysis\.
Figure 1:Distribution of raw LambdaG scores for each corpus adopted from Nini et al\. \(2026\), shown as violin plots\. Colours indicate target label\.A LLR of 4, for example, indicates that the observed evidence is 10,000 times more likely under the prosecution hypothesis than under the defence hypothesis\. However, even with the most extreme values omitted, the remainingλG\\lambda\_\{G\}values in Figure[1](https://arxiv.org/html/2607.09501#S5.F1)still span a notably large range\. Moreover, aside from the remaining extreme values in the tails of the distribution, many values still exceed four, with numerous values between 50 and 100\.Niniet al\.\([2026](https://arxiv.org/html/2607.09501#bib.bib112), p\.15\)demonstrate that these overinflated raw scores are misleading, withλG\\lambda\_\{G\}yielding aCllrC\_\{llr\}of greater than one across each corpus\.
Figure[1](https://arxiv.org/html/2607.09501#S5.F1)also illustrates that, across most corpora, the distributions ofλG\\lambda\_\{G\}for same\-source texts and different\-source texts are approximately sign\-aligned\. True problems tend to obtain a positiveλG\\lambda\_\{G\}value and false problems a negativeλG\\lambda\_\{G\}\. Given thatλG\\lambda\_\{G\}does not need to be shifted, and is already constructed as a LLR, the primary role of calibration may therefore be to scale the scores to mitigate overstated evidence\. Logistic regression calibration may therefore constitute an excessively complex solution for what might be solved by a simpler transformation, such as normalisation\.
### 5\.2The Square Root Correction
The first innovative approach, the Square Root Correction \(Equation[7](https://arxiv.org/html/2607.09501#S5.E7)\), scalesλG\\lambda\_\{G\}using a simple metric based on the number of tokens in the questioned document,N\(Q\)N\(Q\)\. The longer a text is the more likely it is to contain repetitions\(Herdan,[1960](https://arxiv.org/html/2607.09501#bib.bib44); Heaps,[1978](https://arxiv.org/html/2607.09501#bib.bib42)\)\. Therefore, as text length increases, additional tokens are increasingly likely to present evidence that has already been seen\. Subsequent repetitions of the same pattern are unlikely to contribute fully independent evidential weight, as once a stylistic feature has been observed, its recurrence becomes increasingly expected\(Coulthard,[2004](https://arxiv.org/html/2607.09501#bib.bib21); Nini,[2023](https://arxiv.org/html/2607.09501#bib.bib80)\)\. As a result, with increasing tokens, the average contribution of evidential strength for each token is likely to diminish\.
ΛG=λGN\(Q\)\{\\Lambda\_\{G\}=\\frac\{\\lambda\_\{G\}\}\{\\sqrt\{N\(Q\)\}\}\}\(7\)
The selection of a square\-root transformation is consistent with statistical principles, where variance often grows approximately proportionally with the aggregation of weakly\-dependent input features, causing dispersion to increase approximately proportionally toN\(Q\)\\sqrt\{N\(Q\)\}\. Similar forms of square\-root scaling are also used in other domains, such as the scaling factor applied in attention mechanisms within machine learning\(Vaswaniet al\.,[2017](https://arxiv.org/html/2607.09501#bib.bib105)\)\.
Normalising solely by text length would implicitly assume that every token contributes equal evidential value\. In contrast, transforming a score by the square root of text length, captures the diminishing contributions from additional tokens\. The Square Root Correction can therefore moderate the inflated predictions of evidential strength for longer texts, without excessively penalising them\.
### 5\.3Adversarial Manipulation
While the Square Root Correction mitigates score inflation, it does not directly address the underlying independence assumption that may cause such inflation\. To more directly account for repeated tokens providing increasingly redundant stylistic evidence, normalisation must incorporate not only text length but also a measure of repetition, or information redundancy\.
One scenario where scaling by length is likely not sufficient to address the effect of repeated tokens is adversarial manipulation of AV systems, for instance, when impersonation is attempted\. If an adversary is aware of a specific linguistic preference of the individual they are impersonating, then they would likely repeat this known feature several times\.
One example of this is the Starbuck case, where Jamie Starbuck attempted to impersonate his wife, Debbie\(Grant and Grieve,[2022](https://arxiv.org/html/2607.09501#bib.bib71)\)\. Debbie used a high\-frequency of semi\-colons, and having noticed this, when impersonating Debbie, Jamie employed a high\-frequency of semi\-colons comparative to his own typical usage\. The feature was repeated to such an extent that the frequency of semi\-colons in the impersonated texts was higher than had been observed in texts known to be written by Debbie\(Roemling and Grieve,[2024](https://arxiv.org/html/2607.09501#bib.bib70)\)\.
With the current LambdaG system, such impersonation may have the desired effect, because each occurrence would independently increase the score\. The model could not differentiate this artificial inflation from natural language\. As a result, the repeated use of a deliberately imitated stylistic feature may artificially dominate the overallλG\\lambda\_\{G\}of the text\. This means the AV system is vulnerable to manipulation, reducing robustness in forensic settings\.
### 5\.4The Hapax Correction
The second proposed correction, the Hapax Correction \(Equation[8](https://arxiv.org/html/2607.09501#S5.E8)\), aims to mitigate inflated evidential strength in instances of limited stylistic diversity, relative to the length of the text\. To capture the degree of repetition and redundancy in a text, the correction incorporates the number ofhapax legomena,V1\(Q\)V\_\{1\}\(Q\)\- the number of distinct types that only appear once in a text, in this case, the questioned document\.
ΛG=λG×V1\(Q\)N\(Q\)\{\\Lambda\_\{G\}=\\lambda\_\{G\}\\times\\frac\{V\_\{1\}\(Q\)\}\{N\(Q\)\}\}\(8\)
For this correction, hapax legomena are counted in the content\-masked text rather than on the original text, as the masked representation constitutes the actual input to the LambdaG model\. The aim of the normalisation process is to counteract the accumulation of scores arising from repeated encounters with the same tokens, including the parts\-of\-speech introduced through POS\-Noise\. Using masked token counts ensures that normalisation operates over the same set of tokens that contributes to score accumulation, thereby maintaining consistency between the source of bias and the corrective mechanisms\.
As raw counts of hapax legomena are not comparable for texts of different lengths, the Hapax Correction calculates the ratio of hapax legomena to the total token count of the questioned document \(NN\)\. This provides a more reliable measure of linguistic productivity and lexical diversity\. As this value is always between 0 and 1, multiplyingλG\\lambda\_\{G\}by this ratio proportionally scales texts containing a high degree of repetition\. For texts with relatively few hapax legomena, and therefore a greater proportion of repeated tokens, the score is reduced more than for a text with a wide range of stylistic evidence\. This aims to reduce inflated evidential strength resulting from redundant or highly correlated linguistic patterns\.
## 6Methodology
To address the question of whetherΛG\\Lambda\_\{G\}can be derived fromλG\\lambda\_\{G\}without training a calibration model, we evaluated the Hapax Correction and Square Root Correction against logistic regression calibration across fifteen corpora\. For a subset of these corpora, the corrections were tested across a range of text length configurations, from 100\-9,500 tokens for the Q and K texts\. The data preparation and analysis were executed using R\. In particular, the analysis relied on theidiolectpackage\(Nini,[2026](https://arxiv.org/html/2607.09501#bib.bib73)\)\.
### 6\.1Data
#### 6\.1\.1Overview
The data used in this study comprises fifteen different corpora, alongside corresponding problem sets\. Twelve of the corpora were adopted fromNiniet al\.\([2026](https://arxiv.org/html/2607.09501#bib.bib112)\)\. These consisted of academic papers \(ACL\), newspaper articles \(All\-the\-news\), blog posts \(Koppel’s Blogs\), four corpora of reviews \(Amazon,IMDB,TripAdvisorandYelp\), Wikipedia talk pages \(Wiki\), two forum corpora \(The Apricity\) and \(StackExchange\) , emails \(Enron\) and chat logs \(Perverted Justice\)\.
The component corpora represent a range of scenarios encountered in real\-word AV cases\. Some of the complications captured by the dataset include very short texts \(Perverted JusticeandYelp\), texts of varying lengths \(Stack Exchange\), very formal texts \(ACL\), texts with a high frequency of non\-standard lexical items \(Perverted JusticeandAll\-the\-news\), texts written by the same author several years apart \(YelpandACL\), and texts written by the same author on different topics and different authors on the same topic \(Stack Exchange\)\. Further information on each of these corpora, their preparation, and challenges is detailed byNiniet al\.\([2026](https://arxiv.org/html/2607.09501#bib.bib112), Section 6\.1\)\.
Three further corpora were compiled to complement the existing datasets, introducing higher volumes of data per author and previously unrepresented registers\. The two registers on which LambdaG has not previously undergone formal evaluation are tweets \(Twitter\) and text messages \(Bolt\)\. The register of the final corpus employed in this paper comprises emails \(Avocado\)\. This register has been examined before \(theEnroncorpus\), but only with substantially smaller amounts of data for each author \(averaging 876 tokens of disputed data\)\. This analysis ofAvocadoandTwitterimplements approximately 9,500 tokens of disputed data per author, enabling the assessment of the proposed corrections, and logistic regression calibration, across a range of text\-length configurations\.
Authors in each corpus were randomly assigned to either the case data or the calibration dataset\. The case data represents the target AV problems and is the data on which predictions are made, whereas the calibration data is used solely for training the logistic regression model implemented in the baseline approach\. For the corpora adopted fromNiniet al\.\([2026](https://arxiv.org/html/2607.09501#bib.bib112)\), 40% of each corpus was used for calibration data, and the remaining 60% for case data\. Comparatively, for the three new corpora a 50/50 split was implemented\. Each author features in two problems, one same\-source pairing and one different\-source pairing\. In total, there are 13,146 problems\. 4,766 of the problems are calibration data and 8,380 are case data, both are evenly split between same\-source and different\-source pairs\.
#### 6\.1\.2Preparation
The data preparation processes described below were designed to mirror the standard that would be applied to real\-world case data\. However, no standardised framework exists, as preprocessing requirements differ depending on each register and corpus\.
Two of the corpora adopted fromNiniet al\.\([2026](https://arxiv.org/html/2607.09501#bib.bib112)\)underwent additional preprocessing\. InAll\-the\-news, duplicate sentences within the same text were removed, as some texts contained repeated sections\. InEnron, emails were removed if they contained a sequence of 10 or more tokens that also featured in another text, to reduce duplication\. To address outliers in token distributions, an inter\-quartile range\-based filter\(Tukey,[1977](https://arxiv.org/html/2607.09501#bib.bib115)\)with a multiplier of 0\.8 was applied\. This multiplier was selected on the basis of manual validation of the output\. Authors below the lower threshold were removed, while authors above the upper threshold were trimmed to within the accepted range\.
The first of the new corpora,Avocado, consists of 64,858 emails from4040different accounts\. The data originates from the Avocado Research Email Collection\(Oard, Douglaset al\.,[2015](https://arxiv.org/html/2607.09501#bib.bib81)\), a corpus of emails from a defunct informational technology company\. The objective of preprocessingAvocadowas to emulate the initial preprocessing ofEnron\(Niniet al\.,[2026](https://arxiv.org/html/2607.09501#bib.bib112), p\.22\)\. However, asEnronis much smaller, some of this was achieved manually\. ForAvocadoour aim was to automate as much of the process as possible\.
That being said, a small number of accounts identified as non\-human were manually excluded\. As part of automated preprocessing, salutations and signatures were removed\. Duplicate and near\-duplicate texts were identified using MinHash LSH\(Mullen,[2020](https://arxiv.org/html/2607.09501#bib.bib76)\)\(240240minhashes,8080bands,1010\-token n\-grams\); based on manual inspection, any pair with non\-zero Jaccard similarity was removed\. Automated preprocessing also included lowercasing, standardising elongated words, replacing email addresses with a placeholder, removing URLs and non\-ASCII characters, trimming whitespace, and removing texts shorter than three tokens or longer than 1,000 tokens\. Each author inAvocadohas a minimum of 100,000 tokens, of which approximately 35,000 texts are labelled as*unknown*\(Q data\) This was achieved by randomly assigning texts until the token threshold was exceeded\.
The second new corpus,Twitter, consists of 576,602 tweets by6060authors\. From the original dataset\(Grieveet al\.,[2019b](https://arxiv.org/html/2607.09501#bib.bib37)\), a sample of authors with more than 135,000 tokens was taken\. Authors were manually reviewed to exclude automated accounts\. Duplicates were removed, as were tweets beginning*rt :*, as it indicated a retweet\. Within texts, emojis, leading punctuation, non\-printing Unicode characters and redundant white space were also removed\. URLs were replaced with*<url\>*, and tags were replaced with*userid*\. Elongated words were normalised, and consecutive sequences of more than three identical punctuation marks were reduced to three\. Any texts that now had less than three tokens were then removed\. The total word count for each author was then calculated, and this preprocessing was repeated until a set of6060authors with greater than 135,000 tokens was obtained\. Approximately 35,000 tokens were allocated to the*unknown*data for each author, and 100,000 tokens to the known \(*K*\) data\.
Lastly, theBoltCorpus comprises 3716 messages authored across4646individuals\. The data consists of natural text \(SMS\) and chat messages sourced from the BOLT \(Broad Operational Language Translation\) English SMS/Chat corpus\(Chen, Songet al\.,[2018](https://arxiv.org/html/2607.09501#bib.bib18)\)\. The preprocessing of this dataset involved removing URLs, non\-printing Unicode characters, redundant internal and trailing whitespace and non\-ASCII characters, as well as normalising word elongations\. Texts with fewer than three tokens, and texts from authors with fewer than1,0001,000aggregated tokens were removed from the corpus\. For each author, texts were randomly sampled and assigned as*unknown*until the500500token threshold was exceeded\. The same threshold and process was applied for known texts\. Remaining texts were also removed\.
### 6\.2Protocol
LambdaG was applied to each corpus following the workflow outlined byNiniet al\.\([2026](https://arxiv.org/html/2607.09501#bib.bib112)\)\. To prepare the texts for input, content masking was applied using the POSNoise method with the spaCyR part\-of\-speech tagger \(model:en\_core\_web\_sm\)\(Halvani and Graner,[2021](https://arxiv.org/html/2607.09501#bib.bib39); Benoitet al\.,[2023](https://arxiv.org/html/2607.09501#bib.bib8); Nini,[2024](https://arxiv.org/html/2607.09501#bib.bib78)\)\. The texts were also tokenised into sentences using spaCyR \(model:en\_core\_web\_sm\)\(Benoitet al\.,[2023](https://arxiv.org/html/2607.09501#bib.bib8); Nini,[2024](https://arxiv.org/html/2607.09501#bib.bib78)\)\.
Two hyperparameters are defined within LambdaG:rr\(the number of iterations, i\.e\. the number of reference grammar models\) andNN\(the order of the model\)\(Niniet al\.,[2026](https://arxiv.org/html/2607.09501#bib.bib112)\)\. Both hyperparameters were set to the default values provided in the Idiolect package:r=30r=30andN=10N=10\(Nini,[2024](https://arxiv.org/html/2607.09501#bib.bib78)\)\. Nini et al\.\([2026](https://arxiv.org/html/2607.09501#bib.bib112), Figure 3\)found LambdaG is robust to changes in these hyperparameters\.
On the twelve corpora adopted fromNiniet al\.\([2026](https://arxiv.org/html/2607.09501#bib.bib112)\), LambdaG was applied to the full texts\. OnAvocado,TwitterandBoltthe method was evaluated across a range of text length conditions\. A condition refers to the combination of*Q*length and*K*length\. For each of the three corpora, all texts were content masked and tokenised into sentences, before they were concatenated into*Q*and*K*data respectively\. This ensured that text boundaries were maintained in sentence tokenization\.
From the*Q*data \(the concatenated list of tokenised sentences labelled*unknown*\), sentences were randomly sampled until the cumulative token count exceeded the specified*Q*length\. The same procedure was applied to the*K*data, using the corresponding*K*length as the threshold\. LambdaG was then run on each problem in turn\. This process was applied to both the case and the calibration data\.
As a baseline, rawλG\\lambda\_\{G\}scores were calibrated using logistic regression calibration, following the approach byNiniet al\.\([2026](https://arxiv.org/html/2607.09501#bib.bib112)\)\. TheλG\\lambda\_\{G\}values calculated for each same\-source and different\-source problem in the calibration dataset were used to train the logistic regression model\. The model estimates the optimal parameters to transformλG\\lambda\_\{G\}values intoΛG\\Lambda\_\{G\}by aligning predicted probabilities with observed outcomes\.
The proposed corrections \(the Square Root Correction \(Equation[7](https://arxiv.org/html/2607.09501#S5.E7)\) and the Hapax Correction \(Equation[8](https://arxiv.org/html/2607.09501#S5.E8)\)\) provide an alternative approach to approximate these optimal affine transformations\. These are also applied toλG\\lambda\_\{G\}, enabling the production ofΛG\\Lambda\_\{G\}without training a logistic regression model\.
Performance was evaluated usingCllrC\_\{llr\}\. As LambdaG is a stochastic method, each analysis was repeated five times, and the meanCllrC\_\{llr\}was reported\.
## 7Results
### 7\.1Performance Overview
TableLABEL:tbl\-overviewprovides an overview of the results across all fifteen corpora, by outlining the win rate and close loss rate for each pairwise comparison among the three methods\. Each corpus was weighted equally, thus where multiple tests were conducted on a corpus, the win and close loss percentages were first calculated at the corpus level and then incorporated into an overall average across all corpora\.
The threshold for close loss rate was defined as within 5% of the comparator\. A relative threshold was chosen to provide a consistent and interpretable definition of comparable performance across performance results of varying magnitudes\. Support for the threshold of 5% will be outlined in TableLABEL:tbl\-development, which indicates performance differences within this parameter may be attributable to noise\.
Table 2:Pairwise performance comparison of methods across corpora\. Columns show each methods performance against the comparator, with each corpus weighted equally\. Metrics include the average percentage of tests where the method performed better or equal \(lower values indicate better performance\) and the proportion of losses that were within 5% of the comparatorLog Reg vs Square RootLog Reg vs HapaxSquare Root vs HapaxMetricLog RegSquare RootLog RegHapaxSquare RootHapaxCases outperforms comparator \(%\)74\.8025\.2054\.6045\.4021\.8078\.20Losses within 5% of comparator \(%\)40\.1821\.3649\.1356\.9123\.8945\.29On average across the corpora, the Hapax Correction outperformed the performance of logistic regression calibration in 45\.4% of tests\. Where the Hapax Correction was outperformed by logistic regression calibration, its performance remained within 5% of logistic regression calibration in a higher proportion of tests on average \(56\.91%\) than the reverse comparison \(49\.13%\)\. The Square Root Correction was outperformed by each of the other approaches in over 70% of tests \(on average\)\. The close loss rate was 21\.36% when compared to logistic regression, and 23\.89% when compared to the Hapax Correction\.
The following sections will provide a more fine\-grained view of this analysis\. We will detail the comparative performance of the three approaches by corpus and across text length conditions\.
### 7\.2Performance by Corpus
TableLABEL:tbl\-developmentshows the performance of LambdaG on the twelve corpora adopted fromNiniet al\.\([2026](https://arxiv.org/html/2607.09501#bib.bib112)\), when transformed using logistic regression calibration, the Square Root Correction and the Hapax Correction\. Performance is measured usingCllrC\_\{llr\}\. Each analysis was run five times and the mean and standard deviation was calculated\. The results are separated by corpus, for each approach\.
Table 3:Performance evaluation of logistic regression calibration, the Square Root Correction and the Hapax Correction across twelve corpora, measured using Cllr\. Results are reported as mean Cllr with standard deviation across five random seeds\.CorpusLogRegSquareRootHapaxACL0\.888 ± 0\.0030\.756 ± 0\.0080\.777 ± 0\.010All\-the\-news0\.808 ± 0\.0030\.804 ± 0\.0030\.856 ± 0\.006Amazon0\.302 ± 0\.0010\.437 ± 0\.0010\.313 ± 0\.001Enron0\.518 ± 0\.0260\.505 ± 0\.0070\.479 ± 0\.014IMDB0\.658 ± 0\.0020\.706 ± 0\.0020\.657 ± 0\.006Koppel’s Blogs0\.412 ± 0\.0010\.486 ± 0\.0010\.409 ± 0\.001Perverted Justice0\.219 ± 0\.0020\.253 ± 0\.0020\.226 ± 0\.003StackExchange0\.509 ± 0\.0110\.533 ± 0\.0080\.508 ± 0\.012The Apricity0\.369 ± 0\.0090\.451 ± 0\.0020\.354 ± 0\.007TripAdvisor0\.761 ± 0\.0040\.783 ± 0\.0030\.780 ± 0\.008Wiki0\.459 ± 0\.0130\.579 ± 0\.0040\.474 ± 0\.011Yelp0\.713 ± 0\.0040\.783 ± 0\.0010\.739 ± 0\.004The logistic regression calibration results reproduce the results ofNiniet al\.\([2026](https://arxiv.org/html/2607.09501#bib.bib112), p\.15\)Discounting the two datasets for which additional preprocessing was applied \(EnronandAll\-the\-news\), the mean absolute difference across all datasets was 0\.012 \(rounded to three decimal places; all comparisons also usedCllrC\_\{llr\}values rounded to three decimal places\)\. This demonstrates high consistency between the present study andNiniet al\.\([2026](https://arxiv.org/html/2607.09501#bib.bib112)\), and highlights the reproducibility of LambdaG\.
No single approach has superior performance across all corpora, and the three approaches demonstrate similar patterns of performance across each register, for example, all three achieve their lowestCllrC\_\{llr\}values on thePerverted Justicecorpus \(0\.2190\.219,0\.2530\.253and0\.2260\.226\)\.
Across all three approaches, the low standard deviation values indicate that performance is stable across runs\. This suggests low sensitivity to random seed initialisation\. Of the approaches, logistic regression calibration shows slightly higher variability than the proposed corrections\.
The corpus with the greatest standard deviation wasEnron, with a standard deviation of 0\.026 under the logistic regression calibration condition\. This constitutes approximately 5% of the meanCllrC\_\{llr\}obtained for that condition\. This justified the 5% threshold used in TableLABEL:tbl\-overview, as differences within 5% of the meanCllrC\_\{llr\}across conditions may be comparable to intrinsic variability observed within conditions\.
### 7\.3Text Length Analysis
#### 7\.3\.1Avocado
Figure[2](https://arxiv.org/html/2607.09501#S7.F2)presents the results obtained from the analysis ofAvocado\. It shows how the length of the disputed document \(*Q*\) and the known\-author data \(*K*\), impactsCllrC\_\{llr\}across the three different approaches\. The*Q*and*K*lengths ranged from approximately 500 to 9,500 tokens, sampled at 1,000 token intervals\. Every possible pairwise combination of*Q*and*K*length was evaluated\. The meanCllrC\_\{llr\}is represented by the colour gradient, with empty cells representing aCllrC\_\{llr\}\> 1, i\.e\. misleading output\(van Lieropet al\.,[2024](https://arxiv.org/html/2607.09501#bib.bib104)\)\. For the retained datapoints the meanCllrC\_\{llr\}values are provided within the heatmap and rounded to two decimal places; therefore, where a meanCllrC\_\{llr\}is shown as0, this indicates that the true value was less than 0\.005\.

Figure 2:Comparison of three approaches on the Avocado dataset showing the relationship between Q Length, K Length and mean Cllr\. Each panel corresponds to a post\-processing method applied to LambdaG scores: logistic regression calibration \(left\), the Square Root Correction \(middle\) and the Hapax Correction \(right\)\.\.
Figure[2](https://arxiv.org/html/2607.09501#S7.F2)demonstrates that both proposed corrections yield a positive performance, with the lowest observedCllrC\_\{llr\}of 0\.05 for the Hapax Correction and as low as 0\.04 for the Square Root Correction \- both achieved when*Q*and*K*were at their longest tested length \(9,500\)\. However, logistic regression calibration achieved the best peak performance, with values<<0\.005 in certain instances where*K*length≥\\geq3,500\.
For conditions where the logistic regression calibration obtained aCllrC\_\{llr\}\>0\.1, the Hapax Correction performed better or equal in 92\.3% of cases, and the Square Root Correction in 84\.6% of instances; in a further 7\.69% and 10\.3% of cases respectively, the proposed methods were within 5% of the baselineCllrC\_\{llr\}\.
When logistic regression calibration achieved aCllrC\_\{llr\}≤\\leq0\.1, percentage difference is less informative due to small values\. In these cases, the maximum observed difference was 0\.13 for the Hapax Correction and 0\.23 for the Square Root Correction\. However, in just 11\.48% of these conditions, theCllrC\_\{llr\}obtained using the Hapax Correction was within 0\.05 of theCllrC\_\{llr\}achieved by the baseline, and for the Square Root Correction, this was the case in just 8\.2% of cases\. Comparing the proposed methods directly, the Hapax Correction performed better than the Square Root Correction in 70% of cases\. Figure[2](https://arxiv.org/html/2607.09501#S7.F2)also reveals several anomalous results when logistic regression calibration is implemented \(left\), as indicated by multiple empty cells\.
#### 7\.3\.2Twitter
Figure[3](https://arxiv.org/html/2607.09501#S7.F3)shows the same analysis as Figure[2](https://arxiv.org/html/2607.09501#S7.F2), applied to a register on which LambdaG has not yet been formally tested: tweets\. LambdaG demonstrates strong performance when both logistic regression calibration and the proposed corrections are implemented\.
Figure 3:Comparison of three approaches on the Twitter dataset showing the relationship between Q Length, K Length and mean Cllr\. Each panel corresponds to a post\-processing method applied to LambdaG scores: logistic regression calibration \(left\), the Square Root Correction \(middle\), and the Hapax Correction \(right\)\.Figure[3](https://arxiv.org/html/2607.09501#S7.F3)shows the proposed approaches obtained well\-calibrated LLRs even when the case data was very short \-*K*length=500length=500tokenstokens\. This was the worst performing condition, yetCllrC\_\{llr\}values remained consistently low \(≤\\leq0\.39\)\. Contrastingly, for logistic regression calibration, when*K*length = 1,500, and*Q*length≤\\leq1,500, anomalously highCllrC\_\{llr\}values are evidenced\.
Across all the tested conditions, logistic regression calibration achieved the best performance in 67% of cases\. However, the Hapax Correction performs comparably overall\. Where logistic regression calibration obtained aCllrC\_\{llr\}\> 0\.1, the Hapax Correction matched or outperformed it in 65\.2% of conditions, while the Square Root Correction did so in 60\.9%\. In conditions where logistic regression calibration achievedCllrC\_\{llr\}< 0\.1, the Hapax Correction was always within 0\.04, and the Square Root Correction was within 0\.05 in 87\.01% of cases\. Directly comparing the Hapax Correction and the Square Root Correction, the former outperforms the latter across 87% of the tested conditions\. Nevertheless, the performance of the Square Root Correction remained within 5% of the Hapax Correction across 90% of conditions\.
#### 7\.3\.3Bolt
Figure[4](https://arxiv.org/html/2607.09501#S7.F4)illustrates theCllrC\_\{llr\}obtained when using the three approaches to post\-processλG\\lambda\_\{G\}calculated on the text messages inBolt\. The text lengths evaluated ranged from 100 to 500 tokens in increments of 100\. Considering this very limited amount of data, there is a strong performance across all three approaches, particularly with a minimum of 200 tokens of*Q*and*K*data\. Performance tends to improve further as the amount of data increases\. Whilst logistic regression calibration outperforms both the corrections across 80% of conditions, in certain conditions the Square Root Correction achieves the lowestCllrC\_\{llr\}, for example when*Q*length = 500 and*K*length = 500\.
Figure 4:Comparison of three approaches on the Bolt dataset showing the relationship between Q Length, K Length and mean Cllr\. Each panel corresponds to a post\-processing method applied to LambdaG scores: logistic regression calibration \(left\), the Square Root Correction \(middle\), and the Hapax Correction \(right\)\.In contrast to the performance observed on the other registers, on this dataset, the Square Root Correction tends to outperform the Hapax Correction, exhibiting better performance across 84% of the conditions tested\. Figure[4](https://arxiv.org/html/2607.09501#S7.F4)also indicates that, for this dataset, the Hapax Correction requires more than 100 tokens of*K*data to achieve a meaningful output \(CllrC\_\{llr\}below one\)\.
## 8Discussion
This paper has compared two novel alternatives to the calibration ofλG\\lambda\_\{G\}against the conventional approach of logistic regression calibration\(Niniet al\.,[2026](https://arxiv.org/html/2607.09501#bib.bib112)\)\. The objective was to eliminate the need to train inherently data\-dependent, time\-consuming and opaque calibration models in order to produce simple, accessible and well\-calibrated measures of evidential strength for AV\. We believe both the Square Root Correction and the Hapax Correction achieved this, with the Hapax Correction in particular demonstrating a strong performance across the fifteen corpora used for evaluation\.
Although the Hapax Correction achieved a lower win rate than the baseline of logistic regression calibration, it still outperformed this baseline in approximately 45% of the tests, weighted by corpora\. Furthermore, in the cases where logistic regression calibration achieved better performance than the Hapax Correction, the performance gap was more frequently small in magnitude i\.e\. less than 5% relative difference, than the inverse comparison\. This highlights the consistency in the performance of the Hapax Correction, as whilst logistic regression calibration may achieve stronger peak performance, its performance is also more variable\.
The evaluation across varying*Q*and*K*lengths revealed an interaction between input length and method performance\. Logistic regression calibration tends to outperform the Hapax Correction when*Q*and*K*texts are longer, achieving near optimal performance on theAvocadoandTwitterwhen*K*is longer than 5,500 tokens\. This may be due to the increased input supporting more stable parameter estimation\. Under these conditions, the Hapax Correction also demonstrates strong performance, consistently producing aCllrC\_\{llr\}below 0\.1, indicating well\-calibrated LLRs\. Comparatively, on shorter text lengths the Hapax Correction more frequently outperforms logistic regression calibration, demonstrating greater robustness with less linguistic input\. This is of note because these shorter text lengths are more likely to be practically relevant, indicating that the Hapax Correction may offer more reliable performance in real\-world cases\.
There were, however, certain conditions where the Hapax Correction did not produce meaningful LLRs\. This was demonstrated onBoltwhen the*K*sample was approximately 100 tokens\. This is likely because in a very short text the proportion of hapax legomena is extremely high\. This relates closely to Herdan\-Heap’s Law, which shows that vocabulary growth is rapid at first before it slows down considerably\(Herdan,[1960](https://arxiv.org/html/2607.09501#bib.bib44); Heaps,[1978](https://arxiv.org/html/2607.09501#bib.bib42)\)\. This makes the measure noisy and unreliable as it is overestimating uniqueness\. Nonetheless, for most authorship analysis methods, a*K*sample of 100 tokens is likely insufficient\(Stamatatos,[2009](https://arxiv.org/html/2607.09501#bib.bib97); Stamatatoset al\.,[2023](https://arxiv.org/html/2607.09501#bib.bib96)\)\. Both logistic regression calibration and the Square Root Correction also obtainedCllrC\_\{llr\}values greater than 0\.9, indicating comparatively poor calibration relative to the other conditions tested\.
In other tests, logistic regression calibration obtainedCllrC\_\{llr\}values greater than one, indicating that the system is misleading\(van Lieropet al\.,[2024](https://arxiv.org/html/2607.09501#bib.bib104)\)\. The anomalous values were seen in tests onAvocadoandTwitter\. These anomalies are likely due to the small calibration datasets \(4040problems inAvocado;6060inTwitter\), which makes the logistic regression sensitive to noise and prone to overfitting\. These results highlight a key limitation of logistic regression calibration: its reliance on substantial amounts of suitable calibration data, which is often costly and difficult to obtain\. In contrast, the proposed corrections exhibit lower variability, suggesting greater robustness\.
However, despite the greater robustness of the Square Root Correction, its performance was not as competitive overall\. It was outperformed by both the Hapax Correction and logistic regression calibration across the majority of tests, and where it was outperformed, it was typically not within a 5% margin of the other methods\. That said, logistic regression calibration represents a strong baseline, and the performance of the Hapax Correction is particularly effective\. The Square Root Correction is a simple transformation, and despite this simplicity, it is still able to outperform the logistic regression calibration baseline on certain corpora\. It is therefore still a meaningful and competitive approach\.
Furthermore, both the Square Root correction and the Hapax Correction move towards a more interpretable approach compared to logistic regression calibration\. By further aligning the assumptions underlying LambdaG with our theoretical and probabilistic understanding of language productivity, the proposed corrections move forensic AV towards the generative modeling approach seen in state\-of\-the\-art forensic sciences such as DNA profiling\(Balding and Nichols,[1994](https://arxiv.org/html/2607.09501#bib.bib5); Gillet al\.,[2021](https://arxiv.org/html/2607.09501#bib.bib30); Mitchell and Cheney,[2025](https://arxiv.org/html/2607.09501#bib.bib63)\)\. In these approaches, extraneous sources of variability are incorporated directly into the likelihood function\(Tayloret al\.,[2016](https://arxiv.org/html/2607.09501#bib.bib98)\)\. This enables the true probability of the observed DNA evidence to be calculated under each competing hypothesis\. As a result, LLRs are computed directly from the known generative distributions\(P\(x\|y\)\)\(P\(x\|y\)\)\(Collins and Morton,[1994](https://arxiv.org/html/2607.09501#bib.bib20); Buckletonet al\.,[2019](https://arxiv.org/html/2607.09501#bib.bib16)\), rather than derived through separate calibration procedures\.
We hypothesised that repetition leads to artificially inflatedλG\\lambda\_\{G\}values, given the implicit assumption of conditional independence in naïve\-Bayes\-based algorithms\. The usage\-based theories underlying LambdaG suggest that individuals exhibit consistent preferences for certain sequences\(Nini,[2023](https://arxiv.org/html/2607.09501#bib.bib80)\)\. This enables AV, as individuals use the same sequences across different texts\. However, this also suggests that repetitions within a single text may not be conditionally independent\. Therefore, subsequent repetitions of the same pattern should plausibly contribute less incremental evidential weight than their initial occurrence\.
This effect is consistent with established observations in lexical statistics, such as the sub\-linear growth of vocabulary described by Herdan\-Heap’s law\(Herdan,[1960](https://arxiv.org/html/2607.09501#bib.bib44); Heaps,[1978](https://arxiv.org/html/2607.09501#bib.bib42)\)\. This law formalises the decrease in the rate of introduction of new types as text length increases\.
As a result of this increased redundancy, tokens appearing later in a text are less likely to provide as much new stylistic evidence than earlier tokens\. Therefore, whilstλG\\lambda\_\{G\}grows approximately linearly with the number of tokens in the disputed document, the amount of effective new evidence grows sub\-linearly\. This creates a systematic bias, where longer texts are overstated as disproportionately informative under the raw additive score\. This effect is believed to be mitigated by the Square Root Correction\.
When implementing the Hapax Correction,λG\\lambda\_\{G\}is multiplied by the number of hapax legomena in the*Q*text, divided by the total number of tokens\. This constitutes an existing statistical measure of linguistic productivity called Baayen’s𝒫\\mathcal\{P\}\(Baayen,[2001](https://arxiv.org/html/2607.09501#bib.bib3), p\.50\)\. This metric can be used to estimate the slope of Herdan\-Heap’s law\. The Hapax correction could therefore more precisely model the relationship between text length, lexical productivity and inflatedλG\\lambda\_\{G\}values\.
Ultimately, the simplified approach achieved by both normalisations is important not only for experts but to ensure accessibility for fact\-finders\. The proposed corrections are more interpretable and transparent, which in turn facilitates more reliable decision\-making\.
## 9Limitations
The primary limitation of this research is the absence of formal mathematical proof specifying why the proposed corrections achieve such strong performance\. The corrections have some theoretical grounding, as we hypothesis that they proportionally scaleλG\\lambda\_\{G\}values that may be over\-inflated due to repetition in the data\. However, the reason underlying the impressive empirical performance of these specific metrics, as opposed to other measures of lexical diversity, is not yet fully understood\.
Furthering our understanding of these corrections is an avenue for future research and may enhance understanding of individual language productivity and the robustness of AV methods\. Furthermore, It is plausible that both the proposed corrections are approximating the same underlying concept, which, if identifiable, could constitute the optimal approach\.
It is important to note that, as with logistic regression calibration, for each AV case the forensic linguist is required to validate the application of the proposed corrections to that context\. As such, this partial understanding of the proposed corrections does not invalidate their use\.
A second limitation of this study is that the proposed methods have only been validated on English\. As such, another direction for future research is evaluating the proposed corrections on texts in different languages to obtain further insight as to the robustness and underlying assumptions of the normalisations\.
## 10Conclusion
This paper demonstrates that by applying simple transformations to LambdaG scores\(Niniet al\.,[2026](https://arxiv.org/html/2607.09501#bib.bib112)\), it is possible to generate meaningful LLRs without training a separate calibration model\. The transformations consist of dividing the LambdaG score by the square root of the number of tokens in the disputed text, or multiplying it by the ratio of hapax legomena to tokens in the text\.
One of the primary limitations of training a separate calibration model is the high dependency on data, which is often difficult and time consuming to obtain and prepare\. This process also relies on subjective expert decisions regarding data selection and structure \(e\.g\. anchoring\)\. As there is currently limited justification in the literature, these decisions typically require validation for each case\. This further increases the complexity of the method and the difficulty of explaining results to non\-experts\. By removing the need for a separate calibration model, the proposed corrections also eliminate these limitations, as well as issues such as over\- or underfitting, and the inherently opaque nature of the calibration models\.
The proposed corrections applied to LambdaG constitute what is currently the only method of AV that can generate well\-calibrated LLRs without a separate calibration model\. The corrections achieve this whilst maintaining the state\-of\-the\-art performance of LambdaG when calibrated using logistic regression\. This was tested across fifteen different corpora, and text lengths ranging from 100 to 9,500 tokens\. Overall, on average across datasets, the Hapax Correction outperforms logistic regression calibration approximately 45% the time, with comparable performance across the other instances\. The performance of the Square Root Correction was comparatively modest; however, its simplicity makes the results noteworthy\.
Furthermore, the proposed corrections are consistent with usage\-based linguistic theory and well\-established lexical statistics\. Although the reasoning for their effectiveness is not yet fully understood, the normalisations may account for the notion that repeated sequences within a text provide diminishing evidential value with each occurrence\. This aligns with the principles of language productivity on which LambdaG is based and offers greater scientific validity than logistic regression calibration\.
Overall, the proposed methods, particularly the Hapax Correction, achieved calibration performance comparable to logistic regression calibration across text lengths and registers\. The combination of empirical performance and unique practical advantages identified in this paper supports the adoption of the proposed approaches in forensic settings on English\-language texts\.
## References
- S\. Argamon \(2008\)Interpreting Burrows’s Delta: Geometric and Probabilistic Foundations\.Literary and Linguistic Computing23\(2\),pp\. 131–147\.External Links:ISSN 0268\-1145,[Document](https://dx.doi.org/10.1093/llc/fqn003)Cited by:[§3\.2](https://arxiv.org/html/2607.09501#S3.SS2.p2.1)\.
- R\. H\. Baayen \(2001\)Word Frequency Distributions\.Text, Speech and Language Technology,Springer Netherlands,Dordrecht\.External Links:[Document](https://dx.doi.org/10.1007/978-94-010-0844-0)Cited by:[§8](https://arxiv.org/html/2607.09501#S8.p11.3)\.
- D\. J\. Balding and R\. A\. Nichols \(1994\)DNA profile match probability calculation: how to allow for population stratification, relatedness, database selection and single bands\.Forensic Science International64\(2\),pp\. 125–140\.External Links:ISSN 0379\-0738,[Document](https://dx.doi.org/10.1016/0379-0738%2894%2990222-4)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p1.1),[§8](https://arxiv.org/html/2607.09501#S8.p7.1)\.
- K\. Benoit, A\. Matsuo, and J\. Gruber \(2023\)Spacyr: Wrapper to the ’spaCy’ ’NLP’ Library\.External Links:[Document](https://dx.doi.org/10.32614/CRAN.package.spacyr)Cited by:[§6\.2](https://arxiv.org/html/2607.09501#S6.SS2.p1.1)\.
- D\. Biber \(2012\)Register as a predictor of linguistic variation\.External Links:[Document](https://dx.doi.org/10.1515/cllt-2012-0002)Cited by:[§3\.3\.1](https://arxiv.org/html/2607.09501#S3.SS3.SSS1.p2.1)\.
- A\. Biedermann, S\. Bozza, F\. T\. Taroni, and C\. G\. G\. Aitken \(2016\)Reframing the debate: A question of probability, not of likelihood ratio\.Science & Justice56\(5\),pp\. 392–396\.External Links:ISSN 1355\-0306,[Document](https://dx.doi.org/10.1016/j.scijus.2016.05.008)Cited by:[§2](https://arxiv.org/html/2607.09501#S2.p1.3)\.
- G\. Biosa, D\. Giurghita, M\. Vincenti, and T\. Neocleous \(2020\)Evaluation of Forensic Data Using Logistic Regression\-Based Classification Methods and an R Shiny Implementation\.Frontiers in Chemistry8,pp\. 738\.External Links:ISSN 2296\-2646,[Document](https://dx.doi.org/10.3389/fchem.2020.00738)Cited by:[§3\.3](https://arxiv.org/html/2607.09501#S3.SS3.p1.1)\.
- B\. Boenninghoff, R\. Nickel, S\. Zeiler, and D\. Kolossa \(2019\)Similarity Learning for Authorship Verification in Social Media\.InInternational Conference on Acoustics, Speech and Signal Processing,External Links:[Document](https://dx.doi.org/10.1109/ICASSP.2019.8683405)Cited by:[§4\.1](https://arxiv.org/html/2607.09501#S4.SS1.p4.4)\.
- N\. Brümmer and G\. R\. Doddington \(2013\)Likelihood\-ratio calibration using prior\-weighted proper scoring rules\.InInterspeech 2013,External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2013-470)Cited by:[§3\.3](https://arxiv.org/html/2607.09501#S3.SS3.p1.1)\.
- N\. Brümmer and J\. du Preez \(2006\)Application\-independent evaluation of speaker detection\.Computer Speech & Language20\(2\),pp\. 230–275\.External Links:ISSN 0885\-2308,[Document](https://dx.doi.org/10.1016/j.csl.2005.08.001)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p3.1),[§3\.4](https://arxiv.org/html/2607.09501#S3.SS4.p1.1)\.
- J\. S\. Buckleton, J\. Bright, S\. Gittelson, T\. R\. Moretti, A\. J\. Onorato, F\. R\. Bieber, B\. Budowle, and D\. A\. Taylor \(2019\)The Probabilistic Genotyping Software STRmix: Utility and Evidence for its Validity\.Journal of Forensic Sciences64\(2\),pp\. 393–405\.External Links:ISSN 1556\-4029,[Document](https://dx.doi.org/10.1111/1556-4029.13898)Cited by:[§8](https://arxiv.org/html/2607.09501#S8.p7.1)\.
- S\. F\. Chen and J\. Goodman \(1999\)An empirical study of smoothing techniques for language modeling\.Computer Speech & Language13\(4\),pp\. 359–394\.External Links:ISSN 0885\-2308,[Document](https://dx.doi.org/10.1006/csla.1999.0128)Cited by:[§4\.2](https://arxiv.org/html/2607.09501#S4.SS2.p7.1)\.
- Chen, Song, Fore, Dana, Strassel, Stephanie, Lee, Haejoong, and Wright, Jonathan \(2018\)BOLT English SMS/Chat\.Linguistic Data Consortium\.External Links:[Document](https://dx.doi.org/10.35111/HKFC-7865)Cited by:[§6\.1\.2](https://arxiv.org/html/2607.09501#S6.SS1.SSS2.p6.3)\.
- A\. Collins and N\. E\. Morton \(1994\)Likelihood ratios for DNA identification\.Proceedings of the National Academy of Sciences of the United States of America91,pp\. 6007–11\.External Links:[Document](https://dx.doi.org/10.1073/pnas.91.13.6007)Cited by:[§8](https://arxiv.org/html/2607.09501#S8.p7.1)\.
- M\. Coulthard \(2004\)Author Identification, Idiolect, and Linguistic Uniqueness\.Applied Linguistics25\(4\),pp\. 431–447\.External Links:ISSN 0142\-6001,[Document](https://dx.doi.org/10.1093/applin/25.4.431)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p5.1),[§5\.2](https://arxiv.org/html/2607.09501#S5.SS2.p1.2)\.
- A\. P\. Dawid \(1982\)The Well\-Calibrated Bayesian\.Journal of the American Statistical Association77\(379\),pp\. 605–610\.External Links:ISSN 0162\-1459,[Document](https://dx.doi.org/10.1080/01621459.1982.10477856)Cited by:[§3\.1](https://arxiv.org/html/2607.09501#S3.SS1.p1.1)\.
- M\. H\. DeGroot and S\. E\. Fienberg \(1983\)The Comparison and Evaluation of Forecasters\.Journal of the Royal Statistical Society\. Series D \(The Statistician\)32\(1\),pp\. 12–22\.External Links:2987588,ISSN 0039\-0526,[Document](https://dx.doi.org/10.2307/2987588)Cited by:[§3\.1](https://arxiv.org/html/2607.09501#S3.SS1.p1.1)\.
- S\. Evert, T\. Proisl, F\. Jannidis, I\. Reger, S\. Pielström, C\. Schöch, and T\. Vitt \(2017\)Understanding and explaining Delta measures for authorship attribution\.Digital Scholarship in the Humanities32\(2\),pp\. ii4–ii16\.External Links:ISSN 2055\-7671,[Document](https://dx.doi.org/10.1093/llc/fqx023)Cited by:[§2](https://arxiv.org/html/2607.09501#S2.p7.1)\.
- Forensic Science Regulator \(2021\)Codes of Practice and Conduct\.Technical reportTechnical ReportFSR\-C\-118,United Kingdom\.Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p1.1)\.
- P\. Gill, C\. Benschop, J\. Buckleton, Ø\. Bleka, and D\. Taylor \(2021\)A Review of Probabilistic Genotyping Systems: EuroForMix, DNAStatistX and STRmix™\.Genes12\(10\),pp\. 1559\.External Links:ISSN 2073\-4425,[Document](https://dx.doi.org/10.3390/genes12101559)Cited by:[§8](https://arxiv.org/html/2607.09501#S8.p7.1)\.
- J\. Gonzalez\-Rodriguez, P\. Rose, D\. Ramos, D\. T\. Toledano, and J\. Ortega\-Garcia \(2007\)Emulating DNA: Rigorous Quantification of Evidential Weight in Transparent and Testable Forensic Speaker Recognition\.IEEE Transactions on Audio, Speech, and Language Processing15\(7\),pp\. 2104–2115\.External Links:ISSN 1558\-7924,[Document](https://dx.doi.org/10.1109/TASL.2007.902747)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p3.1),[§3\.3](https://arxiv.org/html/2607.09501#S3.SS3.p1.1)\.
- T\. Grant and J\. Grieve \(2022\)The Starbuck Case\.InMethodologies and Challenges in Forensic Linguistic Casework,pp\. 13–28\.External Links:[Document](https://dx.doi.org/10.1002/9781394266661.ch2),ISBN 978\-1\-394\-26666\-1Cited by:[§5\.3](https://arxiv.org/html/2607.09501#S5.SS3.p3.1)\.
- T\. Grant \(2022\)The Idea of Progress in Forensic Authorship Analysis\.Elements in Forensic Linguistics,Cambridge University Press,Cambridge\.External Links:[Document](https://dx.doi.org/10.1017/9781108974714),ISBN 978\-1\-108\-97132\-4Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p1.1)\.
- J\. Grieve, I\. Clarke, E\. Chiang, H\. Gideon, A\. Heini, A\. Nini, and E\. Waibel \(2019a\)Attributing the Bixby Letter using n\-gram tracing\.Digital Scholarship in the Humanities34\(3\),pp\. 493–512\.External Links:ISSN 2055\-7671, 2055\-768X,[Document](https://dx.doi.org/10.1093/llc/fqy042)Cited by:[§2](https://arxiv.org/html/2607.09501#S2.p7.1)\.
- J\. Grieve, C\. Montgomery, A\. Nini, A\. Murakami, and D\. Guo \(2019b\)Frontiers \| Mapping Lexical Dialect Variation in British English Using Twitter\.External Links:[Document](https://dx.doi.org/10.3389/frai.2019.00011)Cited by:[§6\.1\.2](https://arxiv.org/html/2607.09501#S6.SS1.SSS2.p5.2)\.
- C\. Guo, G\. Pleiss, Y\. Sun, and K\. Q\. Weinberger \(2017\)On Calibration of Modern Neural Networks\.ArXiv\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.1706.04599)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p2.1)\.
- O\. Halvani, L\. Graner, and I\. Vogel \(2018\)Authorship verification in the absence of explicit features and thresholds\.External Links:[Document](https://dx.doi.org/10.1007/978-3-319-76941-7%5F34)Cited by:[§4\.1](https://arxiv.org/html/2607.09501#S4.SS1.p2.2)\.
- O\. Halvani and L\. Graner \(2021\)POSNoise: An Effective Countermeasure Against Topic Biases in Authorship Analysis\.InProceedings of the 16th International Conference on Availability, Reliability and Security,ARES ’21,New York, NY, USA,pp\. 1–12\.External Links:[Document](https://dx.doi.org/10.1145/3465481.3470050),ISBN 978\-1\-4503\-9051\-4Cited by:[§4\.2](https://arxiv.org/html/2607.09501#S4.SS2.p2.3),[§6\.2](https://arxiv.org/html/2607.09501#S6.SS2.p1.1),[footnote 1](https://arxiv.org/html/2607.09501#footnote1)\.
- O\. Halvani, C\. Winter, and L\. Graner \(2017\)On the Usefulness of Compression Models for Authorship Verification\.InProceedings of the 12th International Conference on Availability, Reliability and Security,ARES ’17,New York, NY, USA,pp\. 1–10\.External Links:[Document](https://dx.doi.org/10.1145/3098954.3104050),ISBN 978\-1\-4503\-5257\-4Cited by:[§4\.1](https://arxiv.org/html/2607.09501#S4.SS1.p4.4)\.
- O\. Halvani, C\. Winter, and L\. Graner \(2019\)Assessing the applicability of authorship verification methods\.ARES ’19,New York, NY, USA,pp\. 1–10\.External Links:[Document](https://dx.doi.org/10.1145/3339252.3340508),[Link](https://dl.acm.org/doi/10.1145/3339252.3340508)Cited by:[§4\.1](https://arxiv.org/html/2607.09501#S4.SS1.p1.1)\.
- H\. S\. Heaps \(1978\)Information Retrieval, Computational and Theoretical Aspects\.Academic Press\.External Links:ISBN 978\-0\-12\-335750\-2Cited by:[§5\.2](https://arxiv.org/html/2607.09501#S5.SS2.p1.2),[§8](https://arxiv.org/html/2607.09501#S8.p4.1),[§8](https://arxiv.org/html/2607.09501#S8.p9.1)\.
- A\. B\. Hepler, C\. P\. Saunders, L\. J\. Davis, and J\. Buscaglia \(2012\)Score\-based likelihood ratios for handwriting evidence\.Forensic Science International219\(1\),pp\. 129–140\.External Links:ISSN 0379\-0738,[Document](https://dx.doi.org/10.1016/j.forsciint.2011.12.009)Cited by:[§3\.3\.1](https://arxiv.org/html/2607.09501#S3.SS3.SSS1.p8.1)\.
- G\. Herdan \(1960\)Type\-token Mathematics\.Mouton\.Cited by:[§5\.2](https://arxiv.org/html/2607.09501#S5.SS2.p1.2),[§8](https://arxiv.org/html/2607.09501#S8.p4.1),[§8](https://arxiv.org/html/2607.09501#S8.p9.1)\.
- J\. Huertas\-Tato, A\. Martín, and D\. Camacho \(2024\)Understanding writing style in social media with a supervised contrastively pre\-trained transformer\.Knowledge\-Based Systems296,pp\. 111867\.External Links:ISSN 0950\-7051,[Document](https://dx.doi.org/10.1016/j.knosys.2024.111867)Cited by:[§2](https://arxiv.org/html/2607.09501#S2.p8.1),[§4\.1](https://arxiv.org/html/2607.09501#S4.SS1.p4.4)\.
- S\. Ishihara, S\. Tsuge, M\. Inaba, and W\. Zaitsu \(2022\)Estimating the Strength of Authorship Evidence with a Deep\-Learning\-Based Approach\.InProceedings of the 20th Annual Workshop of the Australasian Language Technology Association,P\. Parameswaran, J\. Biggs, and D\. Powers \(Eds\.\),Adelaide, Australia,pp\. 183–187\.Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p1.1),[§3\.3](https://arxiv.org/html/2607.09501#S3.SS3.p1.1)\.
- S\. Ishihara \(2011\)A Forensic Authorship Classification in SMS Messages: A Likelihood Ratio Based Approach Using N\-gram\.InProceedings of the Australasian Language Technology Association Workshop 2011,D\. Molla and D\. Martinez \(Eds\.\),Canberra, Australia,pp\. 47–56\.Cited by:[§3\.2](https://arxiv.org/html/2607.09501#S3.SS2.p4.1)\.
- S\. Ishihara \(2017\)Strength of forensic text comparison evidence from stylometric features: A multivariate likelihood ratio\-based analysis\.International Journal of Speech Language and the Law24,pp\. 67–98\.External Links:[Document](https://dx.doi.org/10.1558/ijsll.30305)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p1.1)\.
- S\. Ishihara \(2021\)Score\-based likelihood ratios for linguistic text evidence with a bag\-of\-words model\.Forensic Science International327\.External Links:ISSN 0379\-0738,[Document](https://dx.doi.org/10.1016/j.forsciint.2021.110980)Cited by:[§2](https://arxiv.org/html/2607.09501#S2.p1.3)\.
- P\. Juola \(2021\)Verifying authorship for forensic purposes: A computational protocol and its validation\.Forensic Science International325\.External Links:ISSN 0379\-0738,[Document](https://dx.doi.org/10.1016/j.forsciint.2021.110824)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p5.1)\.
- D\. Jurafsky and J\. H\. Martin \(2026\)Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition, with Language Models\.3 edition\.Cited by:[§4\.2](https://arxiv.org/html/2607.09501#S4.SS2.p5.5)\.
- Reinnhard\. Kneser and H\. Ney \(1995\)Improved backing\-off for M\-gram language modeling\.In1995 International Conference on Acoustics, Speech, and Signal Processing,Vol\.1,pp\. 181–184 vol\.1\.External Links:ISSN 1520\-6149,[Document](https://dx.doi.org/10.1109/ICASSP.1995.479394)Cited by:[§4\.2](https://arxiv.org/html/2607.09501#S4.SS2.p7.1)\.
- M\. Koppel, J\. Schler, S\. Argamon, and Y\. Winter \(2012\)The “Fundamental Problem” of Authorship Attribution\.English Studies93\(3\),pp\. 284–291\.External Links:ISSN 0013\-838X,[Document](https://dx.doi.org/10.1080/0013838X.2012.668794)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p5.1)\.
- M\. Koppel and Y\. Winter \(2014\)Determining if two documents are written by the same author\.Journal of the Association for Information Science and Technology65\(1\),pp\. 178–187\.External Links:ISSN 2330\-1643,[Document](https://dx.doi.org/10.1002/asi.22954)Cited by:[§2](https://arxiv.org/html/2607.09501#S2.p8.1),[§4\.1](https://arxiv.org/html/2607.09501#S4.SS1.p4.4)\.
- A\. Macarulla Rodriguez, Z\. Geradts, M\. Worring, and L\. Unzueta \(2024\)Improved likelihood ratios for face recognition in surveillance video by multimodal feature pairing\.Forensic Science International: Synergy8,pp\. 100458\.External Links:ISSN 2589\-871X,[Document](https://dx.doi.org/10.1016/j.fsisyn.2024.100458)Cited by:[§3\.3](https://arxiv.org/html/2607.09501#S3.SS3.p1.1)\.
- K\. J\. Mitchell and N\. Cheney \(2025\)The Genomic Code: the genome instantiates a generative model of the organism\.Trends in Genetics41\(6\),pp\. 462–479\.External Links:ISSN 0168\-9525,[Document](https://dx.doi.org/10.1016/j.tig.2025.01.008)Cited by:[§8](https://arxiv.org/html/2607.09501#S8.p7.1)\.
- S\. Mollin \(2009\)“I entirely understand” is a Blairism: The methodology of identifying idiolectal collocations\.External Links:[Document](https://dx.doi.org/10.1075/ijcl.14.3.04mol)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p5.1)\.
- G\. Morrison, F\. Ochoa, and T\. Thiruvaran \(2012\)Database selection for forensic voice comparison\.Proceedings of Odyssey 2012: The Language and Speaker Recognition Workshop\.Cited by:[§3\.3\.1](https://arxiv.org/html/2607.09501#S3.SS3.SSS1.p3.1),[§3\.3\.1](https://arxiv.org/html/2607.09501#S3.SS3.SSS1.p4.1)\.
- G\. S\. Morrison, C\. Zhang, and P\. Rose \(2011\)An empirical estimate of the precision of likelihood ratios from a forensic\-voice\-comparison system\.Forensic Science International208\(1\),pp\. 59–65\.External Links:ISSN 0379\-0738,[Document](https://dx.doi.org/10.1016/j.forsciint.2010.11.001)Cited by:[§3](https://arxiv.org/html/2607.09501#S3.p1.1)\.
- G\. S\. Morrison \(2011\)Measuring the validity and reliability of forensic likelihood\-ratio systems\.Science & Justice51\(3\),pp\. 91–98\.External Links:ISSN 1355\-0306,[Document](https://dx.doi.org/10.1016/j.scijus.2011.03.002)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p1.1),[§3\.4](https://arxiv.org/html/2607.09501#S3.SS4.p1.1),[§3\.4](https://arxiv.org/html/2607.09501#S3.SS4.p3.1)\.
- G\. S\. Morrison \(2013\)Tutorial on logistic\-regression calibration and fusion:converting a score to a likelihood ratio\.Australian Journal of Forensic Sciences45\(2\),pp\. 173–197\.External Links:ISSN 0045\-0618,[Document](https://dx.doi.org/10.1080/00450618.2012.733025)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p2.1),[§1](https://arxiv.org/html/2607.09501#S1.p3.1),[§3\.2](https://arxiv.org/html/2607.09501#S3.SS2.p4.1),[§3\.3\.2](https://arxiv.org/html/2607.09501#S3.SS3.SSS2.p1.1),[§3\.3](https://arxiv.org/html/2607.09501#S3.SS3.p1.1)\.
- G\. S\. Morrison \(2021\)In the context of forensic casework, are there meaningful metrics of the degree of calibration?\.Forensic Science International: Synergy3\.External Links:ISSN 2589\-871X,[Document](https://dx.doi.org/10.1016/j.fsisyn.2021.100157)Cited by:[§3\.3\.1](https://arxiv.org/html/2607.09501#S3.SS3.SSS1.p2.1),[§3\.3\.1](https://arxiv.org/html/2607.09501#S3.SS3.SSS1.p4.1)\.
- G\. S\. Morrison \(2024\)Bi\-Gaussianized calibration of likelihood ratios\.Law, Probability and Risk23\(1\)\.External Links:ISSN 1470\-8396,[Document](https://dx.doi.org/10.1093/lpr/mgae004)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p3.1),[§3\.3\.1](https://arxiv.org/html/2607.09501#S3.SS3.SSS1.p2.1),[§3](https://arxiv.org/html/2607.09501#S3.p1.1)\.
- L\. Mullen \(2020\)Minhash and locality\-sensitive hashing\.Note:https://cran\.r\-project\.org/web/packages/textreuse/vignettes/textreuse\-minhash\.htmlCited by:[§6\.1\.2](https://arxiv.org/html/2607.09501#S6.SS1.SSS2.p4.3)\.
- A\. Nini, O\. Halvani, L\. Graner, S\. Titze, V\. Gherardi, and S\. Ishihara \(2026\)Grammar as a behavioral biometric: using cognitively motivated grammar models for authorship verification\.Humanities and Social Sciences Communications13\(1\),pp\. 455\.External Links:[Document](https://dx.doi.org/10.1057/s41599-025-06340-3),[Link](https://www.nature.com/articles/s41599-025-06340-3)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p1.1),[§1](https://arxiv.org/html/2607.09501#S1.p6.1),[§10](https://arxiv.org/html/2607.09501#S10.p1.1),[§2](https://arxiv.org/html/2607.09501#S2.p8.1),[§3\.2](https://arxiv.org/html/2607.09501#S3.SS2.p2.1),[§4\.1](https://arxiv.org/html/2607.09501#S4.SS1.p4.4),[§4\.2](https://arxiv.org/html/2607.09501#S4.SS2.p1.1),[§4\.2](https://arxiv.org/html/2607.09501#S4.SS2.p17.2),[§4\.2](https://arxiv.org/html/2607.09501#S4.SS2.p2.3),[§5\.1](https://arxiv.org/html/2607.09501#S5.SS1.p1.2),[§5\.1](https://arxiv.org/html/2607.09501#S5.SS1.p2.3),[§6\.1\.1](https://arxiv.org/html/2607.09501#S6.SS1.SSS1.p1.1),[§6\.1\.1](https://arxiv.org/html/2607.09501#S6.SS1.SSS1.p2.1),[§6\.1\.1](https://arxiv.org/html/2607.09501#S6.SS1.SSS1.p4.1),[§6\.1\.2](https://arxiv.org/html/2607.09501#S6.SS1.SSS2.p2.1),[§6\.1\.2](https://arxiv.org/html/2607.09501#S6.SS1.SSS2.p3.1),[§6\.2](https://arxiv.org/html/2607.09501#S6.SS2.p1.1),[§6\.2](https://arxiv.org/html/2607.09501#S6.SS2.p2.4),[§6\.2](https://arxiv.org/html/2607.09501#S6.SS2.p3.1),[§6\.2](https://arxiv.org/html/2607.09501#S6.SS2.p5.4),[§7\.2](https://arxiv.org/html/2607.09501#S7.SS2.p1.1),[§7\.2](https://arxiv.org/html/2607.09501#S7.SS2.p2.1),[§8](https://arxiv.org/html/2607.09501#S8.p1.1)\.
- A\. Nini \(2017\)Register variation in malicious forensic texts\.The International Journal of Speech, Language and the Law24\(1\),pp\. 99–126\.External Links:ISSN 1748\-8885,[Document](https://dx.doi.org/10.1558/ijsll.30173)Cited by:[§3\.3\.1](https://arxiv.org/html/2607.09501#S3.SS3.SSS1.p3.1)\.
- A\. Nini \(2023\)A Theory of Linguistic Individuality for Authorship Analysis\.Elements in Forensic Linguistics,Cambridge University Press,Cambridge\.External Links:[Document](https://dx.doi.org/10.1017/9781108974851),ISBN 978\-1\-108\-97138\-6Cited by:[§3\.2](https://arxiv.org/html/2607.09501#S3.SS2.p2.1),[§5\.2](https://arxiv.org/html/2607.09501#S5.SS2.p1.2),[§8](https://arxiv.org/html/2607.09501#S8.p8.1)\.
- A\. Nini \(2024\)Idiolect: An R package for forensic authorship analysis\.External Links:[Document](https://dx.doi.org/10.32614/CRAN.package.idiolect)Cited by:[§6\.2](https://arxiv.org/html/2607.09501#S6.SS2.p1.1),[§6\.2](https://arxiv.org/html/2607.09501#S6.SS2.p2.4)\.
- A\. Nini \(2026\)Idiolect: An R package for forensic authorship analysis\.11\(119\)\.Cited by:[§6](https://arxiv.org/html/2607.09501#S6.p1.2)\.
- J\. Noecker Jr and M\. Ryan \(2012\)LREC 2012\.N\. Calzolari, K\. Choukri, T\. Declerck, M\. U\. Doğan, B\. Maegaard, J\. Mariani, A\. Moreno, J\. Odijk, and S\. Piperidis \(Eds\.\),Istanbul, Turkey,pp\. 785–789\.External Links:[Link](https://aclanthology.org/L12-1090/)Cited by:[§4\.1](https://arxiv.org/html/2607.09501#S4.SS1.p2.2)\.
- Oard, Douglas, Webber, William, Kirsch, David A\., and Golitsynskiy, Sergey \(2015\)Avocado Research Email Collection\.Linguistic Data Consortium\.External Links:[Document](https://dx.doi.org/10.35111/WQT6-JG60)Cited by:[§6\.1\.2](https://arxiv.org/html/2607.09501#S6.SS1.SSS2.p3.1)\.
- N\. Potha and E\. Stamatatos \(2014\)A Profile\-Based Method for Authorship Verification\.InArtificial Intelligence:Methods and Application,Lecture Notes in Computer Science,pp\. 312–326\.External Links:[Document](https://dx.doi.org/10.1007/978-3-319-07064-3%5F25)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p5.1)\.
- N\. Potha and E\. Stamatatos \(2017\)An Improved Impostors Method for Authorship Verification\.External Links:[Document](https://dx.doi.org/10.1007/978-3-319-65813-1%5F14)Cited by:[§4\.1](https://arxiv.org/html/2607.09501#S4.SS1.p4.4)\.
- Y\. Pull and C\. Hurlin \(2025\)A Bayesian Approach to Probability Default Model Calibration: Theoretical and Empirical Insights on the Jeffreys Test\.SSRN Scholarly Paper,Social Science Research Network,Rochester, NY\.External Links:5291474,[Document](https://dx.doi.org/10.2139/ssrn.5291474)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p2.1)\.
- D\. Ramos and J\. Gonzalez\-Rodriguez \(2013\)Reliable support: Measuring calibration of likelihood ratios\.Forensic Science International230\(1\-3\),pp\. 156–69\.External Links:[Document](https://dx.doi.org/10.1016/j.forsciint.2013.04.014)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p4.1)\.
- S\. Reinders, Y\. Guan, D\. Ommen, and J\. Newman \(2022\)Source\-anchored, trace\-anchored, and general match score\-based likelihood ratios for camera device identification\.Journal of Forensic Sciences67\(3\),pp\. 975–988\.External Links:ISSN 1556\-4029,[Document](https://dx.doi.org/10.1111/1556-4029.14991)Cited by:[§3\.3\.1](https://arxiv.org/html/2607.09501#S3.SS3.SSS1.p7.1)\.
- R\. A\. Rivera\-Soto, O\. E\. Miano, J\. Ordonez, B\. Y\. Chen, A\. Khan, M\. Bishop, and N\. Andrews \(2021\)Learning Universal Authorship Representations\.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing,M\. Moens, X\. Huang, L\. Specia, and S\. W\. Yih \(Eds\.\),Online and Punta Cana, Dominican Republic,pp\. 913–919\.External Links:[Document](https://dx.doi.org/10.18653/v1/2021.emnlp-main.70)Cited by:[§2](https://arxiv.org/html/2607.09501#S2.p8.1)\.
- D\. Roemling and J\. Grieve \(2024\)Forensic Authorship Analysis\.Note:https://crestresearch\.ac\.uk/comment/forensic\-authorship\-analysis/Cited by:[§5\.3](https://arxiv.org/html/2607.09501#S5.SS3.p3.1)\.
- P\. Rose \(2006\)Technical forensic speaker recognition: Evaluation, types and testing of evidence\.Computer Speech & Language20\(2\),pp\. 159–191\.External Links:ISSN 0885\-2308,[Document](https://dx.doi.org/10.1016/j.csl.2005.07.003)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p1.1)\.
- T\. Silva Filho, H\. Song, M\. Perello\-Nieto, R\. Santos\-Rodriguez, M\. Kull, and P\. Flach \(2023\)Classifier calibration: a survey on how to assess and improve predicted class probabilities\.Machine Learning112\(9\),pp\. 3211–3260\.External Links:ISSN 1573\-0565,[Document](https://dx.doi.org/10.1007/s10994-023-06336-7)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p2.1),[§3\.1](https://arxiv.org/html/2607.09501#S3.SS1.p1.1)\.
- E\. Stamatatos, K\. Kredens, P\. Pezik, A\. Heini, J\. Bevendorff, B\. Stein, and M\. Potthast \(2023\)Overview of the Authorship Verification Task at PAN 2023\.InConference and Labs of the Evaluation Forum,Cited by:[§8](https://arxiv.org/html/2607.09501#S8.p4.1)\.
- E\. Stamatatos \(2009\)A survey of modern authorship attribution methods\.Journal of the American Society for Information Science and Technology60\(3\),pp\. 538–556\.External Links:ISSN 1532\-2890,[Document](https://dx.doi.org/10.1002/asi.21001)Cited by:[§8](https://arxiv.org/html/2607.09501#S8.p4.1)\.
- D\. Taylor, J\. Bright, and J\. Buckleton \(2013\)The interpretation of single source and mixed DNA profiles\.Forensic Science International\. Genetics7\(5\),pp\. 516–528\.External Links:[Document](https://dx.doi.org/10.1016/j.fsigen.2013.05.011)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p1.1)\.
- D\. Taylor, J\. Buckleton, and J\. Bright \(2016\)Factors affecting peak height variability for short tandem repeat data\.Forensic Science International: Genetics21,pp\. 126–133\.External Links:ISSN 1872\-4973,[Document](https://dx.doi.org/10.1016/j.fsigen.2015.12.009)Cited by:[§8](https://arxiv.org/html/2607.09501#S8.p7.1)\.
- J\. W\. \(\. W\. Tukey \(1977\)Exploratory data analysis\.Addison\-Wesley Pub\. Co\.,Reading, Massachusetts\.External Links:[Link](http://archive.org/details/exploratorydataa0000tuke_7616)Cited by:[§6\.1\.2](https://arxiv.org/html/2607.09501#S6.SS1.SSS2.p2.1)\.
- D\. van der Vloed \(2024\)Interchangeability of Calibration Audio Datasets for Forensic Automatic Speaker Recognition\.In2024 12th International Workshop on Biometrics and Forensics \(IWBF\),pp\. 1–6\.External Links:[Document](https://dx.doi.org/10.1109/IWBF62628.2024.10593938)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p2.1)\.
- S\. van Lierop, D\. Ramos, M\. Sjerps, and R\. Ypma \(2024\)An overview of log likelihood ratio cost in forensic science – Where is it used and what values can we expect?\.Forensic Science International: Synergy8\.External Links:ISSN 2589\-871X,[Document](https://dx.doi.org/10.1016/j.fsisyn.2024.100466)Cited by:[§3\.4](https://arxiv.org/html/2607.09501#S3.SS4.p1.1),[§3\.4](https://arxiv.org/html/2607.09501#S3.SS4.p6.3),[§7\.3\.1](https://arxiv.org/html/2607.09501#S7.SS3.SSS1.p1.6),[§8](https://arxiv.org/html/2607.09501#S8.p5.3)\.
- A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. ukasz Kaiser, and I\. Polosukhin \(2017\)Attention is All you Need\.InAdvances in Neural Information Processing Systems,Vol\.30\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.1706.03762)Cited by:[§5\.2](https://arxiv.org/html/2607.09501#S5.SS2.p3.1)\.
- P\. Vergeer, Y\. van Schaik, and M\. Sjerps \(2021\)Measuring calibration of likelihood\-ratio systems: A comparison of four metrics, including a new metric devPAV\.Forensic Science International321\.External Links:ISSN 0379\-0738,[Document](https://dx.doi.org/10.1016/j.forsciint.2021.110722)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p4.1)\.
- D\. S\. Wilks \(2009\)Extending logistic regression to provide full\-probability\-distribution MOS forecasts\.Meteorological Applications16\(3\),pp\. 361–368\.External Links:ISSN 1469\-8080,[Document](https://dx.doi.org/10.1002/met.134)Cited by:[§3\.1](https://arxiv.org/html/2607.09501#S3.SS1.p3.1)\.
- D\. Wright \(2017\)Using word n\-grams to identify authors and idiolects: A corpus approach to a forensic linguistic problem\.International Journal of Corpus Linguistics22\(2\),pp\. 212–241\.External Links:ISSN 1384\-6655, 1569\-9811,[Document](https://dx.doi.org/10.1075/ijcl.22.2.03wri)Cited by:[§1](https://arxiv.org/html/2607.09501#S1.p5.1)\.
## Appendix\\thechapter\.ALambdaG Algorithm
Algorithm 1LambdaGInput:
QQ,
KK,
ℝ\\mathbb\{R\},
NN,
rr
QQis the questioned text;KKis the known\-author text;ℝ\\mathbb\{R\}is a set of reference texts \(ℝ\\mathbb\{R\}=R1R^\{1\},R2R^\{2\}, …\);NNis the order of the model;rris the number of repetitions
Function
posnoise\(D\)posnoise\(D\)Returns a POS\-noise version ofDD, where content words are replaced by their part\-of\-speech tags
Function
sent\(D\)sent\(D\)Returns the set of all tokenized sentences in document\(s\)DD
Function
sample\(𝕊,n\)sample\(\\mathbb\{S\},n\)Randomly samplesnntokens from the set𝕊\\mathbb\{S\}
Function
KN\(𝕊,N\)KN\(\\mathbb\{S\},N\)Train an n\-gram language model of order N with the set of sentences𝕊\\mathbb\{S\}using Kneser\-Ney smoothing
Apply POSNoise to al respective documents and construct tokenized sentences
𝕊Q←sent\(posnoise\(Q\)\)\\mathbb\{S\}\_\{Q\}\\leftarrow sent\(posnoise\(Q\)\)
𝕊K←sent\(posnoise\(K\)\)\\mathbb\{S\}\_\{K\}\\leftarrow sent\(posnoise\(K\)\)
𝕊R←sent\(posnoise\(ℝ\)\)\\mathbb\{S\}\_\{R\}\\leftarrow sent\(posnoise\(\\mathbb\{R\}\)\)
Build Grammar Model for the known author
GK←KN\(𝕊K,N\)G\_\{K\}\\leftarrow KN\(\\mathbb\{S\}\_\{K\},N\)
Build reference grammar models
for
i←1i\\leftarrow 1to
rrdo
𝕊i←sample\(𝕊R,\|𝕊K\|\)\\mathbb\{S\}\_\{i\}\\leftarrow sample\(\\mathbb\{S\}\_\{R\},\|\\mathbb\{S\}\_\{K\}\|\)
GR\(i\)←KN\(𝕊i,N\)G\_\{R\}^\{\(i\)\}\\leftarrow KN\(\\mathbb\{S\}\_\{i\},N\)
endfor
Calculate theλG\\lambda\_\{G\}\(the log\-likelihood ratio of the Grammar Model\) over𝕊Q\\mathbb\{S\}\_\{Q\}
λG\(𝕊Q←0\{\\lambda\_\{G\}\}\(\\mathbb\{S\}\_\{Q\}\\leftarrow 0
for
Si∈S\_\{i\}\\in𝕊Q\\mathbb\{S\}\_\{Q\}doDecomposeSiS\_\{i\}into a sequence of tokens
\(t1,t2,…,tz\)←Si\(t\_\{1\},t\_\{2\},\\dots,t\_\{z\}\)\\leftarrow S\_\{i\}
for
j←1j\\leftarrow 1to
zzdoCalculate meanλG\\lambda\_\{G\}fortjt\_\{j\}over reference Grammar Models
for
i←1i\\leftarrow 1to
rrdo
λG\(𝕊Q\)←λG\(𝕊Q\)\+1rlogP\(tj\|t<j;GKP\(tj\|t<j;Gj\\lambda\_\{G\}\(\\mathbb\{S\}\_\{Q\}\)\\leftarrow\\lambda\_\{G\}\(\\mathbb\{S\}\_\{Q\}\)\+\\frac\{1\}\{r\}\\log\\frac\{P\(t\_\{j\}\|t\_\{<j\};G\_\{K\}\}\{P\(t\_\{j\}\|t\_\{<j\};G\_\{j\}\}
endfor
endfor
endfor
return
λG\(𝕊Q\)\\lambda\_\{G\}\(\\mathbb\{S\}\_\{Q\}\)Similar Articles
Fusing Stylometric and Embedding Systems to Estimate Authorship Likelihood Ratios in Japanese
This paper applies the likelihood ratio framework for forensic authorship attribution to Japanese texts, fusing stylometric features with embedding-based systems to improve discrimination and calibration.
Log-Likelihood, Simpson's Paradox, and the Detection of Machine-Generated Text
This paper addresses the degradation of likelihood-based machine-generated text detectors by identifying a Simpson's paradox in token-score aggregation. It proposes a learned local calibration step that significantly improves detection performance across various models and datasets.
READER: Robust Evidence-based Authorship Decoding via Extracted Representations
Introduces READER, a lightweight framework for dynamic black-box LLM provenance that uses a frozen proxy LLM to extract authorship evidence from responses and performs Bayesian evidence accumulation across multiple queries, achieving high accuracy on the Agent500 dataset.
Retrieval-Augmented Linguistic Calibration
This paper proposes Retrieval-Augmented Linguistic Calibration (RALC), a post-hoc pipeline for calibrating confidence signals in LLMs by modeling linguistic confidence as a distribution and using retrieval-augmented rewriting. It introduces Faithfulness Divergence metric and shows significant improvements across benchmarks.
Lightweight Stylistic Consistency Profiling: Robust Detection of LLM-Generated Textual Content for Multimedia Moderation
Proposes LiSCP, a lightweight stylistic consistency profiling method for robust detection of LLM-generated textual content, focusing on feature stability under adversarial manipulation. Achieves superior performance on in-domain and cross-domain detection with notable robustness.