Interpretable, Fairly Evaluated Automated L2 Speaking Assessment that Beats the Single-Human Ceiling and Why Pause Encoding Does Not Change LLM Fluency Scores
Summary
This paper presents an interpretable feature-LLM hybrid system for automated L2 English speaking assessment that outperforms individual human raters and shows that pause encoding does not significantly impact LLM fluency scores.
View Cached Full Text
Cached at: 08/28/26, 09:19 AM
# Interpretable, Fairly Evaluated Automated L2 Speaking Assessment that Beats the Single-Human Ceiling — and Why Pause Encoding Does Not Change LLM Fluency Scores
Source: [https://arxiv.org/html/2608.26137](https://arxiv.org/html/2608.26137)
###### Abstract
Second\-language \(L2\) English learners can rarely rehearse speaking with a partner\. Speaking is also the most anxiety\-laden skill\. These gaps drive a fast\-growing market for automated speaking practice and scoring\. But an automated score is trustworthy only if it is accurate, interpretable, fair, and benchmarked against the*right*human bar\. We build an interpretable feature\-plus\-LLM hybrid for spontaneous L2 dialogue\. We evaluate it without ever fitting to the human labels, against the ICNALE Global Rating Archive:140140speeches rated by∼\\sim8080trained raters on1010analytic criteria\. We score the130130L2 speeches with usable audio\. A deterministic De\-Jong speech\-timing composite reachesρ=0\.764\\rho=0\.764\. Blended with a single text\-LLM fluency judgment, it reaches Spearmanρ=0\.818\\rho=0\.818against the consensus gold\. This agrees with the consensus better than81%81\\%of the8080individual trained raters: above the median rater \(ρ=0\.73\\rho=0\.73\) and near the best, and at∼\\sim83%83\\%of the reliability\-corrected maximum \(κmax=0\.99\\kappa\_\{\\max\}=0\.99\)\. The blend improves on the composite alone by\+0\.054\+0\.054\(paired\-bootstrap95%95\\%CI\[0\.017,0\.108\]\[0\.017,0\.108\], excludes0\); the LLM adds a coarse fluency ranking that the continuous composite refines\. We also report a controllednullon pause encoding, bounded to effects below about±0\.1ρ\\pm 0\.1\\,\\rhoat this sample size\. Holding the LLM and learner words fixed and varying only how pauses are written into the prompt, inline pause*locations*do not beat aggregate pause*statistics*\(−0\.069\-0\.069, CI\[−0\.15,\+0\.08\]\[\-0\.15,\+0\.08\]\), and a grounded mid\-clause criterion gives no reliable gain\. The fluency signal comes from the measured speech\-timing features, not from how pauses are written for the LLM\. We back every claim with two agreeing learner\-isolation methods, paired\-bootstrap CIs, a monologue negative control, per\-feature reproduction of classical measurements, and a per\-L1 fairness audit\.
## 1Introduction
Second\-language acquisition theory says learners advance not only by receiving language but by*producing*and*negotiating*it\. Swain’s output hypothesis came from Canadian French\-immersion data\. It argued that comprehensible input alone is not enough\. Learners must produce “pushed output” to notice gaps, test hypotheses, and develop accuracy and fluency\(Swain,[1985](https://arxiv.org/html/2608.26137#bib.bib27),[1995](https://arxiv.org/html/2608.26137#bib.bib28)\)\. Long’s interaction hypothesis adds that acquisition is driven by interactional adjustments\. These include clarification requests and confirmation and comprehension checks\. They occur during the negotiation of meaning with a conversation partner\(Long,[1996](https://arxiv.org/html/2608.26137#bib.bib20)\)\. The implication is sharp\. The interactive speaking that theory privileges is exactly what a solitary learner cannot self\-supply\.
This matters because oral production is the skill L2 English learners can least rehearse\. In many EFL contexts, classroom talk is teacher\-dominated, classes are large, and out\-of\-class exposure is minimal\. For many learners the coursebook is the only place they meet English\(ERIC EJ1319829,[2021](https://arxiv.org/html/2608.26137#bib.bib11)\)\. Speaking is also the most anxiety\-laden skill\. Horwitz, Horwitz and Cope established foreign language anxiety as a distinct construct\. Its first and dominant component is communication apprehension while speaking\(Horwitz et al\.,[1986](https://arxiv.org/html/2608.26137#bib.bib13); ERIC EJ1305502,[2019](https://arxiv.org/html/2608.26137#bib.bib12)\)\. Cross\-national learner surveys confirm the pattern at scale\. Fear of speaking up and making mistakes is the single most reported barrier to language\-learning success\(Preply,[2025](https://arxiv.org/html/2608.26137#bib.bib25)\)\. AI conversation partners now address this access\-and\-anxiety gap at population scale\. Peer\-reviewed evidence shows that judgment\-free AI practice both raises proficiency and lowers speaking anxiety\(Ding & Yusof,[2025](https://arxiv.org/html/2608.26137#bib.bib10)\)\. The same wave has put automated speaking*scoring*into high\-stakes use such as university admissions\(Isaacs et al\.,[2023](https://arxiv.org/html/2608.26137#bib.bib15)\)\.
But an automated speaking score is only as good as it is trustworthy\. A recent comparison of commercial scoring tools reports correlations with human raters aroundr≈0\.85r\\approx 0\.85, while also finding score inflation and missed detail, and operational practice still routes hard cases to humans\(Chen & Sun,[2025](https://arxiv.org/html/2608.26137#bib.bib6)\)\. Two consequences follow\. First, the right target is human\-rater*agreement under known rater unreliability*, not a single noiseless gold label\. A model should be measured against what one trained human actually achieves, and against the highest score a perfect model could reach once we correct for rater noise \(the reliability\-corrected maximum\)\. Second, a high correlation does not make an opaque system trustworthy\. An interpretable scorer, evaluated*without fitting to the labels*, is the defensible object of study\. Machine speaking scoring must therefore be accurate, interpretable, fair, and honestly evaluated against the right human bar\.
We study whether modern speech\-timing features and large language models \(LLMs\) can score spontaneous L2 dialogue fairly against a uniquely rich human gold standard\. That standard is the ICNALE Global Rating Archive \(GRA\), with140140speeches rated by∼\\sim8080trained raters on1010analytic criteria\. We make a positive and a negative claim and keep them strictly separate\. Our contributions are fourfold:
- •An interpretable, fairly\-evaluated scorer that beats the single\-human ceiling\.A De\-Jong speech\-timing composite, with signs fixed in advance and no fitting to human labels, blended with one LLM fluency judgment, reachesρ=0\.818\\rho=0\.818\. This agrees with the consensus better than81%81\\%of the8080individual trained raters \(above the median rater\) and reaches∼\\sim83%83\\%of the reliability\-corrected maximum \(κmax=0\.99\\kappa\_\{\\max\}=0\.99\)\. The blend improves on the composite alone by\+0\.054\+0\.054\(paired\-bootstrap CI\[0\.017,0\.108\]\[0\.017,0\.108\], excludes0\), mainly by resolving the LLM’s coarse ties\.
- •A controlled null on LLM pause encoding\.We vary only how pauses are written into the prompt\. Inline pause*locations*do not beat aggregate pause*statistics*\(−0\.069\-0\.069, CI\[−0\.15,\+0\.08\]\[\-0\.15,\+0\.08\]\), and a linguistically grounded mid\-clause×\\timesword\-frequency criterion gives no reliable gain\. The three pause encodings perform about the same, and all differences fall within the confidence intervals\.
- •A reliability\-corrected ceiling analysison the full140×80×10140\\times 80\\times 10GRA matrix, reporting single\-rater\-vs\-consensusρ=0\.621\\rho=0\.621, mean pairwise inter\-raterρ=0\.393\\rho=0\.393, and Cronbachα=0\.979\\alpha=0\.979, with principledκmax\\kappa\_\{\\max\}and human\-like target bars\.
- •Extensive verification, the paper’s signature: two independent learner\-isolation methods that agree, paired\-bootstrap CIs on every contrast, a monologue negative control, per\-feature reproduction of classical De\-Jong magnitudes, and a per\-L1 fairness audit\.
The remainder of the paper reviews related work \(§[2](https://arxiv.org/html/2608.26137#S2)\), describes the data \(§[3](https://arxiv.org/html/2608.26137#S3)\) and method \(§[4](https://arxiv.org/html/2608.26137#S4)\), presents results \(§[5](https://arxiv.org/html/2608.26137#S5)\), devotes a prominent section to verification and robustness \(§[6](https://arxiv.org/html/2608.26137#S6)\), discusses interpretation \(§[7](https://arxiv.org/html/2608.26137#S7)\) and limitations \(§[8](https://arxiv.org/html/2608.26137#S8)\), and concludes \(§[9](https://arxiv.org/html/2608.26137#S9)\) with an explicit honesty statement \(§[10](https://arxiv.org/html/2608.26137#S10)\)\.
## 2Related Work
### 2\.1Classical automated speaking and fluency assessment
Automated assessment of L2 speaking has a mature engineering lineage rooted in feature\-based scoring engines\. ETS’s SpeechRater\(Zechner et al\.,[2009](https://arxiv.org/html/2608.26137#bib.bib31)\)pioneered operational spoken\-response scoring\. It extracts interpretable features—fluency, pronunciation, prosody, vocabulary, and grammar—and combines them in a linear model validated against human raters\. The fluency part of such systems builds directly on the utterance\-fluency measures defined byde Jong et al\. \([2012](https://arxiv.org/html/2608.26137#bib.bib7)\)\. They broke speed, breakdown, and repair fluency into countable acoustic measures: speech rate, mean length of run, and silent\- and filled\-pause frequency and duration\. They showed that these measures explain much of the difference in human proficiency judgments\. The predictive strength of these features is robust and well quantified\. The meta\-analysis ofSuzuki et al\. \([2021](https://arxiv.org/html/2608.26137#bib.bib26)\)reports that, across studies,*speech rate*correlates with proficiency atr=\.76r=\.76,*mean length of run*atr=\.72r=\.72, and*pause frequency*negatively atr=−\.59r=\-\.59\. These are not incidental correlations\. They are the empirical backbone of every deployed delivery scorer, and, as we show, they remain difficult to beat\. Our deterministic composite is, deliberately, a faithful re\-implementation of this De\-Jong feature family\. Our per\-feature correlations land in the same direction and broad magnitude on our corpus\. Mean length of run \(ρ=\+\.75\\rho=\+\.75\) and pause ratio \(ρ=−\.72\\rho=\-\.72\) closely track the priors\. Our*speech rate including pauses*\(ρ=\+\.68\\rho=\+\.68\) is a different way of measuring from de Jong’s pause\-excluding speech rate, so we read it as same\-direction agreement rather than a point reproduction of the\+\.76\+\.76prior\.
### 2\.2Self\-supervised speech representations
A second line replaces hand\-engineered features with self\-supervised speech representations such as wav2vec 2\.0\(Baevski et al\.,[2020](https://arxiv.org/html/2608.26137#bib.bib1)\), HuBERT\(Hsu et al\.,[2021](https://arxiv.org/html/2608.26137#bib.bib14)\), and WavLM\(Chen et al\.,[2022](https://arxiv.org/html/2608.26137#bib.bib4)\)\.Bañno & Matassoni \([2022](https://arxiv.org/html/2608.26137#bib.bib2)\)apply wav2vec 2\.0 to L2 proficiency assessment*on ICNALE*, the corpus we also use\. They report that a frozen wav2vec 2\.0 representation reaches77\.9%77\.9\\%CEFR accuracy versus53\.5%53\.5\\%for a BERT text baseline\. This shows the acoustic signal carries proficiency information the transcript alone cannot give\.Liu et al\. \([2023](https://arxiv.org/html/2608.26137#bib.bib17)\)push further with an ASR\-free fluency scorer built on SSL features and clustering\. It reaches0\.7970\.797correlation versus0\.6900\.690for handcrafted features on their data\. These results motivate an acoustic representation, but two caveats bear on our setting\. The largest SSL win is precisely the temporal/fluency signal our De\-Jong features already capture\. SSL encoders also encode L1 and accent\. That creates a fairness risk across the ten Asian L1s in our data\. We therefore treat SSL as a fairness\-gated add\-on rather than the centerpiece\.
### 2\.3Speech\-LLM graders
The most recent wave couples speech encoders to large language models\. These graders score—and often*explain*—L2 proficiency directly from audio\.Ma et al\. \([2025](https://arxiv.org/html/2608.26137#bib.bib21)\)evaluate speech\-LLM graders on L2 proficiency at Interspeech 2025\. The Radboud group\(Parikh et al\.,[2026a](https://arxiv.org/html/2608.26137#bib.bib22),[b](https://arxiv.org/html/2608.26137#bib.bib23),[c](https://arxiv.org/html/2608.26137#bib.bib24)\)systematically builds rubric\-guided, multi\-rater, rationale\-emitting speech\-LLM assessors on speechocean762, with a fine\-tuned Qwen2\-Audio reaching sentence\-level fluency PCC up to0\.850\.85\. A repeated finding across this line is that zero\-shot audio\-LLMs judge delivery badly\. They systematically over\-score\. Qwen2\-Audio scores fluency essentially at chance \(PCC0\.0530\.053\) in zero\-shot use, with its prosody correlation only0\.1400\.140under direct matching\(Parikh et al\.,[2026c](https://arxiv.org/html/2608.26137#bib.bib24)\)\.Bañno et al\. \([2025](https://arxiv.org/html/2608.26137#bib.bib3)\)take an interpretable text\-only route\. They prompt a large LLM with the transcript and CEFR can\-do descriptors \(overall PCC≈0\.76\\approx 0\.76\) but*explicitly exclude*acoustic fluency cues as inaccessible through text\. The acoustic fluency signal these systems either mishandle \(zero\-shot audio\) or punt on \(descriptor\-based text\) is exactly where our work operates\.
### 2\.4Textualizing prosody and pauses as text for LLMs
One line of work renders acoustic structure as*text*inside the LLM prompt, and our method comparison builds on it directly\. SpeechCueLLM\(Wu et al\.,[2025](https://arxiv.org/html/2608.26137#bib.bib30)\)bins acoustic features by threshold and describes them in words for emotion recognition\. This is the cleanest template for “acoustics\-as\-text\.” Closest to us, TextPA \(“Read to Hear”;Chen et al\.,[2025](https://arxiv.org/html/2608.26137#bib.bib5)\) feeds a transcript augmented with IPA, CMU phones, and*inline pause durations*\(e\.g\."D \(0\.12s pause\) G"\) to zero\-shot LLMs for pronunciation and fluency scoring\. It reaches fluency PCC0\.6500\.650on MultiPA and0\.7840\.784when fused with a supervised system\. TextPA establishes the mechanism: textualize pause locations, then let the LLM judge fluency\. It owns that mechanism\. A widely shared idea is that telling an LLM*where*pauses fall should help it judge fluency, more than raw aggregate statistics do\. The L2 psycholinguistics below motivates this\. We test that idea in a controlled comparison of three ways of writing pauses into the prompt\.
### 2\.5LLM\-as\-judge and comparative judgement
Our evaluation design draws on the broader LLM\-as\-judge literature\. MT\-Bench and the associated judge studies\(Zheng et al\.,[2023](https://arxiv.org/html/2608.26137#bib.bib32)\)established both the promise and the biases of LLM judges\. Comparative\-judgement methods have repeatedly outperformed pointwise scoring\. LCES\(Shibata & Miyamura,[2025](https://arxiv.org/html/2608.26137#bib.bib18)\)reports that pairwise comparison markedly improves quadratic weighted kappa \(QWK\) over vanilla pointwise scoring \(e\.g\.0\.6330\.633vs\.0\.0210\.021on TOEFL11 with Llama\-3\.1\-8B\)\. The canonical comparative\-assessment study ofLiusie et al\. \([2024](https://arxiv.org/html/2608.26137#bib.bib19)\)makes the same point across NLG evaluation\. These results come from*essay*and text NLG scoring, not speech\. They inform our reading of a measured defect in pointwise LLM scoring \(coarse binning, frequent ties\), but they are not the primary contribution here\.
### 2\.6Evaluation methodology and the human ceiling
A scorer can only be judged against a target whose own reliability is known\. Classical test theory supplies the correction\.Uto \([2026](https://arxiv.org/html/2608.26137#bib.bib29)\)recently brought it to automated assessment\. He derives the achievable agreement ceiling: the maximum correlation a perfect model could reach given finite rater reliability\. He argues that models should not be penalized for irreducible human noise\. This reframes the success bar from “match the consensus exactly” to “reach the reliability\-corrected ceiling\.” It also reframes single\-rater reliability as the question that matters: can one trained judge do this? We adopt this method directly\. On our140×80140\\times 80rating matrix the 80\-rater Fluency mean is near\-noise\-free \(Cronbachα=0\.979\\alpha=0\.979\)\. This implies single\-rater reliabilityr1=0\.371r\_\{1\}=0\.371and a single\-rater\-vs\-consensus agreement of onlyρ=0\.62\\rho=0\.62, the human ceiling our scorer must clear, against a theoretical maximum of0\.979≈0\.99\\sqrt\{0\.979\}\\approx 0\.99\. The reliability\-corrected human\-like barr1α=0\.60\\sqrt\{r\_\{1\}\\alpha\}=0\.60closely matches the directly measured0\.620\.62\(within0\.020\.02\)\.
### 2\.7Psycholinguistic grounding of L2 pausing
The idea that pause*placement*matters is well grounded\.de Jong \([2016](https://arxiv.org/html/2608.26137#bib.bib9)\)shows that L2 speakers pause disproportionately*within*clauses rather than at clause boundaries\. Both L1 and L2 speakers pause more before*lower\-frequency*words\. This supplies both ingredients of a placement\-sensitive criterion on the production side: clause position and word frequency\.Kahng \([2018](https://arxiv.org/html/2608.26137#bib.bib16)\)uses phonetic manipulation to provide evidence that pause*location*drives*perceived*fluency, with mid\-clause pauses weighing most heavily\. The pause\-detection threshold we use \(250250ms\) followsde Jong & Bosker \([2013](https://arxiv.org/html/2608.26137#bib.bib8)\)\. This literature is sound, and our per\-feature analysis confirms its measures carry real signal\. We use it to build the grounded criterion among the three pause encodings we compare below\.
### 2\.8Positioning
Our closest neighbors partition cleanly\.TextPA owns the mechanism: textualizing inline pauses for an LLM to score fluency is theirs, and we cite it as prior art, not as something we claim\.Uto owns the ceiling method: reliability\-corrected achievable ceilings are his contribution, which we apply to the speaking/fluency case\. We contribute four things that remain unclaimed\. First, an*interpretable, fairly\-evaluated*delivery scorer: a De\-Jong composite blended with an LLM, with no fitting to the human labels, that reachesρ≈0\.82\\rho\\approx 0\.82\. It*beats*the single\-trained\-human ceiling of0\.620\.62and reaches≈83%\\approx 83\\%of the theoretical maximum\. Second, a controlled*null*on pause encoding: with ground\-truth learner isolation, inline pause locations do*not*beat plain pause statistics for an LLM \(ρ0\.71\\rho\\,0\.71vs\.0\.770\.77; difference−0\.069\-0\.069,95%95\\%CI\[−0\.150,\+0\.083\]\[\-0\.150,\+0\.083\]\), and a grounded mid\-clause×\\timesword\-frequency criterion gives no reliable gain\. Third, a reliability\-corrected ceiling analysis on a uniquely rich archive\. Fourth, and as the paper’s signature,*extensive verification*\. We use two independent learner\-isolation methods that agree \(feature rank\-correlation0\.770\.77–0\.890\.89\), paired\-bootstrap confidence intervals on every contrast, a monologue negative control \(broken speaker linkage correctly drives the correlation to∼0\\sim\\\!0\), per\-feature reproduction of classical De\-Jong magnitudes, and a per\-L1 fairness check\. Where TextPA asked whether an LLM*can*read textualized pauses for fluency, we ask whether the encoding changes the score\. Evaluated against a properly measured human ceiling, an interpretable feature composite does the job at least as well\.
## 3Data
#### Corpus\.
We use the ICNALE Global Rating Archive \(GRA\)\. It distributes140140rated L2 English speeches together with the full matrix of analytic ratings\. These rated speeches are the first9090seconds of the*part\-time\-job roleplays*drawn from the ICNALE Spoken Dialogue interviews\. Each is a two\-party interaction between a learner and an interviewer\. They are*not*the clean solo monologues that ICNALE also distributes\. The correct construct is spontaneous interactive dialogue\. This has direct consequences for the pipeline: the learner must be separated from the interviewer\. It also shapes construct validity: we score delivery in interaction, not in a read\-aloud or rehearsed\-monologue style\. The speakers span ten Asian L1 backgrounds—Chinese \(CHN\), Taiwanese \(TWN\), Korean \(KOR\), Japanese \(JPN\), Indonesian \(IDN\), Thai \(THA\), Hong Kong \(HKG\), Malaysian \(MYS\), Pakistani \(PAK\), and Filipino \(PHL\)\. They sit at CEFR levels A2 to B2\. Table[1](https://arxiv.org/html/2608.26137#S3.T1)summarizes the corpus\.
Table 1:Corpus composition of then=130n=130scored ICNALE GRA speeches\. Each gold label is the mean of∼80\{\\sim\}80trained raters on a0–1010scale\. Learner words and speaking time are measured after the learner is isolated from the interviewer\.
#### Scope of scoring\.
We scoren=130n=130speeches\. We exclude the four native\-English control speakers, because the target population is L2 learners, and six speeches whose source YouTube audio was unavailable\. A second, robustness pipeline based on speaker diarization is reported on then=125n=125subset for which diarization succeeded\.
#### Gold standard\.
For each criterion the gold label is the mean of about8080trained raters, each scoring on a0–1010scale\. The archive provides1010analytic criteria \(Fluency, Accuracy, Intelligibility, Involvement, Sophistication, Purposefulness, Logicality, and others\) plus a holistic score\. The ratings are ELF\-referenced \(English as a Lingua Franca\)\. The construct is therefore communicative effectiveness among non\-native conversation partners rather than native\-likeness\. This140×80×10140\\times 80\\times 10design is unusually rich\. It is what makes a reliability\-corrected ceiling analysis possible\.
#### Learner isolation\.
Each clip is a two\-party dialogue, so the learner’s speech must be separated from the interviewer’s before any feature is computed\. We use two methods that agree \(feature rank\-correlation0\.770\.77–0\.890\.89; §[6](https://arxiv.org/html/2608.26137#S6)\)\. The*primary*method uses the official ICNALE Spoken Dialogue “PTJ\_ROL” learner\-only transcripts\. In these, interviewer turns are stripped and participant identifiers match the GRA codes\. We align them to the timed ASR words \(n=130n=130\)\. The*robustness*method uses speaker diarization \(pyannote\-3\.0 plus wespeaker\-resnet34\) followed by a linguistic learner classifier that distinguishes “I/my” narration from “you/your”\-laden questions\. We need the classifier because the two speakers’ voice embeddings collapsed onto each other \(within\-clip cosine0\.800\.80\)\. Diarization alone is then unreliable\. This yieldsn=125n=125\.
#### License note\.
The GRA labels are used for*validation only*\. No model parameter, feature sign, or threshold is fit to the human ratings\. The labels are touched solely to compute correlations and confidence intervals at evaluation time\.
## 4Method
### 4\.1ASR and pipeline
Audio is transcribed with the Parakeet\-TDT\-0\.6B model \(sherpa\-onnx\)\. It provides word\-level timings and a per\-word minimum sub\-token confidence\. After learner isolation \(§[3](https://arxiv.org/html/2608.26137#S3)\), all timing features are computed on the learner’s words only\. Pauses are detected at a250250ms threshold followingde Jong & Bosker \([2013](https://arxiv.org/html/2608.26137#bib.bib8)\)\.
### 4\.2Interpretable De\-Jong feature composite
The deterministic scorer is a faithful re\-implementation of the De\-Jong utterance\-fluency family\(de Jong et al\.,[2012](https://arxiv.org/html/2608.26137#bib.bib7)\)\. It combines five features whose directional signs are fixed in advance from the SLA literature, never fit to the labels:
1. 1\.Pause ratio\(silent time / total time\), sign−\-\(more pausing⇒\\Rightarrowlower fluency\);
2. 2\.Mean length of run\(words per inter\-pause run\), sign\+\+;
3. 3\.Speech rateincluding pauses \(words per total second\), sign\+\+;
4. 4\.Long\-pause rate\(rate of pauses exceeding a long threshold\), sign−\-;
5. 5\.ASR word\-confidenceas an intelligibility proxy, sign\+\+\.
Each feature is converted to azz\-score using the in\-sample mean and standard deviation, multiplied by its fixed sign, and averaged\. Thiszz\-scoring uses only the feature distribution, not the human labels, so the composite remains fairly evaluated\. We deliberately exclude pure articulation rate \(rate excluding pause time\)\. The literature predicts it to be a weak proficiency cue, and we confirm it is null on our data\.
### 4\.3Text\-LLM fluency call
The LLM arm passes the learner’s transcript to DeepSeek\-chat at temperature0with JSON\-structured output, asking for a single pointwise fluency score\. The transcript is augmented with pause information according to the pause encoding below\. The LLM is used zero\-shot with a fixed rubric prompt\. No examples are drawn from the labeled set\.
### 4\.4The blend
The final scorer is a simple combination of the deterministic composite and the LLM fluency score\. The two signals are constructed independently: one is continuous and feature\-derived, the other a coarse discrete LLM judgment\. We evaluate whether their combination adds information via the paired bootstrap \(§[5](https://arxiv.org/html/2608.26137#S5)\)\.
### 4\.5Pause\-encoding settings \(A/B/C\)
To compare ways of writing pauses into the prompt, we run a controlled comparison\. We hold the LLM, the prompt scaffold, and the learner words fixed, and change only the textual encoding of pauses:
- •Statistics \(Arm A\)— transcript plus*aggregate pause statistics*\(overall pause ratio, counts, mean durations\) appended as text;
- •Inline locations \(Arm B\)— transcript with*inline pause locations*, each silence rendered as a token such as\(0\.8s\)at its position in the word stream \(the TextPA\-style mechanism\);
- •Grounded criterion \(Arm C\)— the inline locations of Arm B plus a*grounded*criterion that asks the LLM to weight mid\-clause pauses before low\-frequency words, applying the psycholinguistic placement findings ofde Jong \([2016](https://arxiv.org/html/2608.26137#bib.bib9)\)andKahng \([2018](https://arxiv.org/html/2608.26137#bib.bib16)\)\.
The contrasts B−\-A and C−\-B are the controlled tests of the inline locations and the grounded criterion, respectively\.
### 4\.6Reliability\-corrected ceilings
FollowingUto \([2026](https://arxiv.org/html/2608.26137#bib.bib29)\), we report two principled bars\. The*theoretical maximum*correlation a perfect model could attain against a gold with reliabilityα\\alphaisκmax=α\\kappa\_\{\\max\}=\\sqrt\{\\alpha\}\. The*human\-like*bar is the agreement a single trained rater achieves against the consensus, after correcting for rater noise\. It isr1α\\sqrt\{r\_\{1\}\\,\\alpha\}, wherer1r\_\{1\}is the implied single\-rater reliability\. For Fluency,α=0\.979\\alpha=0\.979givesκmax=0\.979≈0\.99\\kappa\_\{\\max\}=\\sqrt\{0\.979\}\\approx 0\.99and a human\-like bar of0\.371×0\.979=0\.603≈0\.60\\sqrt\{0\.371\\times 0\.979\}=0\.603\\approx 0\.60\. This closely matches the directly measured single\-rater\-vs\-consensusρ=0\.621\\rho=0\.621\(within0\.020\.02\)\. We emphasize that0\.600\.60–0\.620\.62, not the reliability\-corrected0\.830\.83, is the human bar\. The0\.830\.83figure is an upper\-bound*interpretation*of our scorer’s performance \(its reliability\-corrected value\)\. Reading it as a human target would mistake an upper bound for a human score\.
## 5Results
### 5\.1The human ceiling
Table[2](https://arxiv.org/html/2608.26137#S5.T2)reports the ceiling analysis from the full140×80140\\times 80rating matrix\. Across all ten analytic criteria the 80\-rater mean is near\-noise\-free \(Cronbachα\\alphabetween0\.9760\.976and0\.9800\.980\)\. Yet an individual trained rater agrees with the consensus at onlyρ=0\.569\\rho=0\.569–0\.6210\.621\. The mean pairwise inter\-rater agreement is lower still \(ρ=0\.393\\rho=0\.393, Pearson0\.4030\.403\)\. The implied single\-rater reliability for Fluency isr1=0\.371r\_\{1\}=0\.371\. The theoretical maximum isκmax=0\.99\\kappa\_\{\\max\}=0\.99, and the human\-like bar is0\.600\.60\. Together they frame the success criterion\. A scorer that reaches the high\-0\.70\.7s or low\-0\.80\.8s has matched or beaten what one trained human does\. It still leaves headroom to the reliability\-corrected ceiling\. We use the*mean*single\-rater agreement \(0\.6210\.621\) as the primary human bar\. We use the mean because it matches the reliability\-correction step: by construction it lines up withr1α=0\.60\\sqrt\{r\_\{1\}\\alpha\}=0\.60, computed from the rater\-mean reliability\. The classical\-test\-theory correction works on the mean, not the median\. The mean is also the harder choice in only one direction\. The*median*trained rater is stronger, atρ=0\.733\\rho=0\.733, so we report both and check our conclusion against the harder median bar below\.
Table 2:Human ceiling from the140×80140\\times 80GRA matrix\. The single\-raterρ\\rhois the agreement of one trained rater with the 80\-rater consensus;κmax=α\\kappa\_\{\\max\}=\\sqrt\{\\alpha\}is the theoretical maximum; the human\-like bar isr1α\\sqrt\{r\_\{1\}\\alpha\}\. Fluency values are exact; the criterion range spans all ten analytic criteria\.
### 5\.2Per\-feature correlations
Table[3](https://arxiv.org/html/2608.26137#S5.T3)reports the Spearman correlation of each De\-Jong feature with the Fluency gold onn=130n=130under the official isolation\. The magnitudes line up, in direction and broad size, with the classical measurements compiled bySuzuki et al\. \([2021](https://arxiv.org/html/2608.26137#bib.bib26)\)\. Mean length of run \(\+0\.749\+0\.749vs\. prior\+0\.72\+0\.72\) and pause ratio \(−0\.715\-0\.715vs\. prior pause frequency−0\.59\-0\.59\) closely track the priors\. Our speech rate \(\+0\.676\+0\.676\) is in the same direction as the prior\+0\.76\+0\.76\. It is a different way of measuring, rate*including*pauses rather than de Jong’s pause\-excluding speech rate, so we do not read the0\.080\.08gap as a point reproduction\. Long\-pause rate \(−0\.580\-0\.580\) and ASR word\-confidence \(\+0\.595\+0\.595\) also carry strong signal\. Pure articulation rate is null \(\+0\.08\+0\.08\), which confirms the advance decision to exclude it\.111All six correlations in Table[3](https://arxiv.org/html/2608.26137#S5.T3), including the articulation\-rate null \(\+0\.08\+0\.08\), are from the singlen=130n=130official\-isolation run; no per\-feature value is taken from then=125n=125diarization run\.The strongly predictive measures reproduce in direction and broad magnitude on a new corpus, under ground\-truth learner isolation\. That reproduction is itself a verification result\.
Table 3:Per\-feature Spearman correlation with the Fluency gold \(n=130n=130, official isolation\), with fixed sign and the corresponding prior magnitude fromSuzuki et al\. \([2021](https://arxiv.org/html/2608.26137#bib.bib26)\)where available\.
### 5\.3Representation comparison
Table[4](https://arxiv.org/html/2608.26137#S5.T4)reports the controlled representation comparison of §[4\.5](https://arxiv.org/html/2608.26137#S4.SS5)\. We hold the LLM and the learner words fixed and change only how pauses are written into the prompt\. We compare three pause encodings: aggregate statistics, inline locations, and a grounded mid\-clause criterion\. Arm A \(statistics\) reachesρ=0\.774\\rho=0\.774\. Arm B \(inline locations\) reaches0\.7050\.705\. Arm C \(grounded criterion\) reaches0\.6870\.687\. The paired bootstrap \(10k resamples\) gives B−\-A=−0\.069=\-0\.069,95%95\\%CI\[−0\.15,\+0\.08\]\[\-0\.15,\+0\.08\], and C−\-B=−0\.018=\-0\.018, CI\[−0\.10,\+0\.08\]\[\-0\.10,\+0\.08\]\. Both intervals include zero\. Atn=130n=130this comparison can detect only encoding effects larger than about±0\.1\\pm 0\.1–0\.15ρ0\.15\\,\\rho, so we claim no effect in that range, not that no effect could exist at any scale\. The three encodings perform about the same, and none beats simple statistics\. Inline locations give no gain over aggregate statistics\. The grounded criterion gives no reliable improvement over locations; its confidence interval includes zero\.222Under the official isolation every encoding contrast has a95%95\\%CI including zero \(Table[4](https://arxiv.org/html/2608.26137#S5.T4)\); the diarization point estimates are small and likewise show no encoding beating statistics\.The takeaway is plain: the fluency signal comes from the measured speech\-timing features, not from how pauses are written for the LLM\. The deterministic composite \(ρ=0\.764\\rho=0\.764\) comes within0\.010\.01of the best LLM arm \(Arm A,0\.7740\.774\), with no language model\.
Table 4:Pause\-representation comparison \(n=130n=130, official isolation\), same LLM and learner words, varying only the textual encoding of pauses\. Contrasts are paired\-bootstrap differences \(1010k resamples\) with95%95\\%CIs\. The deterministic composite uses no LLM\. The diarization column \(n=125n=125\) is the robustness replication\.
### 5\.4Composite, blend, and beating the ceiling
Blending the deterministic composite \(ρ=0\.764\\rho=0\.764\) with the LLM fluency score yieldsρ=0\.818\\rho=0\.818, an improvement of\+0\.054\+0\.054over the composite alone \(paired\-bootstrap95%95\\%CI\[0\.017,0\.108\]\[0\.017,0\.108\], excludes zero\)\. The mechanism is specific, and we state it plainly\. The pointwise LLM score is coarse: it takes about66distinct values, with27%27\\%of pairs tied\. The continuous composite resolves these ties in a gold\-aligned order, and this tie resolution does most of the lifting\. Blending the composite with the weakest arm \(C,ρ=0\.687\\rho=0\.687alone\) still reaches0\.8150\.815, almost the same as with the strongest arm \(A,0\.7740\.774alone,→0\.818\\to 0\.818\)\. The blend is carried mainly by the composite, with the LLM supplying a coarse complementary ranking\. We therefore read the\+0\.054\+0\.054gain as the LLM adding a coarse ranking the composite refines, not as the LLM supplying a large amount of new information\.
How good is0\.8180\.818against a human? On the same130130speeches we measure each rater’s agreement with the consensus of the other raters\. The blend agrees with the consensus better than81%81\\%of the8080trained raters: above the median rater \(ρ=0\.73\\rho=0\.73\) and the7575th\-percentile rater \(ρ=0\.80\\rho=0\.80\), though below the best \(ρ=0\.88\\rho=0\.88\)\. Corrected for rater noise it reaches≈0\.83\\approx 0\.83, about83%83\\%of theκmax=0\.99\\kappa\_\{\\max\}=0\.99maximum, leaving roughly0\.170\.17of headroom\. This is the paper’s defensible “better”: an interpretable scorer, never fit to the labels, agrees with the consensus as well as a strong individual rater\. We report then=130n=130official\-isolation figures \(composite0\.7640\.764, blend0\.8180\.818\) as primary; then=125n=125diarization run gives a composite of≈0\.72\\approx 0\.72and a blend of≈0\.78\\approx 0\.78, with the same method ranking\. Figure[1](https://arxiv.org/html/2608.26137#S5.F1)places every system against the single\-human band and the maximum, and Figure[2](https://arxiv.org/html/2608.26137#S5.F2)shows the blend’s agreement with the gold\.
Figure 1:Spearmanρ\\rhowith the human Fluency gold \(n=130n=130\); error bars are paired\-bootstrap95%95\\%CIs and the number above each bar is the point estimate\. The shaded band is the single\-human\-rater range \(mean0\.620\.62to median0\.730\.73\); the dash\-dot line is the reliability\-corrected maximum \(κmax=0\.99\\kappa\_\{\\max\}=0\.99\)\. The three LLM pause encodings \(A statistics, B inline locations, C grounded criterion\) perform about the same, and none beats the deterministic composite, which uses no language model\. The blend of the composite and the LLM is best: it sits above the single\-human band and below the maximum\.Figure 2:Blend score against the human Fluency gold \(n=130n=130,ρ=0\.818\\rho=0\.818\)\. The visible horizontal banding reflects the LLM’s coarse, near\-tied pointwise scores, which the continuous composite resolves\.
### 5\.5Construct map across ten criteria
The composite is built only from fluency\-relevant features\. It still correlates with all ten GRA criteria in the range0\.680\.68–0\.770\.77\(Table[5](https://arxiv.org/html/2608.26137#S5.T5)\)\. This is expected, because the GRA criteria move together\. The composite correlates most with the delivery\-proximal criteria: Fluency \(0\.7650\.765\), Accuracy \(0\.7440\.744\), Involvement \(0\.7390\.739\), and Intelligibility \(0\.7360\.736\)\. It correlates least with the content\-oriented criteria: Sophistication \(0\.6820\.682\), Purposefulness \(0\.6840\.684\), and Logicality \(0\.6860\.686\)\. The ordering is broadly the one a delivery\-based measure should show\. The gradient is not perfectly clean\. Complexity \(0\.7130\.713\) and Comprehensibility \(0\.7120\.712\) sit mid\-pack, and Complexity, a range/content criterion, ranks above the three lowest content criteria\. This overlap keeps the full spread narrow \(0\.6820\.682–0\.7650\.765\)\. We read the raw gradient as suggestive of delivery\-proximity rather than a sharp construct separation\.
This overlap can be removed\. When we control for the holistic score \(partial Spearman\), the composite stays positive and strong on Fluency \(\+0\.44\+0\.44\) but turns*negative*on the content criteria \(Sophistication−0\.30\-0\.30, Purposefulness−0\.21\-0\.21, Logicality−0\.17\-0\.17; Table[5](https://arxiv.org/html/2608.26137#S5.T5)\)\. The composite is therefore delivery\-specific, not a general proficiency proxy: once shared proficiency is removed, it tracks fluency and the delivery\-adjacent criteria and actively diverges from content\. This is the strongest construct\-validity evidence we have\.
Table 5:Construct map: correlation of the fluency\-built deterministic composite with each GRA criterion \(n=130n=130\)\.*Raw*Spearman correlations are inflated by overlap across criteria\. The*partial*column controls for the holistic score and exposes a clear delivery\-vs\-content gradient: positive on delivery, negative on content\.
### 5\.6Per\-L1 descriptive fairness
Figure[3](https://arxiv.org/html/2608.26137#S5.F3)reports the composite\-versus\-Fluency correlation within each L1 group\. The cells are small \(n=4n=4–2020\), so these figures are descriptive only and carry wide uncertainty\. The point estimates range from0\.400\.40\(TWN\) to0\.940\.94\(THA\)\. But the per\-group bootstrap confidence intervals are wide and overlap heavily\. TWN, the apparent low, has a95%95\\%CI of\[−0\.16,\+0\.85\]\[\-0\.16,\+0\.85\]\(Figure[3](https://arxiv.org/html/2608.26137#S5.F3)\) and overlaps every other group\. The spread is therefore consistent with sampling noise at these cell sizes, and we cannot conclude differential validity in either direction\. The remaining four L1 cells \(HKG, MYS, PAK, PHL\) have onlyn=4n=4each and returned no stable estimate, so the plotted groups cover six of the ten L1s\. We make no per\-L1 fairness guarantee in either direction, and report this audit as a documented limitation\.
Figure 3:Per\-L1 descriptive fairness: composite\-vs\-Fluencyρ\\rhowithin each L1 group, for the six groups withn≥10n\\geq 10\. Cells are small \(n=18n=18–2020\), so these are descriptive only\. The four smallest L1 cells \(HKG, MYS, PAK, PHL;n=4n=4each\) returned no stable estimate and are omitted\. The dashed line marks the mean single\-rater ceiling \(0\.620\.62\)\. The spread \(0\.400\.40for TWN to0\.940\.94for THA\) is a documented limitation\.
## 6Verification and Robustness
Verification is the signature of this paper\. We make a positive claim \(beating the human ceiling\) and a method finding \(how pauses are encoded does not change the fluency score\)\. Each must survive scrutiny that a single correlation cannot provide\. We report five converging checks\.
#### \(1\) Two independent learner\-isolation methods that agree\.
The central risk in a dialogue corpus is contaminating the learner’s timing features with the interviewer’s speech\. We isolate the learner two ways\. The primary method aligns the official ICNALE learner\-only “PTJ\_ROL” transcripts, with participant ids matching the GRA codes, to the timed ASR words \(n=130n=130\)\. The robustness method combines pyannote\-3\.0 \+ wespeaker\-resnet34 diarization with a linguistic learner classifier \(n=125n=125\)\. Within\-clip speaker embeddings collapsed onto each other \(cosine0\.800\.80\), so diarization alone would be unreliable\. That is why we rely on the transcripts\. The two methods still agree, with feature rank\-correlations of0\.770\.77–0\.890\.89across the De\-Jong features\. The encoding comparison also holds across both methods: the three encodings perform about the same under each, and their small differences are within the confidence intervals \(Table[4](https://arxiv.org/html/2608.26137#S5.T4)\)\.
#### \(2\) Paired\-bootstrap CIs and a small\-nnprotocol\.
Every contrast in the paper has a paired\-bootstrap95%95\\%CI from1010k resamples over then=130n=130speeches\. The bootstrap respects the paired structure, the same speeches scored two ways, and makes no normality assumption\. This is what licenses the equivalence reading: B−\-A=−0\.069=\-0\.069, CI\[−0\.15,\+0\.08\]\[\-0\.15,\+0\.08\]and C−\-B=−0\.018=\-0\.018, CI\[−0\.10,\+0\.08\]\[\-0\.10,\+0\.08\]both include zero, while the blend gain\+0\.054\+0\.054, CI\[0\.017,0\.108\]\[0\.017,0\.108\]excludes it\. The CIs exclude only effects larger than roughly±0\.10\\pm 0\.10–0\.150\.15ρ\\rho\. We therefore claim no*practically useful*difference between encodings, not exact equivalence\.
#### \(3\) A monologue negative control\.
ICNALE also distributes clean solo*monologue*audio\. Its speaker ids are*disjoint*from the dialogue/GRA participants, confirmed four ways, including via the participant survey: matching on \(L1, Sex, Age, Vocabulary\-Size\-Test\) leaves266266of425425dialogue speakers with no monologue counterpart\. We compute features on the unmatched monologues and correlate them with the GRA fluency gold\. The correlation collapses to∼\\sim0\. This is a clean negative control\. When speaker linkage is deliberately broken, the signal correctly vanishes\. So the strong dialogue correlation reflects real, speaker\-matched information, not an artifact of the corpus\.
#### \(4\) Per\-feature reproduction of prior measurements\.
The individual De\-Jong feature correlations \(Table[3](https://arxiv.org/html/2608.26137#S5.T3)\) align in direction and broad magnitude with the meta\-analysis ofSuzuki et al\. \([2021](https://arxiv.org/html/2608.26137#bib.bib26)\)\. Mean length of run\+0\.749\+0\.749\(prior\+0\.72\+0\.72\) and pause ratio−0\.715\-0\.715\(prior pause frequency−0\.59\-0\.59\) track the priors closely\. Our speech rate\+0\.676\+0\.676matches the*direction*of the prior\+0\.76\+0\.76under a different way of measuring \(rate including pauses\)\. The predicted\-null articulation rate is indeed null \(\+0\.08\+0\.08\)\. Independent reproduction of the strongly predictive established measures on a new corpus and pipeline shows that the isolation and feature extraction are sound\.
#### \(5\) Reliability\-corrected ceiling and honest treatment of residual issues\.
We benchmark against a properly measured human bar \(§[5\.1](https://arxiv.org/html/2608.26137#S5.SS1)\) rather than an idealized noiseless gold\. We are explicit that the human bar is0\.600\.60–0\.620\.62, not the reliability\-corrected0\.830\.83\. We surface two soft spots rather than hide them\. The first is the per\-L1 spread \(ρ=0\.40\\rho=0\.40–0\.940\.94onn=4n=4–2020cells; §[5\.6](https://arxiv.org/html/2608.26137#S5.SS6)\), reported as descriptive only\. The second is the coarse\-LLM/tie phenomenon \(∼\\sim66distinct values, with2727–36%36\\%of pairs tied across arms\)\. We treat it both as a measured defect of pointwise LLM scoring and as the mechanism by which the continuous composite resolves the LLM’s ties in the blend\.
## 7Discussion
#### What “beating the single\-human ceiling” means—and does not mean\.
Our blend reachesρ=0\.818\\rho=0\.818\. This beats the agreement that an individual trained rater has with the8080\-rater consensus \(mean single\-raterρ=0\.62\\rho=0\.62; median0\.7330\.733\); it agrees with the consensus better than81%81\\%of the8080individual raters\. The defensible reading is narrow\. Evaluated fairly and without ever fitting to the labels, the scorer is at least as trustworthy as a single trained human\. It clears both the mean and the median single\-rater bar, and reaches∼\\sim83%83\\%of the reliability\-corrected maximum\. It does*not*mean the scorer matches an8080\-rater panel\. The consensus mean is far more reliable \(α=0\.979\\alpha=0\.979\) and is the gold, not the bar\. The reliability\-corrected0\.830\.83does not describe an achieved agreement\. It is an upper\-bound interpretation of the observed0\.8180\.818\. The honest claim is narrow and strong: an interpretable, auditable scorer can do what one trained human does\.
#### Where the fluency signal lives\.
The signal is in the features, not in how pauses are written for the LLM\. The deterministic composite \(0\.7640\.764\) is built directly from the De\-Jong measures and comes within0\.010\.01of the best LLM arm\. We tried three ways of writing pauses into the prompt: aggregate statistics, inline locations, and a grounded mid\-clause criterion\. They perform about the same, and none beats simple statistics\. Inline locations give no gain over statistics, and the grounded criterion gives no reliable improvement over locations\. This result is a clean fact about the method\. The measured speech\-timing features carry the fluency information, and the textual rendering of pauses adds nothing useful\. There is a plausible reason\. Counting silence tokens and weighting them by word frequency is exact arithmetic\. A deterministic feature does this exactly; a language model only approximates it\. The constructive lesson is to route grounding into deterministic features\. The LLM is then best used for the holistic judgment it does well, and the two are combined in the blend rather than the LLM asked to do the timing arithmetic\.
#### Construct validity\.
The criterion map \(Table[5](https://arxiv.org/html/2608.26137#S5.T5)\) shows the composite correlating most with delivery\-proximal criteria \(Fluency, Accuracy, Intelligibility, Involvement\) and least with the most content\-oriented criteria \(Sophistication, Purposefulness, Logicality\)\. Complexity and Comprehensibility sit in between\. This cross\-criterion overlap keeps the raw spread within0\.0830\.083, so the raw ordering is only mild support\. The partial correlations are decisive: controlling for the holistic score, the composite stays strongly positive on Fluency \(\+0\.44\+0\.44\) and turns negative on the content criteria \(Sophistication−0\.30\-0\.30; Table[5](https://arxiv.org/html/2608.26137#S5.T5)\)\. Once shared proficiency is removed, the composite tracks delivery and diverges from content\. It is a delivery\-specific measure, not a general proficiency proxy\.
#### Deployment implications\.
The resulting system is cheap, interpretable, and auditable\. The deterministic composite requires only ASR timings and five transparent features with fixed signs\. The LLM arm is a single zero\-shot text call\. Every component can be inspected, and each feature’s contribution can be explained to a learner or an examiner\. Opaque high\-correlation systems lack these properties\. High\-stakes speaking assessment now needs them\.
## 8Limitations
Several limitations bound our claims\.*Single corpus and L1 skew:*all data are from ICNALE and the ten Asian L1 groups; generalization to other L1s, speech styles, or proficiency bands is unverified\.*Dialogue construct:*we score9090\-second part\-time\-job roleplays, an interactive style that differs from monologue or read\-aloud tasks; the construct is interaction\-embedded delivery\.*ASR error on accented A2 speech:*word timings and confidences degrade on the lowest\-proficiency, most\-accented speech, which may weaken features unevenly across L1s\.*Per\-L1 variation:*predictive accuracy varies fromρ=0\.40\\rho=0\.40to0\.940\.94across L1 cells ofn=4n=4–2020; these are too small to support any per\-L1 fairness guarantee\.*Single text\-LLM:*the LLM results use one model \(DeepSeek\-chat\); a different model might shift the absolute arm values, but the controlled*contrasts*are what carry the encoding finding, and these hold across isolation methods\.*Sample size:*all correlations are onn=130n=130\(official\) orn=125n=125\(diarization\), so the encoding CIs exclude only effects larger than∼\\sim±0\.10\\pm 0\.10–0\.150\.15ρ\\rho\.*Embedding collapse:*within\-clip speaker embeddings collapsed onto each other, which is why diarization is robustness rather than primary\.
## 9Conclusion
L2 English learners face a structural shortage of speaking practice and trustworthy feedback\. This is pushing automated speaking assessment into high\-stakes use\. We have shown that an interpretable feature\-plus\-LLM hybrid, evaluated fairly with no fitting to the human labels, can score spontaneous L2 dialogue atρ=0\.818\\rho=0\.818\. This agrees with the consensus better than81%81\\%of individual trained raters \(above the median,ρ=0\.73\\rho=0\.73\) and reaches∼\\sim83%83\\%of the reliability\-corrected maximum on a uniquely rich140×80×10140\\times 80\\times 10archive\. We have also shown, in a controlled comparison, that how pauses are encoded for the LLM does not change the fluency score\. Aggregate statistics, inline locations, and a grounded placement criterion all perform about the same, and none beats simple statistics\. The signal lies in the measured speech\-timing features\. The LLM is best used for holistic judgment and combined with, not asked to replace, deterministic timing analysis\. We back both claims with two agreeing isolation methods, paired\-bootstrap CIs, a monologue negative control, per\-feature reproduction of classical measurements, and a per\-L1 audit\. The resulting scorer is cheap, interpretable, and auditable\. These properties matter precisely because the assessment it performs is becoming high\-stakes\.
## 10Ethics and Honesty Statement
We state plainly what we do and do not claim\. We donotclaim that writing pauses into the prompt helps LLM\-based fluency scoring\. We find that the three pause encodings perform about the same, and the inline\-location*mechanism*itself is prior art \(TextPA;Chen et al\.,[2025](https://arxiv.org/html/2608.26137#bib.bib5)\), as are our ceiling methodology\(Uto,[2026](https://arxiv.org/html/2608.26137#bib.bib29)\)and rationale\-faithfulness framing\(Parikh et al\.,[2026a](https://arxiv.org/html/2608.26137#bib.bib22)\)\. The encoding finding is bounded by power\. The confidence intervals exclude only effects larger than∼\\sim±0\.10\\pm 0\.10–0\.150\.15ρ\\rho, so we claim*no practically useful difference*, not exact equivalence\. Our “beats the human ceiling” claim is against a*single\-rater*bar\. We use the mean single\-rater agreement \(ρ≈0\.62\\rho\\approx 0\.62\) as the primary bar\. It matches the reliability\-correction step, lining up withr1α\\sqrt\{r\_\{1\}\\alpha\}\. The blend also clears the harder median trained\-rater bar \(0\.7330\.733\), so the claim does not depend on choosing the weaker bar\. The8080\-rater consensus mean is far more reliable \(α=0\.979\\alpha=0\.979\) and is the gold, not the bar\. The reliability\-corrected0\.830\.83is an upper\-bound interpretation, not a measured agreement\. Because in\-clip speaker embeddings collapsed onto each other \(cosine0\.800\.80\), our primary learner isolation relies on the official transcripts, with diarization as robustness only\. The data are9090\-second part\-time\-job roleplays from ten Asian L1s at CEFR A2–B2; generalization beyond this speech style, proficiency band, and L1 set is unverified\. Per\-L1 accuracy varies \(ρ=0\.40\\rho=0\.40–0\.940\.94\) on cells ofn=4n=4–2020and is descriptive only; we make no per\-L1 fairness guarantee\. The motivation citations concerning AI practice \(e\.g\.Ding & Yusof,[2025](https://arxiv.org/html/2608.26137#bib.bib10)\) speak to*demand*for trustworthy scoring, not to the validity of our method, and market and community figures are signals of need, never evidence of scoring validity\. The GRA labels are used for validation only; no model parameter is fit to them\. All reported numbers are computed onn=130n=130\(official isolation\) orn=125n=125\(diarization\)\.
#### Reproducibility\.
We do not release the scoring pipeline as source code\. The validation labels come from the ICNALE Global Rating Archive, distributed by Kobe University under a research\- and education\-only license that does not permit redistribution, so neither the corpus nor any data derived from it is included here; a licensed copy can be obtained directly from the corpus maintainers\. The running system is instead available for inspection and independent verification at[https://www\.aflo\.one](https://www.aflo.one/)\.
## References
- Baevski et al\. \(2020\)Baevski, A\., Zhou, H\., Mohamed, A\., & Auli, M\. \(2020\)\. wav2vec 2\.0: A framework for self\-supervised learning of speech representations\.*Advances in Neural Information Processing Systems \(NeurIPS\)*\.
- Bañno & Matassoni \(2022\)Bañno, S\., & Matassoni, M\. \(2022\)\. Proficiency assessment of L2 spoken English using wav2vec 2\.0\.*Proceedings of IEEE Spoken Language Technology Workshop \(SLT\)*\. arXiv:2210\.13168\.
- Bañno et al\. \(2025\)Bañno, S\., Ma, R\., Qian, M\., et al\. \(2025\)\. Natural Language\-based Assessment of L2 Oral Proficiency using LLMs\.*Proc\. SLaTE 2025*\. arXiv:2507\.10200\.
- Chen et al\. \(2022\)Chen, S\., Wang, C\., Chen, Z\., et al\. \(2022\)\. WavLM: Large\-scale self\-supervised pre\-training for full stack speech processing\.*IEEE Journal of Selected Topics in Signal Processing*, 16\(6\), 1505–1518\.
- Chen et al\. \(2025\)Chen, Y\., Ma, M\., & Hirschberg, J\. \(2025\)\. Read to Hear: A Zero\-Shot Pronunciation Assessment Using Textual Descriptions and LLMs\.*Proceedings of EMNLP 2025*\. arXiv:2509\.14187\.
- Chen & Sun \(2025\)Chen, T\., & Sun, S\. \(2025\)\. Evaluating automated evaluation systems for spoken English proficiency\.*PLOS ONE*, 20\. DOI:10\.1371/journal\.pone\.0320811\.
- de Jong et al\. \(2012\)de Jong, N\. H\., Steinel, M\. P\., Florijn, A\. F\., Schoonen, R\., & Hulstijn, J\. H\. \(2012\)\. Facets of speaking proficiency\.*Studies in Second Language Acquisition*, 34\(1\), 5–34\.
- de Jong & Bosker \(2013\)de Jong, N\. H\., & Bosker, H\. R\. \(2013\)\. Choosing a threshold for silent pauses to measure second language fluency\.*Proceedings of Disfluency in Spontaneous Speech \(DiSS\)*, 17–20\.
- de Jong \(2016\)de Jong, N\. H\. \(2016\)\. Predicting pauses in L1 and L2 speech: The effects of utterance boundaries and word frequency\.*International Review of Applied Linguistics in Language Teaching \(IRAL\)*, 54\(2\), 113–132\.
- Ding & Yusof \(2025\)Ding, D\., & Yusof, A\. M\. B\. \(2025\)\. Investigating the role of AI\-powered conversation bots in enhancing L2 speaking skills and reducing speaking anxiety: A mixed methods study\.*Humanities and Social Sciences Communications*, 12, 1223\. DOI:10\.1057/s41599\-025\-05550\-z\.
- ERIC EJ1319829 \(2021\)Challenges faced by bachelor\-level students while speaking English\. \(2021\)\.*ERIC*EJ1319829\.
- ERIC EJ1305502 \(2019\)Speaking anxiety overwhelms English learners\. \(2019\)\.*ERIC*EJ1305502\.
- Horwitz et al\. \(1986\)Horwitz, E\. K\., Horwitz, M\. B\., & Cope, J\. \(1986\)\. Foreign language classroom anxiety\.*The Modern Language Journal*, 70\(2\), 125–132\. DOI:10\.1111/j\.1540\-4781\.1986\.tb05256\.x\.
- Hsu et al\. \(2021\)Hsu, W\.\-N\., Bolte, B\., Tsai, Y\.\-H\. H\., Lakhotia, K\., Salakhutdinov, R\., & Mohamed, A\. \(2021\)\. HuBERT: Self\-supervised speech representation learning by masked prediction of hidden units\.*IEEE/ACM Transactions on Audio, Speech, and Language Processing*, 29, 3451–3460\.
- Isaacs et al\. \(2023\)Isaacs, T\., Hu, R\., Trenkic, D\., & Varga, J\. \(2023\)\. Examining the predictive validity of the Duolingo English Test: Evidence from a major UK university\.*Language Testing*, 40\(3\)\. DOI:10\.1177/02655322231158550\.
- Kahng \(2018\)Kahng, J\. \(2018\)\. The effect of pause location on perceived fluency\.*Applied Psycholinguistics*, 39\(3\), 569–591\.
- Liu et al\. \(2023\)Liu, W\., et al\. \(2023\)\. An ASR\-free fluency scoring approach with self\-supervised learning\.*Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\)*\. arXiv:2302\.09928\.
- Shibata & Miyamura \(2025\)Shibata, T\., & Miyamura, Y\. \(2025\)\. LCES: Zero\-shot automated essay scoring via pairwise comparisons using LLMs\.*Proceedings of EMNLP 2025*\. arXiv:2505\.08498\.
- Liusie et al\. \(2024\)Liusie, A\., Manakul, P\., & Gales, M\. J\. F\. \(2024\)\. LLM comparative assessment: Zero\-shot NLG evaluation through pairwise comparisons using large language models\.*Proceedings of EACL 2024*\. arXiv:2307\.07889\.
- Long \(1996\)Long, M\. H\. \(1996\)\. The role of the linguistic environment in second language acquisition\. In W\. C\. Ritchie & T\. K\. Bhatia \(Eds\.\),*Handbook of Second Language Acquisition*\(pp\. 413–468\)\. Academic Press\.
- Ma et al\. \(2025\)Ma, R\., Qian, M\., Tang, S\., Bañno, S\., Knill, K\., & Gales, M\. \(2025\)\. Assessment of L2 Oral Proficiency using Speech Large Language Models\.*Proceedings of Interspeech 2025*\. arXiv:2505\.21148\.
- Parikh et al\. \(2026a\)Parikh, A\., et al\. \(2026\)\. A Finetuned SpeechLLM for Joint Multi\-Granular L2 Assessment and Natural\-Language Rationales\.*arXiv preprint*arXiv:2606\.09470\.
- Parikh et al\. \(2026b\)Parikh, A\., et al\. \(2026\)\. Rubric\-guided fine\-tuning of speech\-LLMs for multi\-aspect, multi\-rater L2 reading\-speech assessment\.*Proceedings of LREC 2026*\. arXiv:2603\.16889\.
- Parikh et al\. \(2026c\)Parikh, A\., et al\. \(2026\)\. Zero\-Shot Speech LLMs for Multi\-Aspect Evaluation of L2 Speech: Challenges and Opportunities\.*arXiv preprint*arXiv:2601\.16230\.
- Preply \(2025\)Preply\. \(2025\)\. Language barriers survey \(n=3,608n=3\{,\}608adults across six countries\)\. Preply Blog\.
- Suzuki et al\. \(2021\)Suzuki, S\., Kormos, J\., & Uchihara, T\. \(2021\)\. The relationship between utterance and perceived fluency: A meta\-analysis of correlational studies\.*The Modern Language Journal*, 105\(2\), 435–463\.
- Swain \(1985\)Swain, M\. \(1985\)\. Communicative competence: Some roles of comprehensible input and comprehensible output in its development\. In S\. Gass & C\. Madden \(Eds\.\),*Input in Second Language Acquisition*\(pp\. 235–253\)\. Newbury House\.
- Swain \(1995\)Swain, M\. \(1995\)\. Three functions of output in second language learning\. In G\. Cook & B\. Seidlhofer \(Eds\.\),*Principle and Practice in Applied Linguistics*\(pp\. 125–144\)\. Oxford University Press\.
- Uto \(2026\)Uto, M\. \(2026\)\. Has Automated Essay Scoring Reached Sufficient Accuracy? Deriving Achievable QWK Ceilings from Classical Test Theory\.*Proceedings of AIED 2026*\. arXiv:2604\.19131\.
- Wu et al\. \(2025\)Wu, Z\., Gong, Z\., Ai, L\., et al\. \(2025\)\. Beyond Silent Letters: Amplifying LLMs in Emotion Recognition with Vocal Nuances \(SpeechCueLLM\)\.*Findings of NAACL 2025*\. arXiv:2407\.21315\.
- Zechner et al\. \(2009\)Zechner, K\., Higgins, D\., Xi, X\., & Williamson, D\. M\. \(2009\)\. Automatic scoring of non\-native spontaneous speech in tests of spoken English\.*Speech Communication*, 51\(10\), 883–895\.
- Zheng et al\. \(2023\)Zheng, L\., et al\. \(2023\)\. Judging LLM\-as\-a\-judge with MT\-Bench and Chatbot Arena\.*arXiv:2306\.05685*\.Similar Articles
Evaluating Language Models in Realistic Conversational Contexts
This paper introduces UPHELD, a large benchmark for evaluating human-scale conversational ability in LLMs, and proposes a Mixture-of-Judges framework that improves correlation with human assessments by approximately 30%.
(Towards) Scalable Reliable Automated Evaluation with Large Language Models
This paper proposes a scalable, domain-agnostic framework for automated LLM evaluation that uses pairwise comparisons by multiple LLMs and an Elo rating system to approximate expert judgments, reducing the need for human intervention.
RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization
This paper introduces RLearner-LLM, a framework using Hybrid-DPO to balance logical correctness and fluency in LLM-generated explanations, achieving significant NLI entailment improvements across multiple domains and base models while mitigating the verbosity bias of standard preference signals.
How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework
This paper introduces a register-aware linguistic evaluation framework to assess how human-like large language models (LLMs) are by comparing the distribution of 67 lexico-grammatical features between human and LLM-generated texts using Maximum Mean Discrepancy. Experiments across seven instruction-tuned open-source models and five registers show that no model perfectly matches human baselines, and closeness to human language varies by register rather than model size.
@dair_ai: // Your LLM judge disagrees with the experts // LLM Judges can be tricky to build. Here is an interesting showcasing wh…
The paper presents UPHELD, a benchmark with extensive human annotations for evaluating conversational LLMs, and a Mixture-of-Judges framework that enhances evaluation accuracy by 30%.