Phonological Perception of Sign Language Models

arXiv cs.CL Papers

Summary

This paper evaluates whether Sign Language Recognition models exhibit phonological sensitivity by probing them with minimal pairs of signs, revealing architectural trade-offs and emergent but limited phonological perception.

arXiv:2606.28667v1 Announce Type: new Abstract: Sign languages are compositional systems where meaning arises by combining sublexical phonological parameters, such as handshape, location, and movement. While deep learning models for Sign Language Recognition (SLR) have achieved increased performance on translation benchmarks, it remains unclear whether these models distinguish abstract phonological features or merely rely on low-level statistical correlations. This work evaluates the phonological perception of SLR models trained on American Sign Language (ASL) by probing phonological sensitivity using minimal pairs and evaluating representational alignment with human behavioral data. Our results reveal that SLR models exhibit emergent phonological sensitivity, but with clear architectural trade-offs: pose-based models are sensitive to handshape contrasts, while pixel-based models better capture location changes. Furthermore, pose-based models learn latent representations that correlate with human perceptual similarity judgments (r~0.49). These findings suggest that while SLR models exhibit emergent phonology, current training paradigms are insufficient to scale them beyond their architectural inductive biases.
Original Article
View Cached Full Text

Cached at: 06/30/26, 05:27 AM

# Phonological Perception of Sign Language Models
Source: [https://arxiv.org/html/2606.28667](https://arxiv.org/html/2606.28667)
Jessica Carter Johns Hopkins University Alex X\. Lu Microsoft Research Annemarie Kocab Johns Hopkins University

###### Abstract

Sign languages are compositional systems where meaning arises by combining sublexical phonological parameters, such as handshape, location, and movement\. While deep learning models for Sign Language Recognition \(SLR\) have achieved increased performance on translation benchmarks, it remains unclear whether these models distinguish abstract phonological features or merely rely on low\-level statistical correlations\. This work evaluates the phonological perception of SLR models trained on American Sign Language \(ASL\) by probing phonological sensitivity using minimal pairs and evaluating representational alignment with human behavioral data\. Our results reveal that SLR models exhibit emergent phonological sensitivity, but with clear architectural trade\-offs: pose\-based models are sensitive to handshape contrasts, while pixel\-based models better capture location changes\. Furthermore, pose\-based models learn latent representations that correlate with human perceptual similarity judgments \(r≈0\.49r\\approx 0\.49\)\. These findings suggest that while SLR models exhibit emergent phonology, current training paradigms are insufficient to scale them beyond their architectural inductive biases\.111Code and data:[https://github\.com/kayoyin/sign\-phonology](https://github.com/kayoyin/sign-phonology)\.

## 1Introduction

In signed languages, signs are constructed from sublexical phonological parameters: handshape, location, movement, orientation, and non\-manual markers\(Stokoe,[1960](https://arxiv.org/html/2606.28667#bib.bib25)\)\. While modern Sign Language Recognition \(SLR\) models achieve high translation fidelity\(Camgozet al\.,[2020](https://arxiv.org/html/2606.28667#bib.bib30); Yin and Read,[2020](https://arxiv.org/html/2606.28667#bib.bib11); Guanet al\.,[2025](https://arxiv.org/html/2606.28667#bib.bib28)\), it remains unclear if they acquire robust phonological representations\. Current evaluations focus on translation accuracy on narrow datasets\(Desaiet al\.,[2024](https://arxiv.org/html/2606.28667#bib.bib8)\)\. However, high accuracy does not guarantee linguistic competence; it could mask a limited ability to generalize beyond training distributions\. High scores may instead stem from overfitting to supplementary cues such as mouthing, which—unlike core phonological parameters—is optional and varies across signs and signers\. The model may not generalize to the significant variability in signing across signers, or overfit to the recording conditions in the training data\. This reliance on spurious features poses significant limitations for using SLR models as general\-purpose tools for linguistic analysis, especially for low\-resource sign languages\. Given that distinguishing phonological parameters is critical for human lexical access\(Lieberman and Borovsky,[2020](https://arxiv.org/html/2606.28667#bib.bib38); Williamset al\.,[2017](https://arxiv.org/html/2606.28667#bib.bib39)\), if models are not sensitive to the true underlying phonology, they lack the robust representational quality required for cross\-linguistic transfer or language understanding\.

In this work, we evaluate whether SLR models exhibit phonological sensitivity by probing them withminimal pairsof signs \(Figure[1](https://arxiv.org/html/2606.28667#S1.F1)\)\. These signs differ by one phonological element \(e\.g\.,QUEENvs\.KING, which share the same articulation location and sign movement and differ only in their handshape\)\. This framework allows us to explicitly measure whether the model’s latent space distinguishes signs along their phonological dimensions\. Furthermore, by studying latent representations instead of model outputs, our framework directly compares various architectures and data sources without requiring explicit training on phonological labels\.

![Refer to caption](https://arxiv.org/html/2606.28667v1/x1.png)Figure 1:Auditing phonological sensitivity with minimal pairs\. Feature representations of each sign in a minimal pair \(e\.g\.QUEENvs\.KING\) are extracted from the model’s penultimate layer\. We quantify the model’s sensitivity to the phonological contrast by calculating the cosine similarity between these latent representations\.We apply this framework in two experiments\. First, we assess discriminability in naturalistic American Sign Language \(ASL\) data\. We compare distinct architectures \(pixel\-based vs\. pose\-based\) and training domains \(sign language vs\. general action recognition\) to determine which features facilitate phonological sensitivity\.

Second, we investigate the alignment between model representations and human perception\. We use controlled synthetic data to study how the latent space of models organize handshapes compares to theoretical models of phonology based on human perception data\.

Our results reveal trade\-offs in model sensitivity governed by architectural constraints\. While pixel\-based models better capture broader spatial contrasts like sign location, pose\-based models demonstrate superior sensitivity to fine\-grained handshape distinctions\. We also find that the latent space of pose\-based models significantly aligns with human perceptual judgments, partially reproducing the confusion patterns and hierarchical groupings found in human signers\. These findings suggest that while emergent phonology exists in SLR models, models designed with different visual processing show different phonological sensitivities, highlighting the critical importance of cognitively grounded benchmarks to guide the development of robust computational sign systems\.

## 2Background & Related Work

##### Phonological structure of sign languages\.

Sign languages contain sublexical structure where meaningful signs are constructed from combinatorial parameters\. We adopt the standard five\-parameter model established byBattison \([1978](https://arxiv.org/html/2606.28667#bib.bib24)\)andLiddell and Johnson \([1989](https://arxiv.org/html/2606.28667#bib.bib14)\), which expandsStokoe \([1960](https://arxiv.org/html/2606.28667#bib.bib25)\)’s original classification to include:handshape,location,movement,palm orientation, andnon\-manual markers\. This compositional structure allows us to defineminimal pairs– signs that differ by exactly one parameter – to measure discriminability\. While originally described for ASL, this parameter model also generalizes across different sign languages\(Sandler,[2012](https://arxiv.org/html/2606.28667#bib.bib23)\)\. Psycholinguistic research demonstrates that signers activate these sublexical representations incrementally during lexical access\(Meadeet al\.,[2018](https://arxiv.org/html/2606.28667#bib.bib33)\), while behavioral studies show that phonological similarity influences sign production and recognition speeds\(Carreiraset al\.,[2008](https://arxiv.org/html/2606.28667#bib.bib34); Williamset al\.,[2017](https://arxiv.org/html/2606.28667#bib.bib39)\)\.

##### Sign language recognition\.

SLR research has largely prioritized end\-to\-end translation accuracy, often treating models as “black boxes” whose internal representations remain unexamined\(Camgozet al\.,[2020](https://arxiv.org/html/2606.28667#bib.bib30); Yin and Read,[2020](https://arxiv.org/html/2606.28667#bib.bib11)\)\. Existing attempts to analyze phonology typically rely on explicit supervision, training classifiers to predict specific features\(Tavellaet al\.,[2022](https://arxiv.org/html/2606.28667#bib.bib21); Bilgeet al\.,[2022](https://arxiv.org/html/2606.28667#bib.bib16)\)\. However, this approach is limited by data scarcity for rare features and statistical confounds \(e\.g\., specific handshapes correlate with specific locations, so a classifier can score well by predicting a feature whenever a correlated parameter occurs, rather than by detecting the parameter of interest\)\. To avoid these validity issues, we assessimplicitphonological emergence using a minimal pair strategy\. This method controls for confounding variables and requires no phonological labels, allowing us to probe the model’s latent structure directly\.

##### Phonological representations in spoken language models\.

Analogous work in speech processing demonstrates that in neural networks, phonological structure is emergent without explicit supervision\. Research shows that phonological information is encoded in the lower to middle layers of speech recognition models\(Belinkov and Glass,[2017](https://arxiv.org/html/2606.28667#bib.bib19); Pasadet al\.,[2021](https://arxiv.org/html/2606.28667#bib.bib18)\)and can be retrieved via linear geometry\(Gauthieret al\.,[2025](https://arxiv.org/html/2606.28667#bib.bib10)\)\. We extend this line of inquiry to the visual\-spatial domain, investigating whether sign language models exhibit similar emergent sensitivity to phonological constraints\.

## 3Methodology

To assess the phonological sensitivity of SLR models, we use aminimal pairframework\. Rather than relying on final classification outputs, we analyze the models’ latent representations of sign pairs that differ by a single phonological parameter\. This approach allows us to quantify sensitivity to specific phonological changes in models trained without explicit phonological supervision\.

### 3\.1Minimal pairs and datasets

We curate minimal pairs from three datasets: ASL Citizen\(Desaiet al\.,[2023](https://arxiv.org/html/2606.28667#bib.bib12)\), Sem\-Lex\(Kezaret al\.,[2023](https://arxiv.org/html/2606.28667#bib.bib13)\), and Handshapes in Context Stimuli\(Carteret al\.,[2026](https://arxiv.org/html/2606.28667#bib.bib35)\)\. We adopt the parameter inventory ofBattison \([1978](https://arxiv.org/html/2606.28667#bib.bib24)\)andLiddell and Johnson \([1989](https://arxiv.org/html/2606.28667#bib.bib14)\)for minimal\-pair decisions: each minimal pair contains exactly one change in handshape, location, movement, orientation, or non\-manual markers, verified by a Deaf author\. Our experiments focus on handshape, location, and movement contrasts, as orientation and non\-manual contrasts are too sparse in the available naturalistic corpora to evaluate reliably\. We provide samples from each dataset in Figure[2](https://arxiv.org/html/2606.28667#S3.F2)\.

![Refer to caption](https://arxiv.org/html/2606.28667v1/x2.png)Figure 2:Examples of minimal pair data and model sensitivity metrics\.\(Left, Center\) Naturalistic minimal pairs from ASL Citizen and Sem\-Lex contrasting in Handshape and Location, respectively\. Below:tt\-test results indicate the statistical significance of SLR models’ ability to distinguish these pairs\. \(Right\) A controlled minimal pair from HCS contrasting Handshape\. Below: Cosine similarity scores quantify the distance between the two nonce signs in the models’ latent space \(lower similarity = higher sensitivity\)\.##### ASL Citizen\.

ASL Citizen is a large\-scale, crowd\-sourced dataset featuring 52 deaf or hard\-of\-hearing \(DHH\) signers\. A Deaf author curated 259 minimal pairs spanning 360 unique signs to cover diverse phonological changes:137pairs differ in handshape, 31 in location,73in movement, and18in orientation or non\-manual markers\. To establish a baseline for human perceptual similarity, we enlisted a hearingnovicesigner to annotate the pairs as “very similar” \(51 pairs\), “somewhat similar” \(98 pairs\), or “not similar” \(103 pairs\)\. We use the standard test split, which includes 4,427 videos from 11 signers for the selected signs\.

##### Sem\-Lex\.

As all SLR models in this study were trained on ASL Citizen, to test for generalization, we also curated pairs from the Sem\-Lex Benchmark\. Unlike ASL Citizen, where users mimicked a seed video, Sem\-Lex prompted signers with concepts, which resulted in natural variability in sign production\. We mapped our ASL Citizen minimal pairs to Sem\-Lex by manually verifying that each video sharing a gloss matched the intended sign, and excluding lexical variants whose production differed from its pair partner in more than one parameter\. This ensures our framework evaluates robust phonological sensitivity rather than being confounded by mislabeled videos or lexical variation\. This yielded 178 pairs across 263 signs \(8,934 videos from 42 signers\)\.

##### Handshapes in Context Stimuli \(HCS\)\.

Real\-world vocabularies often lack systematic coverage of phonological combinations\. To address this, we use HCS, a set of controlled experimental stimuli consisting of nonce signs designed to systematically sample variations in handshape and phonological context\. The design of these stimuli was guided by prior phonological inventories\(Laneet al\.,[1976](https://arxiv.org/html/2606.28667#bib.bib36); Stungis,[1981](https://arxiv.org/html/2606.28667#bib.bib32)\), while the specific handshape inventory was expanded from 20 to 36 shapes based on frequency scores from a lexical database\(Sehyret al\.,[2021](https://arxiv.org/html/2606.28667#bib.bib31)\)\. HCS includes 972 videos covering 36 handshapes in 27 phonological contexts \(3 locations × 3 movements × 3 orientations\)\. This systematic design allows for precise minimal pair contrasts even for features rare in natural vocabulary\. Unlike ASL Citizen and Sem\-Lex, HCS features one Deaf signer with one video per nonce sign, recorded in a controlled laboratory setting\.

### 3\.2Models

We evaluate variants of two distinct video classification architectures, I3D and STGCN, which represent the predominant modeling paradigms222I3D and STGCN are the dominant open\-weight representatives of pixel\- and pose\-based SLR at the time of publication\. Although more advanced systems have been proposed in more recent work, these largely do not make their weights openly available, which precludes systematic comparison\.in the current SLR literature \(i\.e\. vision versus pose\-based\)\(Desaiet al\.,[2023](https://arxiv.org/html/2606.28667#bib.bib12)\)\. By comparing these standard backbones, we investigate howtraining domain\(general human action vs\. sign language recognition\) andinput modality\(raw pixels vs\. pose estimations\) influence phonological perception\. The two input modalities are illustrated in Appendix[A\.1](https://arxiv.org/html/2606.28667#A1.SS1)\.

##### I3D \(Pixel\-based\)\.

The Inflated 3D ConvNet \(I3D;Carreira and Zisserman \([2017](https://arxiv.org/html/2606.28667#bib.bib26)\)\) is a two\-stream architecture operating directly on RGB video frames\. We evaluate three variants: \(1\)I3D\-Rand, initialized with random weights to serve as a baseline; \(2\)I3D\-Kine, trained on the Kinetics dataset for general action recognition; and \(3\)I3D\-ASL, trained on ASL Citizen for sign language recognition\.

##### STGCN \(Pose\-based\)\.

The Spatio\-Temporal Graph Convolutional Network \(STGCN;Yanet al\.\([2018](https://arxiv.org/html/2606.28667#bib.bib27)\)\) operates on a skeletal graph representation of the body\. We evaluate: \(1\)STGCN\-Rand, a randomly initialized baseline; and \(2\)STGCN\-ASL, trained on ASL Citizen\. We excluded an STGCN model trained on Kinetics because existing pre\-trained models use action recognition pose graphs that lack fine\-grained hand landmarks\.

##### Strict exclusion of test signs\.

To ensure our analysis measures generalized phonological sensitivity rather than lexical memorization, we re\-trained both ASL\-trained models \(I3D\-ASL and STGCN\-ASL\) on a dataset split that explicitly excludes all signs appearing in our minimal pair test sets\.

Table 1:Model sensitivity to phonological changes on ASL Citizen\.Results are grouped by linguistic feature \(Handshape, Location, Movement\) and human perception of similarity\.tt\-stat represents the meantt\-statistic calculated across all minimal pairs in that category\. % wins \(vs\. Baseline\) denotes the percentage of pairs where the trained model had a significanttt\-statistic \(p<0\.05p<0\.05\) that was higher than its randomly initialized counterpart\. Head\-to\-Head columns show the percentage of pairs where one trained model \(I3D\-ASL or STGCN\-ASL\) significantly outperformed the other \(p<0\.05p<0\.05\)\. Note that Head\-to\-Head percentages do not sum to 100% as pairs with no significanttt\-stat are excluded\. STGCN is generally more sensitive to phonological changes than I3D, except for changes in location\.BaselinePixel\-based \(I3D\)BaselinePose\-based \(STGCN\)Head\-to\-HeadI3D\-RandI3D\-KineticsI3D\-ASLSTGCN\-RandSTGCN\-ASLI3D\-ASLSTGCN\-ASLCategorytt\-stattt\-stat% winstt\-stat% winstt\-stattt\-stat% wins% wins% winsTotal \(N=259N=259\)0\.130\.121\.935\.0069\.500\.236\.0881\.0833\.2056\.37Control \(N=259N=259\)0\.060\.76—10\.98—0\.4412\.66———By Phonological ParameterHandshape \(N=137N=137\)0\.140\.120\.734\.8069\.340\.186\.3086\.1329\.2064\.96Location \(N=31N=31\)0\.021\.136\.457\.7380\.650\.595\.6767\.7464\.5216\.13Movement \(N=73N=73\)0\.190\.201\.374\.7273\.970\.305\.8180\.8238\.3654\.79By Human PerceptionNot similar \(N=103N=103\)0\.150\.350\.975\.0166\.990\.106\.9889\.3228\.1664\.08Somewhat similar \(N=98N=98\)0\.170\.062\.045\.0876\.530\.435\.6479\.5937\.7656\.12Very similar \(N=51N=51\)0\.070\.173\.924\.2656\.860\.034\.6264\.7131\.3743\.14

Table 2:Out\-of\-domain phonological sensitivity results on Sem\-Lex\.Models trained on ASL Citizen were evaluated on minimal pairs from the Sem\-Lex dataset to test domain generalization\. Both models demonstrate a reduced sensitivity to phonological differences in this out\-of\-domain setting compared to ASL Citizen\.

## 4Are Models Sensitive to Phonological Changes?

### 4\.1Experiment

We evaluate whether SLR models learn latent representations that preserve the phonological distinctions found in minimal pairs\. By focusing on the latent feature space rather than final prediction accuracy, we establish a unified evaluation framework that allows for direct comparison across diverse model architectures and training objectives\. Furthermore, we posit that robust phonological separability in the latent space is a strong indicator of a model’s capacity for zero\-shot transfer—a critical capability for analyzing underrepresented sign languages where training data are scarce\.

To analyze these representations, we extract feature vectors from the model’s penultimate layer \(preceding the classification head\) for videos in our naturalistic minimal pair datasets, where there are multiple samples per sign \(ASL Citizen and Sem\-Lex\)\. We operationalize “phonological sensitivity” by comparing intra\-sign similarity against inter\-sign similarity\. Our hypothesis is that if a model is sensitive to phonological change, the cosine similarity between twominimally differentsigns should besignificantly lowerthan two instances of thesamesign \(produced by different signers\)\.

Formally, for each minimal pair of signs\(s1,s2\)\(s\_\{1\},s\_\{2\}\)we construct two distributions of similarity scores\.\(1\) Intra\-sign similarity \(Ds​a​m​eD\_\{same\}\):The cosine similarity between the representation of signs1s\_\{1\}produced by signerAAand an instance of the same signs1s\_\{1\}produced by a different signerBB\.\(2\) Inter\-sign similarity \(Dd​i​f​fD\_\{diff\}\):The cosine similarity between the representation of signs1s\_\{1\}produced by signerAAand the phonologically distinct signs2s\_\{2\}produced by signerBB\.\(3\) Sampling:We repeat these computations for 10 random pairs of signers\(A,B\)\(A,B\)to populate the distributionsDs​a​m​eD\_\{same\}andDd​i​f​fD\_\{diff\}\.

To quantify phonological sensitivity, we perform a two\-tailed independent samples t\-test to determine if the distribution of similarities ofDs​a​m​eD\_\{same\}is statistically significantly higher than that ofDd​i​f​fD\_\{diff\}:

t​\(s1,s2\)=D¯s​a​m​e−D¯d​i​f​fss​a​m​e2/N\+sd​i​f​f2/N,t\(s\_\{1\},s\_\{2\}\)\\;=\\;\\frac\{\\bar\{D\}\_\{same\}\-\\bar\{D\}\_\{diff\}\}\{\\sqrt\{s^\{2\}\_\{same\}/N\+s^\{2\}\_\{diff\}/N\}\},\(1\)whereD¯\\bar\{D\}ands2s^\{2\}denote the sample mean and variance of each similarity distribution, andNNis the number of sampled signer pairs\.

As a control, we repeat this procedure on random non\-minimal sign pairs sampled from the same gloss pool, providing an upper bound on thett\-statistics that trained models should trivially achieve \(Controlrows in Tables[1](https://arxiv.org/html/2606.28667#S3.T1)–[2](https://arxiv.org/html/2606.28667#S3.T2)\)\.

### 4\.2Results

We present the results oftt\-tests on ASL Citizen and Sem\-Lex in Table[1](https://arxiv.org/html/2606.28667#S3.T1)and Table[2](https://arxiv.org/html/2606.28667#S3.T2), respectively\. For each phonological category, we report themean t\-statistic\(magnitude of separation\) and the% wins\(frequency with which the trained model significantly outperforms the randomly initialized baseline,p<0\.05p<0\.05\)\. Additionally, we provide head\-to\-head comparisons indicating how often one ASL\-trained model proved significantly more sensitive than the other\. In Appendix[A\.2](https://arxiv.org/html/2606.28667#A1.SS2), we also provide qualitative examples of model sensitivity results to minimal pairs in ASL\-Citizen\.

First,training on sign language data is prerequisite for phonological sensitivity,which notably deviates from human learning\. Across all categories, models trained on general action recognition \(I3D\-Kinetics\) failed to distinguish minimal pairs, achieving a negligible 1\.93% win rate over random baselines on ASL Citizen\. In contrast, both ASL\-trained models demonstrated significant emergent sensitivity \(I3D\-ASL: 69\.50%; STGCN\-ASL: 81\.08%\)\. TheControlrow confirms that the ASL\-trained models separate random non\-minimal pairs far more strongly than minimal pairs \(e\.g\.,t=12\.66t=12\.66vs\.6\.086\.08for STGCN\-ASL\), while the untrained and Kinetics baselines remain at chance\.

This failure of the general action model is surprising given that phonological sensitivity builds upon general visual capabilities: prior work demonstrates that non\-signers can reliably distinguish phonological contrasts in ASL and there is a strong correlation between signer and non\-signer perceptual confusion\(Stungis,[1981](https://arxiv.org/html/2606.28667#bib.bib32)\)\. Our results suggest that current video models are poor proxies for human perception\. Unlike the human visual system, which can transfer general object and motion recognition to identify sign contrasts, these models require domain\-specific training to acquire phonological sensitivity\. We attribute this to the training signal itself: general action datasets like Kinetics lack the fine\-grained phonological contrasts found in sign language, and so never pressure the model to encode the subtle distinctions that sign\-specific training captures\.

Second,architectures exhibit distinct biases in what aspects of phonology they are sensitive to\.On ASL Citizen, the pose\-based STGCN\-ASL model demonstrates superior sensitivity tohandshapeover I3D\-ASL \(86\.13% vs\. 69\.34%\) andmovement\(80\.82% vs\. 73\.97%\)\. This suggests that explicit skeletal modeling better captures fine\-grained configuration and trajectory than raw pixels\. Conversely, the pixel\-based I3D\-ASL model significantly outperformed STGCN\-ASL onlocationcontrasts \(80\.65% vs\. 67\.74%\), winning the head\-to\-head comparison in 64\.52% of location pairs\. This indicates that pixel\-based models may better retain the spatial context relative to the frame \(e\.g\., forehead vs\. chin\) that graph\-based representations – which are often normalized for position – can lose\.

Third,model sensitivity aligns with human perception\.For both architectures, performance degrades for minimal pairs rated by a human non\-signer as perceptually more similar\. A Kruskal\-Wallis H\-test confirms a statistically significant difference in STGCN\-ASL sensitivity across the “Not / Somewhat / Very Similar” human ratings \(H​\(2\)=12\.92,p=0\.002H\(2\)=12\.92,p=0\.002\), validating that the models struggle most with the distinctions humans may find subtle\. In contrast, this difference was not statistically significant for the I3D\-ASL model or the non\-ASL trained variants\.

Finally,representations are brittle to domain shifts\.When evaluated on Sem\-Lex, which introduces variations in prompts and filming conditions, sensitivity drops substantially \(I3D\-ASL: 45\.07%; STGCN\-ASL: 49\.3%\)\. This suggests that while current models learn phonological distinctions, their representations remain highly brittle to the specific production constraints of their training distribution\.

## 5How Does Model Phonological Sensitivity Compare to Human Perception?

### 5\.1Experiment

Beyond assessing whether modelscandistinguish phonological features, we investigate whether their internal representation space captures meaningful phonological structure\. We use the HCS data to compare model feature distances with subjective human perception and articulatory geometry\. While naturalistic datasets like ASL Citizen offer ecological validity, they suffer from sparse phonological coverage\. HCS allows us to control for non\-phonological variance \(e\.g\., lighting, signer identity\) while systematically sampling the full spectrum of handshape contrasts\.

First, we quantifyhuman perceptionusing confusion matrices collected byStungis \([1981](https://arxiv.org/html/2606.28667#bib.bib32)\)\. This data quantifies how frequently human signers and non\-signers confuse different ASL handshapes\. We treat the confusion frequency as a proxy for perceptual similarity: handshapes that are confused more often are considered more perceptually similar\.

Second, we calculate thehandshape distance \(HD\)between handshape pairs\. HD, a computational measure of articulatory geometry adapted fromYinet al\.\([2024](https://arxiv.org/html/2606.28667#bib.bib37)\), is defined as the mean angular difference between their corresponding joints\. For each pair of handshapes in HCS data, we computed the HD for the handshape pair in 27 phonological contexts and took the median\.

Third, we measuremodel sensitivityto the phonological contrast in each pair of handshapes in HCS by calculating the cosine similarity between their corresponding video representations in the model’s latent space\. We then compute the Pearson correlation coefficient \(rr\) between the vector of human confusion frequencies, handshape distances, and model feature similarities\. We summarize the results of our correlation analysis in Table[3](https://arxiv.org/html/2606.28667#S5.T3)\.

To visualize the structural organization of the learned feature spaces, we performed U\-statistic hierarchical clustering\(D’Andrade,[1978](https://arxiv.org/html/2606.28667#bib.bib40)\)of handshapes according to HD and model feature similarity, following the methodology inStungis \([1981](https://arxiv.org/html/2606.28667#bib.bib32)\)\. In Figure[3](https://arxiv.org/html/2606.28667#S5.F3), we plotted dendrograms for four representations: \(1\) theLane\-Boyes\-Braem II \(LBB2\) Model\(a theoretical model based on human perceptual studies\(Laneet al\.,[1976](https://arxiv.org/html/2606.28667#bib.bib36); Stungis,[1981](https://arxiv.org/html/2606.28667#bib.bib32)\)\), \(2\)Handshape Distance\(our geometric reference metric\), \(3\)I3D\-ASL, and \(4\)STGCN\-ASL\. This allows us to inspect whether models recover high\-level phonological features that humans rely on, or if they use different features for distinguishing handshapes\.

Table 3:Correlation between perceptual similarity reported by humans\(Stungis,[1981](https://arxiv.org/html/2606.28667#bib.bib32)\), handshape distance \(HD\), and cosine similarity of model features of handshapes\.Significant correlations \(p<0\.05p<0\.05\) are bolded\. Models trained on ASL data, especially STGCN, have higher alignment with human perception\.
### 5\.2Results

![Refer to caption](https://arxiv.org/html/2606.28667v1/x3.png)Figure 3:Structural organization of handshape representations\.From top to bottom: the LBB2 Model \(theoretic model based on human perception\), handshape distance \(geometric reference\), I3D\-ASL, STGCN\-ASL\.##### Correlation analysis\.

First, we established reference bounds for our analysis\. Consistent with prior literature, human signers and non\-signers have high agreement \(r=0\.89r=0\.89\)\. The theoretical HD metric showed moderate positive correlation with human perception, suggesting that human perception is grounded in articulatory geometry \(r=0\.53r=0\.53for signers\)\. As expected, models with randomly initialized weights \(Rand\) yielded insignificant correlations, confirming that untrained networks do not inherently encode phonologically meaningful features\.

Second, we find thattraining on ASL data significantly improved alignment with human perception\. For the I3D architecture, the model trained on general human action recognition \(Kine\) has a weak correlation with human signers \(r=0\.19r=0\.19\)\. Training on ASL data increased this correlation tor=0\.31r=0\.31for the I3D model, suggesting that exposure to sign language data drives the feature space closer to human perceptual representations\.

Third, we observe cleararchitectural differences in handshape perceptionbetween models\. STGCN\-ASL is the model with the highest alignment with human signers \(r=0\.49r=0\.49\)\. The correlation between I3D\-ASL and STGCN\-ASL is itself notably low \(r=0\.19r=0\.19\)\. This indicates that while both models successfully distinguish some handshapes, their architectures induce different structures in their latent spaces\. STGCN\-ASL also shows a strong correlation with handshape distance \(r=0\.55r=0\.55\), substantially higher than I3D\-ASL \(r=0\.20r=0\.20\)\. I3D, lacking the skeletal prior of pose\-based models, seems to learn a visual abstraction that correlates less with joint geometry and the human perception of handshape\. We also note that the alignment between STGCN\-ASL and humans is not higher than the alignment between HD and humans, which may suggest that training on sign language allows for similarities already present in input space to trickle through the model, but does not enhance perceptual similarity past what is expected from joint geometry\.

##### Hierarchical clustering analysis\.

The dendrograms reveal distinct topological differences between models that corroborate the correlation results\. TheLBB2 Modelorganizes handshapes primarily by the ‘Compact’ feature \(whether the three middle fingers are closed\)\. Within the ‘\+Compact’ group, it further distinguishes based on the number of bent fingers, while the ‘\-Compact’ group is organized by the number of extended fingers\. Similarly, theHD dendrogramorganizes handshapes by finger count, creating a primary split between handshapes with≤1\\leq 1finger extended \(with the exception of two handshapes that has the thumb and another finger extended\) versus those with\>1\>1\. It further sub\-groups the first cluster by specific articulators \(thumb vs\. index extension\) and the second by the count of extended fingers\.

TheI3D\-ASL dendrogramdiverges the most from the LBB2 Model\. It primarily separates by handshapes that visually resemble a fist from those that do not\. Within the ‘fist’ cluster, it sorts by closure tightness\. The non\-fist cluster separates a group with extended thumbs but leaves the remaining handshapes in sub\-clusters lacking clear patterns\. This suggests I3D relies primarily on global visual texture \(e\.g\., “compact fist” vs\. “open hand with protruding thumb”\) rather than fine\-grained articulatory features\.

In contrast, theSTGCN\-ASL dendrogramclosely recovers the structure of the LBB2 Model\. It organizes handshapes into four primary clusters: \(1\) all three middle fingers are closed, \(2\) 4\-5 extended fingers, \(3\) 2\-3 extended fingers, and \(4\) just one extended finger \(with the exception of the ‘K’ handshape on the far right\)\. STGCN also captures fine\-grained sub\-distinctions within these groups: in the first cluster, it distinguishes thumb/pinky extension; in the second, it separates curved fingers from straight extension\. This structural alignment explains the STGCN’s high correlation with human perception \(LBB2\) and the theoretical HD metric\.

## 6Discussion & Future Work

Our investigation reveals that while phonological sensitivity emerges without supervision in SLR models, it is fractionated by architectural inductive biases: pose\-based STGCN are more sensitive to fine\-grained handshape contrasts, whereas pixel\-based I3D are better at location distinctions\. This functional split suggests that current training methods and architectures are insufficient for holistic sign language processing, and that a more complete model likely requires dual\-stream integration of the skeletal abstraction captured by pose\-based models and the holistic spatial processing captured by pixel\-based models\.

Furthermore, the strong correlation of STGCN and human perception serves as a benchmark for the structural integrity of the latent space in skeletal models\. However, these representations remain brittle to domain shifts\. The significant drop in phonological sensitivity on out\-of\-domain data such as Sem\-Lex exposes a tendency to overfit production variables \(e\.g\., recording conditions\) rather than acquiring robust phonological abstractions\. Thus, scaling current training paradigms on naturalistic data may not be sufficient to induce the robust phonological representations required for real\-world applications\.

While our framework for phonological sensitivity advances SLR evaluation, our study has three primary limitations\. First, data sparsity restricted our analysis to handshape, location, and movement, omitting palm orientation and non\-manual markers\. Second, we used only supervised baselines due to the unavailability of open\-source models; future work should investigate whether self\-supervised objectives induce different representations\. Finally, our focus on isolated signs leaves the effects of co\-articulation and syntax in continuous SLR for future study\.

Beyond these, our approach enables broader exploration of model robustness and generalizability\. Future research can probe compositionality of model features, similarly to how word embeddings enable analogical reasoning, and explore multimodal architectures that bridge the trade\-off between I3D and STGCN models\. The minimal pairs could also be used to improve training itself: they supply hard negatives for contrastive fine\-grained phonological discrimination, and could pre\-train models to first acquire phonological distinctions that then transfer to lower\-resource sign languages\. Future work could also test for cross\-linguistic transfer between sign languages to see whether phonological representations learned from higher\-resource sign languages could generalize to lower\-resources sign languages; testing transfer across structurally unrelated sign languages could further isolate phonological constraints grounded in human body biomechanics from language\-specific learning, which would inform data\-efficient modeling\.

#### Acknowledgements

KY is supported by the Vitalik Buterin Ph\.D\. Fellowship in AI Existential Safety\. AK is supported by NSF STEM\-APWD 136073\-5130403\.

## References

- R\. Battison \(1978\)Lexical borrowing in american sign language\.\.ERIC\.Cited by:[§2](https://arxiv.org/html/2606.28667#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2606.28667#S3.SS1.p1.1)\.
- Y\. Belinkov and J\. Glass \(2017\)Analyzing hidden representations in end\-to\-end automatic speech recognition systems\.Advances in Neural Information Processing Systems30\.Cited by:[§2](https://arxiv.org/html/2606.28667#S2.SS0.SSS0.Px3.p1.1)\.
- Y\. C\. Bilge, R\. G\. Cinbis, and N\. Ikizler\-Cinbis \(2022\)Towards zero\-shot sign language recognition\.IEEE transactions on pattern analysis and machine intelligence45\(1\),pp\. 1217–1232\.Cited by:[§2](https://arxiv.org/html/2606.28667#S2.SS0.SSS0.Px2.p1.1)\.
- N\. C\. Camgoz, O\. Koller, S\. Hadfield, and R\. Bowden \(2020\)Sign language transformers: joint end\-to\-end sign language recognition and translation\.InProceedings of the IEEE/CVF conference on computer vision and pattern recognition,pp\. 10023–10033\.Cited by:[§1](https://arxiv.org/html/2606.28667#S1.p1.1),[§2](https://arxiv.org/html/2606.28667#S2.SS0.SSS0.Px2.p1.1)\.
- J\. Carreira and A\. Zisserman \(2017\)Quo vadis, action recognition? a new model and the kinetics dataset\.Inproceedings of the IEEE Conference on Computer Vision and Pattern Recognition,pp\. 6299–6308\.Cited by:[§3\.2](https://arxiv.org/html/2606.28667#S3.SS2.SSS0.Px1.p1.1)\.
- M\. Carreiras, E\. Gutiérrez\-Sigut, S\. Baquero, and D\. Corina \(2008\)Lexical processing in spanish sign language \(lse\)\.Journal of Memory and Language58\(1\),pp\. 100–122\.Cited by:[§2](https://arxiv.org/html/2606.28667#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Carter, A\. Kocab,et al\.\(2026\)Manuscript in preparation\.Note:Manuscript in preparationCited by:[§3\.1](https://arxiv.org/html/2606.28667#S3.SS1.p1.1)\.
- R\. G\. D’Andrade \(1978\)U\-statistic hierarchical clustering\.Psychometrika43\(1\),pp\. 59–67\.Cited by:[§5\.1](https://arxiv.org/html/2606.28667#S5.SS1.p5.1)\.
- A\. Desai, L\. Berger, F\. O\. Minakov, V\. Milan, C\. Singh, K\. Pumphrey, R\. E\. Ladner, H\. Daumé III, A\. X\. Lu, N\. Caselli, and D\. Bragg \(2023\)ASL citizen: a community\-sourced dataset for advancing isolated sign language recognition\.arXiv preprint arXiv:2304\.05934\.Cited by:[§3\.1](https://arxiv.org/html/2606.28667#S3.SS1.p1.1),[§3\.2](https://arxiv.org/html/2606.28667#S3.SS2.p1.1)\.
- A\. Desai, M\. De Meulder, J\. A\. Hochgesang, A\. Kocab, and A\. X\. Lu \(2024\)Systemic biases in sign language ai research: a deaf\-led call to reevaluate research agendas\.arXiv preprint arXiv:2403\.02563\.Cited by:[§1](https://arxiv.org/html/2606.28667#S1.p1.1)\.
- J\. Gauthier, C\. Breiss, M\. K\. Leonard, and E\. F\. Chang \(2025\)Emergent morpho\-phonological representations in self\-supervised speech models\.InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing,pp\. 28055–28074\.Cited by:[§2](https://arxiv.org/html/2606.28667#S2.SS0.SSS0.Px3.p1.1)\.
- M\. Guan, Y\. Wang, G\. Ma, J\. Liu, and M\. Sun \(2025\)MSKA: multi\-stream keypoint attention network for sign language recognition and translation\.Pattern Recognition165,pp\. 111602\.Cited by:[§1](https://arxiv.org/html/2606.28667#S1.p1.1)\.
- L\. Kezar, J\. Thomason, N\. Caselli, Z\. Sehyr, and E\. Pontecorvo \(2023\)The sem\-lex benchmark: modeling asl signs and their phonemes\.Proceedings of the 25th International ACM SIGACCESS Conference on Computers and Accessibility\.Cited by:[§3\.1](https://arxiv.org/html/2606.28667#S3.SS1.p1.1)\.
- H\. Lane, P\. Boyes\-Braem, and U\. Bellugi \(1976\)Preliminaries to a distinctive feature analysis of handshapes in american sign language\.Cognitive Psychology8\(2\),pp\. 263–289\.Cited by:[§3\.1](https://arxiv.org/html/2606.28667#S3.SS1.SSS0.Px3.p1.1),[§5\.1](https://arxiv.org/html/2606.28667#S5.SS1.p5.1)\.
- S\. K\. Liddell and R\. E\. Johnson \(1989\)American sign language: the phonological base\.Sign language studies64\(1\),pp\. 195–277\.Cited by:[§2](https://arxiv.org/html/2606.28667#S2.SS0.SSS0.Px1.p1.1),[§3\.1](https://arxiv.org/html/2606.28667#S3.SS1.p1.1)\.
- A\. M\. Lieberman and A\. Borovsky \(2020\)Lexical recognition in deaf children learning american sign language: activation of semantic and phonological features of signs\.Language learning70\(4\),pp\. 935–973\.Cited by:[§1](https://arxiv.org/html/2606.28667#S1.p1.1)\.
- G\. Meade, B\. Lee, K\. J\. Midgley, P\. J\. Holcomb, and K\. Emmorey \(2018\)Phonological and semantic priming in american sign language: n300 and n400 effects\.Language, Cognition and Neuroscience33\(9\),pp\. 1092–1106\.Cited by:[§2](https://arxiv.org/html/2606.28667#S2.SS0.SSS0.Px1.p1.1)\.
- A\. Pasad, J\. Chou, and K\. Livescu \(2021\)Layer\-wise analysis of a self\-supervised speech representation model\.In2021 IEEE Automatic Speech Recognition and Understanding Workshop \(ASRU\),pp\. 914–921\.Cited by:[§2](https://arxiv.org/html/2606.28667#S2.SS0.SSS0.Px3.p1.1)\.
- W\. Sandler \(2012\)The phonological organization of sign languages\.Language and linguistics compass6\(3\),pp\. 162–182\.Cited by:[§2](https://arxiv.org/html/2606.28667#S2.SS0.SSS0.Px1.p1.1)\.
- Z\. S\. Sehyr, N\. Caselli, A\. M\. Cohen\-Goldberg, and K\. Emmorey \(2021\)The asl\-lex 2\.0 project: a database of lexical and phonological properties for 2,723 signs in american sign language\.The Journal of Deaf Studies and Deaf Education26\(2\),pp\. 263–277\.Cited by:[§3\.1](https://arxiv.org/html/2606.28667#S3.SS1.SSS0.Px3.p1.1)\.
- Jr\. Stokoe \(1960\)Sign Language Structure: An Outline of the Visual Communication Systems of the American Deaf\.The Journal of Deaf Studies and Deaf Education10\(1\),pp\. 3–37\.External Links:ISSN 1081\-4159,[Document](https://dx.doi.org/10.1093/deafed/eni001),[Link](https://doi.org/10.1093/deafed/eni001),https://academic\.oup\.com/jdsde/article\-pdf/10/1/3/1034248/eni001\.pdfCited by:[§1](https://arxiv.org/html/2606.28667#S1.p1.1),[§2](https://arxiv.org/html/2606.28667#S2.SS0.SSS0.Px1.p1.1)\.
- J\. Stungis \(1981\)Identification and discrimination of handshape in american sign language\.Perception & Psychophysics29\(3\),pp\. 261–276\.Cited by:[§3\.1](https://arxiv.org/html/2606.28667#S3.SS1.SSS0.Px3.p1.1),[§4\.2](https://arxiv.org/html/2606.28667#S4.SS2.p3.1),[§5\.1](https://arxiv.org/html/2606.28667#S5.SS1.p2.1),[§5\.1](https://arxiv.org/html/2606.28667#S5.SS1.p5.1),[Table 3](https://arxiv.org/html/2606.28667#S5.T3.3.1),[Table 3](https://arxiv.org/html/2606.28667#S5.T3.4.1)\.
- F\. Tavella, A\. Galata, and A\. Cangelosi \(2022\)Phonology recognition in american sign language\.InICASSP 2022\-2022 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 8452–8456\.Cited by:[§2](https://arxiv.org/html/2606.28667#S2.SS0.SSS0.Px2.p1.1)\.
- J\. T\. Williams, A\. Stone, and S\. D\. Newman \(2017\)Operationalization of sign language phonological similarity and its effects on lexical access\.The Journal of Deaf Studies and Deaf Education22\(3\),pp\. 303–315\.Cited by:[§1](https://arxiv.org/html/2606.28667#S1.p1.1),[§2](https://arxiv.org/html/2606.28667#S2.SS0.SSS0.Px1.p1.1)\.
- S\. Yan, Y\. Xiong, and D\. Lin \(2018\)Spatial temporal graph convolutional networks for skeleton\-based action recognition\.InProceedings of the AAAI conference on artificial intelligence,Vol\.32\.Cited by:[§3\.2](https://arxiv.org/html/2606.28667#S3.SS2.SSS0.Px2.p1.1)\.
- K\. Yin and J\. Read \(2020\)Better sign language translation with STMC\-transformer\.InProceedings of the 28th International Conference on Computational Linguistics,D\. Scott, N\. Bel, and C\. Zong \(Eds\.\),Barcelona, Spain \(Online\),pp\. 5975–5989\.External Links:[Link](https://aclanthology.org/2020.coling-main.525/),[Document](https://dx.doi.org/10.18653/v1/2020.coling-main.525)Cited by:[§1](https://arxiv.org/html/2606.28667#S1.p1.1),[§2](https://arxiv.org/html/2606.28667#S2.SS0.SSS0.Px2.p1.1)\.
- K\. Yin, T\. Regier, and D\. Klein \(2024\)Pressures for communicative efficiency in american sign language\.InAnnual Conference of the Association for Computational Linguistics \(ACL\),Cited by:[§5\.1](https://arxiv.org/html/2606.28667#S5.SS1.p3.1)\.

## Appendix AAppendix

### A\.1SLR model input modalities

In Figure[4](https://arxiv.org/html/2606.28667#A1.F4), we show examples of the video input modality to the I3D and STGCN models\.

![Refer to caption](https://arxiv.org/html/2606.28667v1/x4.png)Figure 4:Comparison of model input modalities\.\(Left\) Raw RGB video frames processed by the I3D model\. \(Right\) Skeletal pose graph estimation processed by the STGCN model\. Both inputs depict a frame from the sign “HIPPO\.”
### A\.2Qualitative analysis of minimal pair sensitivity

In Table[4](https://arxiv.org/html/2606.28667#A1.T4), we provide qualitative examples oft\-test results for I3D\-ASL and STGCN\-ASL models\. The first three examples show cases where both models fail to represent the phonological change in minimal pairs of varying similarity as judged by humans\. The first pair \(HOSPITAL / PATIENT\) has a change in handshape, where the handshapes primarily differ by the position of the index finger\. The second pair \(5DOLLARS / FIFTH\) has a change in movement, where5DOLLARSis signed by twisting the wrist once, whileFIFTHis signed by twisting the wrist inward\-outward in small movements\. The third pair \(BEAVER / TABLE\) has a change in location, where the elbow of the dominant arm and wrist of the non\-dominant arm stay in contact forBEAVER, but neither elbows and wrists stay in contact during the movements forTABLE\.

The final two examples show the divergence between architectures’ phonological sensitivity\. In example 4, only STGCN fails to distinguishFEEL / HAPPY, where these signs mostly differ in whether the middle finger is bent or straight\. In example5, only I3D fails to distinguishSEVEN / NINE, where the handshapes differ in the finger that is in contact with the thumb\.

Table 4:Qualitative examples of model sensitivity to minimal pairs in ASL Citizen\.

Similar Articles

Phone Segmentation and Recognition through Phonological Activation Mapping

Hugging Face Daily Papers

This paper introduces SPAM (S3M-based Phonological Activation Mapping), a method that leverages self-supervised speech models to perform both phone segmentation and recognition simultaneously using lightweight, gradient-descent-free prediction heads requiring minimal phonetic transcriptions.

Direct Translation between Sign Languages

arXiv cs.CL

This paper introduces a direct sign-to-sign translation model that bypasses intermediate text by using back-translation to create synthetic parallel sign language data, achieving significant improvements in speed and accuracy over cascade methods for ASL, CSL, and DGS.

Emotion Recognition in Sign Language Conversation

arXiv cs.CL

This paper introduces the eJSL Dialog dataset for emotion recognition in sign language conversations, addressing the lack of conversational context in existing datasets. Benchmarking shows a domain gap when applying generic multimodal models, highlighting the need for context-aware visual extractors for sign language.