Isolated Sign Language Recognition for Icelandic Sign Language: 低资源环境下的实验

arXiv cs.CL 论文

摘要

本文介绍了针对Icelandic Sign Language的孤立手语识别的首次实验,使用低资源数据集,表明从其他手语的跨语言迁移显著提高了性能。

arXiv:2609.25862v1 Announce Type: new Abstract: We present the first experiments on isolated sign language recognition (ISLR) for Icelandic Sign Language (\'ITM). We use \'ITM SignWiki, a dataset derived from a bilingual Icelandic--\'ITM online dictionary. It is genuinely low-resource: 1,845 videos cover 849 classes, 86% of which have only two examples, making the full task effectively one-shot recognition across signers. We compare two open-source ISLR frameworks, OpenHands and SPOTER, on three tasks of increasing vocabulary size (22, 117 and 849 classes), and evaluate three pose estimators and two forms of cross-lingual transfer. With \'ITM data alone, SPOTER outperforms OpenHands on all three tasks, and MediaPipe poses give better results than AlphaPose or SDPose. Cross-lingual transfer brings the largest gains: pretraining SPOTER on American Sign Language data before finetuning on \'ITM raises accuracy by 14--24 percentage points, to 72.7%, 47.9% and 22.6% on the three tasks, and multilingual training with data from six other sign languages lifts OpenHands from 1.41% to 28.86% on the full task. Although far from practical use, the results suggest that transfer from better-resourced sign languages is promising for very low-resource ones. We release our adapted versions of both frameworks.
查看原文
查看缓存全文

缓存时间: 2026/09/23 09:19

# Isolated Sign Language Recognition for Icelandic Sign Language: Experiments in a Low-resource Setting
Source: [https://arxiv.org/html/2609.25862](https://arxiv.org/html/2609.25862)
Finnur Ágúst IngimundarsonAffiliation:Department of Computational Linguistics, University of ZurichEmail:[finnuragust\.ingimundarson@uzh\.ch](mailto:)Guðný Björk ÞorvaldsdóttirAffiliation:Communication Centre for the Deaf and Hard of Hearing in IcelandEmail:[gudny\.bjork\.thorvaldsdottir@shh\.is](mailto:)Mathias MüllerAffiliation:Department of Computational Linguistics, University of ZurichEmail:[mmueller@cl\.uzh\.ch](mailto:)Sarah Ebling

###### Abstract

We present the first experiments on isolated sign language recognition \(ISLR\) for Icelandic Sign Language \(ÍTM\)\. We use ÍTM SignWiki, a dataset derived from a bilingual Icelandic–ÍTM online dictionary\. It is genuinely low\-resource: 1,845 videos cover 849 classes, 86% of which have only two examples, making the full task effectively one\-shot recognition across signers\. We compare two open\-source ISLR frameworks, OpenHands and SPOTER, on three tasks of increasing vocabulary size \(22, 117 and 849 classes\), and evaluate three pose estimators and two forms of cross\-lingual transfer\. With ÍTM data alone, SPOTER outperforms OpenHands on all three tasks, and MediaPipe poses give better results than AlphaPose or SDPose\. Cross\-lingual transfer brings the largest gains: pretraining SPOTER on American Sign Language data before finetuning on ÍTM raises accuracy by 14–24 percentage points, to 72\.7%, 47\.9% and 22\.6% on the three tasks, and multilingual training with data from six other sign languages lifts OpenHands from 1\.41% to 28\.86% on the full task\. Although far from practical use, the results suggest that transfer from better\-resourced sign languages is promising for very low\-resource ones\. We release our adapted versions of both frameworks\.

## 1Introduction

Although not always recognized as such, sign languages are full\-fledged natural languages that have their own grammar and vocabulary\.111This paper is based on[Ingimundarson \(2026\)](https://arxiv.org/html/2609.25862#bib.bib2)the first author’s master’s thesis at the University of ZurichContrary to a widespread belief, there is no universal sign language and several hundred sign languages have been documented around the world\. The number of deaf people who use them as a primary means of communication worldwide is estimated to be roughly 70 million, and more than half a million in Europe\([Way et al\., 2024](https://arxiv.org/html/2609.25862#bib.bib26)\)\.

The field of sign language processing \(SLP\) is more recent than that of natural language processing \(NLP\) and has lagged behind, constrained by a lack of technological or computational resources, software, and data\([Yin et al\., 2021](https://arxiv.org/html/2609.25862#bib.bib7);[Müller et al\., 2022](https://arxiv.org/html/2609.25862#bib.bib6)\)\. This situation has gradually improved but SLP remains underrepresented and fundamental design decisions underexplored\([O’Brien et al\., 2026a](https://arxiv.org/html/2609.25862#bib.bib16)\)\. In general, it can also be argued that all sign languages are low\-resource languages when compared to spoken languages\([Joshi et al\., 2020](https://arxiv.org/html/2609.25862#bib.bib3)\)\. Nevertheless, larger sign language communities, such as that of American Sign Language \(ASL\), are comparatively well represented and well resourced, whereas sign languages of smaller communities are low\-resource, which restricts the development and deployment of SLP tools based on them\.

One such low\-resource sign language is Icelandic Sign Language, that will hereafter be abbreviated as ÍTM \(íslenskttáknmál\), which is preferred by the Icelandic Deaf community and ÍTM researchers\. ÍTM is now estimated to have about 300 native users, and in total approximately 2,000 users\. The total number includes L2 users with late onset of hearing loss, children of deaf adults, foreign L1 signers \(ÍTM L2 signers\) and hearing signers with various levels of proficiency, the last group comprising somewhere between 1,000 and 1,500 people\([Koulidobrova and Sverrisdóttir, 2021](https://arxiv.org/html/2609.25862#bib.bib12)\)\.

Automatic processing of ÍTM has not taken place so far and this paper represents one of the first steps in that direction\. The task at hand, isolated sign language recognition \(ISLR\), is not the same as translation, and therefore less practical than a translation system\. Nevertheless, for a low\-resource sign language such as ÍTM, ISLR is a good starting point\. With that in mind, this paper explores how well ISLR can perform in a real low\-resource setting, on limited ÍTM data\.

The contributions of this paper include:

1. 1\.The first\-ever ISLR experiments on ÍTM data
2. 2\.Comparison of SPOTER and OpenHands, two open\-source frameworks for ISLR 1. \(a\)SPOTER yields significantly better baseline results than OpenHands on all three tasks, outperforming it by 36\.36, 14\.53 and 6\.95 in test accuracy 2. \(b\)Multilingual training in OpenHands improves the test performance by 27\.8 and 27\.45 percentage points 3. \(c\)Pretraining on ASL data and finetuning on ÍTM data improves SPOTER test performance by 22\.73, 23\.93 and 14\.25 percentage points 4. \(d\)MediaPipe pose estimates in SPOTER performed better than AlphaPose and SDPose pose estimates

## 2Related Work

Automatic sign language recognition \(SLR\) can be split into two tasks,continuousorisolatedSLR\([Koller, 2020](https://arxiv.org/html/2609.25862#bib.bib4)\)\. In continuous SLR \(CSLR\), the recognition is performed either directly on the continuous stream or with an intermediate segmentation step, where sign boundaries are identified and the recognition is then performed on isolated signs based on the assumption that a single sign/gloss is contained within the segment boundaries\. Isolated SLR \(ISLR\) is performed on single sign input\.

The recognition is then treated as a classification task where information is extracted from signed video input that is processed into a representation suitable for a downstream task\([Way et al\., 2024](https://arxiv.org/html/2609.25862#bib.bib26)\)\. The output labels can be, for instance, words or lemmas in the corresponding spoken language or sign language glosses\.

### 2\.1Video\-based Recognition

As observed by[Way et al\. \(2024\)](https://arxiv.org/html/2609.25862#bib.bib26), in the wake of deep learning, the field of SLR has witnessed a surge of new techniques, models and model architectures\. One of the first end\-to\-end approaches to SLR was the work of[Camgoz et al\. \(2017\)](https://arxiv.org/html/2609.25862#bib.bib25), who drew on speech recognition and the use of Connectionist Temporal Classification \(CTC\) algorithms\. The novel architecture they proposed for sequence\-to\-sequence learning, calledSubUNets, consisted of three layers: a CNN layer to extract spatial features from image input, bidirectional LSTM layers to temporally model those spatial features, and a CTC loss layer on top\. This was followed by theSign Language Recognition Transformer\([Camgoz et al\., 2020](https://arxiv.org/html/2609.25862#bib.bib24)\), a unified model which was trained to jointly learn CSLR and translation\.

### 2\.2Pose\-based Recognition

Pose estimation provides a lower\-dimensional alternative to raw video\. Pose estimators extract human skeletal keypoints, creating signer\-invariant representations that abstract over clothing, skin color, gender etc\., focusing on the most relevant parts of the signal and ignoring irrelevant RGB features in the video\. Despite these abstractions, the poses remain human\-interpretable, but are not entirely anonymous\([Battisti et al\., 2024](https://arxiv.org/html/2609.25862#bib.bib23)\)\. Previous work has predominantly used two pose estimators: OpenPose\([Cao et al\., 2021](https://arxiv.org/html/2609.25862#bib.bib22)\)and MediaPipe\([Lugaresi et al\., 2019](https://arxiv.org/html/2609.25862#bib.bib21)\)\. Neither was developed specifically for SLP, but given their lightweight nature compared to raw videos, they have become widespread in SLP\.

A recent and comprehensive comparison study of pose estimators is the work of[O’Brien et al\. \(2026a\)](https://arxiv.org/html/2609.25862#bib.bib16)\. They compare eight different pose estimators and evaluate them on sign language translation \(SLT\)\. They show that four pose estimators achieve higher translation scores than MediaPipe, and suggest that alternatives to MediaPipe should be more strongly considered for pose\-based SLT\. Well\-performing and interesting alternatives include Sapiens\([Khirodkar et al\., 2024](https://arxiv.org/html/2609.25862#bib.bib15)\)and SDPose\([Liang et al\., 2026](https://arxiv.org/html/2609.25862#bib.bib8)\)– which are the best\-performing estimators, but have significant compute requirements – and AlphaPose\([Fang et al\., 2023](https://arxiv.org/html/2609.25862#bib.bib14)\), which had the third\-best translation performance and the fastest compute speed\.

### 2\.3Publicly Available ISLR Frameworks

#### SPOTER

TheSign Pose\-based Transformer\([Boháček and Hrúz, 2022](https://arxiv.org/html/2609.25862#bib.bib20)\),222[https://github\.com/maty\-bohacek/spoter](https://github.com/maty-bohacek/spoter)proposed using pose estimates from Apple’s Vision API to train a word\-level SLR model\. Along with a novel normalization scheme, they presented new data augmentation methods, namely four different spatial augmentations and adding Gaussian noise, which gave the best results\. In addition, they demonstrated how well the SPOTER model performs on limited training data compared to an I3D model in the same setting\. Their code and data availability has enabled follow\-up works such as[Azad and Rahman \(2026\)](https://arxiv.org/html/2609.25862#bib.bib18)who introduce several modifications to the original SPOTER, use MediaPipe poses instead of Vision API, adapt the implementation to Bengali Sign Language data, and substantially improve the previous results on an existing ISLR dataset\.

[Boháček et al\. \(2022\)](https://arxiv.org/html/2609.25862#bib.bib19)improves the original SPOTER approach by replacing the Vision API poses with MediaPipe Holistic\. The poses were adapted to the landmarks used in the original SPOTER, in addition to performing hyperparameter search over augmentation parameters\. By doing so, they greatly improved the performance on WLASL100, achieving state\-of\-the\-art results for it at the time\. This code is, however, not available\.

SPOTER only implements one model architecture \(Transformer\) and neither pretraining strategies nor multilingual training setups are supported\.

#### OpenHands

Another important contribution in terms of open\-sourced code is the OpenHands library by[Selvaraj et al\. \(2022\)](https://arxiv.org/html/2609.25862#bib.bib17),333[https://github\.com/AI4Bharat/OpenHands](https://github.com/AI4Bharat/OpenHands)which aimed to exploit key low\-resource language insights from traditional NLP for the benefit of sign languages\. The open\-source framework they released is no longer maintained, but includes pose estimates of existing datasets for five sign languages and more than 1,000 hours of Indian Sign Language data for self\-supervised training\.

OpenHands supports four model variants: two sequence\-based models, RNN and Transformer, and two graph\-based models, a spatio\-temporal graph convolutional network \(ST\-GCN\) and a sign language GCN \(SL\-GCN\)\. Based on the results reported by[Selvaraj et al\. \(2022\)](https://arxiv.org/html/2609.25862#bib.bib17), the graph\-based models outperformed the sequence\-based models on all of the datasets where accuracy was reported, with the SL\-GCN performing best overall\.

The framework also provides a multilingual training configuration that allows combining data from multiple sign languages\. Two vocabulary strategies are available: a unified vocabulary, where the original classes from each dataset are “normalized” into English glosses, or a combined vocabulary of the original vocabulary of each dataset\.

In its default configuration, OpenHands supports 11 datasets for 7 sign languages \(American, Chinese, Indian, Turkish, Greek, Argentinian and German\)\.444See[https://openhands\.ai4bharat\.org/en/latest/instructions/ datasets\.html](https://openhands.ai4bharat.org/en/latest/instructions/datasets.html)for an overview of the supported datasets\.However, three of these datasets are not publicly available\. For the other datasets, the authors provide MediaPipe poses\.

## 3Dataset

We use theÍTM SignWikidataset presented in[Ingimundarson \(2026\)](https://arxiv.org/html/2609.25862#bib.bib2)\. It is based on the dictionary component of the Icelandic part ofSignWiki,555[https://is\.signwiki\.org/](https://is.signwiki.org/index.php/Fors%C3%AD%C3%B0a)a web and mobile platform for sign languages and deaf education\. The dictionary is bilingual, where ÍTM sign language videos are mapped to Icelandic words or phrases, and currently contains roughly 13,000 signs and phrases\. Although the dictionary is open for public input, the vast majority of the recordings originate from the Communication Centre for the Deaf and Hard of Hearing in Iceland \(SHH\), contributed by people who have been or are currently employed there, and the recordings date from the early 1990s and up until the present day\.

### 3\.1Dataset Profile

An overview of the dataset is given in Tables[1](https://arxiv.org/html/2609.25862#S3.T1)and[2](https://arxiv.org/html/2609.25862#S3.T2)\. As the numbers illustrate, the dataset is very limited in terms of size\. Instead of having many examples on a limited set of classes, there are limited examples for many classes\. This is in stark contrast to most ISLR datasets, such as the subsets of the MS\-ASL dataset\([Joze and Koller, 2019](https://arxiv.org/html/2609.25862#bib.bib10)\), that range from 100–1,000 classes and have, on average, 25\.5–57\.4 samples per class \(189–222 signers\)\.

Table 1:Overall statistics of ÍTM SignWiki datasetTable 2:Samples per class distribution in the ÍTM SignWiki dataset\. Most classes represented by two samples![Refer to caption](https://arxiv.org/html/2609.25862v1/leftright_example.png)Figure 1:Example of left\-/right\-handed contrast in the ÍTM SignWiki datasetAs far as the signers in the data are concerned, 84\.77% are L1 deaf, 10\.30% are hearing and 4\.93% are L2 deaf\. It is also worth noting that 89\.7% of the signers are female and 10\.3% are male, and all are white, which gives an indication of the overall composition of the data\. As the data has been collected over many years, some of the more frequently featured signers have several different appearances, i\.e\. wearing different clothing, with different hairstyles etc\. This would increase the diversity in a video\-/RGB\-based approach, but is less relevant for pose\-based methods\.

A further characteristic of the dataset is the fact that the second\-most frequently occurring signer in the dataset \(Signer 1\) is left\-handed, the only one out of the 33 signers\. The signer appears in 267 out of 1,845 samples\. Therefore, many of the classes have a left\-/right\-hand contrast in addition to the signer difference between training and test samples\. In total, 1,075 instances are labeled as right\-handed, 202 left\-handed, and 568 samples are two\-handed symmetrical signs\. An example of left\-/right\-handed contrast is shown in Figure[1](https://arxiv.org/html/2609.25862#S3.F1)\.

It is worth noting that a few instances of phrases, as opposed to isolated signs, are included in the data, e\.g\.hvað heitir þú\(‘what is your name’\)\. The dataset also contains multiple instances of other classes that consist of two or even three signs, such as the wordsérkennari‘special needs teacher’, composed of the signs SÉR \(‘special’\), KENNA \(‘teach’\), and PERSÓNA \(‘person’\)\. Furthermore, seven of the 849 classes are fingerspelled\. It must also be noted that the vocabulary of the dataset is unsystematic and contains, for example, signs for the wordsAlgeria,browned potatoes,Lord\(in Christian sense\),featherandhippopotamus\.

### 3\.2Data splits and recognition tasks

Table 3:Recognition tasks based on the ÍTM SignWiki dataset, varying number of classes in the training, validation and test sets#### Data splits

We split the full dataset into a training, validation and test set with 917, 79 and 849 samples, respectively\. The difference in size between the validation and test splits is due to dataset constraints\. For 38 of the 117 classes that have three samples or more, two samples were retained for training instead of having a validation sample\. In many cases, the validation sample is a different recording by the same signer as the training sample, whereas the test sample is drawn from a different signer not represented for that class\. Therefore, the performance on the validation set might be misleading with regard to test set performance\. This is preferable to validating on an unseen signer while testing on a seen signer for the same class\.

#### Tasks

We define three different recognition tasks with increasing difficulty \(varying random chance of success and class balance of training and validation data\), see Table[3](https://arxiv.org/html/2609.25862#S3.T3)\. TheMinimaltask entails recognizing 22 classes \(with≥\\geq4 samples per class\), theTrimmedtask has 117 classes \(with≥\\geq3 samples per class\) and theFulltask has 849 classes \(with≥\\geq2 samples per class\)\. Moving to the next task means adding more, and more challenging, classes, with the Full task effectively being one\-shot recognition over 849 classes\.

## 4Experiments

We train a series of baselines with OpenHands and SPOTER using only the ÍTM SignWiki dataset \(Section[4\.1](https://arxiv.org/html/2609.25862#S4.SS1)\)\. Then we perform additional experiments on multilingual training \(Section[4\.2](https://arxiv.org/html/2609.25862#S4.SS2)\) and a pretraining/finetuning scheme \(Section[4\.3](https://arxiv.org/html/2609.25862#S4.SS3)\) where the training data includes other languages as well\.

### 4\.1Baselines

#### OpenHands

We train systems for all four model variants that OpenHands supports: two sequence\-based models, RNN and Transformer \(BERT\), and two graph\-based models, a spatio\-temporal graph convolutional network \(ST\-GCN\) and a sign language GCN \(SL\-GCN\) \(see Section[2\.3](https://arxiv.org/html/2609.25862#S2.SS3)\)\.

The framework provides precomputed poses for eight ISLR datasets \(listed in Table[8](https://arxiv.org/html/2609.25862#A1.T8)\), but also a pipeline to extract MediaPipe poses from video data\. We used this pipeline on the ÍTM data\. In the native OpenHands setup, the authors then define two different keypoint presets for the pose estimates:Minimal\(27 2D keypoints of the upper\-body, hands and face\) orTop body\(59 keypoints with better coverage of the upper body\)\. Following[Selvaraj et al\. \(2022\)](https://arxiv.org/html/2609.25862#bib.bib17)’s example scripts we only used the minimal preset in all of our experiments\.

In addition to trying different model architectures, we evaluate on all three recognition tasks \(see Section[3\.2](https://arxiv.org/html/2609.25862#S3.SS2)\) and either disable or enable data augmentation\. Taken together, we train4×3×2=244\\times 3\\times 2=24baseline OpenHands models\.

In the native OpenHands setup, the validation accuracy is monitored for early stopping, and we followed this example \(max epochs=500/1000, patience=80 mode=max\)\. Further hyperparameters were the following: CosineAnnealingLR scheduler with Adam, lr =1​e−31e\-3, and batch size 4/8/16 forMinimal/Trimmed/Full\. Models were trained on an A100 or H100 GPU \(this is true for all models in our experiments\)\.

#### SPOTER

The best results with SPOTER were achieved with MediaPipe poses as the input representation \(see Section[2\.3](https://arxiv.org/html/2609.25862#S2.SS3)\)\. For this paper, we adapted the SPOTER code to the binary pose format developed by[Moryossef et al\. \(2021a\)](https://arxiv.org/html/2609.25862#bib.bib11)\. Inspired by[O’Brien et al\. \(2026a\)](https://arxiv.org/html/2609.25862#bib.bib16)we compare three different pose estimators: AlphaPose\([Fang et al\., 2023](https://arxiv.org/html/2609.25862#bib.bib14)\), MediaPipe\([Lugaresi et al\., 2019](https://arxiv.org/html/2609.25862#bib.bib21)\)and SDPose\([Liang et al\., 2026](https://arxiv.org/html/2609.25862#bib.bib8)\), instead of assuming MediaPipe as a fixture of the experiment\.

As described in Section[2](https://arxiv.org/html/2609.25862#S2), MediaPipe and OpenPose have been the two predominant pose estimators used in SLP in recent years\. AlphaPose and SDPose are lesser known in an SLP context\. In[O’Brien et al\. \(2026a\)](https://arxiv.org/html/2609.25862#bib.bib16)’s comparison of pose estimators, evaluated on SLT, AlphaPose ranked third, and was found to be more computationally efficient than the better\-performing estimators\. It is primarily the low\-latency benefit of AlphaPose that makes it appealing for the purposes of this project\. While working with a low\-resource language does not necessarily entail limited computational resources, a lightweight solution with competitive performance would be advantageous\.

To extract the poses, we used thevideo\-to\-poserepository by[O’Brien et al\. \(2026b\)](https://arxiv.org/html/2609.25862#bib.bib5),666[https://github\.com/ZurichNLP/video\-to\-pose](https://github.com/ZurichNLP/video-to-pose)which at the time of writing supports eight pose estimators in total\. In the SPOTER implementation, a total of 54 body landmarks are extracted, including five head landmarks \(eyes, ears, and nose\) and 21 body landmarks that represent body joints\. This results in 54 2D points and an 108 dimensional feature vector for each frame\. Following[Boháček and Hrúz \(2022\)](https://arxiv.org/html/2609.25862#bib.bib20)we train for 350 epochs and add Gaussian noise to the training set \(along with other data augmentation techniques the framework provides\)\.

To summarize, we trained SPOTER baselines for all three recognition tasks \(see Section[3\.2](https://arxiv.org/html/2609.25862#S3.SS2)\) and for three different pose estimators \(3×3=93\\times 3=9baseline SPOTER models\)\.

### 4\.2Multilingual training

We train additional OpenHands models on more datasets in other sign languages \(see Table[8](https://arxiv.org/html/2609.25862#A1.T8)in Appendix[A](https://arxiv.org/html/2609.25862#A1)for the full list of datasets\)\.

#### Vocabulary strategies

We experiment with either keeping the vocabularies of all datasets separate \(original\) or unifying them into a single, normalized vocabulary \(unified\)\. In the first setting, the language code \(ISO\) of each sign language included is prepended to the class label as a one\-hot encoded vector, i\.e\.ice\_\_ for ÍTM\.

In the second setting, the vocabulary of each non\-English/non\-ASL dataset is normalized to English glosses\. Here, the ÍTM subset has only 817 classes as the normalization allows to combine different sign variants of the same sign, of which there are several examples in the dataset, with the same normalized gloss\. See Appendix[A](https://arxiv.org/html/2609.25862#A1)for a more detailed explanation\.

This results in a single normalized vocabulary and predictions are therefore made with normalized glosses, which can then be mapped back to the original vocabulary, but not to separate variants\. As the unified vocabulary experiment effectively is a different version of theFulltask, we keep those results separate and present them in Appendix[A](https://arxiv.org/html/2609.25862#A1)\.

These experiments use the multilingual training feature of OpenHands, which, is neither described in the documentation nor in the paper\. The code is, however, nearly fully implemented, and the experiments were inspired by example configs for multilingual training provided with the framework\.777See[https://github\.com/AI4Bharat/OpenHands/tree/main/\-examples/configs/multilingual](https://github.com/AI4Bharat/OpenHands/tree/main/examples/configs/multilingual)We used the following hyperparameters: CosineAnnealingLR scheduler with Adam, learning rate1​e−31e\-3, batch size 64, max epochs 200, with early stopping on validation accuracy \(max, patience 30\)\.

Due to time and resource constraints, we evaluate these multilingual models only on theFullrecognition task\.

### 4\.3Pretraining/finetuning scheme

SPOTER does not support pretraining out\-of\-the\-box\. For this paper, we adapted the framework to support pretraining \(either with the encoder frozen for a specified number of epochs or unfrozen from the start\) and finetuning\.

For this experiment, a SPOTER model was pre\-trained on the ASL Citizen dataset\([Desai et al\., 2023](https://arxiv.org/html/2609.25862#bib.bib9)\)\. It is a community\-sourced dataset for ASL that has nearly 84,000 videos filmed by 52 signers and 2,731 classes\. One motivation for using it was that MediaPipe poses in pose format for the dataset were already available, though any comparable ISLR dataset for a higher\-resource sign language would have been a viable alternative\. As MediaPipe outperformed the other two pose estimators in the baseline experiments \(see Section[5\.1](https://arxiv.org/html/2609.25862#S5.SS1)\), this setup was exclusively tested with MediaPipe poses\. As the aim here was simply for the encoder to learn general representations of ASL and not to achieve the best results on the dataset, the training was limited to 30 epochs \(test accuracy 40\.28\)\.

For the finetuning we tested two approaches: one where the encoder was frozen for the first 30 epochs and only the decoder and head trained, and another where the full model was finetuned from the beginning\.

Thus the model is pre\-trained on ASL data and finetuned on ÍTM, with the encoder either frozen from the start or unfrozen; exclusively with MediaPipe poses \(3×23\\times 2experiments\)\.

### 4\.4Evaluation Metrics

We report test accuracy and validation accuracy, given that the validation set is somewhat particular and limited in two ways \(see Section[3\.2](https://arxiv.org/html/2609.25862#S3.SS2)\): on the one hand it covers only 79 classes \(compared to the total of 849\) and on the other hand, the validation samples are mostly with the same signer as \(one of\) the training sample\(s\), whereas the signer in the test set is always unseen for that particular sign\.

## 5Results

This section presents the initial ISLR experiments we performed with the two frameworks\. An overview of the best results is shown in Table[7](https://arxiv.org/html/2609.25862#S5.T7)and an additional table for the multilingual results is included in Appendix[A](https://arxiv.org/html/2609.25862#A1)\.

### 5\.1Baselines \(ÍTM Data Only\)

#### OpenHands

Table 4:Accuracy of OpenHands baseline models trained on ÍTM data only \(✓=with augmentation, \-=only normalization\)Table[4](https://arxiv.org/html/2609.25862#S5.T4)shows the performance of all OpenHands baselines\. The graph\-based models generally outperform the LSTM and Transformer models, with a GCN variant achieving the best test accuracy of 13\.64% / 9\.40% / 1\.41% on theMinimal/Trimmed/Fulltasks respectively\. Only one configuration yielded more than 10% accuracy on theMinimaltask; an ST\-GCN model without any data augmentation\. For reference, the best OpenHands baseline model is repeated in Table[7](https://arxiv.org/html/2609.25862#S5.T7)\.

Across the baselines, disabling augmentation yields slightly better test accuracy in 8 of 12 pairs and in general higher validation accuracy\. However, since the test sets are rather small \(22, 117 and 849 examples\) these are in fact minor differences\.

Table 5:Accuracy of SPOTER baseline models trained on ÍTM data only, varying the pose estimation system
#### SPOTER

Table[5](https://arxiv.org/html/2609.25862#S5.T5)shows the performance of all SPOTER baselines\. MediaPipe poses consistently outperforms other estimators, for example the test accuracy of MediaPipe is roughly 5 percentage points higher than AlphaPose on all three tasks\. In general, the SPOTER baselines show considerably higher accuracy on the test set than comparable OpenHands baselines \(see above\)\. For reference, the best SPOTER baseline model \(only MediaPipe results\) is repeated in Table[7](https://arxiv.org/html/2609.25862#S5.T7)\.

### 5\.2Multilingual Training

Multilingual training results with the original vocabulary are shown in Table[7](https://arxiv.org/html/2609.25862#S5.T7), for a direct comparison with baseline scores\. Individual per\-dataset scores are reported in Table[8](https://arxiv.org/html/2609.25862#A1.T8)in Appendix[A](https://arxiv.org/html/2609.25862#A1)\. Multilingual training results only concern theFulltask\. Adding multilingual training outperforms the best OpenHands and SPOTER baselines\. For example, the test accuracy of the best SPOTER baseline is 8\.36, while the test accuracy of the best multilingual model is 28\.86\. Unifying the multilingual vocabulary \(as opposed to keeping separate vocabularies\) yields higher accuracy, 29\.21, but on a slightly smaller vocabulary \(see Section[A\.1](https://arxiv.org/html/2609.25862#A1.SS1)\)\.

### 5\.3Pretraining / finetuning scheme

Table 6:Test accuracy of a SPOTER model pre\-trained on ASL and finetuned on ÍTM compared to the baseline in Table[5](https://arxiv.org/html/2609.25862#S5.T5)\(Frozen=Pre\-trained encoder frozen for first 30 epochs, Unfrozen=Full model finetuned from start\)TaskRandomFrameworkModelSettingValTestTimeMinimal4\.54OpenHandsST\-GCNBaseline \(No Augmentation\)40\.913\.6400:01:54SPOTERTransformerBaseline \(MP\)72\.7250\.0000:02:18SPOTERTransformerASL\-Finetuned \(MP\)\-Frozen95\.4572\.73\*05:22:24Trimmed0\.85OpenHandsSL\-GCNBaseline \(No Augmentation\)59\.59\.4000:07:39SPOTERTransformerBaseline \(MP\)72\.1523\.9300:18:06SPOTERTransformerASL\-Finetuned \(MP\)\-Frozen82\.2747\.86\*05:25:13Full0\.12OpenHandsSL\-GCNBaseline \(Augmentation\)17\.701\.4101:14:32OpenHandsSL\-GCNMultilingual Original68\.228\.8608:15:16SPOTERTransformerBaseline \(MP\)67\.098\.3601:31:25SPOTERTransformerASL\-Finetuned \(MP\)\-Frozen78\.4822\.61\*06:07:20Table 7:Validation and test accuracy on the three ÍTM recognition tasks for OpenHands and SPOTER\. \(Baseline=best baseline score, Random=Random classification accuracy, MP=MediaPipe Holistic poses, Time=Training time as HH:MM:SS, \*=Combined training time \(ASL Citizen pretraining took 05:21:11 hours\)Pretraining / finetuning results are shown in full in Table[6](https://arxiv.org/html/2609.25862#S5.T6)\. Pretraining on ASL data also outperforms the best baselines by at least 10 percentage points in test accuracy\. For instance, on theFulltask, the best SPOTER baseline achieves 8\.36 test accuracy, while the best finetuned model achieves 22\.61 accuracy\. Furthermore the results demonstrate that freezing the pre\-trained encoder at the beginning of finetuning increases the test accuracy by at least 10 percentage points\.

## 6Discussion

#### Baselines: SPOTER vs\. OpenHands

When only using ÍTM data, SPOTER clearly outperforms Openhands on all three tasks: 50\.00 vs\. 13\.64, 23\.93 vs\. 9\.40 and 8\.36 vs\. 1\.41 test accuracy \(see Section[5\.1](https://arxiv.org/html/2609.25862#S5.SS1)\)\. These margins should be read with the size of the test sets in mind\. On theMinimaltask each prediction is worth1/22=4\.541/22=4\.54percentage points, and the gap amounts to 11 vs\. 3 correct predictions; onTrimmedit is 28 vs\. 11 out of 117\. TheFulltask, with 71 vs\. 12 correct predictions out of 849, therefore carries most of the evidence, and the smaller tasks should not be over\-interpreted\.

We emphasize that this is a comparison of two frameworks as they are distributed, not of two architectures\. Input representation, normalization, augmentation, optimization and checkpoint selection all differ at the same time, and our experiments do not isolate these factors\. We therefore offer hypotheses rather than explanations\.

For example, OpenHands and SPOTER reduce the full set of MediaPipe keypoints in different ways\. OpenHands offers 27\-point and 59\-point presets, but our OpenHands models use only the 27 point preset, following[Selvaraj et al\. \(2022\)](https://arxiv.org/html/2609.25862#bib.bib17)\. SPOTER, on the other hand, uses 54 2D keypoints\. This difference in keypoint resolution may in part explain the difference in performance between the frameworks: a coarser hand representation could plausibly mean a disadvantage for OpenHands\. Re\-running OpenHands with the 59\-point preset would test this directly\.

Second, the frameworks normalize differently\. OpenHands applies a single shoulder\-referenced centering and scaling to the whole skeleton, so hand landmarks occupy a small region of the normalized space, whereas SPOTER normalizes body and hands separately, distorting hand keypoints to a lesser degree\. Normalization is intimately tied to generalization; because the test signer is always unseen for a given class \(Section[3\.2](https://arxiv.org/html/2609.25862#S3.SS2)\), our test set specifically measures signer\-invariant generalization\. Consistent with this, SPOTER retains a much larger share of its validation accuracy on the test set \(69%/33%/12% across the three tasks\) than the best OpenHands models do \(33%/16%/8%\)\. Both frameworks overfit to the seen signer; OpenHands does so considerably more\. The importance of preprocessing choices of this kind for pose\-based SLP has been noted before\([Coster et al\., 2023](https://arxiv.org/html/2609.25862#bib.bib1);[O’Brien et al\., 2026a](https://arxiv.org/html/2609.25862#bib.bib16)\)\.

Third, the training and model selection protocols differ\. SPOTER trains for a fixed number of epochs, saves the two checkpoints with highest training and validation accuracy every ten epochs, and evaluates over all of them\. Our OpenHands runs use early stopping and checkpoint selection on validation accuracy, which on theFulltask is computed over 79 of 849 classes, and a single checkpoint was then chosen manually fromkksaved checkpoints\. Model selection is thus both noisier and more weakly related to the target task for OpenHands, and we did not evaluate all saved checkpoints for comparison\.

Fourth, and in our view most informative, the graph\-based models appear to be data\-starved rather than unsuited to the task\. The same SL\-GCN that reaches 1\.41 on theFulltask with ÍTM data alone reaches 28\.86 once other sign languages are added \(Section[5\.2](https://arxiv.org/html/2609.25862#S5.SS2)\)\.[Selvaraj et al\. \(2022\)](https://arxiv.org/html/2609.25862#bib.bib17)report their graph models performing best on datasets with tens of samples per class; ÍTM SignWiki offers one or two\. In this respect our results are in line with[Boháček and Hrúz \(2022\)](https://arxiv.org/html/2609.25862#bib.bib20), who show SPOTER learning effectively from small training sets, although their comparison was against a video\-based I3D model that must first learn general properties of human motion, whereas both frameworks compared here operate on poses\. Our results are therefore consistent with their claim but do not test it under the same conditions\.

Finally, although both frameworks use MediaPipe, the poses were extracted with different pipelines: the extraction script shipped with OpenHands in one case andvideo\-to\-pose\([O’Brien et al\., 2026b](https://arxiv.org/html/2609.25862#bib.bib5)\)in the other\. Differences in version or configuration cannot be ruled out as a contributing factor\.

#### SPOTER: Choice of pose estimator

MediaPipe clearly outperformed AlphaPose and SDPose\. It is worth reiterating that this was not a comparison of the full pose estimates but of the customized \(reduced\) SPOTER format\. However, the format was the same for all three estimators and the training schemes identical\. As mentioned in Section[2](https://arxiv.org/html/2609.25862#S2), a comparison study of pose estimators performed by[O’Brien et al\. \(2026a\)](https://arxiv.org/html/2609.25862#bib.bib16)revealed that three other pose estimators performed better than MediaPipe on SLT\. In our experiments, the downstream task is different, recognition instead of translation, and the set of keypoints is reduced\. Nevertheless, MediaPipe appears better suited to ISLR\. This is in line with findings of the comparative studies of[Coster et al\. \(2023\)](https://arxiv.org/html/2609.25862#bib.bib1)and[Moryossef et al\. \(2021b\)](https://arxiv.org/html/2609.25862#bib.bib13)of pose estimators for SLR, where MediaPipe outperformed OpenPose and MMPose\. Our results are, to the best of our knowledge, the first comparison of MediaPipe, AlphaPose and SDPose for SLR\.

#### Cross\-lingual transfer effects

Both multilingual training and fine\-tuning a pretrained model improved the performance, meaning that both constitute genuine cross\-lingual transfer effects from higher\-resourced sign languages to a low\-resource one\. Multilingual training raises the best OpenHands result on theFulltask from 1\.41 to 28\.86, a twentyfold increase, while ASL pretraining raises the best SPOTER result from 8\.36 to 22\.61, and by 14–24 percentage points across the three tasks\. On theFulltask the multilingual model is ahead \(245 vs\. 192 correct predictions of 849\)\.

These two numbers should not be read as a ranking of the two strategies\. They come from different frameworks, whose baselines already differ by a factor of six; from different source data, eight datasets in six sign languages in one case and a single ASL dataset in the other; and from different label spaces at inference\. The multilingual model predicts over all 5,732 classes of the concatenated dataset, and 137 of its 849 test predictions fall on classes belonging to other sign languages \(see Appendix[B](https://arxiv.org/html/2609.25862#A2)\); these are incorrect by construction, so its ÍTM accuracy is measured under a handicap that the pretrained model does not face\. A controlled comparison would require training SPOTER jointly on ASL Citizen and ÍTM and evaluating both schemes within one framework, and the multilingual configuration on the two smaller tasks, neither of which we were able to do\.

## 7Conclusion

In this paper we explore the performance of two publicly available SLR frameworks in a low\-resource setting, using the first ISLR dataset for ÍTM\. The dataset is very limited with regard to the intended use in a machine learning task; it has many classes \(849\) and few samples per class, only two samples for the majority of classes \(732\)\. The experiments therefore explored how well SLR can perform in a real low\-resource setting\.

We demonstrate that multilingual learning and cross\-lingual transfer can benefit lower\-resource sign languages\. In the OpenHands experiments, the multilingual training scheme achieved the best results on the full dataset\. In the case of SPOTER, we show that SPOTER learns effectively from very small training sets, corroborating the findings of[Boháček and Hrúz \(2022\)](https://arxiv.org/html/2609.25862#bib.bib20)\. Additionally, finetuning on ÍTM after pretraining on ASL Citizen data yielded the best overall results on the original ÍTM vocabulary on two out of three tasks and demonstrates the effectiveness of a strong pre\-trained checkpoint that can be finetuned to different tasks\. The multilingual model with original vocabulary in OpenHands yields the best performance on the full \(most challenging\) task\.

## 8Limitations

#### Framework choice

Both of these frameworks are from 2022 and comparison with at least one newer ISLR method would have been preferable, but no more recent, publicly available code could be found, except extensions of SPOTER\.

#### Error analysis

Beyond general accuracy measures, we present no error analysis\. Given the composition of the dataset, it could for instance be insightful to analyze errors based on sign types, e\.g\. one\-handed vs\. two\-handed, fingerspelled signs and multi\-sign compounds \(see Section[3\.1](https://arxiv.org/html/2609.25862#S3.SS1)\), or per\-signer performance\.

#### Baseline tuning

The discrepancy in performance in the OpenHands baselines with regard to data augmentation, where the performance without augmentation was in general better, was unexpected and would need further inspection\. The augmentation techniques used followed the examples of the OpenHands authors, but could perhaps be adjusted better to the ÍTM dataset\.

#### Variability

We only present one single run for each training configuration\. Training several models with different random seeds would make our results and conclusions drawn from them more robust\.

#### Applicability

Although the experiments yielded meaningful results it should nonetheless be stressed that the results do not yet constitute practically applicable performance\. ISLR systems can be applied to tasks such as looking for signs in videos or dictionaries, but none of the models we trained could be immediately deployed in such a setting\. This would require substantially greater amounts of training data as well as a larger vocabulary\.

## 9Acknowledgements

We would like to thank three anonymous reviewers for useful comments\. MM received funding from the SIGMA project \(grant no\. G\-95017\-01\-07\), supported by the Digital Society Initiative \(DSI\) at the University of Zurich\.

The language of the abstract and discussion sections was refined with a Claude agent\.

## References

- Azad and Rahman \(2026\)S\. I\. Azad and Md\. A\. RahmanBdSL\-SPOTER: A Transformer\-Based Framework for Bengali Sign Language Recognition with Cultural Adaptation\.InAdvances in Visual Computing,pp\. 304–315\.External Links:ISBN 9783032144959,ISSN 1611\-3349,[Link](http://dx.doi.org/10.1007/978-3-032-14495-9_23),[Document](https://dx.doi.org/10.1007/978-3-032-14495-9%5F23)Cited by:[§2\.3](https://arxiv.org/html/2609.25862#S2.SS3.SSS0.Px1.p1.1)\.
- Battistiet al\.\(2024\)A\. Battisti, E\. van den Bold, A\. Göhring, F\. Holzknecht, and S\. EblingPerson Identification from Pose Estimates in Sign Language\.InProceedings of the LREC\-COLING 2024 11th Workshop on the Representation and Processing of Sign Languages: Evaluation of Sign Language Resources,E\. Efthimiou, S\. Fotinea, T\. Hanke, J\. A\. Hochgesang, J\. Mesch, and M\. Schulder \(Eds\.\),Torino, Italia,pp\. 13–25\.External Links:[Link](https://aclanthology.org/2024.signlang-1.2/)Cited by:[§2\.2](https://arxiv.org/html/2609.25862#S2.SS2.p1.1)\.
- Boháčeket al\.\(2022\)M\. Boháček, Z\. Cao, and M\. HrúzCombining Efficient and Precise Sign Language Recognition: Good pose estimation library is all you need\.External Links:2210\.00893,[Link](https://arxiv.org/abs/2210.00893)Cited by:[§2\.3](https://arxiv.org/html/2609.25862#S2.SS3.SSS0.Px1.p2.1)\.
- Boháček and Hrúz \(2022\)M\. Boháček and M\. HrúzSign Pose\-based Transformer for Word\-level Sign Language Recognition\.In2022 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops \(WACVW\),Vol\.,pp\. 182–191\.External Links:[Document](https://dx.doi.org/10.1109/WACVW54805.2022.00024)Cited by:[§2\.3](https://arxiv.org/html/2609.25862#S2.SS3.SSS0.Px1.p1.1),[§4\.1](https://arxiv.org/html/2609.25862#S4.SS1.SSS0.Px2.p3.1),[§6](https://arxiv.org/html/2609.25862#S6.SS0.SSS0.Px1.p6.1),[§7](https://arxiv.org/html/2609.25862#S7.p2.1)\.
- Camgozet al\.\(2017\)N\. C\. Camgoz, S\. Hadfield, O\. Koller, and R\. BowdenSubUNets: End\-to\-End Hand Shape and Continuous Sign Language Recognition\.In2017 IEEE International Conference on Computer Vision \(ICCV\),Vol\.,pp\. 3075–3084\.External Links:[Document](https://dx.doi.org/10.1109/ICCV.2017.332)Cited by:[§2\.1](https://arxiv.org/html/2609.25862#S2.SS1.p1.1)\.
- Camgozet al\.\(2020\)N\. C\. Camgoz, O\. Koller, S\. Hadfield, and R\. BowdenSign Language Transformers: Joint End\-to\-end Sign Language Recognition and Translation\.External Links:2003\.13830,[Link](https://arxiv.org/abs/2003.13830)Cited by:[§2\.1](https://arxiv.org/html/2609.25862#S2.SS1.p1.1)\.
- Caoet al\.\(2021\)Z\. Cao, G\. Hidalgo, T\. Simon, S\. Wei, and Y\. SheikhOpenPose: Realtime Multi\-Person 2D Pose Estimation Using Part Affinity Fields\.IEEE Transactions on Pattern Analysis and Machine Intelligence43\(1\),pp\. 172–186\.External Links:[Document](https://dx.doi.org/10.1109/TPAMI.2019.2929257)Cited by:[§2\.2](https://arxiv.org/html/2609.25862#S2.SS2.p1.1)\.
- Costeret al\.\(2023\)M\. D\. Coster, E\. Rushe, R\. Holmes, A\. Ventresque, and J\. DambreTowards the extraction of robust sign embeddings for low resource sign language recognition\.External Links:2306\.17558,[Link](https://arxiv.org/abs/2306.17558)Cited by:[§6](https://arxiv.org/html/2609.25862#S6.SS0.SSS0.Px1.p4.1),[§6](https://arxiv.org/html/2609.25862#S6.SS0.SSS0.Px2.p1.1)\.
- Desaiet al\.\(2023\)A\. Desai, L\. Berger, F\. O\. Minakov, V\. Milan, C\. Singh, K\. Pumphrey, R\. E\. Ladner, H\. Daumé, A\. X\. Lu, N\. Caselli, and D\. BraggASL Citizen: A Community\-Sourced Dataset for Advancing Isolated Sign Language Recognition\.InProceedings of the 37th International Conference on Neural Information Processing Systems,NIPS ’23,Red Hook, NY, USA\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/f29cf8f8b4996a4a453ef366cf496354-Abstract-Datasets_and_Benchmarks.html)Cited by:[§4\.3](https://arxiv.org/html/2609.25862#S4.SS3.p2.1)\.
- Fanget al\.\(2023\)H\. Fang, J\. Li, H\. Tang, C\. Xu, H\. Zhu, Y\. Xiu, Y\. Li, and C\. LuAlphaPose: Whole\-Body Regional Multi\-Person Pose Estimation and Tracking in Real\-Time\.IEEE Trans\. Pattern Anal\. Mach\. Intell\.45\(6\),pp\. 7157–7173\.External Links:ISSN 0162\-8828,[Link](https://doi.org/10.1109/TPAMI.2022.3222784),[Document](https://dx.doi.org/10.1109/TPAMI.2022.3222784)Cited by:[§2\.2](https://arxiv.org/html/2609.25862#S2.SS2.p2.1),[§4\.1](https://arxiv.org/html/2609.25862#S4.SS1.SSS0.Px2.p1.1)\.
- Ingimundarson \(2026\)F\. Á\. IngimundarsonIsolated Sign Language Recognition for Icelandic Sign Language \(ITM\): Experiments in a Low\-resource Setting\.Master’s Thesis,University of Zurich\.External Links:[Link](https://doi.org/10.5167/uzh-436015)Cited by:[§3](https://arxiv.org/html/2609.25862#S3.p1.1),[footnote 1](https://arxiv.org/html/2609.25862#footnote1)\.
- Joshiet al\.\(2020\)P\. Joshi, S\. Santy, A\. Budhiraja, K\. Bali, and M\. ChoudhuryThe State and Fate of Linguistic Diversity and Inclusion in the NLP World\.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics,D\. Jurafsky, J\. Chai, N\. Schluter, and J\. Tetreault \(Eds\.\),Online,pp\. 6282–6293\.External Links:[Link](https://aclanthology.org/2020.acl-main.560/),[Document](https://dx.doi.org/10.18653/v1/2020.acl-main.560)Cited by:[§1](https://arxiv.org/html/2609.25862#S1.p2.1)\.
- Joze and Koller \(2019\)H\. R\. V\. Joze and O\. KollerMS\-ASL: A Large\-Scale Data Set and Benchmark for Understanding American Sign Language\.External Links:1812\.01053,[Link](https://arxiv.org/abs/1812.01053)Cited by:[§3\.1](https://arxiv.org/html/2609.25862#S3.SS1.p1.1)\.
- Khirodkaret al\.\(2024\)R\. Khirodkar, T\. Bagautdinov, J\. Martinez, S\. Zhaoen, A\. James, P\. Selednik, S\. Anderson, and S\. SaitoSapiens: Foundation for Human Vision Models\.InComputer Vision – ECCV 2024: 18th European Conference, Milan, Italy, September 29–October 4, 2024, Proceedings, Part IV,Berlin, Heidelberg,pp\. 206–228\.External Links:ISBN 978\-3\-031\-73234\-8,[Link](https://doi.org/10.1007/978-3-031-73235-5_12),[Document](https://dx.doi.org/10.1007/978-3-031-73235-5%5F12)Cited by:[§2\.2](https://arxiv.org/html/2609.25862#S2.SS2.p2.1)\.
- Koller \(2020\)O\. KollerQuantitative Survey of the State of the Art in Sign Language Recognition\.External Links:2008\.09918,[Link](https://arxiv.org/abs/2008.09918)Cited by:[§2](https://arxiv.org/html/2609.25862#S2.p1.1)\.
- Koulidobrova and Sverrisdóttir \(2021\)E\. Koulidobrova and R\. SverrisdóttirHow to Ensure Bilingualism/Biliteracy in an Indigenous Context: The Case of Icelandic Sign Language\.Languages6\(2\)\.External Links:[Link](https://www.mdpi.com/2226-471X/6/2/98),ISSN 2226\-471X,[Document](https://dx.doi.org/10.3390/languages6020098)Cited by:[§1](https://arxiv.org/html/2609.25862#S1.p3.1)\.
- Lianget al\.\(2026\)S\. Liang, J\. He, C\. Wang, L\. Liao, G\. Zhang, Y\. Chen, and Y\. YuanSDPose: Exploiting Diffusion Priors for Out\-of\-Domain and Robust Pose Estimation\.External Links:2509\.24980,[Link](https://arxiv.org/abs/2509.24980)Cited by:[§2\.2](https://arxiv.org/html/2609.25862#S2.SS2.p2.1),[§4\.1](https://arxiv.org/html/2609.25862#S4.SS1.SSS0.Px2.p1.1)\.
- Lugaresiet al\.\(2019\)C\. Lugaresi, J\. Tang, H\. Nash, C\. McClanahan, E\. Uboweja, M\. Hays, F\. Zhang, C\. Chang, M\. G\. Yong, J\. Lee, W\. Chang, W\. Hua, M\. Georg, and M\. GrundmannMediaPipe: A Framework for Building Perception Pipelines\.External Links:1906\.08172,[Link](https://arxiv.org/abs/1906.08172)Cited by:[§2\.2](https://arxiv.org/html/2609.25862#S2.SS2.p1.1),[§4\.1](https://arxiv.org/html/2609.25862#S4.SS1.SSS0.Px2.p1.1)\.
- Moryossefet al\.\(2021a\)A\. Moryossef, M\. Müller, and R\. Fahrnipose\-format: Library for viewing, augmenting, and handling \.pose files\.Note:[https://github\.com/sign\-language\-processing/pose](https://github.com/sign-language-processing/pose)Cited by:[§4\.1](https://arxiv.org/html/2609.25862#S4.SS1.SSS0.Px2.p1.1)\.
- Moryossefet al\.\(2021b\)A\. Moryossef, I\. Tsochantaridis, J\. Dinn, N\. C\. Camgöz, R\. Bowden, T\. Jiang, A\. Rios, M\. Müller, and S\. EblingEvaluating the Immediate Applicability of Pose Estimation for Sign Language Recognition\.External Links:2104\.10166,[Link](https://arxiv.org/abs/2104.10166)Cited by:[§6](https://arxiv.org/html/2609.25862#S6.SS0.SSS0.Px2.p1.1)\.
- Mülleret al\.\(2022\)M\. Müller, S\. Ebling, E\. Avramidis, A\. Battisti, M\. Berger, R\. Bowden, A\. Braffort, N\. Cihan Camgöz, C\. España\-bonet, R\. Grundkiewicz, Z\. Jiang, O\. Koller, A\. Moryossef, R\. Perrollaz, S\. Reinhard, A\. Rios, D\. Shterionov, S\. Sidler\-miserez, and K\. TissiFindings of the First WMT Shared Task on Sign Language Translation \(WMT\-SLT22\)\.InProceedings of the Seventh Conference on Machine Translation \(WMT\),P\. Koehn, L\. Barrault, O\. Bojar, F\. Bougares, R\. Chatterjee, M\. R\. Costa\-jussà, C\. Federmann, M\. Fishel, A\. Fraser, M\. Freitag, Y\. Graham, R\. Grundkiewicz, P\. Guzman, B\. Haddow, M\. Huck, A\. Jimeno Yepes, T\. Kocmi, A\. Martins, M\. Morishita, C\. Monz, M\. Nagata, T\. Nakazawa, M\. Negri, A\. Névéol, M\. Neves, M\. Popel, M\. Turchi, and M\. Zampieri \(Eds\.\),Abu Dhabi, United Arab Emirates \(Hybrid\),pp\. 744–772\.External Links:[Link](https://aclanthology.org/2022.wmt-1.71/),[Document](https://dx.doi.org/10.18653/v1/2022.wmt-1.71)Cited by:[§1](https://arxiv.org/html/2609.25862#S1.p2.1)\.
- O’Brienet al\.\(2026a\)C\. O’Brien, G\. Sant, M\. Müller, and S\. EblingEvaluation of Pose Estimation Systems for Sign Language Translation\.InProceedings of the LREC 2026 12th Workshop on the Representation and Processing of Sign Languages: Language in Motion,E\. Efthimiou, S\. Fotinea, T\. Hanke, J\. A\. Hochgesang, J\. Mesch, and M\. Schulder \(Eds\.\),Palma, Mallorca \(Spain\),pp\. 371–386\.External Links:[Link](https://aclanthology.org/2026.signlang-1.39/),[Document](https://dx.doi.org/10.63317/4shxzirxykmm)Cited by:[§1](https://arxiv.org/html/2609.25862#S1.p2.1),[§2\.2](https://arxiv.org/html/2609.25862#S2.SS2.p2.1),[§4\.1](https://arxiv.org/html/2609.25862#S4.SS1.SSS0.Px2.p1.1),[§4\.1](https://arxiv.org/html/2609.25862#S4.SS1.SSS0.Px2.p2.1),[§6](https://arxiv.org/html/2609.25862#S6.SS0.SSS0.Px1.p4.1),[§6](https://arxiv.org/html/2609.25862#S6.SS0.SSS0.Px2.p1.1)\.
- O’Brienet al\.\(2026b\)C\. O’Brien, G\. Sant, and M\. MüllerConvenience code for installing and using several pose estimation systems\.Note:[https://github\.com/ZurichNLP/video\-to\-pose](https://github.com/ZurichNLP/video-to-pose)Cited by:[§4\.1](https://arxiv.org/html/2609.25862#S4.SS1.SSS0.Px2.p3.1),[§6](https://arxiv.org/html/2609.25862#S6.SS0.SSS0.Px1.p7.1)\.
- Selvarajet al\.\(2022\)P\. Selvaraj, G\. Nc, P\. Kumar, and M\. KhapraOpenHands: Making Sign Language Recognition Accessible with Pose\-based Pretrained Models across Languages\.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),S\. Muresan, P\. Nakov, and A\. Villavicencio \(Eds\.\),Dublin, Ireland,pp\. 2114–2133\.External Links:[Link](https://aclanthology.org/2022.acl-long.150/),[Document](https://dx.doi.org/10.18653/v1/2022.acl-long.150)Cited by:[§A\.1](https://arxiv.org/html/2609.25862#A1.SS1.p1.1),[§A\.1](https://arxiv.org/html/2609.25862#A1.SS1.p2.1),[Table 8](https://arxiv.org/html/2609.25862#A1.T8),[Table 8](https://arxiv.org/html/2609.25862#A1.T8.2.2.6),[§2\.3](https://arxiv.org/html/2609.25862#S2.SS3.SSS0.Px2.p1.1),[§2\.3](https://arxiv.org/html/2609.25862#S2.SS3.SSS0.Px2.p2.1),[§4\.1](https://arxiv.org/html/2609.25862#S4.SS1.SSS0.Px1.p2.1),[§6](https://arxiv.org/html/2609.25862#S6.SS0.SSS0.Px1.p3.1),[§6](https://arxiv.org/html/2609.25862#S6.SS0.SSS0.Px1.p6.1)\.
- Wayet al\.\(2024\)A\. Way, L\. Leeson, and D\. ShterionovSign Language Machine Translation\.Springer Nature Switzerland\.External Links:[Link](https://link.springer.com/book/10.1007/978-3-031-47362-3)Cited by:[§1](https://arxiv.org/html/2609.25862#S1.p1.1),[§2\.1](https://arxiv.org/html/2609.25862#S2.SS1.p1.1),[§2](https://arxiv.org/html/2609.25862#S2.p2.1)\.
- Yinet al\.\(2021\)K\. Yin, A\. Moryossef, J\. Hochgesang, Y\. Goldberg, and M\. AlikhaniIncluding Signed Languages in Natural Language Processing\.InProceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing \(Volume 1: Long Papers\),C\. Zong, F\. Xia, W\. Li, and R\. Navigli \(Eds\.\),Online,pp\. 7347–7360\.External Links:[Link](https://aclanthology.org/2021.acl-long.570/),[Document](https://dx.doi.org/10.18653/v1/2021.acl-long.570)Cited by:[§1](https://arxiv.org/html/2609.25862#S1.p2.1)\.

## Appendix AMultilingual Results

Table 8:Per\-dataset test accuracy for two multilingual models, 1\) original vocabulary \(5,732 classes\) and 2\) unified vocabulary \(4,261 classes\), along with the best standalone results \(all SL\-GCN\) reported by[Selvaraj et al\. \(2022\)](https://arxiv.org/html/2609.25862#bib.bib17)### A\.1Multilingual Training

As mentioned in Section[2\.3](https://arxiv.org/html/2609.25862#S2.SS3), the multilingual training provided in the OpenHands framework is neither discussed in the paper nor the documentation\. We are unaware of other published results of this training configuration, but it is possible that they exist\. In Table[8](https://arxiv.org/html/2609.25862#A1.T8), we present the results of the test set evaluation for both vocabulary approaches of the multilingual training, with two different architectures, along with a comparison of the results reported by[Selvaraj et al\. \(2022\)](https://arxiv.org/html/2609.25862#bib.bib17)\.

The multilingual training outperforms the best accuracy reported by[Selvaraj et al\. \(2022\)](https://arxiv.org/html/2609.25862#bib.bib17)on four datasets out of five\. In the OpenHands paper\([Selvaraj et al\., 2022](https://arxiv.org/html/2609.25862#bib.bib17)\), no accuracy is reported for the ASLLVD, MSASL and RWTH\-PHOENIX datasets and the performance on the latter is especially poor\. This is presumably due to the specific nature of that data: the poses have been extracted from low\-resolution image files instead of videos that are furthermore from continuous data\. Regarding the results for ASL datasets, the performance on WLASL is significantly better than the results reported by the authors \(30\.6\), though the result remains far from competitive\. For AUTSL, the performance of the multilingual models is worse than the SL\-GCN results reported by[Selvaraj et al\. \(2022\)](https://arxiv.org/html/2609.25862#bib.bib17)\(that are, however, slightly worse than the state\-of\-the\-art results at the time\), whereas the performance on GSL and INCLUDE is better\.

#### Original vs\. unified vocabulary

As an alternative to a simple concatenation of all sign classes from all languages, we consider aunifiedvocabulary\. As an example from our dataset, two different classes for sign variants ofallt fínt‘all good’ are mapped to the same normalized gloss, ALL\_GOOD\. Including ÍTM in this experiment is admittedly only synthetic, as the original data is not glossed nor are there English glosses available\. To synthesize this, the classes \(words\) in the ÍTM data were machine\-translated to English, manually reviewed, and compared to normalized glosses for the other datasets – with 40\.8% overlap between the translated ÍTM glosses and the other normalized glosses\. Another potential limitation is the fact that it is not always clear where the normalized English glosses in other datasets derive from, if not from the original dataset\.

Our results demonstrate how the unified vocabulary outperforms the original vocabulary and demonstrate the effectiveness of crosslingual transfer\. It must, however, be stressed that the unified vocabulary has fewer labels than the original vocabulary since several classes that have variants in the dataset are collapsed into a single class in the unified vocabulary\. The evaluation is therefore done on a test set with 817 classes instead of the full 849, as in the original vocabulary\. And consequently, the results of the two vocabulary approaches are not fully comparable and the difficulty of the tasks is a confounding variable\.

In a similar vein, regarding the unified vocabulary approach, we emphasize that these results are based on synthetic ÍTM glosses that have not been verified by an ÍTM expert\. Therefore, the results are only an indication of the benefits of this approach, whereas the results with the original vocabulary provide more direct evidence of the benefits of multilingual training\.

## Appendix BMultilingual Prediction Analysis

If the original vocabulary is preserved in the multilingual OpenHands model, then the recognition is performed over a large multilingual label space, and predictions in other languages are possible\. On the full test set there are 137 cases of another sign language being predicted and the overall distribution is shown in Table[9](https://arxiv.org/html/2609.25862#A2.T9)\. This distribution is in line with the number of classes in the included datasets, and ASL has the largest part of the vocabulary in the concatenated dataset\. It is worth considering whether these foreign predictions are in fact correct for the other sign languages\. As an example of this, one might for instance look at the prediction for the ÍTM classfrumskógur‘jungle’, which is the plural form of tree in ASL \(ase\_\_tree\_pl\), or the ASL predictionboyfor the ÍTM class ‘man’\. The similarity at the level of the word label or gloss alone suggests that the ASL predictions may be correct\. However, this might also simply be a coincidence and there are multiple examples of foreign predictions with no label similarity to the ÍTM class, although the underlying signs may nonetheless be phonologically similar\. This would require a closer inspection and comparison of the videos\.

Table 9:Multilingual test prediction distribution with original vocabulary \(ST\-GCN\)

相似文章

手语对话中的情感识别

arXiv cs.CL

本文介绍了用于手语对话情感识别的eJSL Dialog数据集,填补了现有数据集缺乏对话上下文的空白。基准测试表明,应用通用多模态模型时存在领域差距,凸显了针对手语的上下文感知视觉提取器的必要性。

无标注表示学习用于跨数据集手语词汇检测

arXiv cs.CL

本文提出了一种无标注表示学习方法,用于跨数据集手语词汇检测,利用土耳其手语广播中的弱对齐字幕。研究表明,LLM辅助的伪标注规范化能改善时间定位和下游翻译质量。

迈向实时句子级手语翻译

arXiv cs.CL

本文提出一个句子级手语翻译系统,在How2Sign子集上用QLoRA微调,BLEU达到15.9。主要贡献是一个硬件感知的流式处理流水线,使用Raspberry Pi 4B客户端和CPU/GPU后端,平均延迟降低27.71%。