Detecting Alarming Student Verbal Responses using Text and Audio Classifier
Summary
This paper presents a hybrid framework for detecting alarming or distressed student verbal responses by combining a text classifier (content-based) and an audio classifier (prosodic features), aimed at expediting human review in Automated Verbal Response Scoring systems. The approach addresses a safety gap in automated scoring pipelines where at-risk student responses may otherwise go unnoticed.
View Cached Full Text
Cached at: 04/21/26, 07:04 AM
# Detecting Alarming Student Verbal Responses using Text and Audio Classifier
Source: [https://arxiv.org/html/2604.16717](https://arxiv.org/html/2604.16717)
\(Paper to be Presented at the National Council on Measurement in Education Conference on April 10, 2026\)
###### Abstract
This paper addresses a critical safety gap in the use Automated Verbal Response Scoring \(AVRS\)\. We present a novel hybrid framework for troubled student detection that combines a text classifier, trained to detect responses based on their content, and an audio classifier, trained to detect responses using prosodic markers\. This approach overcomes key limitations of traditional AVRS systems by considering both content and prosody of responses, achieving enhanced performance in identifying potentially concerning responses\. This system can expedite the review process by humans, which can be life\-saving particularly when timely intervention may be crucial\.
## 1Introduction
Automated Scoring \(AS\) systems are sophisticated statistical models designed to evaluate student responses, mimicking the grading process of human educators\. When AS are held to high standards, these systems have proven to be cost\-effective alternatives to hand scoring, making them an increasingly attractive option for educational institutions and assessment organizations dealing with large volumes of student work\[[17](https://arxiv.org/html/2604.16717#bib.bib9)\]\. Examples include Automated Essay Scoring \(AES\)\[[14](https://arxiv.org/html/2604.16717#bib.bib70)\], Automated Short Answer Scoring \(ASAS\)\[[15](https://arxiv.org/html/2604.16717#bib.bib68)\], and Automated Verbal Response Scoring \(AVRS\)\[[3](https://arxiv.org/html/2604.16717#bib.bib138)\]\. The benefits of these systems are accompanied by many risks\. AS systems can be vulnerable to gaming\[[9](https://arxiv.org/html/2604.16717#bib.bib270)\], they can be less accurate on certain classes of responses, and they can perpetuate and introduce additional biases beyond the training data\[[11](https://arxiv.org/html/2604.16717#bib.bib92)\]\. These biases can be exacerbated for AVRS systems where ASR can be unreliable for certain subgroups\[[7](https://arxiv.org/html/2604.16717#bib.bib149)\]\.
A critical yet often overlooked risk in taking humans out of the hand\-scoring process is that humans naturally respond with concern when confronted with a student response that indicates that the student is at risk of self\-harm or inflicting harm on others\. These responses, in some hand\-scoring materials, are simply called ”Alerts”, which will be the term we will use in this article\. In traditional human scoring processes, it might take weeks before such Alerts are noticed, delaying crucial interventions\. Recent studies have demonstrated that Large Language Models \(LLMs\) can effectively flag a small percentage of text responses for immediate human review, significantly expediting the process and enabling timely action when necessary\[[12](https://arxiv.org/html/2604.16717#bib.bib272)\]\. We seek to implement a similar pipeline for verbal student responses\.
Traditional AVRS systems typically integrate automatic speech recognition \(ASR\) with ASAS\. However, applying a standard AVRS approach to detect alerts presents two key limitations: ASR systems often struggle with distressed speech and miss crucial vocal indicators\. Conversely, focusing solely on tone can overlook concerning content delivered in a neutral voice\. Our work shows that a hybrid detection framework achieves enhanced performance\. By a audio classifier for vocal characteristics with transcript content evaluation, the system captures both delivery and substance of responses\. This comprehensive approach enables more accurate detection of potentially concerning responses by considering both*w*hat participants say and*h*ow they say it\.
Our paper is organized as follows: We give a summary of the data we used, how each model was trained, the architecture of the system, and how we benchmark the system in §[2](https://arxiv.org/html/2604.16717#S2)\. In §[3](https://arxiv.org/html/2604.16717#S3)we show the effectiveness of this new pipeline over the baseline approach\. In §[4](https://arxiv.org/html/2604.16717#S4), we provide some natural directions and applications for this research\.
## 2Method
### 2\.1System Architecture
At a high level, the system contains three main components; a transcription service, a text scorer, and an audio scorer\. The transcription service converts the audio to text which is used as input into the text\-scorer, while the audio scorer is applied to the audio directly\. The outputs of the text\-scorer and audio scorer are real numbers\. When we apply cut\-offs to the outputs, we obtain two classifications\. Our combined pipeline classifies the audio as an alert if either of our classifiers judges the audio to be an alert\. This flow is presented in Figure[1](https://arxiv.org/html/2604.16717#S2.F1)\.
Content ClassifierProsodic ClassifierorAudioClassificationFigure 1:The system is defined by two parallel processes; a content classification and a prosodic classification\.This means that we have two classifiers; one based on the transcription, which we call the content classifier, and one based directly on the audio, which we call the prosodic classifier\. Together, they classify*w*hat is being said, and*h*ow they say it\.
### 2\.2Data
To define our task, we first reference the guidelines the hand\-scoring team use to identify alerts\. It is worth noting that many testing agencies have differing definitions of what constitutes an alert\. The Smarter Balanced Consortium Hand\-scoring rules identify “Troubled Student Alerts” to include instances of suicide, criminal activity, alcohol or drug use, extreme depression, violence, rape, sexual, or physical abuse, self\-harm or intent to harm others, or neglect\. The hand\-scoring team we used to identify alerts classifies alerts into five general categories; harm to self, harm to others, harm from others, severe depression, and a specific request for help\. The clearest description of these categories can be found in the work of Burkhardt et al\.\[[2](https://arxiv.org/html/2604.16717#bib.bib7)\]\.
As motivation for this study, many of these categories can be judged on the textual content of the speech alone\. The elements of communication that are not covered by transcriptions are known as \(vocal\) prosody, which include elements like stress, tone, rhythm, and tempo\. There are well\-documented changes in vocal prosody that are linked with depression\[[18](https://arxiv.org/html/2604.16717#bib.bib271)\]\. Clearly, any comprehensive system for detecting alerts must take prosody into account\. This motivates our use of a content classifier and a prosodic classifier\.
Given we require a textual classifier, we first specify a corpus of textual responses\. In\[[10](https://arxiv.org/html/2604.16717#bib.bib193)\]and\[[12](https://arxiv.org/html/2604.16717#bib.bib272)\], there is a corpus of appropriate text responses for classifying alarming student responses\. As discussed in\[[12](https://arxiv.org/html/2604.16717#bib.bib272)\], there are approximately one alert for every 8,000 responses, which means that they are incredibly rare\. In order to sample the number of alerts at a reasonable rate for the purposes of training a classifier, the set of alarming student responses is complimented by a set of supplementary responses vetted by a hand\-scoring team as responses that would satisfy the criteria defined by the hand\-scoring rules\. Since these responses have no associated audio, they serve to train the content classifier\. The details for this corpus are presented in Table[1](https://arxiv.org/html/2604.16717#S2.T1)\.
Table 1:This data represents the text training data used in this study\.The prosodic classifier is trained on audio data as described in Table[2](https://arxiv.org/html/2604.16717#S2.T2)\. Along with alarming student response, the responses used for the non\-alert data were drawn from the typical types of items we expect to see alerts from\.
Table 2:The audio data used to build the prosodic classifier and validate the system\.One of the problems we have at this time is that the set of training examples for troubled students is an order of magnitude smaller than the number of training examples for text\. We therefore use only a subset of the available non\-alert data, with 200 sampled audio responses per each of the 49 prompts \(9,800\) for training\. At validation, we use a large \(86,783\) corpus of responses to approximate the distribution of outputs from both the text scorer and the audio scorer\. While this corpus is audio, we need to establish the distribution of the entire pipeline, not just the distribution of scores from text\.
### 2\.3Models
Transformer\-based architectures \(see\[[16](https://arxiv.org/html/2604.16717#bib.bib95)\]\) have proven to be effective in multiple modalities including language\[[5](https://arxiv.org/html/2604.16717#bib.bib81)\], speech\[[13](https://arxiv.org/html/2604.16717#bib.bib273)\], and vision\[[6](https://arxiv.org/html/2604.16717#bib.bib274)\]\. They also prove to be effective in representation learning, allowing for learning to take place between different modalities\[[1](https://arxiv.org/html/2604.16717#bib.bib275)\]\. The architecture described above requires three models; a speech\-to\-text model, as part of an ASR pipeline, a text\-classifier that classifies the text output of the speech\-to\-text model, and an audio\-classifier that classifies the audio directly\. To be more specific, we have three transformer\-based models we chose are versions of the Whisper models from\[[13](https://arxiv.org/html/2604.16717#bib.bib273)\]and a version of the ELECTRA model\[[4](https://arxiv.org/html/2604.16717#bib.bib50)\]\. The transcriber model was not fine\-tuned for the task\. These choices, and basic descriptions of these models, are presented in Table[3](https://arxiv.org/html/2604.16717#S2.T3)\.
Table 3:The transformer\-based models used in the classification pipeline\.The text scorer was trained using the Adam classifier with weight decay\[[8](https://arxiv.org/html/2604.16717#bib.bib61)\]with a learning rate of5×10−65\\times 10^\{\-6\}applied to the cross\-entropy loss function over 2 epochs over the entire dataset\. Similarly, the audio scorer was trained with the same optimizer, loss function, and number of epochs with a learning rate of5×10−65\\times 10^\{\-6\}\. As the loss function applies to the log\-probabilities, the final score is the component associated with the probability of an alert when the log\-probabilities are passed through a softmax function, meaning the final score has an interpretation as the probability of being an alert\.
### 2\.4Benchmarking
The system is designed to classify a particular percentage of all responses for review by a team of humans\. There are mainly cost considerations to take into account when choosing that percentage\. For this study, we report a range of percentage values that make sense between 0\.3% and 4%, however, the we typically chose between 1% and 2%\.
Following the system architecture, the content classifier and prosodic classifier require cut\-off values associated with those percentage\. Since the classifiers are defined by providing cut\-off values to the content and prosodic scorers, the performance of the classifiers can be determined independently\. Suppose thatXXis the set of responses, then we denote the application of the transcription followed by the text scorer function byfc:X→\[0,1\]f\_\{c\}:X\\to\[0,1\]\. The application of the audio scorer is denoted byfp:X→\[0,1\]f\_\{p\}:X\\to\[0,1\]\. The set of validation responses allows us to approximate the percentile function fairly accurately using linear interpolation, which provides us with approximations for the cut\-off values,ccc\_\{c\}andcpc\_\{p\}\. That is to say,
P\(fc\(x\)\>cc\|x∈X\)=p100andP\(fp\(x\)\>cp\|x∈X\)=p100P\\left\(f\_\{c\}\(x\)\>c\_\{c\}\|x\\in X\\right\)=\\frac\{p\}\{100\}\\hskip 28\.45274pt\\textrm\{and\}\\hskip 28\.45274ptP\\left\(f\_\{p\}\(x\)\>c\_\{p\}\|x\\in X\\right\)=\\frac\{p\}\{100\}whereppis one of the values between0\.30\.3and44discussed above\.
In setting the cut\-off values when considering the combination of the content and prosodic classifiers, our key assumption is that the percentage of responses flagged by each classifier,p~\\tilde\{p\}, is the same\. Given thatc~c\\tilde\{c\}\_\{c\}andc~p\\tilde\{c\}\_\{p\}are the associated cut\-off values for a given percentagep~\\tilde\{p\}, for a desired percentage,pp, our goal is to find ap~\\tilde\{p\}such that
g\(p~\)=P\(fc\(x\)\>c~c\|\|fp\(x\)\>c~p\|x∈X\)=p100\.g\(\\tilde\{p\}\)=P\(f\_\{c\}\(x\)\>\\tilde\{c\}\_\{c\}\|\|f\_\{p\}\(x\)\>\\tilde\{c\}\_\{p\}\|x\\in X\)=\\frac\{p\}\{100\}\.We can derivep~\\tilde\{p\}implicitly using numerical root finding applied to the equationg\(p~\)−p100=0g\(\\tilde\{p\}\)\-\\frac\{p\}\{100\}=0\. We used the secant method to define the two appropriate cut\-off values for each chosen percentage value with an initial estimate ofp~=p/2\\tilde\{p\}=p/2, which assumes a negligible intersection between the two classifiers\. Once appropriate cut\-off values are found, we can use the alerts in our test set to estimate the percentage of alerts flagged for review\.
## 3Results
As described above, for any given percentageppof the population flagged for review, we can calculate the number and percentage of alerts that are flagged by each methods: the content classifier alone, the prosodic classifier alone, and both classifiers combined\. The effectiveness of each of these methods using the number of alerts flagged under these assumptions\. For a set of reasonable values between0\.3%0\.3\\%and4%4\\%, we present these numbers in Table[4](https://arxiv.org/html/2604.16717#S3.T4)\.
Table 4:The efficacy results for the audio \(prosodic\) classifier and text \(content\) classifier individually and combined in the hybrid system\.While the validation sample only contains 100 alerts, we clearly see that the combination of content and prosodic classifiers outperforms either of the two approaches alone\. The results demonstrate that combining content and prosodic classifiers significantly improves the detection of concerning student responses compared to using either classifier alone\. At the operationally relevant range of 1\-2% of responses flagged for review, the hybrid system identifies 79\.0\-85\.0% of alerts, compared to 60\.0\-66\.0% for the content classifier and 69\.0\-78\.0% for the prosodic classifier alone\.
## 4Discussion
The results’ substantial improvement in detection rates could mean identifying several additional students in crisis\. The performance gap between the content and prosodic classifiers \(roughly 10 percentage points across most thresholds\) likely reflects two factors\. First, our content classifier benefited from a larger training dataset, including supplementary examples\. Second, some alert categories, such as specific threats, may be more reliably detected through content than prosody alone\. However, the prosodic classifier’s ability to identify alerts not caught by content analysis \(as evidenced by the hybrid system’s superior performance\) validates our hypothesis that vocal characteristics provide crucial complementary information\.
Several limitations of this study suggest natural directions for future research\. The relatively small number of audio training examples with alerts indicates that collecting additional audio data could substantially improve the prosodic classifier’s performance\. Furthermore, while our current architecture treats the two classifiers as independent, future work could explore more sophisticated ways of combining their outputs, potentially using the confidence scores from each classifier to weight their relative contributions\.
The system could also be extended to classify alerts into the specific categories outlined in Table 1, which would help prioritize responses for human review\. Additionally, while we focused on English\-language responses, the framework could be adapted for other languages, though this would require careful consideration of how prosodic markers of distress may vary across cultures\.
From an implementation perspective, the system’s modular architecture allows for straightforward updates as improved models become available\. For instance, the recent rapid advances in ASR and language models suggest that both classifiers could benefit from newer architectures or pre\-trained models as they are released\.
Finally, while this work focuses on automated scoring contexts, the approach could be valuable in other scenarios where early detection of concerning responses is critical, such as mental health hotlines or counseling services\. However, such applications would require careful validation and likely modification of the alert criteria and classifiers for these specific contexts\.
## References
- \[1\]Y\. Bengio, A\. Courville, and P\. Vincent\(2013\-08\)Representation Learning: A Review and New Perspectives\.IEEE Transactions on Pattern Analysis and Machine Intelligence35\(8\),pp\. 1798–1828\.Note:Conference Name: IEEE Transactions on Pattern Analysis and Machine IntelligenceExternal Links:ISSN 1939\-3539,[Link](https://ieeexplore.ieee.org/abstract/document/6472238),[Document](https://dx.doi.org/10.1109/TPAMI.2013.50)Cited by:[§2\.3](https://arxiv.org/html/2604.16717#S2.SS3.p1.1)\.
- \[2\]A\. Burkhardt, S\. Lottridge, and S\. Woolf\(2021\)A Rubric for the Detection of Students in Crisis\.Educational Measurement: Issues and Practice40\(2\),pp\. 72–80\(en\)\.Note:\_eprint: https://onlinelibrary\.wiley\.com/doi/pdf/10\.1111/emip\.12410External Links:ISSN 1745\-3992,[Link](https://onlinelibrary.wiley.com/doi/abs/10.1111/emip.12410),[Document](https://dx.doi.org/10.1111/emip.12410)Cited by:[§2\.2](https://arxiv.org/html/2604.16717#S2.SS2.p1.1)\.
- \[3\]A\. Cahill and K\. Evanini\(2020\)Natural language processing for writing and speaking\.InHandbook of automated scoring: Theory into practice,D\. Yan, A\. Rupp, and P\. Foltz \(Eds\.\),pp\. 69–92\.Cited by:[§1](https://arxiv.org/html/2604.16717#S1.p1.1)\.
- \[4\]K\. Clark, M\. Luong, Q\. V\. Le, and C\. D\. Manning\(2020\-03\)ELECTRA: Pre\-training Text Encoders as Discriminators Rather Than Generators\.Technical reportTechnical ReportarXiv:2003\.10555,arXiv\.Note:arXiv:2003\.10555 \[cs\] type: articleComment: ICLR 2020External Links:[Link](http://arxiv.org/abs/2003.10555),[Document](https://dx.doi.org/10.48550/arXiv.2003.10555)Cited by:[§2\.3](https://arxiv.org/html/2604.16717#S2.SS3.p1.1),[Table 3](https://arxiv.org/html/2604.16717#S2.T3.1.4.2.3),[Table 3](https://arxiv.org/html/2604.16717#S2.T3.1.4.2.4.1.1)\.
- \[5\]J\. Devlin, M\. Chang, K\. Lee, and K\. Toutanova\(2018\)BERT: Pre\-training of Deep Bidirectional Transformers for Language Understanding\.Technical reportTechnical ReportarXiv:1810\.04805,arXiv\.Note:arXiv:1810\.04805 \[cs\] type: articleExternal Links:[Link](http://arxiv.org/abs/1810.04805),[Document](https://dx.doi.org/10.48550/arXiv.1810.04805)Cited by:[§2\.3](https://arxiv.org/html/2604.16717#S2.SS3.p1.1)\.
- \[6\]A\. Dosovitskiy, L\. Beyer, A\. Kolesnikov, D\. Weissenborn, X\. Zhai, T\. Unterthiner, M\. Dehghani, M\. Minderer, G\. Heigold, S\. Gelly, J\. Uszkoreit, and N\. Houlsby\(2021\-06\)An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale\.arXiv\.Note:arXiv:2010\.11929External Links:[Link](http://arxiv.org/abs/2010.11929),[Document](https://dx.doi.org/10.48550/arXiv.2010.11929)Cited by:[§2\.3](https://arxiv.org/html/2604.16717#S2.SS3.p1.1)\.
- \[7\]A\. Kwako, Y\. Wan, J\. Zhao, K\. Chang, L\. Cai, and M\. Hansen\(2022\-07\)Using Item Response Theory to Measure Gender and Racial Bias of a BERT\-based Automated English Speech Assessment System\.InProceedings of the 17th Workshop on Innovative Use of NLP for Building Educational Applications \(BEA 2022\),Seattle, Washington,pp\. 1–7\.External Links:[Link](https://aclanthology.org/2022.bea-1.1),[Document](https://dx.doi.org/10.18653/v1/2022.bea-1.1)Cited by:[§1](https://arxiv.org/html/2604.16717#S1.p1.1)\.
- \[8\]I\. Loshchilov and F\. Hutter\(2019\-01\)Decoupled Weight Decay Regularization\.arXiv\.Note:arXiv:1711\.05101 \[cs, math\]Comment: Published as a conference paper at ICLR 2019External Links:[Link](http://arxiv.org/abs/1711.05101),[Document](https://dx.doi.org/10.48550/arXiv.1711.05101)Cited by:[§2\.3](https://arxiv.org/html/2604.16717#S2.SS3.p2.2)\.
- \[9\]S\. Lottridge, B\. Godek, A\. Jafari, and M\. PatelComparing the Robustness of Deep Learning and Classical Automated Scoring Approaches to Gaming Strategies\.\(en\)\.Cited by:[§1](https://arxiv.org/html/2604.16717#S1.p1.1)\.
- \[10\]C\. M\. Ormerod and A\. E\. Harris\(2018\-09\)Neural network approach to classifying alarming student responses to online assessment\.Technical reportTechnical ReportarXiv:1809\.08899,arXiv\.Note:arXiv:1809\.08899 \[cs, stat\] type: articleComment: 9 pages, 2 figuresExternal Links:[Link](http://arxiv.org/abs/1809.08899),[Document](https://dx.doi.org/10.48550/arXiv.1809.08899)Cited by:[§2\.2](https://arxiv.org/html/2604.16717#S2.SS2.p3.1)\.
- \[11\]C\. M\. Ormerod, A\. Malhotra, and A\. Jafari\(2021\-02\)Automated essay scoring using efficient transformer\-based language models\.arXiv\.Note:Number: arXiv:2102\.13136 arXiv:2102\.13136 \[cs\]Comment: 11 pages, 1 figure, 3 tablesExternal Links:[Link](http://arxiv.org/abs/2102.13136),[Document](https://dx.doi.org/10.48550/arXiv.2102.13136)Cited by:[§1](https://arxiv.org/html/2604.16717#S1.p1.1)\.
- \[12\]C\. M\. Ormerod, M\. Patel, and H\. Wang\(2023\-05\)Using Language Models to Detect Alarming Student Responses\.arXiv\.Note:arXiv:2305\.07709External Links:[Link](http://arxiv.org/abs/2305.07709),[Document](https://dx.doi.org/10.48550/arXiv.2305.07709)Cited by:[§1](https://arxiv.org/html/2604.16717#S1.p2.1),[§2\.2](https://arxiv.org/html/2604.16717#S2.SS2.p3.1)\.
- \[13\]A\. Radford, J\. W\. Kim, T\. Xu, G\. Brockman, C\. McLeavey, and I\. Sutskever\(2022\-12\)Robust Speech Recognition via Large\-Scale Weak Supervision\.arXiv\.Note:arXiv:2212\.04356External Links:[Link](http://arxiv.org/abs/2212.04356),[Document](https://dx.doi.org/10.48550/arXiv.2212.04356)Cited by:[§2\.3](https://arxiv.org/html/2604.16717#S2.SS3.p1.1),[Table 3](https://arxiv.org/html/2604.16717#S2.T3.1.1.4),[Table 3](https://arxiv.org/html/2604.16717#S2.T3.1.3.1.3)\.
- \[14\]M\. D\. Shermis and B\. Hamner\(2013\-04\)Contrasting State\-of\-the\-Art Automated Scoring of Essays\.pp\. 335–368\(en\)\.Note:Publisher: Routledge Handbooks OnlineExternal Links:[Link](https://typeset.io/papers/contrasting-state-of-the-art-automated-scoring-of-essays-29j53mjjzy),[Document](https://dx.doi.org/10.4324/9780203122761.CH19)Cited by:[§1](https://arxiv.org/html/2604.16717#S1.p1.1)\.
- \[15\]M\. D\. Shermis\(2015\-01\)Contrasting State\-of\-the\-Art in the Machine Scoring of Short\-Form Constructed Responses\.Educational Assessment20\(1\),pp\. 46–65\.Note:Publisher: Routledge \_eprint: https://doi\.org/10\.1080/10627197\.2015\.997617External Links:ISSN 1062\-7197,[Link](https://doi.org/10.1080/10627197.2015.997617),[Document](https://dx.doi.org/10.1080/10627197.2015.997617)Cited by:[§1](https://arxiv.org/html/2604.16717#S1.p1.1)\.
- \[16\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, Ł\. Kaiser, and I\. Polosukhin\(2017\)Attention is All you Need\.InAdvances in Neural Information Processing Systems,Vol\.30\.External Links:[Link](https://proceedings.neurips.cc/paper/2017/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html)Cited by:[§2\.3](https://arxiv.org/html/2604.16717#S2.SS3.p1.1)\.
- \[17\]D\. M\. Williamson, X\. Xi, and F\. J\. Breyer\(2012\)A Framework for Evaluation and Use of Automated Scoring\.Educational Measurement: Issues and Practice31\(1\),pp\. 2–13\(en\)\.Note:\_eprint: https://onlinelibrary\.wiley\.com/doi/pdf/10\.1111/j\.1745\-3992\.2011\.00223\.xExternal Links:ISSN 1745\-3992,[Link](https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1745-3992.2011.00223.x),[Document](https://dx.doi.org/10.1111/j.1745-3992.2011.00223.x)Cited by:[§1](https://arxiv.org/html/2604.16717#S1.p1.1)\.
- \[18\]Y\. Yang, C\. Fairbairn, and J\. F\. Cohn\(2013\-04\)Detecting Depression Severity from Vocal Prosody\.IEEE Transactions on Affective Computing4\(2\),pp\. 142–150\.Note:Conference Name: IEEE Transactions on Affective ComputingExternal Links:ISSN 1949\-3045,[Link](https://ieeexplore.ieee.org/abstract/document/6365169),[Document](https://dx.doi.org/10.1109/T-AFFC.2012.38)Cited by:[§2\.2](https://arxiv.org/html/2604.16717#S2.SS2.p2.1)\.Similar Articles
Multimodal Speaker Identification in Classroom Environments
This paper evaluates a multimodal framework for speaker identification in K-12 classrooms by combining acoustic embeddings (ECAPA-TDNN) with LLM-derived semantic context from transcripts, improving accuracy from 39% to 50.3% overall and from 64.9% to 76.9% for longer utterances.
Poster: Exploring the Limits of Audio-Based Detection of Turkish Phone Call Scams
This paper introduces the first public multimodal dataset of 100 Turkish scam and benign phone calls, evaluating seven LLMs under raw audio, ASR transcripts, and human-corrected transcripts. Results show transcript-based inputs outperform direct audio, highlighting the need for inclusive AI safety research in low-resource languages.
Towards a Phonology-Informed Evaluation of Multilingual TTS
This paper proposes a classifier-based framework to audit multilingual TTS systems for phonological faithfulness, using Assamese ATR vowel harmony as a case study. It reveals that Meta's MMS TTS frequently misproduces advanced tongue root vowels, a bias absent in human speech.
Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment
This research paper introduces a framework for zero-shot respiratory sound classification by aligning audio encoders with medical terminology through LLM-synthesized reports, outperforming models like CLAP and Qwen2-Audio in clinical diagnostic tasks.
Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts
The paper proposes a multimodal solution for audio sentiment polarity classification that integrates audio and multilingual text transcripts via cross-modal transformers, and uses knowledge distillation to enhance an audio-only model without computational overhead during inference.