Inference-Time Target Speaker Unlearning in LLM-Based Automatic Speech Recognition
Summary
The paper introduces target-speaker unlearning ASR (TSU-ASR) and proposes a novel Enrollment-Conditioned Gating module for dynamic opt-out of speakers during inference in LLM-based ASR, enhancing privacy in online conferencing.
View Cached Full Text
Cached at: 09/28/26, 09:37 AM
# Inference-time Target Speaker Unlearning in LLM-based Automatic Speech Recognition
Source: [https://arxiv.org/html/2609.30439](https://arxiv.org/html/2609.30439)
###### Abstract
We introduce target\-speaker unlearning ASR \(TSU\-ASR\) task in a fully end\-to\-end framework for multi\-speaker ASR and diarization\. Given a multi\-speaker utterance and a set of opt\-out speakers who do not wish to have their speech transcribed, the task requires an ASR system to transcribe all speakers except the opt\-out ones, while still indicating when those speakers are active\. As a first step towards tackling this task, we introduce a novel, light\-weight Enrollment\-Conditioned Gating \(ECG\) module attachable to a frozen dual\-stream speech LLM that enables ASR for new opt\-out speakers dynamically during inference, even those who were not seen during initial ECG training phase\. Our experiments on both AMI \(English\) and AliMeeting \(Mandarin\) datasets show that speech transcription accuracy for corresponding opt\-out words or characters falls from 72\.3% to 48\.2% and from 73\.6% to 27\.3%, respectively, while retained speakers’ transcription error rates maintain more or less the same\. Our approach provides a practical solution for modern video conferencing platforms, allowing speakers to dynamically opt\-out from automated AI transcriptions without forcefully leaving the meeting sessions, enabling a privacy\-preserving interface for potentially millions of online meetings daily\.
###### Index Terms:
Speech Recognition, Unlearning
††address:Indiana University, Bloomington, USA## 1Introduction
AI transcription for online conferencing is often provided to automatically transcribe spoken conversations into text, often via an automated speech recognition \(ASR\) model\. However, some participants may opt out for such an automated process\[[16](https://arxiv.org/html/2609.30439#bib.bib16)\]or do not consent to have their speech transcribed\. Unfortunately, there is currently no mechanisms from popular conference platforms that allows the AI transcription to “ignore” a few selected speakers\. Post\-recording solutions such as speaker\-specific utterance or audio segments filtering do not solve the problem because not only it is after\-the\-fact, but also it can lead to unintended removal of other speaker’s utterances when there are overlap in speech\. Pre\-deployment solutions such as retraining the models is both computationally expensive and impractical because such training process is costly\[[7](https://arxiv.org/html/2609.30439#bib.bib15),[3](https://arxiv.org/html/2609.30439#bib.bib14),[14](https://arxiv.org/html/2609.30439#bib.bib13),[11](https://arxiv.org/html/2609.30439#bib.bib12)\], especially for recent advanced transcription system that is built upon complex LLM architectures, and opt\-out requests can come not only post\-deployment but also dynamically during inference\.
We formulate this task astarget\-speaker unlearning for ASR\(TSU\-ASR\)\. Unlike target\-speaker ASR, which aims to transcribe only a selected speaker\[[10](https://arxiv.org/html/2609.30439#bib.bib5),[2](https://arxiv.org/html/2609.30439#bib.bib4)\], TSU\-ASR aims to transcribe all speakers except those in an opt\-out list\. Intuitively, given a meeting recording and short voice samples of the speakers to exclude, the system should omit their words while preserving other speakers’ transcripts\. For the sake of integrity of online conferencing, TSU\-ASR would also report who spoke when, known as speaker diarization, so readers can distinguish omitted speech from silence\. For instance, the transcript could mark thatSpeaker Aspoke from 10 to 15 seconds without displaying their words\.
To tackle this problem, one branch of research that is applicable is machine unlearning\. While existing machine unlearning approaches for AI speech models are limited, most of which often address a very different problem: removing the influence of training examples\[[9](https://arxiv.org/html/2609.30439#bib.bib11),[6](https://arxiv.org/html/2609.30439#bib.bib10),[1](https://arxiv.org/html/2609.30439#bib.bib9)\]\. In our case, we aim to exclude a selected speaker’s utterances from the transcription, even if the ASR model has never been trained on that speaker’s voice\. Methods that suppress content from a model’s inference stream such as\[[5](https://arxiv.org/html/2609.30439#bib.bib8),[12](https://arxiv.org/html/2609.30439#bib.bib7),[4](https://arxiv.org/html/2609.30439#bib.bib6)\]are more applicable, but doing so for meeting transcription must also handle changing speakers and overlapping speech, making our proposed problem both practical and non\-trivial to tackle\.
Therefore, in this work, we propose a novel module, calledEnrollment\-Conditioned Gating \(ECG\), that can plug\-and\-play to an existing LLM\-based speech model\[[3](https://arxiv.org/html/2609.30439#bib.bib14)\], which has became increasingly more popular in the ASR literature\. Existing LLM\-based ASR models often process speech content and speaker identity through two separate paths\. Intuitively, ECG then the voice samples to estimate where the selected speakers are speaking and scramble the information passed through the content path at those times, leaving the speaker path unchanged\. After plugged to the model, ECG module is briefly trained for adaptation with the the original, backbone model fixed\. This makes the selection of a new speaker to exclude to require no further training\.
Our novel plug\-and\-play ECG module for TSU\-ASR achieves a significant reduction in correct transcriptions of opt\-out speakers when tested on datasets of meeting recordings in both English and Mandarin while observingnonoticeable degradation on transcription quality of the remaining speakers\.
## 2Problem Formulation
Letxxbe a recording containing a set of speakers𝒮\(x\)\\mathcal\{S\}\(x\)\. Letℱ\\mathcal\{F\}be the forget set or set of opt\-out speakers, to be represented by one short enrollment audioese\_\{s\}per opt\-out speaker\. Letℛ\(x\)=𝒮\(x\)∖ℱ\\mathcal\{R\}\(x\)=\\mathcal\{S\}\(x\)\\setminus\\mathcal\{F\}be the remaining, retained speakers\. A TSU\-ASR system computesy^=f\(x,\{es\}s∈ℱ\)\\hat\{y\}\{=\}f\\big\(x,\\\{e\_\{s\}\\\}\_\{s\\in\\mathcal\{F\}\}\\big\)in a single pass, withℱ\\mathcal\{F\}chosen at inference, under two conditions:\(1\) Opt\-out speakers unlearned:No utterances ofℱ\\mathcal\{F\}appear iny^\\hat\{y\}under any speaker label, while their diarization is retained, or the resulting transcript must still report that someone spoke, and when;
\(2\) Retained speakers preserved:Onℛ\(x\)\\mathcal\{R\}\(x\),y^\\hat\{y\}matches model’s outputs without TSU, including where target and retained speech might overlap\.
The two conditions are independently satisfiable where the opt\-out speaker talks alone, and compete where speech overlaps, since a mechanism acting on the time axis cannot suppress one voice in a frame without touching the other\. This is also an intriguing problem to solve for an end\-to\-end multi\-speaker LLM\-based ASR framework\.
## 3Proposed Method
We proposeEnrollment\-Conditioned Gating \(ECG\), a trainable, plug\-and\-play module for an end\-to\-end multi\-speaker LLM\-based ASR framework\. ECG uses short voice samples from opt\-out speakers to scramble or mask the information used to transcribe their words, while leaving the other speakers’ utterance stream unchanged\.
Figure 1:Overview of TSU\-ASR and the proposed ECG module that reduces information in the content stream during inference when target speakers are likely to be speaking\. The speaker stream remains unchanged and provides information about who spoke when\.Base ASR Model\.Our work targets ASR architecture TagSpeech\[[3](https://arxiv.org/html/2609.30439#bib.bib14)\], which uses two Zipformer encoders: a semantic encoder for speech content and a voice encoder for speaker information\. Each encoder has a projector that converts its output into input features for a Qwen2\.5\-7B LLM\. The LLM produces text, speaker labels, and speaking times in the XML format shown in Fig\.[1](https://arxiv.org/html/2609.30439#S3.F1)\. We use the publicly available checkpoints with all parameters frozen\.
Enrollment\-Conditioned Gating Module\.ECG estimates whether each frame, a short time segment of audio, contains speech from a opt\-out speaker\. First, frame\-level voice features and enrollment embeddings are computed from the same voice encoder before projection, so they share the same feature space\. Each enrollment embedding is obtained by averaging and normalizing features extracted from a opt\-out speaker’s short voice samples\. Then, for each frame, a light\-weight two\-layer multilayer perceptron \(MLP\) computes a matching score from the frame voice steam featurevtv\_\{t\}at time steptt, an enrollment embeddingeke\_\{k\}of target speakerkk, their element\-wise product, and their cosine similarity:
gt,ek\\displaystyle g\_\{t,e\_\{k\}\}=σ\(MLP\[vt;ek;cos\(vt,ek\)\]\),\\displaystyle=\\sigma\\\!\\Big\(\\mathrm\{MLP\}\\big\[v\_\{t\};\\,e\_\{k\};\\,\\cos\(v\_\{t\},e\_\{k\}\)\\big\]\\Big\),\(1\)wherecos\(⋅\)\\cos\(\\cdot\)denotes cosine similarity function, andσ\(⋅\)\\sigma\(\\cdot\)denotes logistic function returning a score in \[0,1\]\. When multiple opt\-out speakers are provided, we use the largest matching score at each frame\. The matching score is then used to suppress the respective semantic streamsts\_\{t\}at time steptt, or:
st←st∗max\{gt,ek\|ek∈ℱ\},\\displaystyle s\_\{t\}\\leftarrow s\_\{t\}\*\\max\\\{g\_\{t,e\_\{k\}\}\\;\|\\;e\_\{k\}\\in\\mathcal\{F\}\\\},\(2\)assuming thatℱ≠∅\\mathcal\{F\}\{\\neq\}\\emptyset\. Thus, higher scores lead to stronger suppression of the content stream\. At the same time, ECG keeps the speaker stream unchanged, retaining the information used for speaker diarization\. The gate uses continuous scores in \[0,1\] rather than binary decisions for easy of training through back\-propagation\. The set of opt\-out speakers can be then changed by providing different voice samples, without further training\.
Training and Inference\.We train only the ECG module while keeping the base model’s parameters frozen\. With probability 0\.5, we select speakers present in the recording as targets and remove their words from the training transcript, while keeping the speaker labels and speaking times unchanged\. Otherwise, no speakers are selected for exclusion and the original transcript is used\. Each training example also includes a voice sample from a speaker absent from the recording as a negative example\.
Our training lossℒ\\mathcal\{L\}combines next\-token cross\-entropy on the desired output y against y\- the reference with the forget speakers’ text removed and frame\-level binary cross\-entropy on the gate logits g, using binary label m indicating speaker\-matching segments as follows:
ℒ=αCE\(y^,y−\)\+βBCE\(g,m\),\\mathcal\{L\}=\\alpha\\mathrm\{CE\}\(\\hat\{y\},\\,y^\{\-\}\)\+\\beta\\mathrm\{BCE\}\(g,\\,m\),\(3\)whereα,β\\alpha,\\betaare coefficients to balance the two loss terms\. The first term helps guide the generated output, while the second term helps optimize speaker matching ability\.
## 4Experiments
### 4\.1Datasets and Experimental Setup
We evaluate ECG on both AMI\[[8](https://arxiv.org/html/2609.30439#bib.bib1)\]\(English, single distant microphone\) and AliMeeting\[[15](https://arxiv.org/html/2609.30439#bib.bib2)\]\(Mandarin, first far\-field channel\)\. Following the base model’s decoding settings, we keep each speech segments lasting 0\.5–80 s\. This leaves 2,607 segments on AMI and 4,373 on AliMeeting\.
For each dataset, we first curate the test split by choosing roughly 10% of all recording hours, resulting a set of 16 speakers for AMI and 60 speakers for AliMeeting\. We use the remaining data as train split, making sure that there are no overlap in speakers between train and test\. This allows us to test the generalizability of ECG beyond speakers seen during training phase\.
For each test split, we randomly sample 25% of the speakers \(4 out of 16 for AMI and 15 out of 60 for AliMeeting\) as set of opt\-out speakers, and the remaining 75% serves as the retain set\. We repeat this sampling 10 times and average the results\. We collect five voice samples of 3–8 s per opt\-out speaker, using intervals where that speaker talks alone\. During testing, we average the embeddings from all five samples to computeeke\_\{k\}for each target speakerkk\.
We report results for four segment types:Retain\-only, containing only retained speakers,Forget\-only, containing only opt\-out speakers, and two mixed types containing both groups\.Mixed\-overlapincludes segments where protected and retained speech overlap for more than 0\.05 s\. The remaining mixed segments areMixed\-nonoverlap\. Table[1](https://arxiv.org/html/2609.30439#S4.T1)reports their sizes and durations for test split\. Mixed\-nonoverlap segments are uncommon because segments are split at pauses\. All results are reported withα←1\.0\\alpha\\leftarrow 1\.0andβ←1\.0\\beta\\leftarrow 1\.0\.
### 4\.2Evaluation Metrics
We evaluate opt\-out speaker content leakage and retained speaker transcription accuracy separately\. Utterances are lowercased and punctuation is removed\. We score AMI in word level and AliMeeting in character level\.
Content Leakage Rate\.CLR measures how much of the opt\-out speakers’ reference utterance remains anywhere in the generated transcript:
CLR=∑xLCS\(r\(x\),h\(x\)\)∑x\|r\(x\)\|,\\displaystyle\\mathrm\{CLR\}=\\frac\{\\sum\_\{x\}\\mathrm\{LCS\}\\big\(r\(x\),h\(x\)\\big\)\}\{\\sum\_\{x\}\|r\(x\)\|\},\(4\)wherer\(x\)r\(x\)is the reference text spoken by the opt\-out speaker in segmentxx, andh\(x\)h\(x\)is the full generated transcript, andLCS\(⋅\)LCS\(\\cdot\)computes the length of the longest common subsequence: it counts the largest number of matching words or characters in the same order, without requiring them to be consecutive\. We reportCLR\-allusing all words or characters andCLR\-rarefor evaluating on only rare words or characters or those with document frequency below 1%\. Restricting the measure to rare units reduces matches caused by common expressions shared across speakers\. Both scores are reported as percentages, with lower values indicating effective suppressing\. Because CLR examines the full transcript, opt\-out content still counts if it is assigned to a retained speaker\.
Table 1:Test segments by type, averaged over ten protected\-speaker sets\. Dur\.: total duration; Ovl\.: overlap between protected and retained speakers \(both in hours\)\.Table 2:Each cell depicts score changes fromwithouttowith our module attached, averaged over ten random draws of the retained speakers\. Lower is better throughout\.Retained\-Speaker Transcription Accuracy\.We report standard word error rate metrics for meeting transcription\[[13](https://arxiv.org/html/2609.30439#bib.bib3)\], includingcpWER\-R andgWER\-R on AMI, and the corresponding character error rates, cpCER\-R and gCER\-R, on AliMeeting\. The suffix R indicates that these scores are calculated for retained speakers\. We compare results with and without ECG to measure changes in transcription accuracy\.
For cpWER\-R, we jointly match all reference speakers, both opt\-out and retained, to predicted speaker labels using the assignment with the lowest total transcription error\. We then score only the outputs matched to retained speakers\. Outputs matched to opt\-out speakers are excluded from this score, and unmatched outputs are ignored\.
We also report retained\-speaker errors alongside CLR: low transcription error alone does not establish that protected content has been removed, while an empty transcript can achieve zero CLR by omitting everyone’s words\.
### 4\.3Results and Analysis
Overall Reuslts\.Table[2](https://arxiv.org/html/2609.30439#S4.T2)summarizes the results, part of which is visualized in Fig\.[2](https://arxiv.org/html/2609.30439#S4.F2)to show how the results vary across group types\. Overall, ECG enables effective suppression of opt\-out speakers during inference even if they were not seen during training, reducing CLR\-rare by25\.4%25\.4\\%and42\.7%42\.7\\%in AMI and AliMeeting, while maintaining similar transcription performance in retained\-speaker measured in cpWER and cpCER \(\+1\.2%1\.2\\%and2\.3%2\.3\\%\)\. Noticeably, both CLR\-rare and CLR\-all consistently decrease across all groups containing opt\-out speech, so the improvement remains when common words are included\.
Group\-level Results\.\(Fig\.[2](https://arxiv.org/html/2609.30439#S4.F2)\) show that group\-level performance depends on which speakers are present in the segments with two main observations\. Overall, the presence of retain speakers in the segments slightly hinders the suppression of opt\-out speakers’ transcriptions for AliMeeting but only marginally for AMI\. Interestingly, opt\-out content still leaks in Forget\-only groups, even though there is no retained speech to preserve\. This shows that overlap is not the only obstacle to removing protected words\. One possible explanation is that enough information still reaches the frozen LLM to support transcription\.
Regarding retrained speakers, their error rates change little on Retain\-only groups\. The loss in accuracy appears when protected and retained speech share a group, where the model must suppress some content while continuing to transcribe the rest\. Errors increase in both Mixed\-overlap and Mixed\-nonoverlap groups, although the small sample size of the latter makes that comparison less significant\.
Findings\.Finding \#1: Low leakage on Forget\-only group alone is insufficient\.An empty transcript achieves zero CLR, so a lower score does not by itself establish successful selective removal\. For Forget\-only groups, omitting all words is the intended result, provided that the speaker information is preserved\. For mixed groups, the same output might also discard words that should remain, making this a very challenging task\.Finding \#2: the proposed ECG is a strong first step towards tackling TSU task,demonstrating its ability to learn to match speaker patterns and translate it to a probabilistic gate that can control the residual semantic stream of the speech LLM architecture\. Its simple implementation potentially allows adoption in other speech applications where LLM\-based models are increasing prevalent and that we need to control the speech AI models dynamically with unseen conditional information such as opt\-out speakers\.
Figure 2:CLR\-rare \(top\) and retained\-speaker transcription error \(bottom\), without and with ECG\.
## 5Conclusions
We introduced a novel target\-speaker unlearning for ASR task \(TSU\-ASR\) that aims to omit suppress speakers’ utterances while preserving those of other speakers\. Our proposed Enrollment\-Conditioned Gating \(ECG\) module suppresses the semantic stream of an advanced multi\-speaker LLM\-based ASR model while leaving its speaker stream unchanged\. Our approach trains only the light\-weight ECG, and require no further training for unseen opt\-out speakers\. Our experiments show significant effectiveness on two datasets AMI and AliMeeting in both English and Mandarin\.
## References
- \[1\]J\. Cheng and H\. Amiri\(2025\)Speech Unlearning\.InInterspeech 2025,pp\. 3209–3213\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2025-2412),ISSN 2958\-1796Cited by:[§1](https://arxiv.org/html/2609.30439#S1.p3.1)\.
- \[2\]Z\. Huang, D\. Raj, P\. García, and S\. Khudanpur\(2022\)Adapting self\-supervised models to multi\-talker speech recognition using speaker embeddings\.External Links:2211\.00482,[Link](https://arxiv.org/abs/2211.00482)Cited by:[§1](https://arxiv.org/html/2609.30439#S1.p2.1)\.
- \[3\]M\. Huo, Y\. Shao, and Y\. Zhang\(2026\)TagSpeech: end\-to\-end multi\-speaker ASR and diarization with fine\-grained temporal grounding\.InProceedings of the 64th Annual Meeting of the Association for Computational Linguistics \(Volume 1: Long Papers\),M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 41847–41862\.External Links:[Link](https://aclanthology.org/2026.acl-long.1938/),[Document](https://dx.doi.org/10.18653/v1/2026.acl-long.1938),ISBN 979\-8\-89176\-390\-6Cited by:[§1](https://arxiv.org/html/2609.30439#S1.p1.1),[§1](https://arxiv.org/html/2609.30439#S1.p4.1),[§3](https://arxiv.org/html/2609.30439#S3.p2.1)\.
- \[4\]M\. Lee, E\. Shin, and J\. Lee\(2026\)Erasing your voice before it’s heard: training\-free speaker unlearning for zero\-shot text\-to\-speech\.External Links:2601\.20481,[Link](https://arxiv.org/abs/2601.20481)Cited by:[§1](https://arxiv.org/html/2609.30439#S1.p3.1)\.
- \[5\]C\. Y\. Liu, Y\. Wang, J\. Flanigan, and Y\. Liu\(2024\)Large language model unlearning via embedding\-corrupted prompts\.InAdvances in Neural Information Processing Systems,A\. Globerson, L\. Mackey, D\. Belgrave, A\. Fan, U\. Paquet, J\. Tomczak, and C\. Zhang \(Eds\.\),Vol\.37,pp\. 118198–118266\.External Links:[Document](https://dx.doi.org/10.52202/079017-3754),[Link](https://proceedings.neurips.cc/paper_files/paper/2024/file/d6359156e0e30b1caa116a4306b12688-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2609.30439#S1.p3.1)\.
- \[6\]Z\. Liu\(2025\)Unlearning LLM\-Based Speech Recognition Models\.InInterspeech 2025,pp\. 3214–3218\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2025-287),ISSN 2958\-1796Cited by:[§1](https://arxiv.org/html/2609.30439#S1.p3.1)\.
- \[7\]Z\. Ma, G\. Yang, Y\. Yang, Z\. Gao, J\. Wang, Z\. Du, F\. Yu, Q\. Chen, S\. Zheng, S\. Zhang, and X\. Chen\(2025\)Speech recognition meets large language model: benchmarking, models, and exploration\.InProceedings of the Thirty\-Ninth AAAI Conference on Artificial Intelligence and Thirty\-Seventh Conference on Innovative Applications of Artificial Intelligence and Fifteenth Symposium on Educational Advances in Artificial Intelligence,AAAI’25/IAAI’25/EAAI’25\.External Links:ISBN 978\-1\-57735\-897\-8,[Link](https://doi.org/10.1609/aaai.v39i23.34666),[Document](https://dx.doi.org/10.1609/aaai.v39i23.34666)Cited by:[§1](https://arxiv.org/html/2609.30439#S1.p1.1)\.
- \[8\]S\. Renals, T\. Hain, and H\. Bourlard\(2007\)Recognition and understanding of meetings the ami and amida projects\.In2007 IEEE Workshop on Automatic Speech Recognition & Understanding \(ASRU\),Vol\.,pp\. 238–247\.External Links:[Document](https://dx.doi.org/10.1109/ASRU.2007.4430116)Cited by:[§4\.1](https://arxiv.org/html/2609.30439#S4.SS1.p1.1)\.
- \[9\]N\. Sarwar, S\. Roy Dipta, Z\. Liu, and V\. Patil\(2026\)Multimodal unlearning across vision, language, video, and audio: survey of methods, datasets, and benchmarks\.InFindings of the Association for Computational Linguistics: ACL 2026,M\. Liakata, V\. P\. Moreira, J\. Zhang, and D\. Jurgens \(Eds\.\),San Diego, California, United States,pp\. 27702–27730\.External Links:[Link](https://aclanthology.org/2026.findings-acl.1379/),[Document](https://dx.doi.org/10.18653/v1/2026.findings-acl.1379),ISBN 979\-8\-89176\-395\-1Cited by:[§1](https://arxiv.org/html/2609.30439#S1.p3.1)\.
- \[10\]J\. Seong, J\. Choi, Y\. Jeoung, I\. Kim, and J\. Chang\(2025\)Enhancing Target\-speaker Automatic Speech Recognition Using Multiple Speaker Embedding Extractors with Virtual Speaker Embedding\.InInterspeech 2025,pp\. 4918–4922\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2025-2486),ISSN 2958\-1796Cited by:[§1](https://arxiv.org/html/2609.30439#S1.p2.1)\.
- \[11\]M\. Shi, X\. Xiao, R\. Fan, S\. Ling, and J\. Li\(2026\)Train short, infer long: speech\-llm enables zero\-shot streamable joint asr and diarization on long audio\.InICASSP 2026 \- 2026 IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),Vol\.,pp\. 17442–17446\.External Links:[Document](https://dx.doi.org/10.1109/ICASSP55912.2026.11464726)Cited by:[§1](https://arxiv.org/html/2609.30439#S1.p1.1)\.
- \[12\]V\. C\. Turani, O\. Parraga, J\. V\. B\. Abitante, K\. K\. Arguello, J\. Pasquali, R\. N\. Barros, F\. du Pin Calmon, C\. Mattjie, R\. C\. Barros, and L\. S\. Kupssinskü\(2026\)Inference\-time machine unlearning via gated activation redirection\.External Links:2605\.12765,[Link](https://arxiv.org/abs/2605.12765)Cited by:[§1](https://arxiv.org/html/2609.30439#S1.p3.1)\.
- \[13\]T\. von Neumann, C\. Boeddeker, M\. Delcroix, and R\. Haeb\-Umbach\(2024\)MeetEval: a toolkit for computation of word error rates for meeting transcription systems\.External Links:2307\.11394,[Link](https://arxiv.org/abs/2307.11394)Cited by:[§4\.2](https://arxiv.org/html/2609.30439#S4.SS2.p3.1)\.
- \[14\]H\. Yin, Y\. Chen, C\. Deng, L\. Cheng, H\. Wang, C\. Tan, Q\. Chen, W\. Wang, and X\. Li\(2025\)SpeakerLM: end\-to\-end versatile speaker diarization and recognition with multimodal large language models\.ArXivabs/2508\.06372\.External Links:[Link](https://api.semanticscholar.org/CorpusID:280561546)Cited by:[§1](https://arxiv.org/html/2609.30439#S1.p1.1)\.
- \[15\]F\. Yu, S\. Zhang, Y\. Fu, L\. Xie, S\. Zheng, Z\. Du, W\. Huang, P\. Guo, Z\. Yan, B\. Ma, X\. Xu, and H\. Bu\(2022\)M2MeT: the icassp 2022 multi\-channel multi\-party meeting transcription challenge\.External Links:2110\.07393,[Link](https://arxiv.org/abs/2110.07393)Cited by:[§4\.1](https://arxiv.org/html/2609.30439#S4.SS1.p1.1)\.
- \[16\]X\. Zhan, G\. Sun, J\. Such, and P\. Woodland\(2026\)Protecting bystander privacy via selective hearing in audio llms\.External Links:2512\.06380,[Link](https://arxiv.org/abs/2512.06380)Cited by:[§1](https://arxiv.org/html/2609.30439#S1.p1.1)\.Similar Articles
Machine Unlearning for Speech Question Answering in Large Audio-Language Models
This paper explores machine unlearning techniques for Large Audio-Language Models to remove sensitive information from speech QA tasks, demonstrating methods that reduce privacy leakage by up to 80% while maintaining performance.
Inference-Time Machine Unlearning via Gated Activation Redirection
This paper introduces GUARD-IT, a training-free method for machine unlearning that uses input-dependent activation steering at inference time to remove targeted knowledge from LLMs without modifying weights, matching or exceeding gradient-based baselines while preserving utility and robustness to quantization.
How to Leverage Synthetic Speech for LLM-Based ASR Systems?
This paper investigates the distributional gap between synthetic and real speech in LLM-based ASR systems, identifies where the LLM separates them, and proposes using layer-selection and RIR augmentation to match real-data baselines with less real data.
LaSR: Context-Aware Speech Recognition via Latent Reasoning
LaSR proposes a latent reasoning training paradigm for context-aware speech recognition, aligning chain-of-thought supervision around acoustic features to improve terminology recognition without added latency, outperforming standard fine-tuning on Fun-Audio-Chat.
Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting
This paper introduces Attention-Shifting (AS), a novel framework for selective machine unlearning in LLMs that balances effective removal of sensitive information while preventing hallucinations and preserving model utility. The method uses importance-aware attention suppression and retention enhancement to achieve up to 15% higher accuracy preservation compared to existing unlearning approaches on standard benchmarks.