Learning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized Guidance

arXiv cs.CL Papers

Summary

This paper introduces randomized guidance as an efficient training method for tandem speech-to-speech models, enabling them to learn natural conversational behavior directly from real conversations without simulating backend LLM behavior.

arXiv:2609.30773v1 Announce Type: new Abstract: Tandem speech-to-speech architectures couple a responsive speech frontend with an asynchronous text backend. In KAME, a large language model (LLM) serves as the backend, supplying candidate responses as guidance to the speech frontend while the user is still speaking. Ordinary conversation recordings capture the eventual response but not the guidance the backend would supply during the user's utterance. Generating the missing guidance with a simulator LLM adds substantial data-preparation overhead when training on real conversations. We propose randomized intermediate guidance, which derives guidance directly from the conversation corpus rather than simulating backend LLM behavior. During training, target responses provide informative guidance, while randomly sampled responses provide potentially irrelevant updates during the utterance. This combination aims to teach the frontend to use backend information selectively. On synthetic dialogues, KAME trained with this recipe achieves response quality comparable to that of the LLM-generated and similarity-based baselines. Training on 3.8k hours of real conversations improves smooth turn-taking and audio-judge naturalness over synthetic-data KAME while retaining a response-quality advantage over Moshi. These results show that randomized guidance offers a practical route to combining the response-quality benefits of tandem models with natural conversational behavior learned from real speech.
Original Article
View Cached Full Text

Cached at: 09/28/26, 09:42 AM

# LEARNING NATURAL CONVERSATIONAL BEHAVIOR IN TANDEM SPEECH-TO-SPEECH MODELS WITH RANDOMIZED GUIDANCE
Source: [https://arxiv.org/html/2609.30773](https://arxiv.org/html/2609.30773)
## LEARNING NATURAL CONVERSATIONAL BEHAVIOR IN TANDEM SPEECH\-TO\-SPEECH MODELS WITH RANDOMIZED GUIDANCEThanks:This work was done during Manato Yaguchi’s internship at Sakana AI\.

###### Abstract

Tandem speech\-to\-speech architectures couple a responsive speech frontend with an asynchronous text backend\. In KAME, a large language model \(LLM\) serves as the backend, supplying candidate responses as guidance to the speech frontend while the user is still speaking\. Ordinary conversation recordings capture the eventual response but not the guidance the backend would supply during the user’s utterance\. Generating the missing guidance with a simulator LLM adds substantial data\-preparation overhead when training on real conversations\. We propose randomized intermediate guidance, which derives guidance directly from the conversation corpus rather than simulating backend LLM behavior\. During training, target responses provide informative guidance, while randomly sampled responses provide potentially irrelevant updates during the utterance\. This combination aims to teach the frontend to use backend information selectively\. On synthetic dialogues, KAME trained with this recipe achieves response quality comparable to that of the LLM\-generated and similarity\-based baselines\. Training on 3\.8k hours of real conversations improves smooth turn\-taking and audio\-judge naturalness over synthetic\-data KAME while retaining a response\-quality advantage over Moshi\. These results show that randomized guidance offers a practical route to combining the response\-quality benefits of tandem models with natural conversational behavior learned from real speech\.

###### Index Terms:

speech\-to\-speech dialogue, full\-duplex dialogue, randomized guidance, real conversations

††address:1Sakana AI, Tokyo, Japan2The University of Tokyo, Tokyo, Japan## 1Introduction

A central challenge in real\-time spoken dialogue is to combine high\-quality responses with natural interaction\. Full\-duplex speech\-to\-speech \(S2S\) models such as Moshi support low\-latency speech generation and flexible turn\-taking\[[4](https://arxiv.org/html/2609.30773#bib.bib5)\], but their knowledge and reasoning capabilities remain limited\. Cascaded systems can draw on powerful text\-based large language models \(LLMs\), but their sequential ASR–LLM–TTS pipeline introduces delays that can disrupt conversational flow\.

To combine the strengths of both approaches, tandem architectures couple a responsive speech frontend with an asynchronous backend\. KAME exemplifies this design by guiding a full\-duplex S2S frontend with responses from an asynchronous text LLM\[[7](https://arxiv.org/html/2609.30773#bib.bib6)\]\. MoshiRAG retrieves external knowledge during full\-duplex speech generation\[[3](https://arxiv.org/html/2609.30773#bib.bib10)\], while ConvFill lets a lightweight Talker respond as a slower Reasoner streams knowledge\[[13](https://arxiv.org/html/2609.30773#bib.bib11)\]\. For such asynchronous interfaces, a desirable capability is to use relevant backend information without relying on irrelevant updates, while maintaining natural conversational timing and delivery\.

To train such tandem S2S systems, supervised fine\-tuning requires not only ordinary dialogue pairs but also the intermediate guidance supplied by the backend\. This makes it difficult to use transcribed dialogue audio as a training dataset for the tandem architecture, because such datasets lack the intermediate guidance supplied by the backend LLM to the speech frontend\. To mitigate this issue, KAME, for example, employs guidance generated by a simulator LLM\[[7](https://arxiv.org/html/2609.30773#bib.bib6)\]\. The simulator LLM is instructed to generate provisional responses periodically while a user utterance is ongoing\. The simulation process is designed so that the generated responses gradually become semantically closer to the target response, i\.e\., what the next speaker actually said, as more of the user’s utterance is observed\. This design is intended to supply increasingly useful guidance as the user utterance unfolds, encouraging the frontend to exploit informative backend updates during response generation\.

This per\-example LLM simulation becomes a data\-preparation bottleneck when scaling tandem training to large real\-conversation corpora, limiting the use of real speech for learning natural conversational behavior\. The original KAME is trained on synthetic dialogues derived from question–answer pairs and converted to speech\[[7](https://arxiv.org/html/2609.30773#bib.bib6)\]\. These data improve response quality but do not directly capture the timing and vocal delivery of spontaneous conversation\. Recent work has explored real conversations for interactivity alignment\[[10](https://arxiv.org/html/2609.30773#bib.bib7)\]and full\-duplex data construction\[[5](https://arxiv.org/html/2609.30773#bib.bib8),[8](https://arxiv.org/html/2609.30773#bib.bib9)\]\. However, these efforts focus on full\-duplex modeling or data construction rather than scaling real\-conversation training in tandem architectures\.

To eliminate this bottleneck, we introduce randomized intermediate guidance\. Instead of faithfully simulating the backend’s intermediate guidance, our proposed method uses a sequence of informative and potentially irrelevant guidance texts\. Following KAME, the target response is supplied as informative guidance, and randomly sampled response texts from other conversations are supplied as potentially irrelevant guidance in place of KAME’s intermediate guidance\. By mixing those two types of guidance, we aim to encourage the frontend model to distinguish and exploit useful information\. This allows all the training data to be constructed directly from conversational data with transcripts, without per\-example LLM\-based simulation\.

We evaluate this approach in two complementary experiments on the KAME architecture\. First, on synthetic dialogues, randomized guidance is shown to be competitive with LLM\-generated guidance and similarity\-based retrieval on spoken MT\-Bench, whereas Target\-only training yields lower scores \(Table[1](https://arxiv.org/html/2609.30773#S3.T1)\)\. We then apply randomized guidance to train KAME on a 3\.8k\-hour corpus of real two\-speaker conversations\. The resulting model achieves smoother turn\-taking and higher audio\-judge naturalness than synthetic\-data KAME, while retaining an MT\-Bench score of 5\.04 versus Moshi’s 1\.96 \(Table[2](https://arxiv.org/html/2609.30773#S3.T2)\)\. Together, these results demonstrate a practical way to combine the response\-quality benefits of tandem S2S with natural interaction learned from real conversations\.

Figure 1:Schematic training\-time guidance streams:pip\_\{i\}, simulator\-LLM\-generated responses;rir\_\{i\}, randomly sampled responses;yy, the target transcript\. The inference\-time backend is unchanged\.
## 2Randomized Intermediate Guidance

### 2\.1LLM\-Generated Guidance

In KAME, an asynchronous LLM is introduced to supply candidate responses to the frontend S2S model\[[7](https://arxiv.org/html/2609.30773#bib.bib6)\]\. At inference time, this asynchronous LLM is queried periodically at 2 Hz to generate candidate responses from the partial user inputs\. We refer to this series of updates as the guidance stream \(called the “oracle” stream in\[[7](https://arxiv.org/html/2609.30773#bib.bib6)\]\)\.

A key difficulty in training the tandem S2S model is that ordinary dialogue data provide user utterances and their target responses, but not the corresponding guidance stream\. In KAME, a simulator LLM is employed to simulate a trace that converges to the target response \(see “LLM\-generated” in Fig\.[1](https://arxiv.org/html/2609.30773#S1.F1)\)\. At successive points in the user utterance, the simulator generates a possible interpolation between the response estimated from the partial input and the target transcriptyy\. This guidance stream is designed to end withyyitself\. With this simulation, the frontend S2S model is optimized to start generating response audio once relevant information is obtained from the guidance stream\. To prepare a training dataset, this procedure has to be repeated for each training dialogue\.

### 2\.2Randomized Intermediate Guidance

To avoid costly LLM simulation in large\-scale corpus preparation, we propose a method that derives the training\-time guidance stream directly from the training corpus itself\. This method is grounded in our hypothesis about the KAME training process: the most critical aspect of KAME fine\-tuning is teaching the model to distinguish informative guidance from irrelevant guidance that arises from premature LLM invocation\. In this method, the guidance stream is designed to contain potentially irrelevant guidance texts randomly sampled from other conversation histories alongside the relevant guidance text copied from the target response \(see “Randomized” in Fig\.[1](https://arxiv.org/html/2609.30773#S1.F1)\)\.

For each response segment, we retain the target transcriptyyat the last designated target\-response update and fill the remaining updates with randomly sampled response texts\. We form a pool of response texts from the training corpus, deduplicated by token sequence, and exclude all response texts appearing in the current dialogue\. We retain candidates containing between half and twice as many tokens asyy\. For each segment, we sample the required response texts uniformly without replacement from the filtered pool\.

### 2\.3Application to Real Conversations

To use randomized guidance with real conversations, we first prepare clean single\-speaker audio segments and time\-aligned transcripts\. We use Silero VAD to detect speech regions in PodcastIndex recordings, then apply the pyannote speaker\-diarization\-3\.1 pipeline\[[1](https://arxiv.org/html/2609.30773#bib.bib1)\]for speaker diarization on each channel\. We project the diarization results onto the VAD segments and discard segments containing multiple speakers or overlapping speech\. We then enhance each retained segment with Sidon\[[9](https://arxiv.org/html/2609.30773#bib.bib14)\], filter it using DNSMOS P\.835\[[12](https://arxiv.org/html/2609.30773#bib.bib15)\]\(OVRL≥3\.0\\geq 3\.0\), and transcribe it with batched faster\-whisper large\-v3\-turbo\. We obtain token\-level timestamps using faster\-whisper’s Whisper\-based alignment routine\[[11](https://arxiv.org/html/2609.30773#bib.bib3)\]\. We further filter transcriptions using Whisper language probability \(≥0\.85\\geq 0\.85\) and average log probability \(≥−0\.4\\geq\-0\.4\), and require non\-empty transcripts\. After this segmentation and filtering process, we construct dialogue examples from eligible segment groups associated with the same source recording and source channel and containing exactly two speaker identities\. We order the selected segments by their original timestamps and concatenate them back\-to\-back\. This preserves the relative chronological order of the retained segments while discarding the original inter\-segment gaps\.

For each response segment, the recorded audio and its aligned transcript serve as the training targets\. The same transcript also serves as target\-response guidanceyy, while texts sampled from other turns in the training corpus provide randomized intermediate guidance\. The corpus thus supplies both the training targets and the guidance stream without per\-example LLM generation\.

## 3Experiments

### 3\.1Response Quality with Randomized Guidance

We compare four guidance\-construction strategies on the same English synthetic dialogues: LLM\-generated, Similarity\-based, Target\-only, and Randomized \(ours\)\.

The first strategy, LLM\-generated, is our baseline and follows the original KAME construction described in Section[2\.1](https://arxiv.org/html/2609.30773#S2.SS1)\[[7](https://arxiv.org/html/2609.30773#bib.bib6)\]\. The second strategy, Similarity\-based, avoids per\-example LLM generation while still approximating the backend’s semantic progression through retrieval\. We first derive similarity trajectories from 30 calibration turns by comparing the backend’s responses to partial user utterances with its response to each complete utterance, using a frozen MiniLM\-based sentence encoder\[[14](https://arxiv.org/html/2609.30773#bib.bib2)\]\. For each training segment, we sample one of these trajectories rather than an averaged curve, and retrieve corpus responses whose cosine similarity to the target transcriptyyapproximately follows the piecewise\-linear interpolation of the sampled trajectory\. These trajectories need not increase monotonically\. We also examine whether the target response alone provides sufficient guidance for competitive response quality\. The third strategy, Target\-only, therefore omits intermediate guidance\. Finally, Randomized \(ours\) samples intermediate guidance from other training dialogues without semantic selection \(Section[2\.2](https://arxiv.org/html/2609.30773#S2.SS2)\)\. This combination is intended to train the frontend to use informative guidance while ignoring irrelevant updates\.

For all four strategies, we fine\-tune all parameters of the same pretrained Moshi checkpoint for one epoch\. We use a global batch size of 32 and learning rates of2×10−62\\times 10^\{\-6\}and4×10−64\\times 10^\{\-6\}for the temporal and depth transformers, respectively\.

We evaluate response quality on a spoken MT\-Bench subset\[[15](https://arxiv.org/html/2609.30773#bib.bib4)\]comprising 30 questions, each with two turns\. All models use GPT\-4\.1 as the inference\-time backend\. We transcribe their generated responses with Whisper large\-v3\[[11](https://arxiv.org/html/2609.30773#bib.bib3)\]and, following the original MT\-Bench evaluation protocol, score the transcripts with a GPT\-4 judge on a 1–10 scale\. For each model, we repeat judging three times with the audio and transcripts held fixed, and report the mean and standard error across the three 60\-turn averages\.

Table 1:Response quality of four guidance\-construction strategies for KAME on the 30\-question spoken MT\-Bench subset\. Mean±\\pmstandard error over three GPT\-4 judge repetitions using fixed audio and transcripts\.Table 2:Response quality, conversational behavior, and audio\-judge naturalness for three models; higher is better\. MTR\-derived success rates average three generation seeds\. MT\-Bench and naturalness report mean±\\pmstandard error\. MT\-Bench and naturalness use three judge repetitions on fixed audio; naturalness uses the common 16\-question subset\.Table[1](https://arxiv.org/html/2609.30773#S3.T1)shows that Randomized achieves competitive response quality, scoring 6\.09 compared with 6\.14 for Similarity\-based and 5\.54 for LLM\-generated \(KAME\)\. It also scores higher than Target\-only \(5\.12\)\. One possible explanation for the improvement over LLM\-generated guidance is that the simulated intermediate guidance in KAME follows a heuristically designed progression toward the target response and may not fully reflect the variability of inference\-time backend updates\. In contrast, randomized guidance can be viewed as a form of negative sampling: by exposing the frontend to both informative target\-response guidance and potentially irrelevant intermediate updates, it may encourage the model to rely on guidance only when it is useful\. Together, these results support randomized guidance as a practical alternative to reconstructing intermediate backend predictions\.

### 3\.2Training on Real Conversations

Randomized guidance allows KAME to be trained on a large\-scale real dataset without per\-example LLM generation\. We prepare 3\.8k hours of English two\-speaker conversations following Section[2\.3](https://arxiv.org/html/2609.30773#S2.SS3), approximately2\.8×2\.8\\timesthe audio duration of the 1\.36k\-hour synthetic corpus used for the synthetic\-data KAME baseline\. The central question is whether learning from these recordings improves conversational timing and delivery while retaining KAME’s response\-quality advantage over Moshi\. We address this question by comparing the real\-data model with pretrained Moshi and KAME \(synthetic data\), the LLM\-generated baseline from Section[3\.1](https://arxiv.org/html/2609.30773#S3.SS1)\.

We fine\-tune all parameters of the same pretrained Moshi checkpoint as in Section[3\.1](https://arxiv.org/html/2609.30773#S3.SS1)\. With batch size and learning rates unchanged, we evaluate the 3,000\-update checkpoint, denoted KAME \(real data, ours\)\.

We assess response quality with MT\-Bench and conversational behavior with three tasks derived from MTR\-DuplexBench\[[6](https://arxiv.org/html/2609.30773#bib.bib12)\]: smooth turn\-taking, interruption handling, and pause handling\. These tasks examine whether a model responds after the user finishes speaking, yields when interrupted, and waits through pauses within a user turn\. For each task, we average success over 200 ten\-round dialogues and three generation seeds, including the interruption task’s initial setup round\. Both KAME models receive identical precomputed GPT\-4\.1 guidance streams from user\-transcript prefixes at 0\.5\-second intervals, with a fixed simulated backend delay of six frames \(0\.48 s\)\. We use local acoustic/ASR\-based scoring; these rates are not directly comparable to published benchmark scores\.

To assess conversational delivery, we use Gemini 3\.1 Pro \(high thinking level\) as an audio judge, following prior work on audio\-aware judging\[[2](https://arxiv.org/html/2609.30773#bib.bib13)\]\. It scores naturalness \(1–10\) from response\-only clips of up to ten seconds, starting at question end, for both turns of 30 MT\-Bench questions\. The evaluation is designed to focus on the audible naturalness of the response: no question text, transcripts, or model identities are provided\. The prompt assesses ordinary, unperformed conversational delivery without requiring polish or expressiveness\. Deductions require audible evidence; answer content, turn\-taking, leading silence, and clip\-boundary artifacts are excluded\. Clips with insufficient speech are unassessable\. We report mean and standard error across three repetition\-level averages on fixed audio, using the 16 questions with both turns assessable for every model in every repetition\.

The results in Table[2](https://arxiv.org/html/2609.30773#S3.T2)show that real\-data KAME narrows the naturalness gap to Moshi while preserving a substantial response\-quality advantage\. Audio\-judge naturalness rises from the synthetic baseline’s 4\.05 to 4\.79, approaching Moshi’s 5\.29\. At the same time, its MT\-Bench score of 5\.04 remains within 0\.50 points of synthetic\-data KAME and well above Moshi’s 1\.96\. Smooth turn\-taking success also increases from 40\.62% to 50\.88%, extending the gains beyond vocal delivery to turn\-taking\. Pause handling remains similar, while interruption handling remains a limitation\. One possible reason is that our training objective does not explicitly target stopping an ongoing response and responding again after the user finishes speaking\.

## 4Conclusion

In this paper, we introduced randomized intermediate guidance, which constructs the missing training\-time guidance stream directly from the corpus by combining target\-response guidance with randomly sampled response texts, avoiding per\-example LLM simulation\. Experiments on synthetic dialogues show that randomized guidance achieves response quality comparable to the LLM\-generated and similarity\-based baselines\. Applying this approach to 3\.8k hours of real conversations yields smoother turn\-taking and higher audio\-judge naturalness than synthetic\-data KAME, while retaining a substantial MT\-Bench advantage over Moshi\. Thus, we show that randomized intermediate guidance supports effective KAME training and enables more natural interaction when combined with real conversational data\.

## References

- \[1\]H\. Bredin\(2023\)pyannote\.audio 2\.1 speaker diarization pipeline: principle, benchmark, and recipe\.InProc\. Interspeech 2023,pp\. 1983–1987\.External Links:[Document](https://dx.doi.org/10.21437/Interspeech.2023-105),[Link](https://www.isca-archive.org/interspeech_2023/bredin23_interspeech.html)Cited by:[§2\.3](https://arxiv.org/html/2609.30773#S2.SS3.p1.1)\.
- \[2\]C\. Chiang, X\. Wang, C\. Lin, K\. Lin, L\. Li, R\. Kopetz, Y\. Qian, Z\. Wang, Z\. Yang, H\. Lee, and L\. Wang\(2025\)Audio\-aware large language models as judges for speaking styles\.InFindings of the Association for Computational Linguistics: EMNLP 2025,Suzhou, China,pp\. 467–480\.External Links:[Document](https://dx.doi.org/10.18653/v1/2025.findings-emnlp.25),[Link](https://aclanthology.org/2025.findings-emnlp.25/)Cited by:[§3\.2](https://arxiv.org/html/2609.30773#S3.SS2.p4.1)\.
- \[3\]C\. Chien, M\. Orsini, E\. Kharitonov, N\. Zeghidour, K\. Livescu, and A\. Défossez\(2026\)MoshiRAG: asynchronous knowledge retrieval for full\-duplex speech language models\.arXiv preprint arXiv:2604\.12928\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2604.12928),[Link](https://arxiv.org/abs/2604.12928)Cited by:[§1](https://arxiv.org/html/2609.30773#S1.p2.1)\.
- \[4\]A\. Défossez, L\. Mazaré, M\. Orsini, A\. Royer, P\. Pérez, H\. Jégou, E\. Grave, and N\. Zeghidour\(2024\)Moshi: a speech\-text foundation model for real\-time dialogue\.arXiv preprint arXiv:2410\.00037\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2410.00037),[Link](https://arxiv.org/abs/2410.00037)Cited by:[§1](https://arxiv.org/html/2609.30773#S1.p1.1)\.
- \[5\]R\. Y\. He, B\. Cao, C\. Xu, Y\. Liu, and T\. Chen\(2026\)ConversationalVoice: full\-duplex speech data from real conversations through source\-faithful reconstruction and conversation\-grounded expansion\.arXiv preprint arXiv:2609\.08147\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2609.08147),[Link](https://arxiv.org/abs/2609.08147)Cited by:[§1](https://arxiv.org/html/2609.30773#S1.p4.1)\.
- \[6\]Z\. He, W\. Cui, H\. Xu, X\. Li, L\. Zhu, H\. Bai, M\. Shaohua, and I\. King\(2026\)MTR\-DuplexBench: towards a comprehensive evaluation of multi\-round conversations for full\-duplex speech language models\.InFindings of the Association for Computational Linguistics: ACL 2026,San Diego, California, United States,pp\. 5334–5351\.External Links:[Document](https://dx.doi.org/10.18653/v1/2026.findings-acl.263),[Link](https://aclanthology.org/2026.findings-acl.263/)Cited by:[§3\.2](https://arxiv.org/html/2609.30773#S3.SS2.p3.1)\.
- \[7\]S\. Kuroki, Y\. Kubo, T\. Akiba, and Y\. Tang\(2026\)KAME: tandem architecture for enhancing knowledge in real\-time speech\-to\-speech conversational AI\.InProc\. IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),External Links:[Link](https://www.cmsworkshops.com/ICASSP2026/view_paper.php?PaperNum=11317)Cited by:[§1](https://arxiv.org/html/2609.30773#S1.p2.1),[§1](https://arxiv.org/html/2609.30773#S1.p3.1),[§1](https://arxiv.org/html/2609.30773#S1.p4.1),[§2\.1](https://arxiv.org/html/2609.30773#S2.SS1.p1.1),[§3\.1](https://arxiv.org/html/2609.30773#S3.SS1.p2.1)\.
- \[8\]W\. Nakata, Y\. Saito, and H\. Saruwatari\(2026\)DuplexChat: constructing speaker\-separated full\-duplex dialogue speech at scale for spoken dialogue language modeling\.arXiv preprint arXiv:2607\.04941\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2607.04941),[Link](https://arxiv.org/abs/2607.04941)Cited by:[§1](https://arxiv.org/html/2609.30773#S1.p4.1)\.
- \[9\]W\. Nakata, Y\. Saito, Y\. Ueda, and H\. Saruwatari\(2026\)Sidon: fast and robust open\-source multilingual speech restoration for large\-scale dataset cleansing\.InProc\. IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 17617–17621\.Cited by:[§2\.3](https://arxiv.org/html/2609.30773#S2.SS3.p1.1)\.
- \[10\]A\. Ohashi, N\. Zeghidour, A\. Défossez, and E\. Kharitonov\(2026\)Multi\-faceted interactivity alignment in full\-duplex speech models\.arXiv preprint arXiv:2606\.11167\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2606.11167),[Link](https://arxiv.org/abs/2606.11167)Cited by:[§1](https://arxiv.org/html/2609.30773#S1.p4.1)\.
- \[11\]A\. Radford, J\. W\. Kim, T\. Xu, G\. Brockman, C\. McLeavey, and I\. Sutskever\(2023\)Robust speech recognition via large\-scale weak supervision\.InProceedings of the 40th International Conference on Machine Learning,Proceedings of Machine Learning Research, Vol\.202,pp\. 28492–28518\.External Links:[Link](https://proceedings.mlr.press/v202/radford23a.html)Cited by:[§2\.3](https://arxiv.org/html/2609.30773#S2.SS3.p1.1),[§3\.1](https://arxiv.org/html/2609.30773#S3.SS1.p4.1)\.
- \[12\]C\. K\. Reddy, V\. Gopal, and R\. Cutler\(2022\)DNSMOS P\.835: a non\-intrusive perceptual objective speech quality metric to evaluate noise suppressors\.InProc\. IEEE International Conference on Acoustics, Speech and Signal Processing \(ICASSP\),pp\. 886–890\.Cited by:[§2\.3](https://arxiv.org/html/2609.30773#S2.SS3.p1.1)\.
- \[13\]V\. Srinivas, Z\. Englhardt, V\. Iyer, and S\. Patel\(2026\)Thinking while speaking: inference\-time knowledge transfer for responsive and intelligent conversational voice agents\.arXiv preprint arXiv:2511\.07397\.External Links:[Document](https://dx.doi.org/10.48550/arXiv.2511.07397),[Link](https://arxiv.org/abs/2511.07397)Cited by:[§1](https://arxiv.org/html/2609.30773#S1.p2.1)\.
- \[14\]W\. Wang, F\. Wei, L\. Dong, H\. Bao, N\. Yang, and M\. Zhou\(2020\)MiniLM: deep self\-attention distillation for task\-agnostic compression of pre\-trained transformers\.InAdvances in Neural Information Processing Systems,Vol\.33\.External Links:[Link](https://proceedings.neurips.cc/paper/2020/hash/3f5ee243547dee91fbd053c1c4a845aa-Abstract.html)Cited by:[§3\.1](https://arxiv.org/html/2609.30773#S3.SS1.p2.1)\.
- \[15\]L\. Zheng, W\. Chiang, Y\. Sheng, S\. Zhuang, Z\. Wu, Y\. Zhuang, Z\. Lin, Z\. Li, D\. Li, E\. P\. Xing, H\. Zhang, J\. E\. Gonzalez, and I\. Stoica\(2023\)Judging LLM\-as\-a\-judge with MT\-Bench and Chatbot Arena\.InAdvances in Neural Information Processing Systems,Vol\.36\.External Links:[Link](https://proceedings.neurips.cc/paper_files/paper/2023/hash/91f18a1287b398d378ef22505bf41832-Abstract-Datasets_and_Benchmarks.html)Cited by:[§3\.1](https://arxiv.org/html/2609.30773#S3.SS1.p4.1)\.

Similar Articles

Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?

Hugging Face Daily Papers

The paper introduces Box² Bench to measure whether language models can selectively rely on external guidance, finding that capable models remain vulnerable to misleading workflows, and shows that counterfactual SFT plus RL trained only on bad workflows can teach models to resist unreliable guidance and better exploit helpful ones.

Learning Agentic Policy from Action Guidance

arXiv cs.CL

The paper proposes ActGuide-RL, a method for training agentic policies in LLMs by using human action data as guidance to overcome exploration barriers in reinforcement learning without extensive supervised fine-tuning.

How to Guide Your Language Flow

arXiv cs.LG

This paper introduces probe guidance, a new method for flow matching models in continuous diffusion language models, which achieves state-of-the-art performance on unconditional generation and improves multiple choice question answering benchmarks while providing insights into autoguidance.

Boosting Visual Instruction Tuning with Self-Supervised Guidance

Hugging Face Daily Papers

This paper proposes augmenting visual instruction tuning in multimodal language models with self-supervised tasks expressed as natural language instructions, improving vision-centric reasoning without additional architecture or annotations. By reformulating classical self-supervised pretext tasks as image-instruction-response triplets, the method achieves consistent performance improvements across multiple benchmarks by injecting only 3-10% visually grounded instructions into the training data.