Tag
This paper introduces randomized guidance as an efficient training method for tandem speech-to-speech models, enabling them to learn natural conversational behavior directly from real conversations without simulating backend LLM behavior.