Tag
Kyutai Labs announces that their audio-to-MIDI model MuScriptor now detects tempo, allowing direct drag-and-drop of MIDI into a DAW without manual tempo matching.
Kyutai released Pocket TTS, a text-to-speech model capable of cloning a voice from just 5 seconds of audio, running on CPU and released under the MIT license. It was benchmarked against Kokoro, Supertonic, and Inflect-Nano for English TTS.
Kyutai presents Hibiki-Zero, a real-time speech-to-speech translation model, at ICML 2026 in Seoul, with an oral presentation scheduled for July 8.
This paper introduces Continuous Audio Language Models (CALM), which generate audio using continuous frames instead of discrete tokens to improve fidelity and reduce computational cost in speech and music generation.