speech-translation

Tag

Cards List
#speech-translation

When to Use Extra Context: Evidence-Grounded Terminology Adaptation for Simultaneous Speech Translation

arXiv cs.CL · 2026-07-21 Cached

This paper proposes EGTA, an Evidence-Grounded Terminology Adaptation framework for simultaneous speech translation that selectively uses document-specific terminology to improve translation of rare terms, achieving significant gains in named-entity and acronym recall without full-model fine-tuning.

0 favorites 0 likes
#speech-translation

@t0m1ab: Heading to ICML 2026 in Seoul next week with @romfbr31 to present Hibiki-Zero[https://kyutai.org/blog/2026-02-12-hibiki…

X AI KOLs Following · 2026-06-30 Cached

Kyutai presents Hibiki-Zero, a real-time speech-to-speech translation model, at ICML 2026 in Seoul, with an oral presentation scheduled for July 8.

0 favorites 0 likes
#speech-translation

PiDA: Phonetically-Informed Data Augmentation for Robust Vietnamese Speech Translation

arXiv cs.CL · 2026-06-12 Cached

This paper presents PiDA, a phonetically-informed data augmentation method for Vietnamese speech translation that improves robustness by generating ASR-like corruptions using phonetic word embeddings, achieving up to +2.04 BLEU on noisy outputs.

0 favorites 0 likes
#speech-translation

Krisp Voice Translation API

Product Hunt · 2026-06-05

Krisp launches a real-time speech-to-speech translation API designed for high accuracy.

0 favorites 0 likes
#speech-translation

OpenSTBench: Beyond Semantic Evaluation for Speech Translation

Hugging Face Daily Papers · 2026-05-29

OpenSTBench is a unified multidimensional evaluation framework for speech translation systems that jointly assesses translation quality, speech quality, speaker preservation, emotion fidelity, and latency across both S2TT and S2ST systems in offline and streaming settings. The framework addresses the gap left by fragmented evaluation protocols and provides a reproducible benchmark for comparing heterogeneous speech translation systems.

0 favorites 0 likes
#speech-translation

MoVE: Translating Laughter and Tears via Mixture of Vocalization Experts in Speech-to-Speech Translation

Hugging Face Daily Papers · 2026-04-19 Cached

MoVE proposes a Mixture-of-LoRA-Experts architecture that preserves laughter and crying in speech-to-speech translation, achieving 76% NV retention with only 30 minutes of curated data.

0 favorites 0 likes
← Back to home

Submit Feedback