@alamin_ai_: OMG, guys, this is unbelievable Please listen to the Levantine Arabic and the seamless code-switching with english, a 7…
Summary
A significant breakthrough in Levantine Arabic speech synthesis and English code-switching, achieving a 76% improvement using a single RTX 3060 in an evening.
View Cached Full Text
Cached at: 07/07/26, 02:13 AM
OMG, guys, this is unbelievable
Please listen to the Levantine Arabic and the seamless code-switching with english, a 76% improvement, love research man And all of this was done with a single RTX 3060 in anevening
THANK YOU, @TheAhmadOsman all my savings FORbuying more GPUs now https://t.co/8XHsXPdCCI
al’amin ai (@alamin_ai_): This is a breakthrough, run 3 finished, the deterministic duration head, evaluated at natural pacing (length_scale≈1, ratio 0.99), gave a massive improvement, because the old 3× stretch was wrecking both the smoothness and the intelligibility? see this table
Similar Articles
Benchmarking Commercial ASR Systems on Code-Switching Speech: Arabic, Persian, and German
This paper presents a benchmark evaluating five commercial ASR systems on code-switching speech across Arabic-English, Persian-English, and German-English pairs, using a two-stage pipeline to select 300 samples per pair and assessing performance with WER and BERTScore. ElevenLabs Scribe v2 achieves the lowest overall WER (13.2%) and highest BERTScore (0.936), with public dataset available.
CohereLabs/cohere-transcribe-arabic-07-2026
Cohere releases an open-source 2B parameter Arabic ASR model optimized for Arabic dialect performance and Arabic-English code-switching, based on the Conformer encoder-decoder architecture.
Dziri Voicebot: An End-to-End Low-Resource Speech-to-Speech Conversational System for Algerian Dialect
This paper presents a modular end-to-end speech-to-speech conversational system for the low-resource Algerian Dialect, integrating ASR, NLU, RAG, and TTS with dedicated datasets and fine-tuned models.
@YichiZ03: https://x.com/YichiZ03/status/2078588932191895976
MOSS-TD, a speaker-aware ASR system, is optimized within the SGLang-Omni serving stack, enabling 38-minute multi-speaker audio to be transcribed in about 49 seconds on a single H100, with concurrent processing of 16 meetings.
I swapped the TTS in my voice agent and it cut the lag people actually feel more than anything else
The author shares their experience swapping the TTS in their voice agent to a custom model (Banter 1) designed for bilingual Arabic-English conversations, which significantly reduced perceived lag.