Tag
This paper addresses the uncontrolled latent variable of transcription style in ASR models by introducing a method using coverage-aware decoder task tokens to enable controllable verbatim transcription with accurate word-level timing, achieving high disfluency detection F1 via zero-shot cross-lingual transfer.