Tag
This paper introduces a method for word-level timestamping in Speech Large Language Models using relative time intervals and masked training to enhance prediction accuracy and robustness against noisy real-world annotations.