Tag
The tweet highlights VoiceArena's evaluation method for voice models, which separates Task Completion and Naturalness via blind pairwise voting, and announces Jarvis Bench v0.5 as a conversational agent benchmark.