Tag
This paper introduces HEAR, a benchmark for evaluating speaker-attributed reasoning in speech language models, and presents A2R, a 30B model optimized with counterfactual data to improve performance on multi-speaker tasks.