Tag
PD-GS introduces a phoneme-driven 3D Gaussian Splatting approach for audio-driven talking heads, using a Linguistic Fusion Module to improve lip articulation and reduce closure violations.
This paper investigates the effects of phonetic versus character targets and selective state-space models (Mamba) for intracortical brain-to-text decoding, finding that a GRU-based recurrent decoder remains the strongest performer on the Brain-to-Text '25 benchmark.
This paper evaluates demographic and accent biases in phoneme-based ASR systems, specifically WhisperIPA and ZIPA, using phoneme error rate and a new Soft PER metric, revealing persistent disparities across languages and groups.