Tag
The paper introduces a transparent framework that maps acoustic speech features to DSM-5 depression indicators for interpretable detection, running locally on commodity hardware to preserve privacy.
This paper analyzes spontaneous dyadic Zoom conversations using multimodal features (acoustic, facial, turn-taking) to identify markers of perceived conversational success, finding that entrainment in speech and facial movements correlates with higher interaction quality.