Shifting Relational Paradigms for Affective Computing: Affective Resonance, Vitality Affects, and Vocal Interaction Fields
Summary
The paper proposes a relational framework for affective computing that shifts focus from individual emotion states to interactional fields in vocal dynamics, supported by an empirical study using self-supervised speech representations to detect directional expressive coupling.
View Cached Full Text
Cached at: 09/11/26, 08:41 AM
# Shifting Relational Paradigms for Affective Computing: Affective Resonance, Vitality Affects, and Vocal Interaction Fields
Source: [https://arxiv.org/html/2609.09864](https://arxiv.org/html/2609.09864)
Gorman Yao
###### Abstract
Affective computing has largely followed an individual\-state paradigm, extracting discrete emotion labels or arousal/valence from isolated speakers\. We argue this framing is incomplete for interaction\. Drawing on affective resonance and vitality\-contour accounts, we propose a relational framework in which the primary unit of affective analysis is the interactional field constituted within vocal dynamics\. As a proof of concept, we present a preliminary empirical study using continuous self\-supervised speech representations to detect directional expressive coupling in multi\-party conversation\. Coupling is regime\-specific, concentrated at sub\-second timescales, and collapses under exclusive\-speech negative controls, consistent with a relational account of affective dynamics\. We introduce design frameworks for Artificial Affective Resonance Intelligence grounded in Affective Resonance Dynamic Ontologies, supported by null\-calibrated directional coupling analyses across interaction regimes\.
###### keywords
affective resonance, vitality affects, vitality contours, vocal interaction field, ARDO, AARI, affective computing, HCI design, enactivism, pre\-semantic affect, multi\-party speech, Granger causality, neo\-cybernetics, WavLM, participatory sense\-making, affective attunement, social robotics
††address:1Nurobodi, Australia††email:info@nurobodi\.com## 1Introduction
Affective computing — the endeavour of building systems that recognise, model, and respond to human emotion — has largely proceeded by asking what emotion a person is likely experiencing within a given interactional context\. Dominant pipelines map acoustic, linguistic, physiological, or biometric signals to discrete emotion categories or continuous arousal/valence dimensions\[[1](https://arxiv.org/html/2609.09864#bib.bib1),[2](https://arxiv.org/html/2609.09864#bib.bib2)\]\. This paradigm remains productive, but recent advances in cybernetic architectures, multimodal sensing, and interaction design open a complementary modelling focus: treating relational dynamics as a primary site of affective organisation\. The relational field is not restricted to any single expressive modality; here, we examine it through vocal dynamics as a tractable and theoretically motivated domain\. As Kappas and Gratch\[[3](https://arxiv.org/html/2609.09864#bib.bib3)\]observe, affective computing lacks robust models of how affective behaviour unfolds relationally over time\. Mühlhoff\[[4](https://arxiv.org/html/2609.09864#bib.bib4)\]proposes that affective resonance is not the outcome of individuals converging toward similar internal states, but an emergent property of the interactional field they co\-constitute — a dynamic individual\-convergence models cannot capture because they treat affect as an attribute of persons rather than a property of interaction\. Acoustic entrainment research demonstrates convergence in vocal features over time\[[5](https://arxiv.org/html/2609.09864#bib.bib5),[6](https://arxiv.org/html/2609.09864#bib.bib6),[7](https://arxiv.org/html/2609.09864#bib.bib7)\], including via specific interactional events such as backchannels\[[8](https://arxiv.org/html/2609.09864#bib.bib8)\], but typically operationalises coupling as growing feature similarity between speaker trajectories rather than directionality, interactional configuration, or field\-level emergence\. Regime\-dependent coupling that collapses under exclusive speech would indicate that interactional structure, not individual state, governs expressive dynamics — a pattern existing frameworks cannot detect\. Granger\-causal approaches to speech coupling remain comparatively rare and are typically bivariate and single\-feature\[[9](https://arxiv.org/html/2609.09864#bib.bib9),[10](https://arxiv.org/html/2609.09864#bib.bib10)\], with no prior work, to our knowledge, combining conditional multi\-speaker models with interaction\-regime decomposition\. Emotion Recognition in Conversations extends affective computing to multi\-party settings but aggregates individual\-state inferences rather than modelling speakers’ relational dynamic\[[3](https://arxiv.org/html/2609.09864#bib.bib3)\]\. Pre\-semantic vocal resonance across interacting speakers remains underexplored as a computational target despite its role as a direct marker of vitality dynamics\[[11](https://arxiv.org/html/2609.09864#bib.bib11),[12](https://arxiv.org/html/2609.09864#bib.bib12)\]\. We propose a relational reorientation toward field\-level models of human interaction, grounded in Mühlhoff’s account of affective resonance\[[4](https://arxiv.org/html/2609.09864#bib.bib4)\], Stern’s theory of vitality affects and contours\[[11](https://arxiv.org/html/2609.09864#bib.bib11)\], as extended by Køppe et al\[[12](https://arxiv.org/html/2609.09864#bib.bib12)\], and enactivist accounts of participatory sense\-making\[[13](https://arxiv.org/html/2609.09864#bib.bib13),[14](https://arxiv.org/html/2609.09864#bib.bib14)\]\. We ground this position theoretically, demonstrate it through directional coupling analysis of vocal interaction, and develop design implications — Affective Resonance Dynamic Ontologies \(ARDO\) and Artificial Affective Resonance Intelligence \(AARI\) — for systems oriented toward interactional dynamics rather than individual state inference\. Our empirical contribution comprises: \(1\) a null\-calibrated directional coupling framework correcting for size inflation in short\-window autoregressive testing; \(2\) an interaction\-regime decomposition enabling within\-corpus empirical controls; \(3\) a multi\-speaker conditional analysis reducing spurious dyadic coupling from group\-level confounds; and \(4\) results on the AMI Meeting Corpus showing coupling concentrated at sub\-second timescales under mutual co\-presence, while macro\-scale effects are provisional\. This work directly addresses the Interspeech 2026 theme, Speaking Together, by treating speech as interactional coordination rather than individual production\.
## 2Theoretical foundations
### 2\.1Affective resonance as emergent field property
Mühlhoff\[[4](https://arxiv.org/html/2609.09864#bib.bib4)\]proposes that affective resonance names a field of phenomena in which interactional dynamic flux constitutes affective experience, rather than transmitting pre\-existing inner states between independently existing individuals\. He identifies three constitutive axioms: \(1\) affective resonance is a dynamical entanglement of moving and being\-moved in relation; \(2\) it is primarily experienced as an immanent force arising within relational interplay; and \(3\) it is an emergent, creative dynamic, producing its own vectors of movement\-in\-relation, rather than operating within pre\-formed affective state spaces defined by pre\-categorised emotion labels\. This reflects an ontological commitment distinct from convergence\-based paradigms where entrainment research typically operationalises coupling as convergence in feature values\[[5](https://arxiv.org/html/2609.09864#bib.bib5),[6](https://arxiv.org/html/2609.09864#bib.bib6)\], treating interaction as alignment between individual trajectories rather than as a field\-level dynamic with its own organisation\. By contrast, Mühlhoff’s account gives ontological primacy to the relational dynamic — subjectivity and individual affective experience arising from life\-long histories of being\-in\-relation — rather than to pre\-formed individuals who subsequently interact\. Affective resonance is therefore distinct from imitation, synchrony, or mimicry; it does not require congruence of form, but the immanent connectivity of forces that constitutes the field as an emergent property of interaction\. Understanding this temporal\-dynamic texture requires a finer\-grained account of the expressive substrate through which resonance becomes perceptible\.
### 2\.2Vitality affects, vitality contours and pre\-semantic voice
Stern\[[11](https://arxiv.org/html/2609.09864#bib.bib11)\]introduced vitality affects as dynamic, non\-categorical qualities and later described vitality contours—surging, fading, rushing, bursting—as their temporal form\. Køppe, Harder, and Væver\[[12](https://arxiv.org/html/2609.09864#bib.bib12)\]clarify three central dimensions: temporality, intensity, and form\. These dynamics are amodal and cross\-modal; here we focus on voice as both a theoretical and practical carrier\. Vocalics \(e\.g\., timbre, formants, harmonic ratios, jitter, shimmer\) form a continuous energetic substrate, while prosody organises it through intonation, stress, rhythm, and pitch contour\. We use*prosodic*broadly to emphasise the pre\-semantic organisation of temporal\-affective contours, potentially underpinned by microtonal and resonant features\. Their continuous co\-modulation traces the “temporal feeling shapes” through which vitality becomes interactionally available\[[11](https://arxiv.org/html/2609.09864#bib.bib11)\]\. This motivates retaining continuous hidden\-state representations rather than discrete tokens\. WavLM\[[15](https://arxiv.org/html/2609.09864#bib.bib15)\], HuBERT\[[16](https://arxiv.org/html/2609.09864#bib.bib16)\], and wav2vec 2\.0\[[17](https://arxiv.org/html/2609.09864#bib.bib17)\]encode paralinguistic and prosodic information without task annotation, as demonstrated by SUPERB\[[18](https://arxiv.org/html/2609.09864#bib.bib18)\]and speech\-emotion studies\[[19](https://arxiv.org/html/2609.09864#bib.bib19),[20](https://arxiv.org/html/2609.09864#bib.bib20)\]\. Continuous representations preserve temporal dynamics that clustering can suppress: discrete speech tokens tend toward phonetic invariance at the expense of prosodic correlation\[[21](https://arxiv.org/html/2609.09864#bib.bib21),[22](https://arxiv.org/html/2609.09864#bib.bib22),[23](https://arxiv.org/html/2609.09864#bib.bib23)\]\. This matters because vitality is upstream of categorical emotion\. Emotional prosody elicits neural responses at roughly 200 ms\[[24](https://arxiv.org/html/2609.09864#bib.bib24)\], before semantic processing around 300–400 ms\[[25](https://arxiv.org/html/2609.09864#bib.bib25)\]\. Category\-based prosody and emotion systems\[[2](https://arxiv.org/html/2609.09864#bib.bib2),[20](https://arxiv.org/html/2609.09864#bib.bib20),[26](https://arxiv.org/html/2609.09864#bib.bib26)\]therefore model derived representations rather than the force, movement, directionality, and aliveness through which affective experience is organised\. Computational access is most consequential at this pre\-semantic level, where the interaction field’s temporal, intensive, and formal modulations are directly expressed\.
### 2\.3Enactivism and the relational subject
The enactivist tradition\[[13](https://arxiv.org/html/2609.09864#bib.bib13)\]grounds cognition in meaning brought forth through embodied interaction rather than internal representation\. De Jaegher and Di Paolo\[[14](https://arxiv.org/html/2609.09864#bib.bib14)\]extend this through participatory sense\-making: meaning is jointly constituted in interaction itself, not carried by pre\-formed individuals\. The relational subject emerges within interaction rather than prior to it — a claim that directly shapes what it means to design a system that participates in an interactional field rather than observes one\.
### 2\.4ARDO and AARI: interaction field modelling
These foundations motivate a working design distinction\. Affective Resonance Dynamic Ontologies \(ARDO\) names interaction\-level organisation that becomes salient when sustained vocal dynamics are modelled as a field rather than as independent states\. Following Mühlhoff\[[4](https://arxiv.org/html/2609.09864#bib.bib4)\], discrete states attributed to people or an AI are downstream products of resonance, not its preconditions\. This commitment shapes feature selection, training objectives, and outputs, and makes ARDO a frame for testing whether interaction dynamics contain organisation irreducible to individual trajectories\. Artificial Affective Resonance Intelligence \(AARI\) names systems built within that frame: they map the field’s vitality trajectories and generate responses calibrated to its modulation rather than classify individual affect\. Conventional categorical or dimensional systems\[[2](https://arxiv.org/html/2609.09864#bib.bib2)\]expose confidence scores over fixed labels; useful as interfaces, these presuppose the categories whose emergence a relational account seeks to explain\. ARDO instead targets continuous tension, volatility, and momentum that may later organise into recognisable emotions; AARI is the corresponding architectural programme\.
## 3Empirical illustration
### 3\.1Rationale
The framework predicts that directional expressive coupling should appear when mutually oriented speakers constitute an interactional field and weaken when that condition is removed\. Tier A captures simultaneous co\-presence; Tier B pools non\-overlapping single\-speaker activity, which may include sequential engagement\[[14](https://arxiv.org/html/2609.09864#bib.bib14)\]; Tier C separates that activity by speaker direction\. We test this through interaction\-regime decomposition, providing within\-corpus empirical controls\. The study is a conceptual proof of principle; complete implementation details are reserved for subsequent work\.
### 3\.2Method
Data:We analyse four\-speaker meeting audio from the AMI Meeting Corpus\[[27](https://arxiv.org/html/2609.09864#bib.bib27)\]using close\-talking headset channels, resampled to 16 kHz and processed in non\-overlapping 60\-second windows\. Voice activity detection is applied per speaker and mapped to the WavLM feature\-frame grid; the effective frame hop is determined by the model’s temporal downsampling \(≈\\approx20 ms at 16 kHz\), and we align masks by window\-length\-to\-T mapping to ensure exact synchrony; all results reported here are AMI\-only, with generalisation to other corpora reserved for future work\.
Feature extraction:We extract continuous hidden\-state representations from the 12 transformer layers of WavLM\-Base\+\[[15](https://arxiv.org/html/2609.09864#bib.bib15)\]— a self\-supervised speech model pre\-trained on large\-scale audio\. \(The HuggingFace API also returns an initial embedding state; in our notation we usehidden\_states\[1\.\.12\]as the transformer layers\.\) Crucially, we do not apply quantisation or K\-means clustering, which discard prosodic and affective information in favour of phonetic invariance\[[21](https://arxiv.org/html/2609.09864#bib.bib21),[22](https://arxiv.org/html/2609.09864#bib.bib22)\]\. We retain continuous floating\-point hidden states, which preserve the dense entanglement of acoustic, prosodic, and paralinguistic information across representational levels — the representational substrate ARDO requires for tracking vitality dynamics\. Our primary feature — expressiveness — is a cross\-layer activation dispersion statistic derived from WavLM hidden\-state dynamics, intended to capture moment\-to\-moment richness and complexity of vocal expression beyond overall magnitude\.
Multi\-dimensional proxies and energy control:Beyond a scalar energy proxy, the same representation pass yields aligned series for cross\-representational expressiveness, high\-level semantic magnitude, and frame\-to\-frame topic drift\. All share the same feature\-frame axis, windowing, and regime decomposition\. A regime\-contiguous episode is a maximal uninterrupted sequence of frames within a window for which a selected pre\-specified dyadic VAD mask remains true\. To test whether expressiveness merely shadows the energy proxy, we residualise it against energy within each regime\-contiguous episode before directional testing and null calibration\. This targets temporal\-dynamic expressive richness, consistent with vitality contours \(Sec\. 2\.2\), rather than categorical affect\.
Interaction condition decomposition:For each speaker pair within each window, we define three conditions via binary voice activity masks, with minor temporal smoothing applied to bridge brief within\-speaker gaps\. Tier A denotes overlapping speech\. Tier B pools non\-overlapping single\-speaker activity across both active\-speaker directions\. Tier C separates that non\-overlapping activity according to which speaker is active\. These conditions compare coupling across overlapping, pooled non\-overlapping, and direction\-specific non\-overlapping activity\. Tier C provides an AMI\-specific empirical control by separating non\-overlapping activity according to which speaker is active\.
Figure 1:Interaction conditions and expected coupling visibility\. Tier A denotes overlapping speech; Tier B pools non\-overlapping single\-speaker activity; Tier C separates that activity by active\-speaker direction and is used as an AMI\-specific empirical control\.Causal analysis:We test directional Granger causality\[[28](https://arxiv.org/html/2609.09864#bib.bib28)\]using bivariate and 4\-speaker conditional vector autoregression \(VAR\) models\[[10](https://arxiv.org/html/2609.09864#bib.bib10),[28](https://arxiv.org/html/2609.09864#bib.bib28)\]\. In the conditional model, each dyadic test conditions on the other two speakers to reduce group\-level confounds such as shared laughter, room events, and topic shifts\. Granger tests are used here to assess linear predictive dependence within a VAR framework, not to claim that affective interaction dynamics are themselves linear; nonlinear extensions such as transfer entropy remain a future direction\. Within 60 s windows, VAR/Granger is applied to regime\-contiguous episodes; their short duration makes lag\-selection size inflation a central concern\. We therefore compare observed rejection behaviour with matched circular\-shift nulls and report null\-quantile calibrated excess rejections \(OBS–NULL, percentage points; Fig\. 2\)\. Here,nobsn\_\{\\mathrm\{obs\}\}counts episode\-level tests and can exceed the number of contributing windows; negative values mean fewer rejections than the calibrated null\. Window\-level directionality is summarised by a signed Directionality Support Score \(DSS∈\{−1,0,\+1\}\\in\\\{\-1,0,\+1\\\}\) after false\-discovery\-rate control\[[29](https://arxiv.org/html/2609.09864#bib.bib29)\], using both Window\-BH and conservative Global\-FDR\. We call lags≤1\\leq 1s micro and those\>1\>1s macro; exact lag grids and gates are implementation\-specific\.
Table 1:Directionality Support Score \(DSS\) distribution \(% of windows with DSS=−1,0,\+1=\-1,0,\+1\) under within\-condition correction \(Window\-BH\) and conservative global correction \(Global\-FDR\), for micro\- and macro\-scale lags\.nndenotes the number of evaluated windows per tier\.
### 3\.3Results
Directional coupling was reliably detectable at sub\-second lags under interaction regimes involving mutual orientation\. Window\-level directionality is sparse but structured: micro\-scale overlap \(Tier A\) shows the highest non\-zero support mass, with directional wins that are near\-symmetric across directions \(Table 1\)\. Tier B \(pooled non\-overlapping activity\) and Tier C \(directional exclusive activity\) shift strongly toward zero support, consistent with weaker or absent dyadic coupling outside mutual co\-presence\. Matched circular\-shift results confirm short\-series size inflation \(Table 2\), motivating calibrated excess rejection rates \(Fig\. 2\)\. Atq=5%q=5\\%, Tier A micro shows approximately \+6\.6 percentage points under the bivariate model and \+7\.4 points under energy\-controlled conditional analysis\. At the stricterq=1%q=1\\%operating point \(not shown\), effects attenuate but remain regime\-specific: Tier A is strongest, Tier B weaker, and Tier C approaches the null\. Coupling therefore cannot be reduced to synchrony in the energy proxy alone and is concentrated at short temporal offsets relevant to an AARI system\. Non\-zero directionality is feature\-selective across regime and timescale; small but structured cross\-feature effects, tested bivariately, support a multidimensional account of vitality dynamics beyond a single scalar energy mechanism
Table 2:Shift\-null rejection rates under nominal thresholds for the full shift\-null grid \(nnulln\_\{\\mathrm\{null\}\}shown\), illustrating size inflation in short\-series lag selection settings and motivating null calibration\.Figure 2:Calibrated excess significance \(OBS−\-NULL,q=5%q=5\\%, percentage points\) across representative operating points\. Filled markers denote bivariate tests; open markers denote 4\-speaker control; values<0<0indicate fewer rejections than the null baseline\.
### 3\.4Limitations
The expressiveness proxy measures modulation in latent vocal\-production dynamics, not valence or discrete emotion; this is deliberate\. Granger analysis captures linear predictive dependence but does not distinguish convergent from divergent coupling, for which VAR coefficient signs are a principled next step\. Results are AMI\-only, limiting immediate generalisation, although they establish regime\- and scale\-conditioned coupling within formal four\-speaker AMI meeting interaction\. Cross\-corpus analysis has been completed and full results will be reported in subsequent work\. Tier B macro analysis is underpowered \(n=18 gated windows\), and the DSS depends on FDR and minimum run\-fraction parameters; conservative Global\-FDR provides only partial robustness\. Future work should expand sensitivity analysis, corpora, and features targeting temporality, intensity, and form\[[12](https://arxiv.org/html/2609.09864#bib.bib12)\], including multimodal extensions warranted by the amodal scope of vitality affects\.
## 4Design implications
The empirical results described above demonstrate that the interactional field leaves measurable traces: coupling is constituted by conversational configuration rather than carried into interaction by individual speakers — a finding that individual\-state pipelines are structurally unable to exploit\. This opens the design question of how to build systems responsive to the interactional field as such, rather than to individual trajectories that only approximate it\. This is where ARDO and AARI move from theoretical proposals to concrete design constraints\. A system designed to participate in a resonance field rather than observe it from outside must make fundamentally different computational choices\. Conventional affective computing pipelines infer individual affective states and respond to those inferences\. By contrast, an AARI system is situated within the interactional field: its outputs function as contributions to the evolving relational dynamic rather than as responses to a diagnosis\. This reorientation has consequences for training objectives\. Rather than minimising classification error against discrete emotion labels, systems oriented by ARDO are evaluated in terms of the quality of relational coupling they sustain— whether they maintain asymmetric, responsive attunement to the user’s affective tempo, analogous to the dynamics described in Stern’s infant–caregiver model\[[11](https://arxiv.org/html/2609.09864#bib.bib11)\]\. Emotion labels remain useful as diagnostic and benchmark interfaces but are not the primary object of ARDO\. The vocalic layer is where this participation becomes concrete\. Treating vocalics as a design medium rather than solely as a diagnostic signal opens possibilities that categorical\-output systems foreclose\. Juslin and Laukka\[[30](https://arxiv.org/html/2609.09864#bib.bib30)\]show that vocal expression and musical performance share acoustic cue patterns for communicating emotion, suggesting that expressive dynamics in vocal interaction may be analysable using formal tools developed for musical structure\. Concepts such as tension and release, consonance and dissonance can, in principle, be applied to moment\-to\-moment contours of multi\-party vocal interaction\. The human voice exhibits a harmonic series and shifts in the distribution of energy across that series—and in relation to another voice—carry affect\-relevant information that categorical labels compress\. We treat this as a working hypothesis consistent with evidence that affective processing precedes semantic integration in auditory cortex\[[24](https://arxiv.org/html/2609.09864#bib.bib24),[25](https://arxiv.org/html/2609.09864#bib.bib25)\]\. The PIAT/ARTIST model\[[31](https://arxiv.org/html/2609.09864#bib.bib31)\]provides a precedent for architectures that internalise tonal structure from unsupervised exposure, suggesting self\-supervised learning of vitality contour structure without categorical emotion supervision as a plausible design direction\. Extended toward generative systems, this points toward Human\-AI cybernetic loops in which a synthetic vocal output participates in, rather than merely observes, the shared resonance field\. ARDO formalises, at the level of design, implications already articulated across Mühlhoff’s relational ontology\[[4](https://arxiv.org/html/2609.09864#bib.bib4)\], enactivist participatory sense\-making\[[14](https://arxiv.org/html/2609.09864#bib.bib14)\], and Stern’s developmental account of affective attunement\[[11](https://arxiv.org/html/2609.09864#bib.bib11)\]\. What ARDO adds is an evaluative standard: whether a computational approach models upstream resonance dynamics or only their downstream categorical interfaces\. AARI foregrounds this distinction architecturally\. Compassion\-focused technology frameworks\[[32](https://arxiv.org/html/2609.09864#bib.bib32)\]and relational ethics accounts of data\-centric systems\[[33](https://arxiv.org/html/2609.09864#bib.bib33)\]converge on the same insight: that systems operating within human affective life require attunement as a design condition, not an ethical afterthought\. An AARI system instantiates this orientation at the level of design: responding to vitality dynamics within a resonance field rather than extracting affective signals for classificatory or commercial purposes\. Entrainment research models convergence between individually defined speakers\[[5](https://arxiv.org/html/2609.09864#bib.bib5),[6](https://arxiv.org/html/2609.09864#bib.bib6),[7](https://arxiv.org/html/2609.09864#bib.bib7),[8](https://arxiv.org/html/2609.09864#bib.bib8)\]; AARI treats the interaction field as the primary object of modelling, with individual trajectories as observable traces of field\-level dynamics\.
## 5Conclusions
Affective computing requires a reorientation from modelling individual affective states to modelling interactional fields as primary sites of affective organisation\. Mühlhoff’s account of affective resonance\[[4](https://arxiv.org/html/2609.09864#bib.bib4)\], Stern and Køppe’s theory of vitality affects\[[11](https://arxiv.org/html/2609.09864#bib.bib11),[12](https://arxiv.org/html/2609.09864#bib.bib12)\], and enactivist approaches to participatory sense\-making\[[13](https://arxiv.org/html/2609.09864#bib.bib13),[14](https://arxiv.org/html/2609.09864#bib.bib14)\]provide the philosophical substrate for this shift\. Our Empirical Illustration supports this reframing by showing that vocal coupling is regime\-conditioned and temporally localised, collapsing when mutual orientation is removed — consistent with a relational account of affective resonance rather than individual\-state convergence\. ARDO and AARI, introduced above, articulate the design implications of this reorientation: treating the interaction field, rather than individual states, as the primary object of modelling, and the vocalic layer as a site of expressive responsiveness, not a diagnostic channel\.
## 6Acknowledgments
The authors thank CSIRO, including the Innovate to Grow and ON Prime programme teams, for their support of Nurobodi’s research and development, and ongoing research translation for industry\. We also thank Dr\. Haytham Fayek for his valuable guidance and feedback, and for his ongoing encouragement to pursue this area of affective computing research\.
## 7Generative AI Use Disclosure
Generative AI tools \(including large language models\) were used to assist with editing, structural clarity, structural refinement, analysis and summarisation\. AI tools were also used to support documentation of analysis workflows and preparation of visualisation scripts\. All theoretical framing, experimental design and rationale, and implementation were produced by the authors, who retain full intellectual ownership and responsibility for the content of this work\.
## References
- \[1\]R\. W\. Picard, Affective Computing\. Cambridge, MA: MIT Press, 1997\.
- \[2\]A\. S\. Cowen et al\., ”The primacy of categories in the recognition of 12 emotions in speech prosody across two cultures,” Nature Human Behaviour, vol\. 3, no\. 4, pp\. 369\-381, 2019\.
- \[3\]A\. Kappas and J\. Gratch, ”These aren’t the droids you are looking for: Promises and challenges for the intersection of affective science and robotics/AI,” Affective Science, vol\. 4, pp\. 580\-585, 2023\.
- \[4\]R\. Mühlhoff, ”Affective resonance and social interaction,” Phenomenology and the Cognitive Sciences, vol\. 14, no\. 4, pp\. 1001– 1019, 2015\. doi: 10\.1007/s11097\-014\-9394\-7
- \[5\]C\. J\. Wynn and S\. A\. Borrie, ”Classifying conversational entrainment of speech behavior: An expanded framework and review,” Journal of Phonetics, vol\. 94, p\. 101173, 2022\.
- \[6\]R\. Levitan and J\. Hirschberg, ”Measuring acoustic\-prosodic entrainment with respect to multiple levels and dimensions,” in Proc\. Interspeech 2011, pp\. 3081\-3084\.
- \[7\]J\. Kejriwal, S\. Benus, and L\. M\. Rojas\-Barahona, ”Unsupervised auditory and semantic entrainment models with deep neural networks,” in Proc\. Interspeech 2023, pp\. 2628\-2632\.
- \[8\]K\. Ochi, K\. Inoue, D\. Lala, and T\. Kawahara, ”Entrainment analysis and prosody prediction of subsequent interlocutor’s backchannels in dialogue,” in Proc\. Interspeech 2024, pp\. 462\-466\.
- \[9\]A\. K\. Seth, ”A MATLAB toolbox for Granger causal connectivity analysis,” Journal of Neuroscience Methods, vol\. 186, no\. 2, pp\. 262\-273, 2010\.
- \[10\]K\. J\. Blinowska, R\. Kus, and M\. Kaminski, ”Granger causality and information flow in multivariate processes,” Physical Review E, vol\. 70, no\. 5, p\. 050902, 2004\.
- \[11\]D\. N\. Stern, The Interpersonal World of the Infant\. New York: Basic Books, 1985; Forms of Vitality\. Oxford: Oxford University Press, 2010\.
- \[12\]S\. Køppe, S\. Harder, and M\. Væver, ”Vitality affects,” International Forum of Psychoanalysis, vol\. 17, no\. 3, pp\. 169–179, Sep\. 2008\. doi: 10\.1080/08037060701650453
- \[13\]F\. J\. Varela, E\. Thompson, and E\. Rosch, The Embodied Mind: Cognitive Science and Human Experience\. Cambridge, MA: MIT Press, 1991\.
- \[14\]H\. De Jaegher and E\. Di Paolo, ”Participatory sense\-making,” Phenomenology and the Cognitive Sciences, vol\. 6, no\. 4, pp\. 485\-507, 2007\.
- \[15\]S\. Chen, C\. Wang, Z\. Chen et al\., ”WavLM: Large\-scale self\-supervised pre\-training for full stack speech processing,” IEEE JSTSP, vol\. 16, no\. 6, pp\. 1505\-1518, 2022\.
- \[16\]W\.\-N\. Hsu, B\. Bolte, Y\.\-H\. H\. Tsai, K\. Lakhotia, R\. Salakhutdinov, and A\. Mohamed, ”HuBERT: Self\-supervised speech representation learning by masked prediction of hidden units,” IEEE/ACM TASLP, vol\. 29, pp\. 3451\-3460, 2021\.
- \[17\]A\. Baevski, Y\. Zhou, A\. Mohamed, and M\. Auli, ”wav2vec 2\.0: A framework for self\-supervised learning of speech representations,” in Proc\. NeurIPS, vol\. 33, pp\. 12449\-12460, 2020\.
- \[18\]S\.\-W\. Yang, P\.\-H\. Chi, Y\.\-S\. Chuang et al\., ”SUPERB: Speech processing universal performance benchmark,” in Proc\. Interspeech 2021, pp\. 1194–1198\.
- \[19\]E\. Morais, R\. Hoory, W\. Zhu et al\., ”Speech emotion recognition using self\-supervised features,” in Proc\. IEEE ICASSP, 2022, pp\. 6922\-6926\.
- \[20\]Z\. Ma, Z\. Zheng, J\. Ye et al\., ”emotion2vec: Self\-supervised pre\-training for speech emotion representation,” in Findings of ACL, Aug\. 2024, pp\. 15747\-15760\.
- \[21\]Y\. Guo, Z\. Li, H\. Wang et al\., ”Recent advances in discrete speech tokens: A review,” arXiv preprint arXiv:2502\.06490, 2025\.
- \[22\]K\. Choi, A\. Pasad, T\. Nakamura et al\., ”Self\-supervised speech representations are more phonetic than semantic,” in Proc\. Interspeech 2024, pp\. 4578\-4582\.
- \[23\]D\. Wells, H\. Tang, and K\. Richmond, ”Phonetic analysis of self\-supervised representations of English speech,” in Proc\. Interspeech 2022, pp\. 3583\-3587\.
- \[24\]S\. Paulmann and S\. A\. Kotz, ”Early emotional prosody perception based on different speaker voices,” Neuroreport, vol\. 19, no\. 3, pp\. 209\-213, 2008\.
- \[25\]S\. A\. Kotz and T\. Schwartze, ”Cortical speech processing unplugged: a timely subcortico\-cortical framework,” Trends in Cognitive Sciences, vol\. 14, no\. 9, pp\. 392\-399, 2010\.
- \[26\]K\. R\. Scherer, ”Vocal affect expression: A review and a model for future research,” Psychological Bulletin, vol\. 99, no\. 2, pp\. 143\-165, 1986\.
- \[27\]J\. Carletta et al\., ”The AMI meeting corpus: A pre\-announcement,” in Machine Learning for Multimodal Interaction, Springer, 2006, pp\. 28– 39\. \[Online\]\. Available:[https://groups\.inf\.ed\.ac\.uk/ami/corpus/](https://groups.inf.ed.ac.uk/ami/corpus/)\(CC BY 4\.0\)\.
- \[28\]C\. W\. J\. Granger, ”Investigating causal relations by econometric models and cross\-spectral methods,” Econometrica, vol\. 37, no\. 3, pp\. 424\-438, 1969\.
- \[29\]Y\. Benjamini and Y\. Hochberg, ”Controlling the false discovery rate: A practical and powerful approach to multiple testing,” Journal of the Royal Statistical Society: Series B, vol\. 57, no\. 1, pp\. 289\-300, 1995\.
- \[30\]P\. N\. Juslin and P\. Laukka, ”Communication of emotions in vocal expression and music performance: Different channels, same code?” Psychological Bulletin, vol\. 129, no\. 5, pp\. 770\-814, 2003\.
- \[31\]F\. G\. P\. Piat, ”ARTIST: Adaptive Resonance Theory to Internalize the Structure of Tonality,” Ph\.D\. dissertation, University of Texas at Dallas, 1999\.
- \[32\]J\. Day, J\. C\. Finkelstein, B\. A\. Field et al\., ”Compassion\-focused technologies: Reflections and future directions,” Frontiers in Psychology, vol\. 12, art\. 603618, 2021\.
- \[33\]J\. Zigon, ”Can machines be ethical? On the necessity of relational ethics and empathic attunement for data\-centric technologies,” Social Research: An International Quarterly, vol\. 86, no\. 4, pp\. 1001–1022, 2019\. doi: 10\.1353/sor\.2019\.0046Similar Articles
Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models
This research paper explores emotion-sensitive neurons in multimodal foundation models, revealing shared affective mechanisms between speech and facial emotion recognition through causal interventions and cross-modal analysis.
Caring Without Feeling: Affective Dynamics as the Control Layer of Human-AI Agent Collaboration
This review synthesizes computational and interactional mechanisms of affective dynamics in human-AI agent collaboration, proposing that affective cues function as a coordination layer for trust, delegation, and oversight rather than internal emotions.
AtmosERC: Modeling Dialogue-Level Affective Atmosphere for Emotion Recognition in Conversation
This paper introduces AtmosERC, a model that models dialogue-level affective atmosphere to enhance emotion recognition in conversations.
AffectOmni: RL-Verifiable People-Centric Grounded Affective Reasoning for Social and Art-Related Scenes
AffectOmni is a GRPO-trained framework for verifiable affective reasoning in multimodal large language models, introducing People Focus and Temporal Order rewards to enhance people-centric evidence selection and temporally structured reasoning, with experiments showing improvements over 7B scale baselines.
Synthetic Resonance: A Framework for Growth-Oriented Human-AI Relationships
This paper introduces 'synthetic resonance,' a framework for understanding meaningful human-AI relationships without attributing subjective experience to AI, and calls for further research.