Computational Measurement of Team-Process Phase Dynamics in Collaborative Virtual Reality

arXiv cs.LG Papers

Summary

This paper presents a computational framework for detecting and interpreting dynamic team-process phases in collaborative virtual reality using late chunking and change-point detection to analyze temporal changes in team communication.

arXiv:2608.18660v1 Announce Type: new Abstract: Collaborative virtual reality (VR) environments make team communication observable as it unfolds, but conventional transcript analyses often summarize entire trials or divide them into fixed temporal windows. Such approaches can obscure changes in team communication and coordination over time. This article presents a computational framework for detecting and interpreting dynamic team-process phases from timestamped dialogue in a collaborative VR game. The framework uses late chunking to generate context-aware transcript representations, aggregates them into temporal chunks, and applies penalized Gaussian-kernel change-point detection to identify semantic transitions in team communication. After boundary detection, term frequency--inverse document frequency (TF-IDF), non-negative matrix factorization (NMF), and representative transcript segments provide structured evidence for phase interpretation. A locally deployed large language model (LLM) uses in-context learning to generate initial interpretations that are subsequently reviewed by humans. Independently recorded interaction logs are then aligned with the detected phases to examine corresponding task-action patterns. The evaluation compares representations, pooling strategies, segmentation methods, parameter settings, reviewed phase interpretations, and phase-aligned interaction profiles. The results show that the framework identifies coherent and interpretable phase structures while preserving traceability to the underlying transcript evidence. The correspondence between transcript-derived phases and interaction behavior further supports their relevance for analyzing collaborative activity. The framework therefore offers a transparent and transferable approach for studying temporal changes in teamwork from timestamped transcripts across collaborative task settings.
Original Article
View Cached Full Text

Cached at: 08/20/26, 10:32 AM

# Computational Measurement of Team-Process Phase Dynamics in Collaborative Virtual Reality
Source: [https://arxiv.org/html/2608.18660](https://arxiv.org/html/2608.18660)
Jianing Zhangand Pooja PolThanks:Qing Huang, Jianing Zhang, and Pooja Pol are with the School of Business, Technical University of Applied Sciences Augsburg, 86161 Augsburg, Germany, and also with the Data Science und Autonome Systeme Technologietransferzentrum \(TTZ\), 86899 Landsberg am Lech, Germany \(e\-mail: qing\.huang@tha\.de; jianing\.zhang@tha\.de; pooja\.pol@tha\.de\)\. Corresponding author: Qing Huang\.

###### Abstract

Collaborative virtual reality \(VR\) environments make team communication observable as it unfolds, but conventional transcript analyses often summarize entire trials or divide them into fixed temporal windows\. Such approaches can obscure changes in team communication and coordination over time\. This article presents a computational framework for detecting and interpreting dynamic team\-process phases from timestamped dialogue in a collaborative VR game\. The framework uses late chunking to generate context\-aware transcript representations, aggregates them into temporal chunks, and applies penalized Gaussian\-kernel change\-point detection to identify semantic transitions in team communication\. After boundary detection, term frequency–inverse document frequency \(TF–IDF\), non\-negative matrix factorization \(NMF\), and representative transcript segments provide structured evidence for phase interpretation\. A locally deployed large language model \(LLM\) uses in\-context learning to generate initial interpretations that are subsequently reviewed by humans\. Independently recorded interaction logs are then aligned with the detected phases to examine corresponding task\-action patterns\. The evaluation compares representations, pooling strategies, segmentation methods, parameter settings, reviewed phase interpretations, and phase\-aligned interaction profiles\. The results show that the framework identifies coherent and interpretable phase structures while preserving traceability to the underlying transcript evidence\. The correspondence between transcript\-derived phases and interaction behavior further supports their relevance for analyzing collaborative activity\. The framework therefore offers a transparent and transferable approach for studying temporal changes in teamwork from timestamped transcripts across collaborative task settings\.

###### Index Terms:

virtual reality collaboration, team communication, team process, change point detection, late chunking, large language models, in\-context learning

## IIntroduction

Collaborative virtual reality \(VR\) training and teamwork platforms provide shared task environments in which communication, coordination, and interaction can be examined as they unfold\[[24](https://arxiv.org/html/2608.18660#bib.bib24),[26](https://arxiv.org/html/2608.18660#bib.bib25),[15](https://arxiv.org/html/2608.18660#bib.bib26)\]\. Because speech, task actions, and system events can be recorded as temporally aligned traces, VR also offers a useful setting for computational teamwork research\[[6](https://arxiv.org/html/2608.18660#bib.bib15),[16](https://arxiv.org/html/2608.18660#bib.bib9)\]\. Team members must establish shared understanding, coordinate their actions, monitor the task situation, and resolve communication problems during collaboration\[[18](https://arxiv.org/html/2608.18660#bib.bib18),[2](https://arxiv.org/html/2608.18660#bib.bib1),[14](https://arxiv.org/html/2608.18660#bib.bib14)\]\. The key analytical challenge is therefore to identify how these communicative and coordination processes change over time rather than merely summarizing the overall amount of interaction\[[24](https://arxiv.org/html/2608.18660#bib.bib24),[13](https://arxiv.org/html/2608.18660#bib.bib12)\]\.

Team\-process research describes teamwork as a sequence of changing activities, including task interpretation, shared\-reference formation, action coordination, and problem response\[[18](https://arxiv.org/html/2608.18660#bib.bib18),[2](https://arxiv.org/html/2608.18660#bib.bib1)\]\. Recent temporal perspectives similarly emphasize that team processes and emergent states should be examined as evolving trajectories rather than static aggregates\[[13](https://arxiv.org/html/2608.18660#bib.bib12),[5](https://arxiv.org/html/2608.18660#bib.bib16),[17](https://arxiv.org/html/2608.18660#bib.bib17),[4](https://arxiv.org/html/2608.18660#bib.bib13)\]\. However, raw transcripts do not directly reveal where one collaborative phase ends and another begins\. Trial\-level summaries remove temporal structure, while fixed temporal windows may split coherent episodes or combine intervals with different communicative functions\. This motivates methods that detect variable\-length semantic phases from changes in the communication trajectory\[[9](https://arxiv.org/html/2608.18660#bib.bib2)\]\.

A further challenge is that short VR utterances often depend on surrounding discourse\. Expressions such as “there,” “the green one,” or “wait” cannot be interpreted reliably without preceding dialogue and task context\. Representing each segment independently may therefore remove information needed to distinguish its local function, which motivates context\-preserving approaches such as late chunking\[[8](https://arxiv.org/html/2608.18660#bib.bib3),[19](https://arxiv.org/html/2608.18660#bib.bib20)\]\. Phase interpretation must also remain traceable to the underlying evidence\. Rather than asking an LLM to summarize an entire trial, structured phase\-specific evidence can be used to generate candidate interpretations that are subsequently reviewed by humans\[[3](https://arxiv.org/html/2608.18660#bib.bib7),[25](https://arxiv.org/html/2608.18660#bib.bib5),[12](https://arxiv.org/html/2608.18660#bib.bib6),[23](https://arxiv.org/html/2608.18660#bib.bib10)\]\.

This article presents a computational framework for detecting and interpreting semantic phases from timestamped team transcripts\. As shown in Fig\.[1](https://arxiv.org/html/2608.18660#S4.F1), the framework combines late\-chunked contextual embeddings, temporal chunk pooling, kernel\-based change\-point \(CP\) detection, structured evidence extraction, LLM\-assisted interpretation with in\-context learning, human review, and phase\-aligned interaction analysis\[[8](https://arxiv.org/html/2608.18660#bib.bib3),[11](https://arxiv.org/html/2608.18660#bib.bib4),[6](https://arxiv.org/html/2608.18660#bib.bib15)\]\. Semantic boundaries are determined before interpretation so that the labeling process cannot influence their placement\. The evaluation examines contextual representation, pooling strategy, segmentation method, and parameter sensitivity, and then assesses the resulting phases through reviewed phase\-process labels, non\-negative matrix factorization \(NMF\) topic evidence, and aligned task\-action indicators\. The framework provides a traceable approach for studying how team communication and coordination develop during collaborative activity\.

## IIRelated Work and Research Motivation

Research on teamwork emphasizes that team behavior should be analyzed as an evolving process rather than as a static aggregate score\[[18](https://arxiv.org/html/2608.18660#bib.bib18),[2](https://arxiv.org/html/2608.18660#bib.bib1)\]\. This view is increasingly important in digital, multimodal, and virtual settings, where emergent states, coordination patterns, and resilience can change over multiple timescales\[[13](https://arxiv.org/html/2608.18660#bib.bib12),[5](https://arxiv.org/html/2608.18660#bib.bib16),[17](https://arxiv.org/html/2608.18660#bib.bib17),[4](https://arxiv.org/html/2608.18660#bib.bib13),[7](https://arxiv.org/html/2608.18660#bib.bib22),[20](https://arxiv.org/html/2608.18660#bib.bib21),[14](https://arxiv.org/html/2608.18660#bib.bib14),[16](https://arxiv.org/html/2608.18660#bib.bib9),[6](https://arxiv.org/html/2608.18660#bib.bib15)\]\. For human–machine systems, training feedback and after\-action review require interpretable information about when and how these processes change\.

Existing computational analyses often operate either at the transcript\-segment level or at the trial level\. Segment\-level labels can be too local to describe sustained team\-process episodes, whereas trial\-level summaries collapse meaningful transitions\. The needed intermediate unit is a semantically coherent phase that is longer than a single transcript segment, shorter than a full task episode, and variable in duration according to the team’s communication dynamics\.

Collaborative VR dialogue differs from well\-formed documents because it contains short turns, deictic expressions, simultaneous player perspectives, task\-specific object references, and local repair sequences\. Late chunking is relevant because it encodes a longer context first and then derives local segment representations, allowing short segments to inherit discourse information\[[8](https://arxiv.org/html/2608.18660#bib.bib3),[19](https://arxiv.org/html/2608.18660#bib.bib20)\]\. For boundary estimation, kernel change\-point detection is suitable because it evaluates changes in the distribution of a semantic trajectory rather than relying only on adjacent local similarity\[[10](https://arxiv.org/html/2608.18660#bib.bib19),[9](https://arxiv.org/html/2608.18660#bib.bib2),[11](https://arxiv.org/html/2608.18660#bib.bib4)\]\.

Interpretation remains separate from boundary detection\. A boundary set does not specify what communicative function a phase serves\. Recent work on in\-context learning and LLM\-assisted communication analysis suggests a useful division of labor in which computational methods extract structured evidence, LLMs convert that evidence into reviewable candidate labels, and human reviewers assess whether the labels are supported by the evidence\[[3](https://arxiv.org/html/2608.18660#bib.bib7),[1](https://arxiv.org/html/2608.18660#bib.bib8),[25](https://arxiv.org/html/2608.18660#bib.bib5),[12](https://arxiv.org/html/2608.18660#bib.bib6),[23](https://arxiv.org/html/2608.18660#bib.bib10),[22](https://arxiv.org/html/2608.18660#bib.bib11)\]\. The present study follows this evidence\-first logic and treats LLM use as LLM\-assisted evidence\-grounded annotation within a human\-reviewed workflow\.

## IIIResearch Questions and Claim Boundaries

The evaluation follows four questions aligned with the measurement pipeline\. RQ1 examines whether late chunking improves semantic separation relative to standard segment embeddings\. RQ2 examines whether CP\-kernel segmentation produces higher semantic separation than fixed\-width and random\-boundary reference partitions\. RQ3 examines sensitivity to the embedding model, chunk pooling, kernel bandwidth, and change\-point penalty\. RQ4 examines whether reviewed phase\-process labels can be assigned to the detected phases, whether these labels are supported by interpreted NMF topic evidence, and whether the phases can be aligned with VR interaction logs for cross\-session comparisons of communication patterns and task\-action profiles\.

Claim boundary\.The paper evaluates transcript\-semantic structure, reviewed interpretability, and alignment with task\-action and performance measures\. The detected phases are analytic units for examining team communication and coordination in collaborative VR rather than direct measurements of latent psychological states\. External validation through trainer ratings, participant self\-reports, and independent behavioral coding remains future work\.

## IVData and Preprocessing

### IV\-AVR Game Context and Transcript Data

The study was conducted in a collaborative VR training game designed around communication and coordination rather than individual dexterity\. Participants entered a shared virtual environment as visually similar avatars and had to move task\-relevant resources to target locations under time pressure\. The task required team members to identify player\-specific attributes, establish mutual visibility, localize resources, coordinate object manipulation, and repair misunderstandings during the ongoing activity\. Each team completed two VR\-game trials\. Between trials, the group left VR for a facilitated reflection and strategy discussion, then returned for a second attempt\. This paired design makes it possible to compare an initial encounter with the task and a later attempt after explicit team reflection\. All participants provided written informed consent before participation, including consent for the recording and analysis of audio, transcript, and VR interaction data\.

The data consist of timestamped German speech transcripts and timestamped VR interaction logs from these trials\. Color terms in this article refer only to VR\-game properties, such as object colors or color\-coded player roles, and should not be interpreted as demographic, personal, or social descriptors\. The dataset contains 15 teams, each of which completed two trials, resulting in 30 team\-session transcripts\. Player\-level audio was transcribed with a Whisper\-based automatic speech recognition \(ASR\) pipeline\[[21](https://arxiv.org/html/2608.18660#bib.bib23)\]\. A*transcript segment*is the minimal textual unit used in the analysis and consists of one timestamped ASR output segment from a player\-level audio recording\. Because each participant had a separate recording channel, the player identifier was inherited from the source channel rather than inferred through diarization\. Theiith segment is

σi=\(si,ei,ci,xi\),\\sigma\_\{i\}=\(s\_\{i\},e\_\{i\},c\_\{i\},x\_\{i\}\),\(1\)wheresis\_\{i\}andeie\_\{i\}denote the start and end timestamps,cic\_\{i\}denotes the player identifier, andxix\_\{i\}denotes the transcript text\. A team trial is represented as the ordered sequence

S=\{σ1,σ2,…,σN\},S=\\\{\\sigma\_\{1\},\\sigma\_\{2\},\\ldots,\\sigma\_\{N\}\\\},\(2\)whereNNis the number of usable timestamped ASR segments\.

Empty, non\-speech, and unusable ASR fragments were removed, and silent intervals were not converted into pseudo\-segments\. Transcript text and timestamp alignment were manually checked\. Because the focus of this study is semantic phase detection from timestamped transcripts rather than speech\-to\-text performance, ASR quality is not evaluated as a separate research objective, and the effect of transcription errors on the downstream results is not examined here\.

Fig\. 1:Workflow for semantic phase detection, evidence extraction, and phase interpretation\.
### IV\-BContextual Segment Embeddings and Temporal Chunks

Three analysis units are used in the framework\. A*transcript segment*is the timestamped automatic speech recognition \(ASR\) unit defined above\. A*contextual segment embedding*represents the semantic content of that segment together with its surrounding discourse\. A*temporal chunk*aggregates nearby contextual segment embeddings for change\-point detection\. Transcript segments retain the original textual evidence, while temporal chunks provide a more stable representation for boundary estimation\.

For late chunking, neighboring transcript segments are concatenated in temporal order into a*transcript block*up to the encoder context limit\. Let

T\(b\)=\(τ1,τ2,…,τL\)T^\{\(b\)\}=\(\\tau\_\{1\},\\tau\_\{2\},\\ldots,\\tau\_\{L\}\)\(3\)denote the token sequence of transcript blockbb, whereτℓ\\tau\_\{\\ell\}is theℓ\\ellth token andLLis the number of tokens in the block\. A long\-context encoderfθf\_\{\\theta\}maps this sequence to contextual token representations\.

\(ϑ1,ϑ2,…,ϑL\)=fθ​\(T\(b\)\)\.\(\\boldsymbol\{\\vartheta\}\_\{1\},\\boldsymbol\{\\vartheta\}\_\{2\},\\ldots,\\boldsymbol\{\\vartheta\}\_\{L\}\)=f\_\{\\theta\}\(T^\{\(b\)\}\)\.\(4\)Here,ϑℓ\\boldsymbol\{\\vartheta\}\_\{\\ell\}denotes the contextual representation of tokenτℓ\\tau\_\{\\ell\}, andθ\\thetadenotes the parameters of the embedding model\. The embedding of transcript segmentσi\\sigma\_\{i\}is obtained by pooling the contextual token representations aligned with that segment\.

𝐱i\\displaystyle\\mathbf\{x\}\_\{i\}=MeanPooltoken⁡\(\{ϑℓ∣τℓ∈σi\}\),\\displaystyle=\\operatorname\{MeanPool\}\_\{\\mathrm\{token\}\}\\bigl\(\\\{\\boldsymbol\{\\vartheta\}\_\{\\ell\}\\mid\\tau\_\{\\ell\}\\in\\sigma\_\{i\}\\\}\\bigr\),\(5\)𝐱~i\\displaystyle\\tilde\{\\mathbf\{x\}\}\_\{i\}=𝐱i‖𝐱i‖2\.\\displaystyle=\\frac\{\\mathbf\{x\}\_\{i\}\}\{\\\|\\mathbf\{x\}\_\{i\}\\\|\_\{2\}\}\.In these expressions,𝐱i\\mathbf\{x\}\_\{i\}is the pooled embedding of transcript segmentσi\\sigma\_\{i\}, and𝐱~i\\tilde\{\\mathbf\{x\}\}\_\{i\}is itsL2L\_\{2\}\-normalized form\. Token representations are combined using mean pooling\. Normalization places all segment embeddings on a comparable scale for cosine\-based evaluation and kernel\-based distance computation\.

When a transcript exceeds the encoder context limit, it is divided into overlapping transcript blocks\. If a segment appears in more than one block, its resulting embeddings are combined before normalization\. This procedure preserves broader conversational context while maintaining alignment with the original ASR timestamps\. The standard embedding condition processes each transcript segment independently and serves as the comparison condition for evaluating the effect of late chunking\.

Change\-point detection is not applied directly to individual transcript segments\. These segments are often short, unevenly distributed over time, and semantically ambiguous without local context\. Direct segmentation at this level would therefore be sensitive to local transcription and timing fluctuations\. Instead, nearby contextual segment embeddings are aggregated into temporal chunks of widthΔ\\Delta\.

Ik\\displaystyle I\_\{k\}=\[t0\+\(k−1\)Δ,t0\+kΔ\),\\displaystyle=\[t\_\{0\}\+\(k\-1\)\\Delta,\\;t\_\{0\}\+k\\Delta\),\(6\)Ck\\displaystyle C\_\{k\}=\{σi:si∈Ik\},\\displaystyle=\\\{\\sigma\_\{i\}:s\_\{i\}\\in I\_\{k\}\\\},𝐮k\\displaystyle\\mathbf\{u\}\_\{k\}=Poolchunk⁡\(\{𝐱~i∣σi∈Ck\}\),\\displaystyle=\\operatorname\{Pool\}\_\{\\mathrm\{chunk\}\}\\bigl\(\\\{\\tilde\{\\mathbf\{x\}\}\_\{i\}\\mid\\sigma\_\{i\}\\in C\_\{k\}\\\}\\bigr\),𝐳k\\displaystyle\\mathbf\{z\}\_\{k\}=𝐮k‖𝐮k‖2,∥𝐮k∥2\>0\.\\displaystyle=\\frac\{\\mathbf\{u\}\_\{k\}\}\{\\\|\\mathbf\{u\}\_\{k\}\\\|\_\{2\}\},\\qquad\\\|\\mathbf\{u\}\_\{k\}\\\|\_\{2\}\>0\.Here,IkI\_\{k\}denotes thekkth temporal interval,t0t\_\{0\}is the trial start time, andΔ\\Deltais the temporal chunk width\. The setCkC\_\{k\}contains all transcript segments whose start timestamps fall withinIkI\_\{k\}\. The vector𝐮k\\mathbf\{u\}\_\{k\}is the pooled chunk representation and𝐳k\\mathbf\{z\}\_\{k\}is itsL2L\_\{2\}\-normalized form\. Thus, after mean, maximum, minimum, or mean–maximum pooling, each nonzero chunk embedding is divided by its Euclidean norm so that its length equals one\. This normalization is applied consistently across all chunk\-pooling strategies before kernel computation\. The ordered sequence

𝐙=\{𝐳1,𝐳2,…,𝐳M\}\\mathbf\{Z\}=\\\{\\mathbf\{z\}\_\{1\},\\mathbf\{z\}\_\{2\},\\ldots,\\mathbf\{z\}\_\{M\}\\\}\(7\)forms the semantic trajectory used for boundary estimation, whereMMis the number of non\-empty temporal chunks\. Chunks without usable transcript segments are excluded rather than imputed\. After the boundaries are detected, phase interpretation and human review return to the original transcript segments\. Temporal chunking therefore stabilizes boundary estimation while preserving access to the underlying textual evidence\.

## VMethod

For each trial, the computation follows the steps below\.

1. 1\.Construct the ordered transcript sequenceS=\{σ1,…,σN\}S=\\\{\\sigma\_\{1\},\\ldots,\\sigma\_\{N\}\\\}from player\-level ASR segments\.
2. 2\.Compute one normalized embedding𝐱~i\\tilde\{\\mathbf\{x\}\}\_\{i\}for each transcript segment using either standard segment embedding or late\-chunked contextual embedding\.
3. 3\.Aggregate the segment embeddings into temporal chunk vectors𝐙=\{𝐳1,…,𝐳M\}\\mathbf\{Z\}=\\\{\\mathbf\{z\}\_\{1\},\\ldots,\\mathbf\{z\}\_\{M\}\\\}using the selected pooling rule\.
4. 4\.Compute the Gaussian kernel matrix over𝐙\\mathbf\{Z\}and the interval costsC^​\(s,e\)\\widehat\{C\}\(s,e\)for admissible chunk intervals\.
5. 5\.Estimate the CP\-kernel boundaries𝝉^\\widehat\{\\boldsymbol\{\\tau\}\}by minimizing the penalized objective with exact dynamic programming\.
6. 6\.Map the detected phase intervals back to the original transcript segments\.
7. 7\.Convert each fixed phase interval into an evidence packet for LLM\-assisted interpretation, human review, and phase\-aligned interaction analysis\.

### V\-AGaussian\-Kernel Change\-Point Objective

Each trial is treated as an ordered trajectory of chunk vectors𝐳k\\mathbf\{z\}\_\{k\}\. The objective is to identify internally coherent intervals separated by changes in the embedding distribution\. A Gaussian radial basis function \(RBF\) kernel is therefore used to compare chunks\.

k⁡\(𝐳i,𝐳j\)=exp⁡\(−γ​‖𝐳i−𝐳j‖22\)\.k\(\\mathbf\{z\}\_\{i\},\\mathbf\{z\}\_\{j\}\)=\\exp\\left\(\-\\gamma\\\|\\mathbf\{z\}\_\{i\}\-\\mathbf\{z\}\_\{j\}\\\|\_\{2\}^\{2\}\\right\)\.\(8\)Hereγ\\gammacontrols the sensitivity of the Gaussian kernel to distances between chunk embeddings and∥⋅∥2\\\|\\cdot\\\|\_\{2\}is the Euclidean norm\. The Gaussian kernel captures changes in the geometry of the embedding distribution, including shifts in location, dispersion, and local density\. Such changes can reflect transitions from exploratory discussion to focused coordination even when task vocabulary remains similar\.

For a candidate interval\[s,e\]\[s,e\], the empirical kernel cost is

C^​\(s,e\)\\displaystyle\\widehat\{C\}\(s,e\)=∑t=sek⁡\(𝐳t,𝐳t\)\\displaystyle=\\sum\_\{t=s\}^\{e\}k\(\\mathbf\{z\}\_\{t\},\\mathbf\{z\}\_\{t\}\)\(9\)−1e−s\+1∑i=se∑j=sek\(𝐳i,𝐳j\)\.\\displaystyle\-\\frac\{1\}\{e\-s\+1\}\\sum\_\{i=s\}^\{e\}\\sum\_\{j=s\}^\{e\}k\(\\mathbf\{z\}\_\{i\},\\mathbf\{z\}\_\{j\}\)\.HereC^​\(s,e\)\\widehat\{C\}\(s,e\)is the cost of representing chunksssthrougheeas one phase, ande−s\+1e\-s\+1is the number of chunks in the interval\. A coherent interval has a lower cost because the within\-interval kernel similarities are high\. An interval that spans a semantic shift is more heterogeneous and therefore receives a higher cost\.

The estimated change points minimize the penalized objective

𝝉^\\displaystyle\\widehat\{\\boldsymbol\{\\tau\}\}=arg⁡min𝝉⁡J⁡\(𝝉\),\\displaystyle=\\arg\\min\_\{\\boldsymbol\{\\tau\}\}J\(\\boldsymbol\{\\tau\}\),\(10\)J⁡\(𝝉\)\\displaystyle J\(\\boldsymbol\{\\tau\}\)=∑r=1K\+1C^​\(τr−1\+1,τr\)\+β​K,\\displaystyle=\\sum\_\{r=1\}^\{K\+1\}\\widehat\{C\}\(\\tau\_\{r\-1\}\+1,\\tau\_\{r\}\)\+\\beta K,𝝉\\displaystyle\\boldsymbol\{\\tau\}:0=τ0<⋯<τK<τK\+1=M\.\\displaystyle:\\;0=\\tau\_\{0\}<\\cdots<\\tau\_\{K\}<\\tau\_\{K\+1\}=M\.In this objective,𝝉^\\widehat\{\\boldsymbol\{\\tau\}\}is the estimated boundary sequence,J⁡\(𝝉\)J\(\\boldsymbol\{\\tau\}\)is the total segmentation cost,KKis the number of detected change points,K\+1K\+1is the number of resulting phases,τr\\tau\_\{r\}is the end chunk of phaserr,MMis the number of temporal chunks in the trial, andβ\\betapenalizes additional boundaries\. A minimum phase length in chunks prevents uninterpretable micro\-phases\. The default kernel scale is derived from the median pairwise squared distance,

γ0=1dmed,\\gamma\_\{0\}=\\frac\{1\}\{d\_\{\\mathrm\{med\}\}\},\(11\)wheredmedd\_\{\\mathrm\{med\}\}is the median of the nonzero pairwise squared distances between chunk embeddings\. Sensitivity analysis uses

γ=α​γ0,\\gamma=\\alpha\\gamma\_\{0\},\(12\)whereα\\alphais the gamma multiplier reported in the sensitivity plots\.

Sensitivity analysis uses late\-chunkedjina\-v2\-de, mean pooling,α=1\\alpha=1,β=1\\beta=1, andmmin=1m\_\{\\min\}=1as the reference configuration\. Embedding model, embedding mode, pooling rule, gamma multiplier, and penalty are varied separately\. Bothα\\alphaandβ\\betaare evaluated over\{0\.01,0\.05,0\.1,0\.5,1,5,10\}\\\{0\.01,0\.05,0\.1,0\.5,1,5,10\\\}\. For each condition, boundaries are recomputed and evaluated with the same semantic\-separation score\. LLM labels and interaction logs are not used for parameter selection\.

In the reported experiments, the penalized objective in \([10](https://arxiv.org/html/2608.18660#S5.E10)\) is solved by exact dynamic programming over the ordered chunk sequence\. LetF⁡\(t\)F\(t\)denote the minimum segmentation cost for the prefix ending at chunktt\. With minimum phase lengthmminm\_\{\\min\}and admissible predecessor set𝒜\(t\)=\{s:0≤s<t,t−s≥mmin\}\\mathcal\{A\}\(t\)=\\\{s:0\\leq s<t,\\ t\-s\\geq m\_\{\\min\}\\\}, the Bellman recursion is

F⁡\(t\)\\displaystyle F\(t\)=mins∈𝒜⁡\(t\)\[F\(s\)\+C^\(s\+1,t\)\\displaystyle=\\min\_\{s\\in\\mathcal\{A\}\(t\)\}\\Bigl\[F\(s\)\+\\widehat\{C\}\(s\+1,t\)\(13\)\+β𝕀\(s\>0\)\],\\displaystyle\+\\beta\\,\\mathbb\{I\}\(s\>0\)\\Bigr\],F⁡\(0\)\\displaystyle F\(0\)=0\.\\displaystyle=0\.Here𝕀⁡\(s\>0\)\\mathbb\{I\}\(s\>0\)denotes an indicator that equals one only when the candidate segment does not start at the beginning of the trial\. Thus, the first phase is penalty\-free, whereas every additional boundary contributesβ\\betato the objective\. For each endpointtt, the algorithm evaluates all admissible predecessor indicesss, stores the best predecessor, and reconstructs the optimal boundary sequence by backtracking fromt=Mt=M\. This exact dynamic\-programming formulation is feasible because temporal chunking keeps the number of analysis units per trial moderate\.

### V\-BSegmentation Methods and Evaluation Metric

The empirical comparison considers three segmentation methods applied to the same semantic chunk trajectory\. The proposed method is CP\-kernel segmentation, which minimizes the penalized kernel objective in \([10](https://arxiv.org/html/2608.18660#S5.E10)\) and places boundaries where the semantic distribution of the chunk sequence changes\. It is therefore the only method that uses the semantic objective directly for boundary placement\.

The first reference method is fixed\-width segmentation\. It uses the same number of phases as the CP\-kernel solution,P=K\+1P=K\+1, but replaces adaptive boundary placement with an equal\-width partition\.

τrfw=⌊r​MP⌋,r=1,…,P−1,\\tau\_\{r\}^\{\\mathrm\{fw\}\}=\\left\\lfloor\\frac\{rM\}\{P\}\\right\\rfloor,\\qquad r=1,\\dots,P\-1,\(14\)whereτrfw\\tau\_\{r\}^\{\\mathrm\{fw\}\}is therrth fixed\-width boundary,MMis the number of temporal chunks, andPPis the number of phases\. This baseline tests whether a regular temporal grid is sufficient to recover interpretable semantic phases\.

The second reference method is random\-boundary segmentation\. It preserves the same number of change points and the same minimum phase\-length constraint as the CP\-kernel solution, but it discards the semantic objective used for boundary placement\. A random boundary sequence satisfies

0=τ0rand<τ1rand<⋯<τP−1rand<τPrand=M,0=\\tau^\{\\mathrm\{rand\}\}\_\{0\}<\\tau^\{\\mathrm\{rand\}\}\_\{1\}<\\cdots<\\tau^\{\\mathrm\{rand\}\}\_\{P\-1\}<\\tau^\{\\mathrm\{rand\}\}\_\{P\}=M,\(15\)and every resulting phase must satisfy

τrandr−τrandr−1≥mmin,r=1,…,P\.\\tau^\{\\mathrm\{rand\}\}\_\{r\}\-\\tau^\{\\mathrm\{rand\}\}\_\{r\-1\}\\geq m\_\{\\min\},\\qquad r=1,\\ldots,P\.\(16\)Here,mminm\_\{\\min\}is the minimum allowed phase length in chunks\. For each session, 1000 admissible random boundary sequences are sampled under this constraint using a fixed pseudo\-random generator with seed 42\. Each random segmentation is matched to the CP\-kernel output in both the number of change points and the minimum phase\-length constraint\. The random\-reference result is summarized by the mean semantic\-separation score across the 1000 repetitions together with an empirical 95% confidence interval\. Using the same seeded random\-number generator makes the baseline reproducible for identical inputs\. This baseline therefore functions as a chance\-level control that preserves segmentation granularity while removing optimized boundary placement\.

These three methods are compared with the same semantic\-separation score\. Because boundary estimation is performed on temporal chunks, semantic separation is evaluated primarily on the chunk embeddings used by the segmentation procedure\. LetCpC\_\{p\}denote the set of temporal chunks assigned to phasepp\. For phases containing at least two chunks, within\-phase similarity is the average pairwise cosine similarity among chunk embeddings inCpC\_\{p\}\.

Wpchunk=2\|Cp\|​\(\|Cp\|−1\)​∑i,j∈Cpi<jcos⁡\(𝐳i,𝐳j\),W\_\{p\}^\{\\mathrm\{chunk\}\}=\\frac\{2\}\{\|C\_\{p\}\|\(\|C\_\{p\}\|\-1\)\}\\sum\_\{\\begin\{subarray\}\{c\}i,j\\in C\_\{p\}\\\\ i<j\\end\{subarray\}\}\\cos\(\\mathbf\{z\}\_\{i\},\\mathbf\{z\}\_\{j\}\),\(17\)whereWpchunkW\_\{p\}^\{\\mathrm\{chunk\}\}is the within\-phase chunk similarity for phaseppandcos⁡\(𝐳i,𝐳j\)\\cos\(\\mathbf\{z\}\_\{i\},\\mathbf\{z\}\_\{j\}\)is the cosine similarity between the normalized chunk embeddings\.

Between\-phase similarity compares the chunks inCpC\_\{p\}with all chunks outside that phase\.

Bpchunk=1\|Cp\|​\|C¬p\|​∑i∈Cp∑j∈C¬pcos⁡\(𝐳i,𝐳j\),B\_\{p\}^\{\\mathrm\{chunk\}\}=\\frac\{1\}\{\|C\_\{p\}\|\\,\|C\_\{\\neg p\}\|\}\\sum\_\{i\\in C\_\{p\}\}\\sum\_\{j\\in C\_\{\\neg p\}\}\\cos\(\\mathbf\{z\}\_\{i\},\\mathbf\{z\}\_\{j\}\),\(18\)whereC¬pC\_\{\\neg p\}denotes all temporal chunks in the same session that are not assigned to phasepp\. The corresponding chunk\-level phase separation score is

Dpchunk=Wpchunk−Bpchunk\.D\_\{p\}^\{\\mathrm\{chunk\}\}=W\_\{p\}^\{\\mathrm\{chunk\}\}\-B\_\{p\}^\{\\mathrm\{chunk\}\}\.\(19\)A larger value indicates that chunks within the same phase are more semantically similar to one another than to chunks outside that phase\.

If a detected phase contains only one temporal chunk, the pairwise term inWpchunkW\_\{p\}^\{\\mathrm\{chunk\}\}is undefined\. In this case, the score is computed from the original transcript\-segment embeddings assigned to that phase rather than discarding the session\. LetSpS\_\{p\}denote the set of transcript segments mapped to phasepp\. The fallback within\-phase and between\-phase similarities are

Wpseg=2\|Sp\|​\(\|Sp\|−1\)​∑i,j∈Spi<jcos⁡\(𝐱~i,𝐱~j\),W\_\{p\}^\{\\mathrm\{seg\}\}=\\frac\{2\}\{\|S\_\{p\}\|\(\|S\_\{p\}\|\-1\)\}\\sum\_\{\\begin\{subarray\}\{c\}i,j\\in S\_\{p\}\\\\ i<j\\end\{subarray\}\}\\cos\(\\tilde\{\\mathbf\{x\}\}\_\{i\},\\tilde\{\\mathbf\{x\}\}\_\{j\}\),\(20\)Bpseg=1\|Sp\|​\|S¬p\|​∑i∈Sp∑j∈S¬pcos⁡\(𝐱~i,𝐱~j\),B\_\{p\}^\{\\mathrm\{seg\}\}=\\frac\{1\}\{\|S\_\{p\}\|\\,\|S\_\{\\neg p\}\|\}\\sum\_\{i\\in S\_\{p\}\}\\sum\_\{j\\in S\_\{\\neg p\}\}\\cos\(\\tilde\{\\mathbf\{x\}\}\_\{i\},\\tilde\{\\mathbf\{x\}\}\_\{j\}\),\(21\)with

Dpseg=Wpseg−Bpseg\.D\_\{p\}^\{\\mathrm\{seg\}\}=W\_\{p\}^\{\\mathrm\{seg\}\}\-B\_\{p\}^\{\\mathrm\{seg\}\}\.\(22\)The final phase\-level score is therefore

Dp=\{Dpchunk,\|Cp\|≥2,Dpseg,\|Cp\|=1​and​\|Sp\|≥2\.D\_\{p\}=\\begin\{cases\}D\_\{p\}^\{\\mathrm\{chunk\}\},&\|C\_\{p\}\|\\geq 2,\\\\ D\_\{p\}^\{\\mathrm\{seg\}\},&\|C\_\{p\}\|=1\\text\{ and \}\|S\_\{p\}\|\\geq 2\.\\end\{cases\}\(23\)The session score is the mean of all valid phase\-level scores\. If a session collapses to a single detected phase, its score is set to zero because no between\-phase contrast exists\. For the reference methods, the phase count is matched to the CP\-kernel solution for the same embedding generation method\. If a valid reference segmentation cannot be constructed under this constraint, the session score is also set to zero\.

### V\-CEvidence Extraction and Reviewed Interpretation

After boundaries are fixed, the pipeline returns to transcript evidence inside each phase\. This step is deliberately separated from boundary estimation\. Evidence extraction and labeling do not move, add, or remove change points\. TF–IDF terms and NMF topics summarize recurrent lexical patterns, while representative transcript segments preserve direct qualitative evidence\. The NMF model provides topic weights and topic\-specific lexical evidence\. The labels reported later are interpreted NMF topic labels, not raw lists of high\-weight topic terms\. For phasepp, the phase\-term matrix is approximated as

𝐕\(p\)\\displaystyle\\mathbf\{V\}^\{\(p\)\}≈𝐖\(p\)​𝐇\(p\),\\displaystyle\\approx\\mathbf\{W\}^\{\(p\)\}\\mathbf\{H\}^\{\(p\)\},\(24\)𝐖\(p\),𝐇\(p\)\\displaystyle\\mathbf\{W\}^\{\(p\)\},\\mathbf\{H\}^\{\(p\)\}≥0,\\displaystyle\\geq 0,where𝐕\(p\)\\mathbf\{V\}^\{\(p\)\}is the TF–IDF\-weighted matrix,𝐖\(p\)\\mathbf\{W\}^\{\(p\)\}contains segment\-topic weights, and𝐇\(p\)\\mathbf\{H\}^\{\(p\)\}contains topic\-term weights\. High\-weight terms and representative segments form the evidence base for phase interpretation\.

A locally deployedqwen3\.5:9bmodel converts each structured evidence packet into an initial phase interpretation\. The LLM serves as an annotation assistant rather than an autonomous evaluator, reducing the need to formulate every label from scratch\. In\-context examples provide the VR task context, expected abstraction level, and previous\-phase information, thereby producing more consistent, task\-grounded labels than unconstrained summarization would\.

The interpretation record contains a predefined phase\-process label for cross\-session comparison and interpreted NMF topic labels as supporting semantic evidence\. The original transcript evidence is in German, while the reported labels are provided in English for readability\. The LLM generates an initial phase\-process label and interpreted NMF topic labels from the structured evidence packet\. Human review then verifies whether these candidate labels are supported by the phase\-specific transcript evidence and task context\. Reviewers may retain, revise, or replace the proposed labels, but they do not modify the detected boundaries\. The reported outputs are therefore human\-reviewed labels from an LLM\-assisted annotation process that reduces the need to formulate every label from scratch\.

Figure[2](https://arxiv.org/html/2608.18660#S5.F2)shows the prompt template used for the initial LLM\-assisted phase interpretation\. The prompt combines task\-level context with phase\-specific evidence, including TF–IDF terms, NMF topic evidence, representative utterances, and information about the preceding phase\. It explicitly instructs the model to interpret the complete detected phase rather than individual utterances and to leave the previously detected phase boundaries unchanged\. The team\-process function is selected from the predefined codebook, while the phase\-specific focus provides a more concrete description of the communication occurring within the phase\. Requiring a short evidence\-based justification and multiple candidate labels makes the initial interpretation easier to inspect during subsequent human review\.

Prompt template for context\-grounded phase labelingRole and task\.You are an expert in German\-language team communication analysis, VR collaboration, and qualitative topic labeling\. You will receive evidence from one detected communication phase in a VR\-based team collaboration task\. Phase boundaries were detected before this labeling step\. Your task is to generate concise, academically usable English labels that describe the dominant communicative function of the complete phase\. Do not label individual utterances and do not create or modify phase boundaries\.Context inputs\.•Scenario context:\{scenario\_context\}•Team\-process codebook and inter\-label clarifications:\{team\_process\_codebook\}•Previous phase context:\{previous\_phase\_summary\}Current phase evidence\.•Phase ID and time interval:\{phase\_id\}; \{phase\_start\}–\{phase\_end\}•Top TF–IDF terms and n\-grams:\{phase\_top\_terms\}•NMF topic evidence:\{nmf\_topics\}•Representative utterances:\{representative\_utterances\}Labeling criteria\.1\.Evidence grounding:base the interpretation on the representative utterances and topic evidence; use NMF as within\-phase evidence, not as separate phase labels\.2\.Coverage:label the dominant communicative function of the complete phase, rather than one salient utterance, object, colour, player, or location\.3\.Discrimination:distinguish the phase from adjacent phases; if the same label is retained, state the concrete difference in focus\.4\.Specificity and style:avoid generic labels such as “communication”, “teamwork”, “discussion”, or “interaction”\. Use natural English suitable for a scientific report; prefer labels of 2–5 words\.Required output format\.Team\-process function: One codebook category that best captures the dominant function\. Phase\-specific focus: One short English phrase describing the concrete focus of this phase\. Reasoning: One short sentence linking the label to transcript evidence\. Candidate labels: \.\.\. Final label: \.\.\.

Fig\. 2:Context\-grounded prompt template for LLM\-assisted phase labeling using task context and phase\-specific transcript evidence\.

## VIResults

The results examine representation choice, pooling strategy, parameter sensitivity, segmentation baselines, detected phase structure, and evidence\-grounded interpretation\. All reported analyses use 30\-second temporal chunks\. In each sensitivity analysis, only the parameter under examination is varied, while all other parameters remain fixed at their final selected values\.

### VI\-ARepresentation Choice and the Effect of Late Context

Figure[3](https://arxiv.org/html/2608.18660#S6.F3)compares late\-chunked and standard embeddings across the tested models using mean pooling and the same semantic\-separation metric\. Across all tested models, late chunking produces higher semantic separation than standard embedding, indicating that contextualization has a stronger influence on the representation than model choice alone\. Within the late\-chunked condition,jina\-v2\-deachieves the highest semantic separation in both trials\. These results show that preserving conversational context is the primary factor, while selecting a language\-appropriate embedding model provides an additional benefit for downstream phase detection\.

Fig\. 3:Semantic separation across embedding models and embedding modes using mean pooling\.
### VI\-BPooling Strategy and Chunk\-Trajectory Stability

Figure[4](https://arxiv.org/html/2608.18660#S6.F4)compares chunk\-pooling strategies under the late\-chunkedjina\-v2\-derepresentation using the same downstream change\-point kernel configuration\. Mean pooling produces the highest semantic separation in both trials\. A temporal chunk can contain several complementary transcript segments that jointly express the local communication function\. Mean pooling captures the central semantic pattern of this exchange, whereas max and min pooling can overemphasize individual segments and mean–max pooling can retain part of this sensitivity\. Because temporal pooling is used only for boundary estimation and interpretation returns to the original transcript segments, mean pooling stabilizes the semantic trajectory without allowing a single segment to dominate it\.

Fig\. 4:Semantic separation across temporal chunk\-pooling strategies using late\-chunkedjina\-v2\-deembeddings\.
### VI\-CSensitivity to the Gaussian\-Kernel Scale and Change\-Point Penalty

Figure[5](https://arxiv.org/html/2608.18660#S6.F5)presents the sensitivity of semantic separation and mean phase count to the gamma multiplier and change\-point penalty under late\-chunkedjina\-v2\-dewith mean pooling\. Very small gamma multipliers make the Gaussian kernel less sensitive to embedding differences, causing chunks to appear more similar and reducing the number of detected transitions\. Larger gamma multipliers make local semantic differences more visible and increase semantic separation, but overly large values can fragment transcripts into many short phases\. The penalty controls the opposing tendency\. Small penalties permit more boundaries and can increase apparent semantic separation, whereas large penalties suppress boundaries and can collapse the phase structure\. The final setting ofα=1\\alpha=1andβ=1\\beta=1provides clear semantic separation without excessive fragmentation or phase collapse\. Both trials exhibit the same overall trade\-off\.

Fig\. 5:Semantic separation and mean number of detected phases across gamma multipliers and change\-point penalties\.
### VI\-DBaseline Comparison of Adaptive and Reference Boundaries

Under the final configuration using late\-chunkedjina\-v2\-de, mean pooling, a gamma multiplier ofα=1\\alpha=1, and a penalty ofβ=1\\beta=1, Fig\.[6](https://arxiv.org/html/2608.18660#S6.F6)compares change\-point kernel segmentation with fixed\-width and random\-boundary references using the same semantic\-separation metric\. For each embedding method, the phase count produced by change\-point kernel segmentation defines the target number of phases for the reference methods, ensuring that all methods are compared under matched segmentation granularity\. Change\-point kernel segmentation achieves the highest semantic separation under late chunking, indicating greater internal semantic separation under the evaluation metric used here when contextual representation and adaptive boundary placement are combined\. Fixed\-width and random\-boundary references preserve the same number of phases but do not use semantic changes to determine boundary locations\.

Fig\. 6:Semantic separation across segmentation methods and embedding modes\.
### VI\-EPhase Count, Timelines, and Interaction Profiles

The combined phase\-structure and interaction results reveal descriptive differences in teamwork between the two trials\. As shown in Fig\.[7](https://arxiv.org/html/2608.18660#S6.F7), Trial 1 is dominated by four\-phase solutions, with such solutions observed in 11 of the 15 sessions, whereas three\-phase solutions are most frequent in Trial 2, occurring in seven sessions\.

Fig\. 7:Distribution of detected phase counts across Trials 1 and 2\.The timelines in Figs\.[8](https://arxiv.org/html/2608.18660#S6.F8)and[9](https://arxiv.org/html/2608.18660#S6.F9)show the temporal order and relative duration of the detected phases for each team\. Colors indicate ordinal phase positions from Phase 1 to Phase 5 within each session\. The color\-coded positions visualize the sequence of semantic changes and do not imply that phases with the same number have the same communicative or coordination meaning across teams\. The specific phase meanings are examined in Section[VII](https://arxiv.org/html/2608.18660#S7)\. The team\-specific boundary locations further show that the detected phases do not follow a fixed temporal template\. Together, the phase\-count and timeline results show that Trial 2 sessions more often contain fewer major changes in the semantic organization of team communication\. Combined with the stronger task engagement observed at earlier ordinal phase positions in the interaction profiles, this pattern is consistent with more stable collaboration in Trial 2 after the initial task experience and interim reflection\.

Fig\. 8:Normalized detected\-phase timelines for Trial 1 sessions\.Fig\. 9:Normalized detected\-phase timelines for Trial 2 sessions\.To examine whether the transcript\-derived phase structure corresponds to task behavior, independently recorded interaction events are aligned with the detected phase intervals after segmentation\. These events are not used to estimate the phase boundaries\.

For phaseppin sessionii, the manipulation count is defined as

Mi​p=Gi​p\+Di​p,M\_\{ip\}=G\_\{ip\}\+D\_\{ip\},\(25\)whereGi​pG\_\{ip\}andDi​pD\_\{ip\}denote the numbers of object\-grab and object\-drop events assigned to that phase\.

The manipulation rate is

Ri​pmanip=Mi​pdi​p,R^\{\\mathrm\{manip\}\}\_\{ip\}=\\frac\{M\_\{ip\}\}\{d\_\{ip\}\},\(26\)wheredi​pd\_\{ip\}is the phase duration in minutes\. This measure represents task\-action intensity\.

The score rate is

Ri​pscore=Qi​pdi​p,R^\{\\mathrm\{score\}\}\_\{ip\}=\\frac\{Q\_\{ip\}\}\{d\_\{ip\}\},\(27\)whereQi​pQ\_\{ip\}is the number of scoring events in the phase\. This measure represents scoring progress per minute\.

Score per manipulation is

Ei​p=Qi​pMi​p,Mi​p\>0\.E\_\{ip\}=\\frac\{Q\_\{ip\}\}\{M\_\{ip\}\},\\qquad M\_\{ip\}\>0\.\(28\)This measure represents how effectively object manipulations translate into successful scoring outcomes\. When no manipulation event occurs within a phase,Ei​pE\_\{ip\}is treated as undefined rather than assigned a value of zero\.

The interaction profiles in Fig\.[10](https://arxiv.org/html/2608.18660#S6.F10)indicate stronger task engagement at earlier ordinal phase positions in Trial 2\. Manipulation rate rises quickly and peaks in Phase 3, whereas Trial 1 shows a more gradual increase across the commonly observed phase positions\. Score rate follows a similar pattern\. In Trial 2, the score rate increases sharply by Phase 2 and remains comparatively high through Phases 3 and 4, whereas the score rate in Trial 1 increases more gradually\. Score per manipulation also reaches a relatively high and stable level at earlier ordinal phase positions in Trial 2, suggesting that teams convert object manipulations into successful outcomes more consistently in the earlier detected phases\.

Taken together, the three measures suggest stronger task engagement, scoring progress, and more stable manipulation efficiency at earlier ordinal phase positions in Trial 2\. This pattern is consistent with teams becoming familiar with the virtual environment, task procedure, and available objects during the first trial and applying that experience more effectively in the second\.

The correspondence between the transcript\-derived phase structure in Figs\.[7](https://arxiv.org/html/2608.18660#S6.F7)–[9](https://arxiv.org/html/2608.18660#S6.F9)and the independently recorded interaction profiles in Fig\.[10](https://arxiv.org/html/2608.18660#S6.F10)provides complementary behavioral evidence for the relevance of the detected phases\. Because phase boundaries are estimated exclusively from transcript semantics and interaction measures are aligned only afterward, the observed correspondence indicates that the transcript\-derived phases are associated with differences in collaborative task activity\.

Fig\. 10:Phase\-aligned interaction measures by ordinal phase position in Trials 1 and 2\.

## VIIEvidence\-Grounded Phase Interpretation

After CP\-kernel segmentation, each fixed phase is interpreted using its time span, TF–IDF terms, NMF topic evidence, representative German transcript segments, and previous\-phase context\. A locally deployed LLM generates an initial English interpretation that is subsequently reviewed by humans\. Each phase receives a reviewed phase\-process label and interpreted NMF topic labels, while boundary placement remains unaffected by the interpretation step\.

Table[I](https://arxiv.org/html/2608.18660#S7.T1)summarizes the most frequent phase\-process labels by ordinal position and trial\. Tables[II](https://arxiv.org/html/2608.18660#S7.T2)and[III](https://arxiv.org/html/2608.18660#S7.T3)show the phase sequences and supporting NMF topic evidence for all 15 teams in Trial 1 and Trial 2, allowing within\-team changes across trials to be examined\.

In Trial 1, Object/Resource Coordination and Player\-Attribute Assignment are most frequent in the early phases, while Spatial Orientation and Localization becomes more prominent in Phases 3 and 4\. At the aggregate level, this distribution is consistent with a shift from establishing player roles, identifying resources, and clarifying task requirements toward spatially grounded execution, and it corresponds to the gradual increase in manipulation and scoring shown in Fig\.[10](https://arxiv.org/html/2608.18660#S6.F10)\.

In Trial 2, Player\-Attribute Assignment is most frequent in Phase 1, followed by Mutual Visibility Check and Spatial Orientation and Localization in Phases 2 and 3\. Together with the stronger task engagement observed at earlier ordinal phase positions in Fig\.[10](https://arxiv.org/html/2608.18660#S6.F10), this pattern suggests that role establishment, shared visibility, and spatial relations are more prominent in the earlier detected phases of Trial 2 following the initial task experience and interim reflection\.

Phase 5 is excluded from the cross\-trial interpretation because it occurs in only one Trial 1 session and two Trial 2 sessions\. The example sequences also show that teams do not follow a single fixed trajectory and that the same phase\-process label can be supported by different NMF topic evidence\.

TABLE I:Most frequent reviewed phase\-process labels at each ordinal phase position in Trials 1 and 2\.TABLE II:Reviewed team\-process functions and supporting NMF topic labels across detected phases in Trial 1\.TABLE III:Reviewed team\-process functions and supporting NMF topic labels across detected phases in Trial 2\.
## VIIIDiscussion, Limitations, Conclusion, and Ethics

This study presents a phase\-centered framework for detecting and interpreting dynamic team\-process phases from timestamped collaborative VR dialogue\. It combines context\-aware transcript representations, adaptive semantic segmentation, evidence\-grounded LLM\-assisted interpretation, and phase\-aligned interaction analysis\. Because boundaries are fixed before interpretation, the assigned labels cannot influence their placement\. Among the tested configurations, the configuration combining late\-chunkedjina\-v2\-de, mean pooling, and CP\-kernel segmentation achieves the highest semantic separation\. The main contribution is the detection of variable\-length semantic phases without relying on fixed temporal windows or trial\-level summaries\. Each phase remains traceable to the original transcript segments, TF–IDF terms, and NMF topic evidence, allowing the resulting labels to be reviewed against their underlying evidence\. The LLM supports this process by generating initial candidate labels, while final interpretations remain subject to human review\.

The detected trajectories show how team communication shifts across grounding, orientation, coordination, and task execution\. Independently recorded interaction and performance logs are aligned only after segmentation, allowing the transcript\-derived phases to be examined in relation to corresponding task activity\. This may support future after\-action review and phase\-specific feedback in collaborative training\.

Several limitations remain\. Semantic separation is an internal evaluation metric because boundary detection and evaluation use the same embedding representation\. In contrast, the independently recorded interaction and performance logs provide complementary behavioral evidence that is aligned only after segmentation\. The reviewed labels should nevertheless be understood as evidence\-grounded interpretations rather than direct measurements of latent psychological states or replacements for independent behavioral coding\. Future work should compare the detected phases with trainer ratings, participant self\-reports, independent behavioral annotations, physiological signals, and broader performance outcomes\.

The workflow also raises data\-protection requirements because audio, transcripts, timestamps, and speaker information may contain identifiable data\. Relevant safeguards include anonymization, data minimization, secure storage, local processing where possible, and continued human oversight of LLM\-assisted interpretation\. The framework is intended for research and training analysis rather than for ranking or diagnosing individual participants\.

In conclusion, the framework transforms timestamped VR dialogue into interpretable semantic phase trajectories that remain traceable to transcript evidence and can be examined alongside task\-action profiles\. It provides a reproducible basis for analyzing how team communication and coordination change over time in collaborative human–machine systems\.

## Conflict of Interest

The authors declare that they have no conflicts of interest\.

## References

- \[1\]K\. Chang, M\. Hsu, S\. Li, and H\. Lee\(2024\)Exploring in\-context learning of textless speech language model for speech classification tasks\.InProceedings of Interspeech 2024,pp\. 4713–4717\.Cited by:[§II](https://arxiv.org/html/2608.18660#S2.p4.1)\.
- \[2\]N\. J\. Cooke, J\. C\. Gorman, C\. W\. Myers, and J\. L\. Duran\(2013\)Interactive team cognition\.Cognitive Science37\(2\),pp\. 255–285\.Cited by:[§I](https://arxiv.org/html/2608.18660#S1.p1.1),[§I](https://arxiv.org/html/2608.18660#S1.p2.1),[§II](https://arxiv.org/html/2608.18660#S2.p1.1)\.
- \[3\]Q\. Dong, L\. Li, D\. Dai, C\. Zheng, J\. Ma, R\. Li, H\. Xia, J\. Xu, Z\. Wu, B\. Chang, X\. Sun, L\. Li, and Z\. Sui\(2024\)A survey on in\-context learning\.InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing,Y\. Al\-Onaizan, M\. Bansal, and Y\. Chen \(Eds\.\),Miami, Florida, USA,pp\. 1107–1128\.Cited by:[§I](https://arxiv.org/html/2608.18660#S1.p3.1),[§II](https://arxiv.org/html/2608.18660#S2.p4.1)\.
- \[4\]B\. Fyhn, V\. Schei, and T\. E\. Sverdrup\(2023\)Taking the emergent in team emergent states seriously: a review and preview\.Human Resource Management Review33,pp\. 100928\.External Links:[Document](https://dx.doi.org/10.1016/j.hrmr.2022.100928)Cited by:[§I](https://arxiv.org/html/2608.18660#S1.p2.1),[§II](https://arxiv.org/html/2608.18660#S2.p1.1)\.
- \[5\]E\. Georganta, C\. S\. Burke, S\. Merk, and F\. Mann\(2021\)Understanding how team process\-sequences emerge over time and their relationship to team performance\.Team Performance Management27\(3/4\),pp\. 159–174\.Cited by:[§I](https://arxiv.org/html/2608.18660#S1.p2.1),[§II](https://arxiv.org/html/2608.18660#S2.p1.1)\.
- \[6\]B\. Goldberg, R\. Spain, A\. Sinatra, K\. Brawner, and R\. Sottilare\(2024\)Towards a multimodal data\-driven framework for adaptive coaching in collaborative simulation\-based training\.InWorkshop on Artificial Intelligence in Support of Guided Experiential Learning,Note:Metadata transcribed from uploaded manuscriptCited by:[§I](https://arxiv.org/html/2608.18660#S1.p1.1),[§I](https://arxiv.org/html/2608.18660#S1.p4.1),[§II](https://arxiv.org/html/2608.18660#S2.p1.1)\.
- \[7\]D\. A\. P\. Grimm, J\. C\. Gorman, N\. J\. Cooke, M\. Demir, and N\. J\. McNeese\(2023\)Dynamical measurement of team resilience\.Journal of Cognitive Engineering and Decision Making17\(4\),pp\. 351–382\.External Links:[Document](https://dx.doi.org/10.1177/15553434231199729)Cited by:[§II](https://arxiv.org/html/2608.18660#S2.p1.1)\.
- \[8\]M\. G”unther, I\. Mohr, D\. J\. Williams, B\. Wang, and H\. Xiao\(2024\)Late chunking: contextual chunk embeddings using long\-context embedding models\.arXiv preprint arXiv:2409\.04701\.Cited by:[§I](https://arxiv.org/html/2608.18660#S1.p3.1),[§I](https://arxiv.org/html/2608.18660#S1.p4.1),[§II](https://arxiv.org/html/2608.18660#S2.p3.1)\.
- \[9\]J\. L\. Harrison, S\. A\. Jain, T\. Dunbar, J\. C\. Gorman, and S\. Varma\(2022\)Toward automated detection of phase changes in team collaboration\.InProceedings of the 44th Annual Conference of the Cognitive Science Society,pp\. 2357–2363\.Cited by:[§I](https://arxiv.org/html/2608.18660#S1.p2.1),[§II](https://arxiv.org/html/2608.18660#S2.p3.1)\.
- \[10\]M\. A\. Hearst\(1997\)TextTiling: segmenting text into multi\-paragraph subtopic passages\.Computational Linguistics23\(1\),pp\. 33–64\.Cited by:[§II](https://arxiv.org/html/2608.18660#S2.p3.1)\.
- \[11\]M\. Jia and J\. Diaz\-Rodriguez\(2026\)Unsupervised text segmentation via kernel change\-point detection on sentence embeddings\.arXiv preprint arXiv:2601\.18788\.Cited by:[§I](https://arxiv.org/html/2608.18660#S1.p4.1),[§II](https://arxiv.org/html/2608.18660#S2.p3.1)\.
- \[12\]T\. Khandelwal\(2025\)Using llm\-based approaches to enhance and automate topic labeling\.arXiv preprint arXiv:2502\.18469\.Cited by:[§I](https://arxiv.org/html/2608.18660#S1.p3.1),[§II](https://arxiv.org/html/2608.18660#S2.p4.1)\.
- \[13\]F\. Klonek, M\. Twemlow, M\. Tims, and S\. K\. Parker\(2025\)It’s about time\! understanding the dynamic team process\-performance relationship using micro\- and macroscale time lenses\.Group & Organization Management50\(5–6\),pp\. 1660–1702\.External Links:[Document](https://dx.doi.org/10.1177/10596011241278556)Cited by:[§I](https://arxiv.org/html/2608.18660#S1.p1.1),[§I](https://arxiv.org/html/2608.18660#S1.p2.1),[§II](https://arxiv.org/html/2608.18660#S2.p1.1)\.
- \[14\]J\. N\. Lane, P\. M\. Leonardi, N\. S\. Contractor, and L\. A\. DeChurch\(2024\)Teams in the digital workplace: technology’s role for communication, collaboration, and performance\.Small Group Research55\(1\),pp\. 139–183\.External Links:[Document](https://dx.doi.org/10.1177/10464964231200015)Cited by:[§I](https://arxiv.org/html/2608.18660#S1.p1.1),[§II](https://arxiv.org/html/2608.18660#S2.p1.1)\.
- \[15\]M\. Lehmann, J\. Mikulasch, H\. Poimann, J\. Backhaus, S\. König, and T\. Mühling\(2025\)Training and assessing teamwork in interprofessional virtual reality\-based simulation using the teamstepps framework: protocol for a randomized pre\-post intervention study\.JMIR Research Protocols14,pp\. e68705\.External Links:[Document](https://dx.doi.org/10.2196/68705)Cited by:[§I](https://arxiv.org/html/2608.18660#S1.p1.1)\.
- \[16\]N\. Lehmann\-Willenbrock and H\. Hung\(2024\)A multimodal social signal processing approach to team interactions\.Organizational Research Methods27\(3\),pp\. 477–515\.External Links:[Document](https://dx.doi.org/10.1177/10944281231202741)Cited by:[§I](https://arxiv.org/html/2608.18660#S1.p1.1),[§II](https://arxiv.org/html/2608.18660#S2.p1.1)\.
- \[17\]J\. Li and C\. Li\(2025\)Understanding the change trajectories of team transition and action processes over time: a regulatory focus perspective\.Journal of Organizational Behavior46,pp\. 850–866\.External Links:[Document](https://dx.doi.org/10.1002/job.2878)Cited by:[§I](https://arxiv.org/html/2608.18660#S1.p2.1),[§II](https://arxiv.org/html/2608.18660#S2.p1.1)\.
- \[18\]M\. A\. Marks, J\. E\. Mathieu, and S\. J\. Zaccaro\(2001\)A temporally based framework and taxonomy of team processes\.Academy of Management Review26\(3\),pp\. 356–376\.Cited by:[§I](https://arxiv.org/html/2608.18660#S1.p1.1),[§I](https://arxiv.org/html/2608.18660#S1.p2.1),[§II](https://arxiv.org/html/2608.18660#S2.p1.1)\.
- \[19\]C\. Merola and J\. Singh\(2025\)Reconstructing context: evaluating advanced chunking strategies for retrieval\-augmented generation\.InKnowledge\-Enhanced Information Retrieval: Second International Workshop, KEIR 2025, Lucca, Italy, April 10, 2025, Revised Selected Papers,Berlin, Heidelberg,pp\. 3–18\.External Links:ISBN 978\-3\-032\-02898\-3Cited by:[§I](https://arxiv.org/html/2608.18660#S1.p3.1),[§II](https://arxiv.org/html/2608.18660#S2.p3.1)\.
- \[20\]C\. Peifer, A\. Pollak, O\. Flak, A\. Pyszka, M\. A\. Nisar, M\. T\. Irshad, M\. Grzegorzek, B\. Kordyaka, and B\. Kożusznik\(2021\)The symphony of team flow in virtual teams: using artificial intelligence for its recognition and promotion\.Frontiers in Psychology12,pp\. 697093\.External Links:[Document](https://dx.doi.org/10.3389/fpsyg.2021.697093)Cited by:[§II](https://arxiv.org/html/2608.18660#S2.p1.1)\.
- \[21\]A\. Radford, J\. W\. Kim, T\. Xu, G\. Brockman, C\. McLeavey, and I\. Sutskever\(2023\)Robust speech recognition via large\-scale weak supervision\.InProceedings of the 40th International Conference on Machine Learning,Vol\.202,pp\. 28492–28518\.Cited by:[§IV\-A](https://arxiv.org/html/2608.18660#S4.SS1.p2.1)\.
- \[22\]A\. Raut, P\. Paromita, S\. Begerowski, S\. Bell, and T\. Chaspari\(2025\)Assessing the feasibility of Large Language Models for detecting micro\-behaviors in team interactions during space missions\.InInterspeech,pp\. 5453–5457\.External Links:ISSN 2958\-1796Cited by:[§II](https://arxiv.org/html/2608.18660#S2.p4.1)\.
- \[23\]R\. Spain, W\. Min, V\. Kumaran, J\. Pande, J\. Saville, and J\. Lester\(2025\)Applying large language models to enhance dialogue and communication analysis for adaptive team training\.International Journal of Artificial Intelligence in Education35,pp\. 2534–2568\.External Links:[Document](https://dx.doi.org/10.1007/s40593-025-00479-5)Cited by:[§I](https://arxiv.org/html/2608.18660#S1.p3.1),[§II](https://arxiv.org/html/2608.18660#S2.p4.1)\.
- \[24\]D\. Sparks, R\. Begum, F\. Aqlan, J\. Saleem, and M\. DeCaro\(2025\)Exploring team dynamics in virtual reality environments\.IISE Transactions on Occupational Ergonomics and Human Factors,pp\. 1–15\.External Links:[Document](https://dx.doi.org/10.1080/24725838.2025.2461481)Cited by:[§I](https://arxiv.org/html/2608.18660#S1.p1.1)\.
- \[25\]S\. Wanna, N\. Solovyev, R\. Barron, M\. E\. Eren, M\. Bhattarai, K\. Ø\. Rasmussen, and B\. S\. Alexandrov\(2024\)TopicTag: automatic annotation of NMF topic models using chain of thought and prompt tuning with LLMs\.InProceedings of the ACM Symposium on Document Engineering 2024,DocEng ’24\.External Links:ISBN 9798400711695Cited by:[§I](https://arxiv.org/html/2608.18660#S1.p3.1),[§II](https://arxiv.org/html/2608.18660#S2.p4.1)\.
- \[26\]J\. Wolfartsberger, J\. Zenisek, and N\. Wild\(2020\)Supporting teamwork in industrial virtual reality applications\.Procedia Manufacturing42,pp\. 2–7\.External Links:[Document](https://dx.doi.org/10.1016/j.promfg.2020.02.016)Cited by:[§I](https://arxiv.org/html/2608.18660#S1.p1.1)\.

Similar Articles

Searching for Synergy in Shared Workspace Human-AI Collaboration

arXiv cs.AI

This paper studies human-AI team coordination in shared workspaces using the Collaborative Gym and DiscoveryBench tasks, finding that adding collaborators can lower performance without proper structure. Scaffolding with shared group memory and human-in-the-loop gates improves performance, especially in three-person teams.