HB-PVI: A Hierarchical Bayesian Personalization and Value-of-Information Framework for Complex Activity Recognition

arXiv cs.LG Papers

Summary

HB-PVI is a hierarchical Bayesian framework that optimizes personalization decisions in complex activity recognition by balancing gains and costs, demonstrating that a population-first deployment policy can reduce labeling expenses while maintaining performance.

arXiv:2609.05582v1 Announce Type: new Abstract: Personalization can improve activity-recognition performance, but participant-specific gains are heterogeneous, and every additional calibration label has an acquisition cost. This study presents HB-PVI, a hierarchical Bayesian personalization and value-of-information framework jointly modeling participant heterogeneity, the benefit and harm of four personalization mechanisms, and the economic value of an additional label, for the 47-participant MUSIC-CAR complex-activity cohort. A leakage-safe, leave-one-participant-out evaluation combines a sequential-Monte-Carlo participant-effect updater with a Student-$t$ hierarchical gain model and a one-step expected-value-of-sample-information (EVSI) stopping rule. Adapter personalization produced small positive mean F1 gains, growing from 0.00099 at one label to 0.00198 at ten, while adapter-plus-head and prototype-residual personalization were negative on average. Under the primary practical-benefit threshold ($\Delta_{\min}=0.01$) and cost setting, one-step EVSI was zero at every decision state, so the policy purchased no labels and retained population inference for all 47 participants, matching always-stop exactly (region-of-practical-equivalence probability $=1$). Relative to fixed ten-shot adapter personalization, this reduced labeling by 100\% while keeping the posterior mean F1 loss at 0.00217 (95\% credible interval, 0.00048 to 0.00389), with posterior probability 0.9992 of remaining below the 0.005 tolerance. HB-PVI was utility-optimal in 199 of 216 cost-threshold settings and in every setting at or above the primary label cost. These results argue for a population-first deployment policy whenever personalization gains are small relative to labeling, computation, and harm costs, and show that value-of-information reasoning, not raw predictive accuracy, should drive personalization decisions in health-sensing applications.
Original Article
View Cached Full Text

Cached at: 09/10/26, 08:22 AM

# HB-PVI: A Hierarchical Bayesian Personalization and Value-of-Information Framework for Complex Activity Recognition
Source: [https://arxiv.org/html/2609.05582](https://arxiv.org/html/2609.05582)
Hammed A\. Olayinka[https://orcid.org/0000-0002-9796-5276](https://orcid.org/0000-0002-9796-5276)††thanks:H\. A\. Olayinka is with the Department of Mathematical Sciences, Worcester Polytechnic Institute, Worcester, MA 01609 USA \(e\-mail: haolayinka@wpi\.edu\)\.

###### Abstract

Personalization can improve activity\-recognition performance, but participant\-specific gains are heterogeneous and every additional calibration label has an acquisition cost\. This study presents HB\-PVI, a hierarchical Bayesian personalization and value\-of\-information framework that jointly models participant heterogeneity, the benefit and harm of four personalization mechanisms, and the economic value of an additional label, for the MUSIC\-CAR complex\-activity cohort of 47 participants\. A leakage\-safe, leave\-one\-participant\-out evaluation combines a sequential\-Monte\-Carlo participant\-effect updater with a Student\-tthierarchical gain model and a one\-step expected\-value\-of\-sample\-information \(EVSI\) stopping rule\. Adapter personalization produced small positive mean F1 gains that grew from 0\.00099 at one label to 0\.00198 at ten labels, whereas joint adapter\-plus\-head and prototype\-residual personalization were negative on average\. Under the primary practical\-benefit threshold \(Δmin=0\.01\\Delta\_\{\\min\}=0\.01\) and cost setting, one\-step EVSI was exactly zero at every decision state, so the policy purchased no labels and retained population inference for all 47 participants, matching always\-stop exactly \(region\-of\-practical\-equivalence probability=1=1\)\. Relative to fixed ten\-shot adapter personalization, this reduced labeling by 100% while keeping the posterior mean F1 loss at 0\.00217 \(95% credible interval, 0\.00048 to 0\.00389\), with posterior probability 0\.9992 of remaining below the 0\.005 tolerance\. HB\-PVI was utility\-optimal in 199 of 216 cost\-threshold settings and in every setting at or above the primary label cost\. These results argue for a population\-first deployment policy whenever personalization gains are small relative to labeling, computation, and harm costs, and they show why value\-of\-information reasoning, not raw predictive accuracy, should drive personalization decisions in health\-sensing applications\.

###### Index Terms:

Adaptive label acquisition, Decision analysis, Health sensing, Human activity recognition, Multi\-label classification, Participant heterogeneity, Uncertainty quantification\.

## IIntroduction

Passive monitoring of complex, health\-indicative activities of daily living \(ADLs\) from smartphone and wearable sensors offers a scalable alternative to infrequent, rater\-dependent clinical assessment\[[1](https://arxiv.org/html/2609.05582#bib.bib1)\]\. Wearable\-sensor human activity recognition \(HAR\) has matured substantially over the last decade\[[2](https://arxiv.org/html/2609.05582#bib.bib22)\], yet unlike simple ambulatory activities, complex activities \(CAs\) are composed of simple activities performed sequentially, concurrently, or interleaved\[[3](https://arxiv.org/html/2609.05582#bib.bib3)\], and the same CA can be executed in markedly different ways by different people\[[4](https://arxiv.org/html/2609.05582#bib.bib2)\]\. Population models trained across many participants therefore often generalize poorly to a specific new user\[[1](https://arxiv.org/html/2609.05582#bib.bib1)\], motivating*personalization*: adapting a population model with a small number of participant\-specific labels\. Personalization is not free, however\. Adapting on unrepresentative or noisy calibration windows can degrade performance for participants whom the population model already served well\[[1](https://arxiv.org/html/2609.05582#bib.bib1)\], and every calibration label carries an acquisition cost in patient or clinician time\. A deployable personalization policy must therefore decide, for each participant and at each point in a calibration sequence,*which*personalization mechanism \(if any\) to apply and*whether*to purchase another label at all\.

Prior work on this cohort established that a personalized transformer \(P\-HART\) with fixed 6\-shot user adapters raised mean F1 from 88\.0% to 92\.6% on average, though personalization improved performance for 41 of 47 participants and slightly reduced it for the remaining six, who already had high non\-personalized performance\[[1](https://arxiv.org/html/2609.05582#bib.bib1)\], and that complex\-activity recognition itself is substantially harder than simple\-activity recognition, with baseline machine\-learning F1 of only 58\.8% on the underlying MUSIC\-CAR dataset\[[4](https://arxiv.org/html/2609.05582#bib.bib2)\]\. Neither study modeled participant\-level uncertainty in the personalization decision itself or priced the value of an additional calibration label against its cost\. This paper addresses this gap with HB\-PVI \(Hierarchical Bayesian Personalization with Value of Information\): a fully Bayesian inferential and decision layer, conditioned on deterministic cross\-fitted sensor representations, that \(1\) propagates population and participant uncertainty into calibrated activity predictions, \(2\) models the participant\-specific benefit and harm distribution of four personalization mechanisms, \(3\) selects the mechanism maximizing posterior expected utility net of labeling, computation, and harm costs, and \(4\) uses a one\-step expected value of sample information \(EVSI\) rule\[[5](https://arxiv.org/html/2609.05582#bib.bib18),[6](https://arxiv.org/html/2609.05582#bib.bib21),[7](https://arxiv.org/html/2609.05582#bib.bib14)\]to decide whether a further label is worth acquiring\. Framing personalization as a sequential label\-purchase decision, rather than a one\-shot model\-selection problem, distinguishes HB\-PVI from few\-shot and meta\-learning approaches to personalization\[[8](https://arxiv.org/html/2609.05582#bib.bib24)\]that optimize predictive accuracy without pricing the cost of the adaptation data itself\.

HB\-PVI builds on a substantial Bayesian and decision\-theoretic literature\. Hierarchical partial pooling and regularized shrinkage priors\[[9](https://arxiv.org/html/2609.05582#bib.bib4),[10](https://arxiv.org/html/2609.05582#bib.bib12)\]let participant\-level effects borrow strength across a modest 47\-participant cohort; Hamiltonian Monte Carlo with the No\-U\-Turn Sampler\[[11](https://arxiv.org/html/2609.05582#bib.bib5),[12](https://arxiv.org/html/2609.05582#bib.bib16),[13](https://arxiv.org/html/2609.05582#bib.bib17)\]provides asymptotically exact posterior inference rather than a variational or Laplace approximation\[[14](https://arxiv.org/html/2609.05582#bib.bib10),[15](https://arxiv.org/html/2609.05582#bib.bib11)\]; and value\-of\-information and Bayesian experimental\-design theory\[[5](https://arxiv.org/html/2609.05582#bib.bib18),[7](https://arxiv.org/html/2609.05582#bib.bib14),[16](https://arxiv.org/html/2609.05582#bib.bib9),[17](https://arxiv.org/html/2609.05582#bib.bib15)\]formalizes when to stop collecting labels, echoing expected value of perfect/sample information \(EVPI/EVSI\) methods long used in medical decision\-making and health\-technology assessment\[[18](https://arxiv.org/html/2609.05582#bib.bib28),[19](https://arxiv.org/html/2609.05582#bib.bib29)\], applied here at the level of an individual participant’s calibration sequence rather than a population\-level trial design\. This study uses the region\-of\-practical\-equivalence \(ROPE\) framework\[[20](https://arxiv.org/html/2609.05582#bib.bib13)\]to test practical equivalence to simpler comparators, and adopts convergence and calibration diagnostics standard in applied Bayesian workflows\[[21](https://arxiv.org/html/2609.05582#bib.bib6),[22](https://arxiv.org/html/2609.05582#bib.bib7)\]\. The personalization mechanisms \(lightweight adapters, classification\-head fine\-tuning, and their combination\) parallel parameter\-efficient transfer learning\[[23](https://arxiv.org/html/2609.05582#bib.bib25),[24](https://arxiv.org/html/2609.05582#bib.bib23)\]applied to the attention\-based sequence backbone used for complex\-activity recognition\[[25](https://arxiv.org/html/2609.05582#bib.bib26),[26](https://arxiv.org/html/2609.05582#bib.bib27),[1](https://arxiv.org/html/2609.05582#bib.bib1)\]\.

HB\-PVI also sits alongside just\-in\-time adaptive interventions and contextual\-bandit algorithms in mobile health, which decide in real time which intervention to deliver given a person’s current context\[[27](https://arxiv.org/html/2609.05582#bib.bib30),[28](https://arxiv.org/html/2609.05582#bib.bib31)\]; that literature typically treats context as free, whereas HB\-PVI addresses the upstream question of whether the context itself is worth purchasing\.

This study’s contributions are: \(i\) a fully Bayesian personalization\-benefit model that reports calibrated, participant\-specific probabilities of meaningful benefit and harm for four personalization mechanisms; \(ii\) a validated sequential\-Monte\-Carlo participant\-effect updater that maintains at least 400 distinct population\-posterior ancestors while incorporating adaptation labels one at a time; \(iii\) a one\-step EVSI stopping rule, evaluated over a sensitivity grid of 216 practical\-benefit\-threshold and cost combinations, that determines whether personalization is worth its price; and \(iv\) a complete leakage\-safe evaluation on 47 participants showing that, under primary costs, no further labels are justified and a population\-first policy preserves F1 within a primary tolerance while eliminating calibration labeling entirely\.

## IIMethods

### II\-ACohort, design, and leakage control

This study used the MUSIC\-CAR cohort of 47 participants performing sequential, concurrent, and interleaved complex ADLs \(nine activity labels,k=1,…,9k=1,\\dots,9\) from wrist\- and pocket\-mounted smartphone sensors\[[4](https://arxiv.org/html/2609.05582#bib.bib2)\]\. The outer evaluation followed leave\-one\-participant\-out logic: for held\-out participantp⋆p^\{\\star\}, all representation training, normalization, and hyperparameter selection excluded that participant\. Windows were ordered prospectively into a 10\-window adaptation block, two 12\-window guard gaps, a 128\-window calibration block, and a prospective test block, so a decision at shotnn\(n=0,…,10n=0,\\dots,10\) used only the firstnnpurchased labels\. Because adjacent 60\-second, 5\-second\-stride windows overlap by approximately 92%, the primary likelihood used a deterministic, label\-blind non\-overlapping subsample to limit pseudo\-replication\.

### II\-BDeterministic representation and cross\-fitting

A nine\-member deep ensemble of P\-HART\-style backbones\[[1](https://arxiv.org/html/2609.05582#bib.bib1)\]\(an attention encoder\[[25](https://arxiv.org/html/2609.05582#bib.bib26)\]with a bidirectional GRU\[[26](https://arxiv.org/html/2609.05582#bib.bib27)\]temporal head\) was trained on nested\-training participants only, with fold\-specific normalization and training\-only principal\-components analysis \(PCA\)\. One deterministic member supplied theqq\-dimensional pooled representation𝐳p​i∈ℝq\\mathbf\{z\}\_\{pi\}\\in\\mathbb\{R\}^\{q\}\(participantpp, windowii\) used to condition every downstream Bayesian model, so a single, fixed latent coordinate system was shared across all analyses; the remaining members were used only to compute the predictive\-uncertainty summaries \(mutual information, entropy, predictive variance\) entering the Model II sequential state \(Section[II\-D](https://arxiv.org/html/2609.05582#S2.SS4)\), not as an approximate posterior\. Every unknown quantity downstream of𝐳p​i\\mathbf\{z\}\_\{pi\}was assigned a prior and sampled with Hamiltonian Monte Carlo / NUTS\[[11](https://arxiv.org/html/2609.05582#bib.bib5),[13](https://arxiv.org/html/2609.05582#bib.bib17),[12](https://arxiv.org/html/2609.05582#bib.bib16)\], so inference is fully Bayesian conditional on leakage\-safe representations, without variational or Laplace approximation\[[14](https://arxiv.org/html/2609.05582#bib.bib10),[15](https://arxiv.org/html/2609.05582#bib.bib11)\]\.

### II\-CModel I: hierarchical multi\-label prediction and sequential updating

For participantpp, windowii, and activitykk, letyp​i​k∈\{0,1\}y\_\{pik\}\\in\\\{0,1\\\}denote the observed label\. Conditional on the representation𝐳p​i\\mathbf\{z\}\_\{pi\},

yp​i​k∼Bernoulli⁡\(σ⁡\(ηp​i​k\)\),ηp​i​k=αk\+𝐳p​i𝖳​𝜷k\+bp​k,y\_\{pik\}\\sim\\mathrm\{Bernoulli\}\(\\sigma\(\\eta\_\{pik\}\)\),\\quad\\eta\_\{pik\}=\\alpha\_\{k\}\+\\mathbf\{z\}\_\{pi\}^\{\\mathsf\{T\}\}\\bm\{\\beta\}\_\{k\}\+b\_\{pk\},\(1\)whereσ⁡\(⋅\)\\sigma\(\\cdot\)is the logistic function;αk\\alpha\_\{k\}is an activity\-specific intercept;𝜷k\\bm\{\\beta\}\_\{k\}is an activity\-specific representation\-coefficient vector; andbp​kb\_\{pk\}is a participant\-specific intercept for activitykk, collected across activities as𝐛p=\(bp​1,…,bp​9\)𝖳\\mathbf\{b\}\_\{p\}=\(b\_\{p1\},\\dots,b\_\{p9\}\)^\{\\mathsf\{T\}\}\. Because theqq\-dimensional representation is high\-dimensional relative to the per\-fold training data, each entry of𝜷k\\bm\{\\beta\}\_\{k\}follows a regularized horseshoe prior\[[10](https://arxiv.org/html/2609.05582#bib.bib12)\]with activity\-specific local and global shrinkage scales and a finite slab scale of 0\.5 \(full specification in Supplementary Section S1\); the fitted primary specification used this horseshoe on𝜷k\\bm\{\\beta\}\_\{k\}alone, with no additional residual cross\-activity factor term\. Participant intercepts follow𝐛p∼𝒩9​\(𝟎,diag⁡\(𝝉b\)​𝛀b​diag⁡\(𝝉b\)\)\\mathbf\{b\}\_\{p\}\\sim\\mathcal\{N\}\_\{9\}\(\\mathbf\{0\},\\operatorname\{diag\}\(\\bm\{\\tau\}\_\{b\}\)\\,\\bm\{\\Omega\}\_\{b\}\\,\\operatorname\{diag\}\(\\bm\{\\tau\}\_\{b\}\)\), where𝝉b\\bm\{\\tau\}\_\{b\}are activity\-specific participant\-effect scales and𝛀b\\bm\{\\Omega\}\_\{b\}is a correlation matrix with an LKJ prior\[[29](https://arxiv.org/html/2609.05582#bib.bib8)\]\. Let𝚯\\bm\{\\Theta\}collect all population\-level parameters \(αk\\alpha\_\{k\},𝜷k\\bm\{\\beta\}\_\{k\},𝝉b\\bm\{\\tau\}\_\{b\},𝛀b\\bm\{\\Omega\}\_\{b\}, and the analogous Model II parameters introduced below\)\. For an unseen participantp⋆p^\{\\star\}, afternnpurchased adaptation windows𝒜p⋆,n\\mathcal\{A\}\_\{p^\{\\star\},n\}\(the labels from the firstnnchronological adaptation windows\), and writing𝒟−p⋆\\mathcal\{D\}\_\{\-p^\{\\star\}\}for all training\-side data excluding participantp⋆p^\{\\star\}, the participant\-effect posterior is

p\(𝐛p⋆,𝚯∣𝒟−p⋆,𝒜p⋆,n\)\\displaystyle p\(\\mathbf\{b\}\_\{p^\{\\star\}\},\\bm\{\\Theta\}\\mid\\mathcal\{D\}\_\{\-p^\{\\star\}\},\\mathcal\{A\}\_\{p^\{\\star\},n\}\)∝p⁡\(𝒜p⋆,n∣𝐛p⋆,𝚯\)​p​\(𝐛p⋆∣𝚯\)​p​\(𝚯∣𝒟−p⋆\)\.\\displaystyle\\quad\\propto p\(\\mathcal\{A\}\_\{p^\{\\star\},n\}\\mid\\mathbf\{b\}\_\{p^\{\\star\}\},\\bm\{\\Theta\}\)\\,p\(\\mathbf\{b\}\_\{p^\{\\star\}\}\\mid\\bm\{\\Theta\}\)\\,p\(\\bm\{\\Theta\}\\mid\\mathcal\{D\}\_\{\-p^\{\\star\}\}\)\.\(2\)This posterior is maintained sequentially, one purchased label at a time, by an adaptive tempered resample\-move sequential Monte Carlo algorithm: each of 16,000 retained population\-posterior draws seeds four conditionally independent participant\-effect particles \(64,000 joint particles\); systematic resampling triggers only when effective sample size \(ESS\) falls below 25% of 64,000; random\-walk Metropolis\-Hastings rejuvenates the nine\-dimensional participant effect after every resampling step, with incremental likelihoods recomputed post\-rejuvenation\. Because population\-level parameters are never rejuvenated, every retained state was required to preserve at least 400 distinct population\-posterior ancestors in addition to a final particle ESS of at least 400; all 47 folds passed both gates \(full algorithm and validation\-gate detail in Supplementary Section S1\)\.

### II\-DPersonalization actions, gain, and Model II

Four personalization mechanismsa∈𝒜a\\in\\mathcal\{A\}=\{adapter,head,adapter\+head,prototype\-residual\} were compared against a zero\-gain population baseline\[[23](https://arxiv.org/html/2609.05582#bib.bib25)\]: the adapter mechanism fine\-tunes a small low\-rank bottleneck inserted into the backbone; the head mechanism fine\-tunes only the final classification layer; adapter\-plus\-head fine\-tunes both jointly \(roughly doubling the number of participant\-adapted parameters relative to either mechanism alone\); and prototype\-residual personalization adjusts predictions using a participant\-specific residual from a nearest\-prototype reference, without gradient fine\-tuning\. For actionaa, shotnn, and participantpp, gain isgp​a​n=F​1p​a​na−F​1ppopulationg\_\{pan\}=F1\_\{pan\}^\{a\}\-F1\_\{p\}^\{\\mathrm\{population\}\}\(sample\-level F1 for actionaaminus the population\-only baseline for the same participant\), evaluated on a held\-out set of47×4×10=1,88047\\times 4\\times 10=1\{,\}880fold\-action\-shot observations\. A robust hierarchical benefit model was fit as

100​gi\\displaystyle 100\\,g\_\{i\}∼tν​\(αai\+cpi​ai\+𝐱i𝖳​𝜸ai,σai\),\\displaystyle\\sim t\_\{\\nu\}\(\\alpha\_\{a\_\{i\}\}\+c\_\{p\_\{i\}a\_\{i\}\}\+\\mathbf\{x\}\_\{i\}^\{\\mathsf\{T\}\}\\bm\{\\gamma\}\_\{a\_\{i\}\},\\ \\sigma\_\{a\_\{i\}\}\),\(3\)𝐜p\\displaystyle\\mathbf\{c\}\_\{p\}∼𝒩⁡\(𝟎,diag⁡\(𝝉\)​𝛀​diag⁡\(𝝉\)\),\\displaystyle\\sim\\mathcal\{N\}\(\\mathbf\{0\},\\operatorname\{diag\}\(\\bm\{\\tau\}\)\\,\\bm\{\\Omega\}\\,\\operatorname\{diag\}\(\\bm\{\\tau\}\)\),\(4\)whereiiindexes a fold\-action\-shot observation with actionaia\_\{i\}and participantpip\_\{i\};αai\\alpha\_\{a\_\{i\}\}is an action\-specific intercept;𝐱i\\mathbf\{x\}\_\{i\}is the corresponding sequential\-state predictor vector \(defined below\);𝜸ai\\bm\{\\gamma\}\_\{a\_\{i\}\}is an action\-specific coefficient vector \(distinct from Model I’s𝜷k\\bm\{\\beta\}\_\{k\}\), assigned a fixed\-scale, weakly informative Gaussian prior rather than a sparsity\-inducing one, since with only 19 predictors this was judged sufficient \(unlike Model I’s higher\-dimensional𝜷k\\bm\{\\beta\}\_\{k\}; full priors for both models in Supplementary Section S1\);cp​ac\_\{pa\}is a participant\-by\-action random intercept, collected as𝐜p=\(cp,adapter,…\)𝖳\\mathbf\{c\}\_\{p\}=\(c\_\{p,\\text\{adapter\}\},\\dots\)^\{\\mathsf\{T\}\}with scales𝝉\\bm\{\\tau\}and correlation matrix𝛀\\bm\{\\Omega\};σai\\sigma\_\{a\_\{i\}\}is an action\-specific residual scale; andν\\nuis the Student\-ttdegrees of freedom \(bounded below at 2 for finite variance\), which downweights outlying participant\-fold gains relative to a Gaussian likelihood\. The predictor vector𝐱i\\mathbf\{x\}\_\{i\}contains 19 sequential\-state predictors spanning purchased\-label prevalence and diversity, ensemble uncertainty, posterior participant\-effect summaries, calibration probability, predictive entropy, shot\-to\-shot change, and representation shift \(Supplementary Table S3\)\. Action\-specific coefficient vectors allow each state predictor to have a different association with gain for each personalization mechanism; Model II did not use a horseshoe prior or an additional interaction layer beyond these action\-specific coefficients\. Four chains were fit per held\-out fold with convergence gates \(rank\-normalizedR^<1\.01\\widehat\{R\}<1\.01\[[21](https://arxiv.org/html/2609.05582#bib.bib6)\], zero divergences, minimum bulk/tail ESS\>400\>400\); all 47 folds passed with larger ESS\(\>1000\)\(\>1000\)and smallerR^<1\.01\\widehat\{R\}<1\.01\.

### II\-EUtility, Bayes action, and EVSI

Letℋn\\mathcal\{H\}\_\{n\}denote the sequential state available afternnpurchased labels \(the posterior summaries and purchased\-label history described above\)\. Cumulative realized utility for actionaaafternnlabels with gainggis

U⁡\(a,n,g\)=g−n​cL−cC​\(a\)−cR​𝕀​\(g<0\),U\(a,n,g\)=g\-nc\_\{L\}\-c\_\{C\}\(a\)\-c\_\{R\}\\,\\mathbb\{I\}\(g<0\),\(5\)wherecLc\_\{L\}is the per\-label acquisition cost,cC​\(a\)c\_\{C\}\(a\)is the action\-specific computation cost,cRc\_\{R\}is a harm penalty applied when realized gain is negative, and𝕀⁡\(⋅\)\\mathbb\{I\}\(\\cdot\)is the indicator function\. The Bayes action maximizes posterior expected utility,an⋆=arg⁡maxa⁡𝔼⁡\[U⁡\(a,n,g\)∣ℋn\]a\_\{n\}^\{\\star\}=\\arg\\max\_\{a\}\\mathbb\{E\}\[U\(a,n,g\)\\mid\\mathcal\{H\}\_\{n\}\], subject to the conservative eligibility ruleP⁡\(g\>Δmin∣ℋn\)≥qP\(g\>\\Delta\_\{\\min\}\\mid\\mathcal\{H\}\_\{n\}\)\\geq qandP⁡\(g<0∣ℋn\)≤rP\(g<0\\mid\\mathcal\{H\}\_\{n\}\)\\leq r, whereΔmin\\Delta\_\{\\min\}is the minimum practically meaningful gain, andq=0\.90q=0\.90,r=0\.10r=0\.10are the primary minimum\-benefit and maximum\-harm probabilities\. To avoid charging label cost twice when comparing the value of one more label, EVSI uses the cost\-free utilityU′​\(a,g\)=g−cC​\(a\)−cR​𝕀​\(g<0\)U^\{\\prime\}\(a,g\)=g\-c\_\{C\}\(a\)\-c\_\{R\}\\,\\mathbb\{I\}\(g<0\); writingYn\+1Y\_\{n\+1\}for the hypothetical outcome of the next purchased label,

EVSIn\+1=\\displaystyle\\mathrm\{EVSI\}\_\{n\+1\}=𝔼Yn\+1\|ℋn\[maxa𝔼\{U′\(a,g\)∣ℋn,Yn\+1\}\]\\displaystyle\\ \\mathbb\{E\}\_\{Y\_\{n\+1\}\\mid\\mathcal\{H\}\_\{n\}\}\\Big\[\\max\_\{a\}\\mathbb\{E\}\\\{U^\{\\prime\}\(a,g\)\\mid\\mathcal\{H\}\_\{n\},Y\_\{n\+1\}\\\}\\Big\]−maxa⁡𝔼⁡\{U′​\(a,g\)∣ℋn\}\.\\displaystyle\-\\max\_\{a\}\\mathbb\{E\}\\\{U^\{\\prime\}\(a,g\)\\mid\\mathcal\{H\}\_\{n\}\\\}\.\(6\)Acquisition continues only whileEVSIn\+1\>cL\\mathrm\{EVSI\}\_\{n\+1\}\>c\_\{L\}, evaluated with 256 hypothetical next\-label draws per state along the fixed chronological window order; active \(non\-chronological\) window selection was deferred to future work\. The sensitivity analysis combinedΔmin∈\{0,0\.005,0\.01\}\\Delta\_\{\\min\}\\in\\\{0,0\.005,0\.01\\\},cL∈\{0,0\.0005,0\.001,0\.0025,0\.005,0\.01\}c\_\{L\}\\in\\\{0,0\.0005,0\.001,0\.0025,0\.005,0\.01\\\},cC∈\{0,0\.0005,0\.001\}c\_\{C\}\\in\\\{0,0\.0005,0\.001\\\}, andcR∈\{0,0\.005,0\.01,0\.02\}c\_\{R\}\\in\\\{0,0\.005,0\.01,0\.02\\\}\(72 cost combinations×\\times3 thresholds=216=216settings\), with primary valuescL=0\.001c\_\{L\}=0\.001,cC=0\.0005c\_\{C\}=0\.0005,cR=0\.01c\_\{R\}=0\.01\.

### II\-FStatistical analysis and study objectives

Participant\-level paired utility differences between HB\-PVI and each comparator were modeled with Student\-ttposteriors, reporting the posterior mean paired differenceμd\\mu\_\{d\}, 95% credible intervals,P⁡\(μd\>0\)P\(\\mu\_\{d\}\>0\), and region\-of\-practical\-equivalence probabilitiesP⁡\(\|μd\|<0\.01\)P\(\|\\mu\_\{d\}\|<0\.01\)\[[20](https://arxiv.org/html/2609.05582#bib.bib13)\]\. Held\-out prediction was summarized by mean absolute error \(MAE\), root\-mean\-squared error \(RMSE\), proper Brier scores\[[30](https://arxiv.org/html/2609.05582#bib.bib20)\], and 95% predictive\-interval coverage using five\-bin reliability curves \(bins with<10<10observations omitted\)\. The label\-efficiency analysis evaluated whether HB\-PVI reduced label use while maintaining a posterior probability greater than 0\.90 that mean F1 loss was below 0\.005\. This analysis used both a participant\-level Bayesian Student\-ttmodel and an independent, non\-parametric 200,000\-resample participant bootstrap\[[31](https://arxiv.org/html/2609.05582#bib.bib19)\]to assess sensitivity to the likelihood\. Detailed analytical objectives are reported in Supplementary Table S12\.

## IIIResults

### III\-AAction\-specific personalization outcomes

Table[I](https://arxiv.org/html/2609.05582#S3.T1)summarizes gain at the first and tenth purchased label for each mechanism \(full ten\-shot trajectories in Fig\.[1](https://arxiv.org/html/2609.05582#S3.F1)\)\. Adapter personalization was the only mechanism with a positive mean gain at every shot, rising from\+0\.00099\+0\.00099to\+0\.00198\+0\.00198; 72\.3% of participants had positive adapter gain at 10 labels, with meaningful benefit \(g\>0\.01g\>0\.01\) in 10\.6% and meaningful harm \(g<−0\.01g<\-0\.01\) in 4\.3%\. Head personalization was near zero on average and became more heterogeneous with more labels \(meaningful benefit rising from 14\.9% to 19\.1%, meaningful harm from 8\.5% to 10\.6%\)\. Adapter\-plus\-head and prototype\-residual personalization were negative on average at every shot, with adapter\-plus\-head showing both the largest benefit subgroup \(up to 21\.3%\) and the largest harm subgroup \(up to 31\.9%\) of any mechanism\.

TABLE I:Action\-specific personalization gain at 1 and 10 purchased labels \(percentages of 47 participants\)- Gain, sample\-level F1 for the personalization action minus population\-only F1; Pos\., percentage of participants with gain\>0\>0; Ben\., percentage with gain exceeding the meaningful\-benefit threshold \(\>0\.01\>0\.01\); Harm, percentage below the meaningful\-harm threshold \(<−0\.01<\-0\.01\)\.

![Refer to caption](https://arxiv.org/html/2609.05582v1/figure1_personalization_trajectories_revised.png)Fig\. 1:Action\-specific personalization outcomes across purchased labels\. \(A\) Mean sample\-level F1 gain relative to population inference, with 95% intervals\. \(B\) Percentage of participants with meaningful benefit \(g\>0\.01g\>0\.01\)\. \(C\) Percentage with meaningful harm \(g<−0\.01g<\-0\.01\)\.
### III\-BPopulation baseline strength and the ceiling on personalization value

The population\-only ensemble member \(no participant\-specific label\) achieved a mean held\-out sample\-level F1 of 0\.9042 across the 47 participants – the reference against which every gain in Table[I](https://arxiv.org/html/2609.05582#S3.T1)and Fig\.[1](https://arxiv.org/html/2609.05582#S3.F1)is measured\. This figure is informative when read against two related studies that share this cohort or sensor modality \(Table[II](https://arxiv.org/html/2609.05582#S3.T2)\)\. On the same MUSIC\-CAR data, five traditional feature\-based machine\-learning baselines reached a best mean F1 of only 58\.8%\[[4](https://arxiv.org/html/2609.05582#bib.bib2)\], underscoring that complex\-activity recognition from raw statistical features is substantially harder than the simple\-activity recognition problem those baselines were designed for\. On the same 47\-participant cohort, a non\-personalized P\-HART transformer reached F1 of 88\.0%, rising to 92\.6% after few\-shot user\-adapter personalization\[[1](https://arxiv.org/html/2609.05582#bib.bib1)\]\. The population\-only backbone used by HB\-PVI exceeded the previously reported non\-personalized P\-HART result by approximately 0\.024 F1, narrowing, but not eliminating, the 0\.046 F1 gap between the previously reported non\-personalized and personalized P\-HART results\. Because the studies differed in training, ensemble configuration, windowing, and evaluation protocol, this comparison is contextual rather than controlled\. Fig\.[2](https://arxiv.org/html/2609.05582#S3.F2)plots each participant’s population\-only F1 against adapter F1 at 1, 6, and 10 labels; most points sit neary=xy=x, and a few participants account for most of the departure\. The lowest\-baseline participant \(F1=0\.566=0\.566\) shows neither benefit nor harm at 1 or 6 shots but crosses into meaningful benefit by 10, showing classification can still change late in the sequence\. The largest\-harm participant \(F1≈0\.79\\approx 0\.79\) remains meaningful harm at all three shots, so not every unfavorable case is transient\.

TABLE II:Population and personalized mean F1 across related studies on the same cohort or sensor modality- aHB\-PVI used a P\-HART\-style backbone within a nine\-model deep ensemble; one deterministic member supplied the representations used in the downstream Bayesian analyses, while the remaining members contributed predictive\-uncertainty summaries\. The MUSIC\-CAR and P\-HART F1 values were taken from the cited publications and were not recomputed here; they are provided for context only, not a controlled head\-to\-head comparison\.

Table[II](https://arxiv.org/html/2609.05582#S3.T2)should be read as context rather than a controlled comparison: the studies differ in ensemble configuration, contrastive\-loss training, cross\-validation protocol, and evaluation windowing\. It nonetheless offers one candidate explanation for why the gains identified by the Model II gain model were both small in absolute F1 units and inconsistent in sign across mechanisms: when the population baseline already captures most of the recoverable signal, the marginal information content of a calibration label about*which*action will help a specific participant is itself small – precisely the condition under which one\-step EVSI is expected to be low\. This mean result concealed substantial participant\-level heterogeneity: realized adapter gain ranged from−0\.0219\-0\.0219to\+0\.0208\+0\.0208at 1 shot,−0\.0249\-0\.0249to\+0\.0193\+0\.0193at 6 shots, and−0\.0265\-0\.0265to\+0\.0168\+0\.0168at 10 shots \(Supplementary Table S6\), even though the population\-averaged gain never exceeded\+0\.00198\+0\.00198\. Exact participant\-specific F1 scores are reported in Supplementary Table S7\.

### III\-CHeld\-out gain prediction and calibration

Table[III](https://arxiv.org/html/2609.05582#S3.T3)report held\-out prediction accuracy\. Prototype\-residual gain was easiest to predict \(MAE 0\.00325, 95\.3% coverage\) and adapter\-plus\-head hardest \(MAE 0\.01396, 88\.9% coverage\); the two more aggressive mechanisms had both larger prediction error and below\-nominal predictive coverage, supporting the conservative eligibility rule used for action selection\. Harm\-specific discrimination was weaker still: comparing each action’s harm Brier score against a trivial baseline that always predicts the action’s empirical harm rate, the fitted model was less accurate than this trivial baseline for every one of the four mechanisms \(e\.g\., adapter: model 0\.269 versus trivial 0\.229; prototype\-residual: model 0\.234 versus trivial 0\.165\)\. Full reliability curves showing where each mechanism’s calibration breaks down \(rather than just the summary coverage and Brier statistics in Table[III](https://arxiv.org/html/2609.05582#S3.T3)\) are provided in Supplementary Fig\. S1\. The gain model therefore did not demonstrate prospective participant\-level discrimination of personalization harm beyond the base rate \(Section[III\-H](https://arxiv.org/html/2609.05582#S3.SS8)\)\.

TABLE III:Held\-out gain\-model prediction and calibration by action- MAE, mean absolute error \(F1 units\); RMSE, root\-mean\-squared error \(F1 units\); Cov\., empirical coverage of the nominal 95% posterior predictive interval; Brierg\>0g\>0andg\>\.01g\>\.01, Brier scores for predicting positive gain and gain exceeding the meaningful\-benefit threshold, respectively \(lower is better; 0 is perfect\)\.

![Refer to caption](https://arxiv.org/html/2609.05582v1/figureS1_population_vs_adapter_f1.png)Fig\. 2:Population\-only F1 versus adapter F1 at \(A\) 1, \(B\) 6, and \(C\) 10 purchased labels, all 47 participants \(Table\)\. Points scatter closely around they=xy=xline at every shot count, and meaningful\-benefit \(green\) and meaningful\-harm \(red\) cases both occur at low and high baseline F1 alike, illustrating that realized gain is not predictable from population\-only performance \(r=−0\.06r=\-0\.06at 10 shots\.
### III\-DEVSI, stopping, and the primary deployment decision

Under the primary cost setting, one\-step EVSI was exactly zero for every state transition from shot 1 through shot 10 \(Fig\.[3](https://arxiv.org/html/2609.05582#S3.F3)A\), so none of the 470 participant\-state decisions continued acquisition\. AtΔmin∈\{0\.005,0\.01\}\\Delta\_\{\\min\}\\in\\\{0\.005,0\.01\\\}the policy stopped immediately, purchased zero labels for all 47 participants \(Fig\.[3](https://arxiv.org/html/2609.05582#S3.F3)B\), and retained population inference \(Fig\.[3](https://arxiv.org/html/2609.05582#S3.F3)C\); realized mean gain and cumulative utility were both exactly zero and no participant was exposed to personalization harm\. AtΔmin=0\\Delta\_\{\\min\}=0the policy purchased a mean of 0\.213 labels \(95% bootstrap interval 0\.000–0\.638\), selecting head personalization for one of the 47 participants after 10 labels and population inference for the rest \(mean gain\+0\.000284\+0\.000284, mean utility\+0\.000061\+0\.000061\)\.

![Refer to caption](https://arxiv.org/html/2609.05582v1/figure3_evsi_and_stopping_revised.png)Fig\. 3:Fixed\-order one\-step EVSI and stopping decisions\. \(A\) Mean EVSI under the primary setting; no state\-level decision continued\. \(B\) Labels purchased across the three practical\-gain thresholds\. \(C\) Action selected at stopping under the primaryΔmin=0\.01\\Delta\_\{\\min\}=0\.01setting\.
### III\-EOne\-step EVSI ceiling across the sensitivity grid

Section III\-C reports EVSI at the primary setting; here we characterize its largest observed value anywhere in the primary design\. Across all 47 participants and 216 cost\-threshold settings \(10,152 policy decisions\), a strictly positive one\-step EVSI occurred in only 72 decisions \(0\.7%\), exclusively atΔmin=0\\Delta\_\{\\min\}=0; atΔmin∈\{0\.005,0\.01\}\\Delta\_\{\\min\}\\in\\\{0\.005,0\.01\\\}no decision, at any label cost, had positive EVSI\. All 72 positive\-EVSI decisions involved the same single participant – the one participant for whom head personalization was ultimately selected atΔmin=0\\Delta\_\{\\min\}=0\(Section III\-C\)\. At label costs at or below the primary value \(cL≤0\.001c\_\{L\}\\leq 0\.001\), this participant’s EVSI exceeded cost at every state, so the policy purchased labels through the full ten\-shot horizon \(mean EVSI per continued step, 0\.0041\)\. At higher label costs \(cL≥0\.0025c\_\{L\}\\geq 0\.0025\), EVSI exceeded cost only immediately after the first label – its peak value anywhere in the study, 0\.01012, narrowly exceeding the largest primary label cost of 0\.01 – and then fell below cost, so the policy purchased exactly one label before stopping\.

### III\-FUtility comparisons, paired differences, and cost sensitivity

Table[IV](https://arxiv.org/html/2609.05582#S3.T4)combines realized comparator utility with the posterior paired\-difference analysis\. Fixed adapter personalization raised mean F1 but had negative cumulative utility once labeling, computation, and harm costs were charged \(e\.g\.,−0\.01129\-0\.01129at 10 shots versus\+0\.00198\+0\.00198mean gain alone\)\. Because HB\-PVI selected population inference for every participant, it was exactly equivalent to always\-stop \(all 47 paired differences=0=0; ROPE probability=1=1\)\. Relative to fixed adapter personalization, the posterior probability that HB\-PVI’s utility was strictly greater was≥0\.989\\geq 0\.989at every shot count tested and reached1\.00001\.0000at 6 and 10 shots\.

TABLE IV:Realized comparator utility and posterior paired\-utility difference \(HB\-PVI minus comparator\)- Util\., mean realized cumulative utility \(F1 units, net of costs\) for the comparator policy;Δ\\Deltamean and 95% CrI, posterior mean and credible interval of the paired utility difference \(HB\-PVI minus comparator\);P⁡\(Δ\>0\)P\(\\Delta\>0\), posterior probability that HB\-PVI’s utility was strictly greater\. The always\-stop comparator is omitted: under the primary setting, HB\-PVI purchased zero labels and selected population inference for all 47 participants, so its decisions and realized utility are identical to always\-stop by construction \(paired difference exactly 0 for every participant; ROPE probability=1=1; see Section III\-D\)\.

Across all 216 cost\-threshold settings, HB\-PVI was utility\-optimal in 199 \(fixed adapter@10 in 14, fixed adapter@1 in 3\)\. Table[V](https://arxiv.org/html/2609.05582#S3.T5)shows why: at zero label cost, HB\-PVI was optimal in only 61\.1% of computation\-and\-harm\-penalty settings, but optimality reached 100% at and above the primary label cost ofcL=0\.001c\_\{L\}=0\.001, across every value ofΔmin\\Delta\_\{\\min\}tested \(Fig\.[4](https://arxiv.org/html/2609.05582#S3.F4)A\)\.

TABLE V:Percentage of the 36 computation\-and\-harm\-penalty settings, per label costcLc\_\{L\}, in which HB\-PVI was utility\-optimal \(across allΔmin\\Delta\_\{\\min\}\)- cLc\_\{L\}, per\-label acquisition cost \(F1\-equivalent units; Section II\-E\)\. Each column aggregates the 36 combinations of computation costcCc\_\{C\}, harm penaltycRc\_\{R\}, and practical\-gain thresholdΔmin\\Delta\_\{\\min\}at that label cost\.

![Refer to caption](https://arxiv.org/html/2609.05582v1/figure4_cost_sensitivity_and_equivalence_revised.png)Fig\. 4:Cost sensitivity and F1 preservation with zero calibration labels\. \(A\) Percentage of computation\-cost and harm\-penalty settings in which HB\-PVI was utility\-optimal, by label cost and practical\-gain threshold; the dotted vertical line marks the primary label cost\. \(B\) Bayesian and bootstrap estimates of mean F1 loss from purchasing zero labels rather than using fixed adapter personalization at 1, 6, or 10 shots\. The dashed vertical line marks the F1\-loss tolerance of 0\.005, and the annotations report the posterior probability that the mean loss remained below this tolerance\.
### III\-GLabel efficiency and F1 preservation

Relative to fixed 10\-shot adapter personalization, HB\-PVI reduced labeling by 100% while the Bayesian posterior mean F1 loss was 0\.00217 \(95% CrI, 0\.00048–0\.00389;P⁡\(μloss<0\.005∣D\)=0\.9992P\(\\mu\_\{\\mathrm\{loss\}\}<0\.005\\mid D\)=0\.9992\); a 200,000\-resample bootstrap gave a similar point estimate \(0\.00198\) but a wider interval that included zero \(95% CI,−0\.00011\-0\.00011to0\.003900\.00390\), reflecting its lack of hierarchical shrinkage across heterogeneous participant\-level losses\. Both methods agreed the loss stayed below the 0\.005 tolerance with probability≥0\.999\\geq 0\.999, including1\.00001\.0000and0\.99990\.9999against the 1\- and 6\-shot references \(Fig\.[4](https://arxiv.org/html/2609.05582#S3.F4)B\)\.

### III\-HHarm avoidance without demonstrated harm identification

At the primary setting, HB\-PVI’s decision rule selected population inference for all 47 participants, so no participant was exposed to realized personalization harm in this study\. This reflects the conservative decision rule under the specified costs, however, rather than a demonstrated ability to identify in advance which participants would be harmed: the fitted gain model’s harm\-specific Brier score was*worse*than a trivial baseline that always predicts each mechanism’s empirical harm rate, for all four personalization mechanisms \(Section[III\-C](https://arxiv.org/html/2609.05582#S3.SS3)\)\. Harm was therefore avoided in this study by declining to personalize under the specified costs, not by successfully flagging high\-risk participants prospectively – a distinction we return to in the Discussion\.

### III\-ISummary of study objectives

In summary, HB\-PVI eliminated calibration labeling while preserving mean F1 within tolerance, was practically equivalent to \(though not demonstrably better than\) always\-stop, and did not establish prospective discrimination of harmful personalization; participant\-level heterogeneity nevertheless changed the deployment recommendation toward a population\-first strategy\. Full definitions and findings for each study objective are in Supplementary Table S12\.

## IVDiscussion

In a leakage\-safe, 47\-participant outer evaluation, personalization produced heterogeneous but generally small changes in sample\-level F1\. Adapter personalization offered the most favorable risk\-benefit profile of the four mechanisms tested, with small, monotonically increasing mean gains and comparatively low harm; adapter\-plus\-head amplified both benefit and harm; and prototype\-residual was reliably unfavorable\. This matches the intuition that more heavily parameterized mechanisms can both help and hurt a given participant more, reflected directly in the calibration results, where the two more aggressive mechanisms had larger prediction error and below\-nominal coverage\.

The decision\-analytic layer changed the interpretation of these predictive results\. Under the primary practical\-benefit, safety, and cost assumptions, the expected value of one additional label never exceeded its cost, at any point in the ten\-shot sequence, for any participant\. The resulting policy purchased zero labels and retained population inference for everyone – not because personalization was ineffective on average, but because its economic value, once labeling, computation, and harm costs were priced in, was smaller than the price of finding out whether it would help a given participant\. This distinction between predictive benefit and decision\-relevant value is the central methodological point of HB\-PVI: average F1 improvement, the outcome most commonly reported in the HAR personalization literature\[[1](https://arxiv.org/html/2609.05582#bib.bib1)\], is not by itself sufficient evidence for a deployment decision\.

The cost\-sensitivity analysis qualifies rather than reverses this conclusion\. HB\-PVI was utility\-optimal in the overwhelming majority of the primary cost\-threshold grid and in every setting at or above the primary label cost; fixed personalization was preferable only in a minority of settings with very low label cost and low harm penalty\. This indicates the population\-first recommendation is a property of the specific cost structure assumed here, not a universal claim that personalization is never worthwhile, and argues for reporting decision surfaces over a cost grid rather than a single point estimate whenever costs are only approximately known\.

### IV\-AImplications for deployment

Translated into an operational recommendation, these results argue against building a routine per\-user calibration step into a MUSIC\-CAR\-style deployment under the primary cost assumptions: a new participant can be served directly by the population model, with no adaptation labels collected, while expected F1 loss relative to a hypothetical adapter\-personalized deployment stays within tolerance with posterior probability 0\.9992\. This recommendation is conditional on the representation\-learning pipeline, cohort, and cost weights used here, and should be revisited whenever any of three things changes materially\. First, if the population backbone is retrained on a smaller or differently sourced cohort and no longer reaches comparably strong F1, the headroom available to personalization enlarges mechanically \(Section III\-B\) and could reverse the decision\. Second, if the relative costs of a label, computation, and harm differ from the values used here, Table[V](https://arxiv.org/html/2609.05582#S3.T5)shows the recommendation is not robust at very low label cost\. Third, applications placing asymmetric value on avoiding subgroup\-specific harm are not directly targeted by the aggregate mean\-utility criterion used here and may warrant a subgroup\-specific harm constraint instead\. None of this requires re\-deriving the EVSI framework itself; only the cost inputs to Table[V](https://arxiv.org/html/2609.05582#S3.T5)and the decision\-region grid \(Supplementary Table S9\) would need recomputing\.

### IV\-BLimitations

The cohort contained 47 participants, limiting power for rare subgroups and higher\-order interactions\. Calibration was weakest for the two more heavily parameterized mechanisms, and harm could not be prospectively identified \(Section[III\-H](https://arxiv.org/html/2609.05582#S3.SS8)\)\. The primary EVSI analysis evaluated only the next chronological window, and costs used a common grid rather than mechanism\-specific runtime; these weights should be re\-elicited for deployment\. The label\-efficiency result concerns the mean F1 difference, not every individual’s loss\. Finally, the tested mechanisms were parameter\-efficient adaptations of a frozen, population\-trained backbone, so larger gains from untested full retraining cannot be ruled out\.

## VConclusion

HB\-PVI provides a fully Bayesian route from participant\-level predictive uncertainty to personalization and label\-acquisition decisions for complex activity recognition\. Under the primary conditions, no further calibration labels were justified for the MUSIC\-CAR cohort, and a population\-first policy preserved F1 within tolerance while eliminating calibration labeling entirely\. Relative to fixed 10\-shot adapter personalization, this policy achieved a 100% reduction in calibration labels, with posterior probability 0\.9992 that the mean F1 loss remained below the 0\.005 tolerance\. These findings demonstrate that small improvements in average predictive performance do not necessarily provide sufficient decision\-relevant value to justify participant\-specific adaptation\. The framework and its EVSI stopping rule generalize to other wearable\-sensing personalization problems in which the benefit of adaptation must be weighed against the cost of acquiring it\.

## Data and Code Availability

The MUSIC\-CAR sensor dataset is available from the CARE Lab at Worcester Polytechnic Institute \(WPI\) at[https://carelab\.wpi\.edu/music\_car/](https://carelab.wpi.edu/music_car/)\. Data preprocessing and analysis code will be made available on GitHub upon publication\.

## Funding

The author received no financial support for the research\.

## Acknowledgment

The author thanks the MUSIC\-CAR data collection team at the CARE Lab at WPI\. Results in this paper were obtained using a high\-performance computing system acquired through NSF MRI grant DMS\-1337943 to WPI\.

## References

- \[1\]\(2026\)A personalized transformer neural network for accurate recognition of health\-indicative complex activities from smartphone sensors\.InProc\. IEEE Int\. Conf\. Biomedical and Health Informatics \(BHI\),Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p1.1),[§I](https://arxiv.org/html/2609.05582#S1.p2.1),[§I](https://arxiv.org/html/2609.05582#S1.p3.1),[§II\-B](https://arxiv.org/html/2609.05582#S2.SS2.p1.1),[§III\-B](https://arxiv.org/html/2609.05582#S3.SS2.p1.1),[TABLE II](https://arxiv.org/html/2609.05582#S3.T2.2.3.1.1),[TABLE II](https://arxiv.org/html/2609.05582#S3.T2.2.4.1.1),[§IV](https://arxiv.org/html/2609.05582#S4.p2.1)\.
- \[2\]A\. Bulling, U\. Blanke, and B\. Schiele\(2014\)A tutorial on human activity recognition using body\-worn inertial sensors\.ACM Comput\. Surv\.46\(3\),pp\. 1–33\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p1.1)\.
- \[3\]S\. Ranasinghe, F\. Al Machot, and H\. C\. Mayr\(2016\)A review on applications of activity recognition systems with regard to performance and evaluation\.Int\. J\. Distrib\. Sensor Netw\.12\(8\),pp\. 1550147716665520\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p1.1)\.
- \[4\]K\. Chandrasekaran, L\. Buquicchio, E\. Agu, and E\. Rundensteiner\(2025\)MUSIC\-CAR: a multi\-label complex health\-related activity recognition dataset from smartphone sensors\.InProc\. IEEE Int\. Conf\. Digital Health \(ICDH\),Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p1.1),[§I](https://arxiv.org/html/2609.05582#S1.p2.1),[§II\-A](https://arxiv.org/html/2609.05582#S2.SS1.p1.1),[§III\-B](https://arxiv.org/html/2609.05582#S3.SS2.p1.1),[TABLE II](https://arxiv.org/html/2609.05582#S3.T2.2.2.1.1)\.
- \[5\]R\. A\. Howard\(1966\)Information value theory\.IEEE Trans\. Syst\. Sci\. Cybern\.2\(1\),pp\. 22–26\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p2.1),[§I](https://arxiv.org/html/2609.05582#S1.p3.1)\.
- \[6\]H\. Raiffa and R\. Schlaifer\(1961\)Applied statistical decision theory\.Harvard Business School\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p2.1)\.
- \[7\]K\. Chaloner and I\. Verdinelli\(1995\)Bayesian experimental design: a review\.Statist\. Sci\.10\(3\),pp\. 273–304\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p2.1),[§I](https://arxiv.org/html/2609.05582#S1.p3.1)\.
- \[8\]C\. Finn, P\. Abbeel, and S\. Levine\(2017\)Model\-agnostic meta\-learning for fast adaptation of deep networks\.InProc\. Int\. Conf\. Mach\. Learn\. \(ICML\),pp\. 1126–1135\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p2.1)\.
- \[9\]A\. Gelman, J\. B\. Carlin, H\. S\. Stern, D\. B\. Dunson, A\. Vehtari, and D\. B\. Rubin\(2013\)Bayesian data analysis\.3rd edition,CRC Press\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p3.1)\.
- \[10\]J\. Piironen and A\. Vehtari\(2017\)Sparsity information and regularization in the horseshoe and other shrinkage priors\.Electron\. J\. Stat\.11\(2\),pp\. 5018–5051\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p3.1),[§II\-C](https://arxiv.org/html/2609.05582#S2.SS3.p1.2)\.
- \[11\]M\. D\. Hoffman and A\. Gelman\(2014\)The no\-u\-turn sampler: adaptively setting path lengths in hamiltonian monte carlo\.J\. Mach\. Learn\. Res\.15,pp\. 1593–1623\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p3.1),[§II\-B](https://arxiv.org/html/2609.05582#S2.SS2.p1.1)\.
- \[12\]M\. Betancourt\(2017\)A conceptual introduction to Hamiltonian Monte Carlo\.arXiv preprint arXiv:1701\.02434\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p3.1),[§II\-B](https://arxiv.org/html/2609.05582#S2.SS2.p1.1)\.
- \[13\]B\. Carpenter, A\. Gelman, M\. D\. Hoffman, D\. Lee, B\. Goodrich, M\. Betancourt, M\. Brubaker, J\. Guo, P\. Li, and A\. Riddell\(2017\)Stan: a probabilistic programming language\.J\. Stat\. Softw\.76\(1\),pp\. 1–32\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p3.1),[§II\-B](https://arxiv.org/html/2609.05582#S2.SS2.p1.1)\.
- \[14\]B\. Lakshminarayanan, A\. Pritzel, and C\. Blundell\(2017\)Simple and scalable predictive uncertainty estimation using deep ensembles\.InProc\. Adv\. Neural Inf\. Process\. Syst\. \(NeurIPS\),Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p3.1),[§II\-B](https://arxiv.org/html/2609.05582#S2.SS2.p1.1)\.
- \[15\]E\. Daxbergeret al\.\(2021\)Laplace redux: effortless Bayesian deep learning\.InProc\. Adv\. Neural Inf\. Process\. Syst\. \(NeurIPS\),Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p3.1),[§II\-B](https://arxiv.org/html/2609.05582#S2.SS2.p1.1)\.
- \[16\]S\. E\. Chick, M\. Forster, and P\. Pertile\(2017\)A Bayesian decision theoretic model of sequential experimentation with delayed response\.J\. R\. Stat\. Soc\. Ser\. B79\(5\),pp\. 1439–1462\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p3.1)\.
- \[17\]B\. Settles\(2009\)Active learning literature survey\.Technical reportTechnical Report1648,Univ\. of Wisconsin–Madison\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p3.1)\.
- \[18\]J\. C\. Felli and G\. B\. Hazen\(1998\)Sensitivity analysis and the expected value of perfect information\.Med\. Decis\. Making18\(1\),pp\. 95–109\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p3.1)\.
- \[19\]A\. E\. Ades, G\. Lu, and K\. Claxton\(2004\)Expected value of sample information calculations in medical decision modeling\.Med\. Decis\. Making24\(2\),pp\. 207–227\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p3.1)\.
- \[20\]J\. K\. Kruschke and T\. M\. Liddell\(2018\)The Bayesian new statistics: hypothesis testing, estimation, meta\-analysis, and power analysis from a Bayesian perspective\.Psychon\. Bull\. Rev\.25\(1\),pp\. 178–206\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p3.1),[§II\-F](https://arxiv.org/html/2609.05582#S2.SS6.p1.1)\.
- \[21\]A\. Vehtari, A\. Gelman, D\. Simpson, B\. Carpenter, and P\.\-C\. Bürkner\(2021\)Rank\-normalization, folding, and localization: an improvedR^\\widehat\{R\}for assessing convergence of MCMC\.Bayesian Anal\.16\(2\),pp\. 667–718\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p3.1),[§II\-D](https://arxiv.org/html/2609.05582#S2.SS4.p1.2)\.
- \[22\]S\. Talts, M\. Betancourt, D\. Simpson, A\. Vehtari, and A\. Gelman\(2018\)Validating Bayesian inference algorithms with simulation\-based calibration\.arXiv preprint arXiv:1804\.06788\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p3.1)\.
- \[23\]N\. Houlsby, A\. Giurgiu, S\. Jastrzebski, B\. Morrone, Q\. de Laroussilhe, A\. Gesmundo, M\. Attariyan, and S\. Gelly\(2019\)Parameter\-efficient transfer learning for NLP\.InProc\. Int\. Conf\. Mach\. Learn\. \(ICML\),pp\. 2790–2799\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p3.1),[§II\-D](https://arxiv.org/html/2609.05582#S2.SS4.p1.1)\.
- \[24\]S\. J\. Pan and Q\. Yang\(2010\)A survey on transfer learning\.IEEE Trans\. Knowl\. Data Eng\.22\(10\),pp\. 1345–1359\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p3.1)\.
- \[25\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, L\. Kaiser, and I\. Polosukhin\(2017\)Attention is all you need\.InProc\. Adv\. Neural Inf\. Process\. Syst\. \(NeurIPS\),Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p3.1),[§II\-B](https://arxiv.org/html/2609.05582#S2.SS2.p1.1)\.
- \[26\]K\. Cho, B\. van Merriënboer, C\. Gulcehre, D\. Bahdanau, F\. Bougares, H\. Schwenk, and Y\. Bengio\(2014\)Learning phrase representations using RNN encoder\-decoder for statistical machine translation\.InProc\. Conf\. Empirical Methods in Natural Language Processing \(EMNLP\),pp\. 1724–1734\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p3.1),[§II\-B](https://arxiv.org/html/2609.05582#S2.SS2.p1.1)\.
- \[27\]I\. Nahum\-Shani, S\. N\. Smith, B\. J\. Spring, L\. M\. Collins, K\. Witkiewitz, A\. Tewari, and S\. A\. Murphy\(2018\)Just\-in\-time adaptive interventions \(JITAIs\) in mobile health: key components and design principles for ongoing health behavior support\.Ann\. Behav\. Med\.52\(6\),pp\. 446–462\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p4.1)\.
- \[28\]A\. Tewari and S\. A\. Murphy\(2017\)From ads to interventions: contextual bandits in mobile health\.InMobile Health: Sensors, Analytic Methods, and Applications,J\. M\. Rehg, S\. A\. Murphy, and S\. Kumar \(Eds\.\),pp\. 495–517\.Cited by:[§I](https://arxiv.org/html/2609.05582#S1.p4.1)\.
- \[29\]D\. Lewandowski, D\. Kurowicka, and H\. Joe\(2009\)Generating random correlation matrices based on vines and extended onion method\.J\. Multivariate Anal\.100\(9\),pp\. 1989–2001\.Cited by:[§II\-C](https://arxiv.org/html/2609.05582#S2.SS3.p1.2)\.
- \[30\]T\. Gneiting and A\. E\. Raftery\(2007\)Strictly proper scoring rules, prediction, and estimation\.J\. Amer\. Statist\. Assoc\.102\(477\),pp\. 359–378\.Cited by:[§II\-F](https://arxiv.org/html/2609.05582#S2.SS6.p1.1)\.
- \[31\]B\. Efron and R\. J\. Tibshirani\(1993\)An introduction to the bootstrap\.Chapman & Hall\.Cited by:[§II\-F](https://arxiv.org/html/2609.05582#S2.SS6.p1.1)\.

## Supplementary Material for: HB\-PVI: A Hierarchical Bayesian Personalization and Value\-of\-Information Framework for Complex Activity Recognition

This supplement provides tables and methodological detail that support the main manuscript\. Section numbers, table numbers, and figure numbers in this document are prefixed with “S” and are distinct from those in the main text\. Every table below is accompanied by a short statement of what it shows beyond the corresponding main\-text result\.

## S1Sequential participant\-effect updater

The main text summarizes the adaptive tempered resample\-move sequential Monte Carlo \(SMC\) algorithm used to update each participant’s effect posterior as adaptation labels are purchased\. This section gives its full specification, the priors used in Model I and Model II, and the validation gates satisfied on all 47 outer folds\.

### S1\.1Population posterior used for new\-participant updating

For each outer fold, the Model I hierarchical prediction model was estimated using four NUTS chains, 2,000 warmup iterations per chain, and 4,000 retained draws per chain, yielding 16,000 population\-posterior draws containing the activity intercepts, representation coefficients, participant\-effect scales, and Cholesky factor of the participant\-effect correlation matrix\. The outer and validation participants were excluded from population\-model fitting\.

### S1\.2Particle initialization and target distribution

Let𝚯\(s\)\\bm\{\\Theta\}^\{\(s\)\}denote retained population\-posterior draws=1,…,16,000s=1,\\ldots,16\{,\}000\. For each𝚯\(s\)\\bm\{\\Theta\}^\{\(s\)\}, four conditionally independent new\-participant activity\-effect particles are initialized,

𝐳p⋆\(s,m\)∼𝒩9​\(𝟎,𝐈\),𝐛p⋆\(s,m\)=diag⁡\(𝝉b\(s\)\)​𝐋Ωb\(s\)​𝐳p⋆\(s,m\),\\mathbf\{z\}\_\{p^\{\\star\}\}^\{\(s,m\)\}\\sim\\mathcal\{N\}\_\{9\}\(\\mathbf\{0\},\\mathbf\{I\}\),\\qquad\\mathbf\{b\}\_\{p^\{\\star\}\}^\{\(s,m\)\}=\\operatorname\{diag\}\(\\bm\{\\tau\}\_\{b\}^\{\(s\)\}\)\\mathbf\{L\}\_\{\\Omega\_\{b\}\}^\{\(s\)\}\\mathbf\{z\}\_\{p^\{\\star\}\}^\{\(s,m\)\},\(1\)form=1,…,4m=1,\\ldots,4, so every fold begins with 64,000 joint particles retaining 16,000 distinct population\-posterior ancestors\. At staten∈\{0,…,10\}n\\in\\\{0,\\ldots,10\\\}the target distribution is proportional to the population posterior, the conditional Gaussian prior for the new\-participant effect, and the likelihood of exactly the firstnnchronological purchased adaptation windows; no future adaptation label is used at statenn\.

### S1\.3Adaptive tempered resample\-move algorithm

For each newly purchased window, the algorithm uses adaptive likelihood tempering: if the complete incremental likelihood would reduce particle effective sample size \(ESS\) below 25% of 64,000, bisection selects an intermediate likelihood temperature\. Systematic resampling is performed only at the 25% ESS threshold\. Following resampling, random\-walk Metropolis\-Hastings moves rejuvenate the nine\-dimensional new\-participant activity effect, with the target including the conditional Gaussian prior and the cumulative adaptation likelihood through the current shot; incremental likelihoods are recomputed after every rejuvenation\. Population\-level posterior parameters are never rejuvenated; consequently every state was required to retain at least 400 distinct population\-draw ancestors in addition to a final particle ESS of at least 400\. A violation of either gate places the fold on hold and excludes the state from downstream utility modeling\. For every state, the algorithm stores posterior means, standard deviations, and 95% credible intervals for the nine participant effects; posterior predictive activity\-probability means and standard deviations; binary predictive entropy; shot\-to\-shot changes; particle ESS; maximum normalized particle weight; tempering and resampling counts; Metropolis\-Hastings acceptance; and the number of distinct population\-posterior ancestors\.

### S1\.4Prior specifications for Model I and Model II

Table[S1](https://arxiv.org/html/2609.05582#as1_S1.T1)gives the complete prior specification for both models, matching the fitted specification described in the main text\.

Table S1:Prior specifications for Model I and Model II- 𝒩\+\\mathcal\{N\}^\{\+\}denotes a half\-normal distribution \(support≥0\\geq 0\)\.kkindexes activities \(Model I\) or actions \(Model II, denotedaa\);qqindexes representation dimensions\.

Model I’s representation coefficients𝜷k\\bm\{\\beta\}\_\{k\}use a regularized horseshoe because theqq\-dimensional representation is high\-dimensional relative to the per\-fold training data, and sparsity\-inducing shrinkage helps prevent overfitting the population model to any one representation dimension\. Model II’s 19\-predictor coefficients𝜸a\\bm\{\\gamma\}\_\{a\}use a much simpler fixed\-scale Gaussian prior with no shrinkage: with 1,880 held\-out predictions across 47 folds and only 19 predictors per action, this ratio does not require sparsity\-inducing regularization, and a weakly informative fixed\-scale prior was judged sufficient\.

### S1\.5Validation gates

This configuration \(four participant\-effect particles per population draw, 25% ESS resampling threshold, tempered likelihood incorporation\) was applied uniformly to all 47 outer folds\. Table[S2](https://arxiv.org/html/2609.05582#as1_S1.T2)summarizes the validation gates and the result on the full 47\-fold run\.

Table S2:Validation gates for the sequential updater, evaluated on all 47 outer foldsAll gates passed on every fold, with no fold requiring exclusion or a fold\-specific configuration override\. The sequential\-updater approximations \(bounded likelihood tempering, a fixed 25% resampling threshold\) therefore did not measurably compromise posterior\-ancestry diversity anywhere in the cohort, so the Model II gain\-model inputs derived from these states \(Section[S2](https://arxiv.org/html/2609.05582#as1_S2)\) are not confounded by degenerate particle sets for any participant\.

### S1\.6Enriched sequential\-state dataset

The state representation used by the Model II gain model combines the information available at each shot\. The all\-fold dataset contains 517 unique states \(47 folds×\\times11 shots\)\. Representation shift is calculated from the firstnnouter\-adaptationqq\-dimensional principal\-component\-analysis \(PCA\) representations relative to the nested\-training reference for that fold; at shot 0, representation\-shift values are structural zeros with an availability indicator of zero\. Calibration labels, outer\-test arrays, and utility outcomes are not joined into this state\. Adaptation loss is excluded from the primary state because no complete, consistent, leakage\-safe action\-by\-shot definition was identified for all mechanisms and shots\.

## S2Sequential\-state predictor dictionary

Table[S3](https://arxiv.org/html/2609.05582#as1_S2.T3)lists the 19 predictors used in the Model II hierarchical gain model, together with their information domain\.

Table S3:Sequential\-state predictor dictionary used in Model II- All predictors are standardized using training\-fold\-only moments within each outer split\. Three additional columns \(population sample F1, action sample F1, gain\) are retained as outcomes, not predictors\.

These 19 predictors span five distinct information domains rather than shot count alone, which is what allows the Model II gain model \(main text, Section II\-D\) to condition personalization\-benefit predictions on each participant’s evolving uncertainty, calibration, and representation\-shift trajectory\. As Table[S1](https://arxiv.org/html/2609.05582#as1_S1.T1)shows, Model II’s coefficients on these predictors use a fixed\-scale, weakly informative Gaussian prior rather than a sparsity\-inducing one; with 19 predictors and 1,880 held\-out predictions, the predictor\-to\-observation ratio is modest enough that this was judged sufficient\.

## S3EVSI cost specification

Table[S4](https://arxiv.org/html/2609.05582#as1_S3.T4)reproduces the cost specification used throughout the sensitivity analysis, fixed before any EVSI results were inspected\.

Table S4:EVSI cost specificationThis grid is deliberately wide relative to the primary setting so that Table V and Fig\. 4 of the main text can show*how far*the population\-first recommendation extends before it is reversed, rather than reporting robustness at a single arbitrary cost\. The primary values were fixed before any EVSI result was computed, so the primary\-setting conclusion in the main text is not the result of selecting a favorable cost combination after the fact\.

## S4Full ten\-shot action\-specific outcomes, including mean F1

Table[S5](https://arxiv.org/html/2609.05582#as1_S4.T5)extends Table I of the main text to all ten purchased\-label counts for all four personalization mechanisms, reporting the realized mean sample\-level F1 \(with its 95% interval\) alongside gain relative to the population baseline \(F1=0\.9042=0\.9042\)\.

Table S5:Action\-specific mean F1 and gain by shot \(all 10 shots; percentages of 47 participants\)- Population\-only reference: mean F1=0\.9042=0\.9042\(constant across shots and mechanisms by construction, since gain is defined relative to this fixed baseline\)\.

Read alongside Table I of the main text, this table shows that adapter’s advantage over the other three mechanisms is present, if small, at every one of the ten shots, while adapter\-plus\-head’s disadvantage widens steadily from shot 1 through roughly shot 8 before leveling off\. None of the four mechanisms’ mean F1 ever separates from the 0\.9042 population baseline by more than about 1\.5 percentage points in either direction at the group level – the large individual\-level swings reported in Section[S5](https://arxiv.org/html/2609.05582#as1_S5)occur despite, not because of, the group\-level trajectories shown here\.

## S5Participant\-level heterogeneity in realized adapter gain and exact F1

Table[S6](https://arxiv.org/html/2609.05582#as1_S5.T6)reports the full distribution of realized adapter gain across the 47 participants at 1, 6, and 10 purchased labels\.

Table S6:Distribution of realized adapter gain across 47 participants, by shot count- Values are realized gain \(action F1 minus population F1\) for the fixed\-adapter comparator at each shot count, at the primary cost and threshold setting \(Δmin=0\.01\\Delta\_\{\\min\}=0\.01,cL=0\.001c\_\{L\}=0\.001,cC=0\.0005c\_\{C\}=0\.0005,cR=0\.01c\_\{R\}=0\.01\);n=47n=47at each shot count\.

Table[S7](https://arxiv.org/html/2609.05582#as1_S5.T7)reports exact participant\-specific scores taken directly from the source population and action F1 columns; no participant score is reconstructed from the cohort mean\. The population baseline is participant\-specific, and the identityF​1p,a,n−F​1p,population=gp,a,nF1\_\{p,a,n\}\-F1\_\{p,\\mathrm\{population\}\}=g\_\{p,a,n\}was verified for every action\-shot row\.

Table S7:Exact participant\-specific population and adapter F1 scores, all 47 participantsFoldPopulation F1Adapter F1@1Adapter F1@6Adapter F1@10Gain@100000\.56640\.56060\.56210\.5777\+0\.01130010\.74620\.75050\.74710\.7482\+0\.00200020\.68040\.68110\.67720\.6699\-0\.01050030\.78830\.78870\.79020\.7885\+0\.00020040\.79060\.76870\.76690\.7641\-0\.02650050\.92030\.91580\.91830\.9203\+0\.00000060\.93070\.92390\.92830\.9288\-0\.00180070\.97460\.96870\.96790\.9716\-0\.00300080\.90790\.91020\.90430\.9013\-0\.00660090\.94960\.94650\.94500\.9468\-0\.00290100\.95680\.95990\.95900\.9594\+0\.00260110\.94540\.94490\.94780\.9480\+0\.00260120\.92070\.91930\.92090\.9209\+0\.00010130\.91630\.91950\.92110\.9211\+0\.00480140\.88480\.90560\.90410\.9011\+0\.01630150\.83840\.84680\.84570\.8486\+0\.01030160\.88490\.88300\.88590\.8833\-0\.00160170\.95360\.93260\.92870\.9476\-0\.00600180\.97760\.97880\.97760\.9779\+0\.00030190\.96170\.95740\.95900\.9620\+0\.00030200\.87490\.88800\.88470\.8849\+0\.01000210\.95510\.95660\.95840\.9584\+0\.00330220\.93860\.94030\.94070\.9421\+0\.00350230\.91920\.92670\.92450\.9264\+0\.00720240\.94750\.95070\.95270\.9521\+0\.00460250\.95610\.95910\.95930\.9582\+0\.00220260\.94790\.94750\.94880\.9495\+0\.00150270\.85350\.86520\.86810\.8703\+0\.01680280\.96340\.96550\.96050\.9637\+0\.00030290\.93250\.93990\.93720\.9359\+0\.00330300\.89800\.90240\.90350\.9035\+0\.00540310\.85520\.85600\.85390\.8538\-0\.00140320\.96040\.96020\.95990\.9602\-0\.00020330\.95380\.95720\.95640\.9547\+0\.00100340\.90020\.90260\.89860\.9054\+0\.00520350\.91880\.91220\.91770\.9179\-0\.00090360\.98610\.98580\.98310\.9831\-0\.00300370\.95820\.95520\.96310\.9631\+0\.00500380\.95560\.94910\.95080\.9479\-0\.00770390\.96380\.96650\.96550\.9642\+0\.00040400\.95690\.95740\.96070\.9626\+0\.00570410\.89510\.89480\.89790\.9009\+0\.00580420\.96420\.96460\.96460\.9646\+0\.00040430\.96110\.96690\.96340\.9642\+0\.00310440\.95130\.95740\.95850\.9599\+0\.00860450\.83040\.84510\.84330\.8394\+0\.00900460\.80520\.80920\.82050\.8170\+0\.0118Fold, anonymized participant identifier \(outer\-fold index\); Population F1, sample\-level F1 with no personalization; Adapter F1@1/@6/@10, sample\-level F1 after adapter personalization at 1, 6, and 10 purchased labels; Gain@10, Adapter F1@10 minus Population F1\.The exact participant\-level results show that the small mean adapter gain at 10 shots \(\+0\.00198\+0\.00198\) masks substantial heterogeneity\. Five participants had gains above\+0\.01\+0\.01, two had gains below−0\.01\-0\.01, and the remaining participants were distributed around smaller positive and negative changes\. Because baseline F1 varied substantially across participants \(0\.566 to 0\.986\), exact participant\-specific scores are more informative than adding gains to the cohort mean, and are the source of the participant\-level statements and the population\-vs\-personalized F1 figure in the main text\.

## S6Held\-out gain\-model calibration diagnostics

Table III of the main text summarizes held\-out prediction error and calibration in numeric form\. Fig\.[S1](https://arxiv.org/html/2609.05582#as1_S6.F1)shows the underlying reliability curves and per\-action breakdown, which reveal*where*each mechanism’s miscalibration concentrates rather than only its overall magnitude\.

![Refer to caption](https://arxiv.org/html/2609.05582v1/figure2_gain_model_calibration_revised.png)Figure S1:Held\-out gain\-model prediction and calibration\. \(A\) Prediction error \(MAE, RMSE\) by action\. \(B–C, E\) Reliability curves comparing predicted and observed frequencies for positive gain, gain\>0\.01\>0\.01, and negative gain \(five probability bins per action\-event combination; bins with<10<10held\-out observations omitted\)\. \(D\) Empirical coverage of the nominal 95% predictive interval by action, with the 95% reference line\.Two patterns stand out beyond the aggregate numbers in Table III\. First, in panel \(C\), adapter’s meaningful\-benefit reliability curve stays well below the diagonal across almost its entire range, meaning the model consistently assigns a higher probability of meaningful benefit than is actually observed for this mechanism, even though adapter has the lowest overall prediction error of the four actions\. Second, and more consequential, panel \(E\) shows that adapter\-plus\-head’s harm\-probability curve is not merely miscalibrated but inverts at the extreme: the bin with the*highest*predicted harm probability \(≈0\.88\\approx 0\.88\) has the*lowest*observed harm frequency \(≈0\.13\\approx 0\.13\) of any bin for that action – the opposite of what a well\-calibrated model would show\. This is the clearest single illustration of the harm\-identification result in the main text \(Section III\-H\): even the most confident harm predictions for the most heavily parameterized mechanism cannot be trusted directionally, not only imprecisely\.

## S7Full paired\-comparison posterior diagnostics

Table[S8](https://arxiv.org/html/2609.05582#as1_S7.T8)extends Table IV of the main text with the underlying MCMC diagnostics for the participant\-level Student\-ttpaired\-utility\-difference models \(HB\-PVI minus each comparator, primary setting\)\.

Table S8:Paired\-comparison posterior diagnostics \(HB\-PVI minus comparator\)- aAll 47 paired differences are exactly zero; MCMC not applicable \(handled analytically\)\.
- Div\., number of divergent transitions; Max treedepth, number of iterations hitting the maximum treedepth;R^\\widehat\{R\}, rank\-normalized potential scale reduction factor; ESS, effective sample size\. All non\-degenerate comparisons pass the convergence gates \(R^<1\.01\\widehat\{R\}<1\.01, zero divergences, zero max\-treedepth hits, bulk/tail ESS\>400\>400\)\.

All three non\-degenerate models converged cleanly with no divergences andR^≤1\.0012\\widehat\{R\}\\leq 1\.0012, so the posterior paired\-utility\-difference estimates in Table IV of the main text \(and the correspondingP⁡\(μd\>0\)P\(\\mu\_\{d\}\>0\)values\) are not confounded by sampler pathologies; the 10\-shot comparison had the smallest ESS among the three, but its minimum ESS remained above 3,400, with no divergences and maximumR^<1\.01\\widehat\{R\}<1\.01\.

## S8Decision\-region breakdown by threshold and label cost

Table[S9](https://arxiv.org/html/2609.05582#as1_S8.T9)gives the full 18\-row breakdown \(three practical\-gain thresholds×\\timessix label costs, each aggregating 12 computation\-cost/harm\-penalty combinations\) underlying Fig\. 4A and Table V of the main text\.

Table S9:HB\-PVI utility\-optimality by practical\-gain threshold and label cost- •Each row aggregates 12 computation\-cost/harm\-penalty settings\.

The 100\.0% rows are the state underlying the primary\-setting result in the main text: once label cost reaches 0\.001, HB\-PVI wins in every one of the 12 computation\-cost/harm\-penalty combinations at every practical\-gain threshold tested, not just on average across them\. The only settings in which a fixed\-shot comparator ever wins are at label cost 0 or 0\.0005, an order of magnitude below the primary value\. This is consistent with Section[S10](https://arxiv.org/html/2609.05582#as1_S10)’s finding that the largest one\-step EVSI observed anywhere in the study \(0\.01012\) narrowly exceeds only the two smallest label costs in the grid, and that positive EVSI was confined throughout toΔmin=0\\Delta\_\{\\min\}=0and a single participant\. Fig\.[S2](https://arxiv.org/html/2609.05582#as1_S8.F2)resolves this label\-cost margin into its two underlying cost dimensions by also varying the harm penaltycRc\_\{R\}\.

![Refer to caption](https://arxiv.org/html/2609.05582v1/figureS2_decision_region_heatmap.png)Figure S2:HB\-PVI utility\-optimality \(%\) across the full label\-cost \(cLc\_\{L\}\) and harm\-penalty \(cRc\_\{R\}\) grid, averaged over the three computation\-cost \(cCc\_\{C\}\) settings, for each practical\-gain threshold\. HB\-PVI is never preferred only in the single cell where both costs are zero; it is preferred in a minority of settings only when exactly one of the two costs is zero and the other is at its lowest tested nonzero level; everywhere else in the grid, HB\-PVI is optimal in 100% of cases\.Fig\.[S2](https://arxiv.org/html/2609.05582#as1_S8.F2)shows that the reversal region identified in Table[S9](https://arxiv.org/html/2609.05582#as1_S8.T9)is smaller than the label\-cost margin alone suggests: HB\-PVI loses to a fixed\-shot comparator in only three of the 24 cells in each panel \(0%, 66\.7%, and either 66\.7% or 33\.3% depending onΔmin\\Delta\_\{\\min\}, all confined tocL≤0\.0005c\_\{L\}\\leq 0\.0005combined withcR≤0\.005c\_\{R\}\\leq 0\.005\), and is optimal in 100% of settings as soon as either cost moves to its next tested level\. Harm penalty and label cost therefore behave as substitutes in this decision: a moderate increase in either one is enough to restore the population\-first recommendation even while the other remains at its lowest tested value\.

## S9Label\-efficiency Bayesian and bootstrap diagnostics

Table[S10](https://arxiv.org/html/2609.05582#as1_S9.T10)reports the posterior degrees\-of\-freedom and scale parameters and convergence diagnostics for the Bayesian label\-efficiency model underlying Section III\-G of the main text\.

Table S10:Label\-efficiency posterior diagnostics \(Bayesian Student\-ttmodel\)- Post\.τ\\tauand post\.ν\\nuare the posterior\-mean scale and degrees\-of\-freedom of the participant\-level Student\-ttF1\-loss model; other columns as in Table[S8](https://arxiv.org/html/2609.05582#as1_S7.T8)\.

Convergence was clean for all three reference comparisons \(zero divergences,R^≤1\.0011\\widehat\{R\}\\leq 1\.0011\), so the posterior probabilities reported in the main text \(e\.g\.,P⁡\(μloss<0\.005∣D\)=0\.9992P\(\\mu\_\{\\mathrm\{loss\}\}<0\.005\\mid D\)=0\.9992against the 10\-shot reference\) rest on well\-mixed chains rather than an under\-explored posterior\. The low degrees\-of\-freedom estimates \(ν≈3\.7\\nu\\approx 3\.7–5\.85\.8\) indicate a genuinely heavy\-tailed distribution of participant\-level F1 losses, consistent with the extreme individual cases identified in Table[S7](https://arxiv.org/html/2609.05582#as1_S5.T7), and justify the use of a Student\-ttrather than Gaussian likelihood for this comparison\.

## S10Stopping\-decision and action\-selection frequencies

Across the complete 47\-fold×\\times216\-setting stopping\-decision dataset \(10,152 fold\-setting decisions\), 10,116 \(99\.6%\) stopped because one\-step EVSI did not exceed the label cost, and the remaining 36 \(0\.4%\) stopped only because the ten\-label horizon was reached \(all occurring atΔmin=0\\Delta\_\{\\min\}=0withcL=0c\_\{L\}=0\)\. Population inference was selected in 10,080 of 10,152 decisions \(99\.3%\); head personalization was selected in the remaining 72 decisions \(0\.7%\), all atΔmin=0\\Delta\_\{\\min\}=0\. Adapter, adapter\-plus\-head, and prototype\-residual personalization were never selected as the terminal action under any cost\-threshold setting, consistent with the primary\-setting result reported in the main text\.

Table[S11](https://arxiv.org/html/2609.05582#as1_S10.T11)details the 72 decisions with strictly positive one\-step EVSI referenced in Section III\-E of the main text, grouped by label cost \(each group aggregates 12 computation\-cost/harm\-penalty combinations, all atΔmin=0\\Delta\_\{\\min\}=0and all attributable to the same single participant\)\.

Table S11:Positive one\-step EVSI decisions, grouped by label cost \(all atΔmin=0\\Delta\_\{\\min\}=0;n=12n=12decisions per row\)The same one participant accounts for every positive\-EVSI decision in the entire 10,152\-decision grid\. Their behavior under cost changes is itself informative: at low label cost the model judges it worthwhile to keep purchasing all the way to the ten\-label horizon \(mean EVSI per continued step, 0\.0041\), but once label cost exceeds about 0\.002 the model buys exactly one label and stops, because EVSI drops sharply after the first purchase\. This is the clearest illustration in the study of EVSI behaving as intended – responding to both the participant’s own trajectory and the price of information – even though, in aggregate, this single case is not enough to change the population\-first recommendation for the cohort as a whole\.

## S11Study objectives and outcomes

The study’s five objectives were specified before any Model I/II fitting or EVSI computation\. Table[S12](https://arxiv.org/html/2609.05582#as1_S11.T12)summarizes each objective and its outcome, drawing on results reported throughout the main text and this supplement\.

Table S12:Study objectives and outcomes- ROPE, region of practical equivalence \(main text, Section II\-F\)\.

Objective 1 was not met because HB\-PVI was exactly equivalent to, rather than better than, always\-stop under the primary cost setting \(main text, Section III\-F\) – a finding that itself supports objective 4 \(practical equivalence, met exactly with ROPE probability 1\)\. Objective 2 was met with high posterior confidence by both a Bayesian and an independent bootstrap analysis \(Section[S9](https://arxiv.org/html/2609.05582#as1_S9)above; main text, Section III\-G\)\. Objective 3 was not established: although the primary decision rule avoided realized harm by selecting population inference for everyone, the underlying harm\-probability predictions did not discriminate risk better than a trivial baseline \(main text, Sections III\-C and III\-H\)\. Objective 5 was supported by the substantial participant\-level heterogeneity documented in Tables[S6](https://arxiv.org/html/2609.05582#as1_S5.T6)and[S7](https://arxiv.org/html/2609.05582#as1_S5.T7)above, which shows that a small subset of participants would benefit or be harmed considerably more than the population\-averaged effect suggests, even though this heterogeneity could not be predicted early enough to justify acting on it under the specified costs\.

Similar Articles

Hierarchical Advantage Weighting for Online RL Fine-Tuning of VLAs from Sparse Episode Outcomes

Hugging Face Daily Papers

This paper proposes Hierarchical Advantage-Weighted Behavior Cloning (HABC) for fine-tuning Vision-Language-Action (VLA) policies using online reinforcement learning with sparse binary episode outcomes. HABC separates viability and efficiency objectives via adaptive critic heads and intervention-aware credit assignment, significantly improving success rates on contact-rich bimanual manipulation tasks.

PAPA: Online Personalized Active Preference Alignment

arXiv cs.LG

This paper introduces Personalized Active Preference Alignment (PAPA), a method for fine-tuning diffusion models using real-time user feedback without a parameterized reward model, enhancing efficiency in personalized tasks like recommendations and image generation.