Adaptive data selection improves wearable prediction under low baseline performance
Summary
This paper evaluates adaptive data selection strategies for wearable health prediction, finding they significantly improve AUROC for participants with low baseline performance but offer limited gains for strong baselines.
View Cached Full Text
Cached at: 06/02/26, 03:39 PM
# Adaptive data selection improves wearable prediction under low baseline performance Source: [https://arxiv.org/abs/2606.00141](https://arxiv.org/abs/2606.00141) [View PDF](https://arxiv.org/pdf/2606.00141) > Abstract:Adaptive sensing strategies that selectively sample data are increasingly used in wearable health systems to improve prediction performance under limited data budgets, yet their benefits across individuals remain poorly understood\. Here, we evaluate adaptive selection of time windows for model training under fixed measurement budgets across multiple sensing modalities, including heart rate, activity, and ecological momentary assessment \(EMA\), in a longitudinal wearable dataset\. We quantify performance gains relative to random sampling using both area under the receiver operating characteristic curve \(AUROC\) and F1 score\. Adaptive strategies yield substantial improvements in AUROC for participants with low baseline performance \(with gains up to 0\.7\), while offering limited or negative gains for participants with strong baselines\. Across modalities, adaptive gain is strongly inversely correlated with baseline performance \(Pearson r = \-0\.67; Spearman p = \-0\.62\)\. At the participant level, most individuals benefit in AUROC \(60\-80% across modalities\), although improvements in F1 are smaller and less consistent\. These findings show that adaptive sensing is not uniformly beneficial, but instead provides the greatest value in underperforming settings\. Our results support selective deployment strategies that tailor adaptive sensing based on baseline performance to improve efficiency in wearable health monitoring\. ## Submission history From: Ali Kargarandehkordi \[[view email](https://arxiv.org/show-email/56952cd3/2606.00141)\] **\[v1\]**Fri, 29 May 2026 00:10:44 UTC \(799 KB\)
Similar Articles
@rohanpaul_ai: New Google paper shows that wearable data becomes far more useful when AI learns the person behind the signals. It's is…
Google researchers propose SensorFM, a foundation model trained on over 1 trillion minutes of unlabeled wearable data from 5 million people, which learns general physiological patterns and outperforms engineered features on 34 of 35 health prediction tasks.
Aligning Data-Driven Predictors with Allocation: A Decision-Focused Approach to Survival Analysis
This paper introduces a decision-focused learning approach for survival analysis that aligns predictive models with downstream allocation decisions, using NDCG optimization. Applied to US heart transplant data, it improves ranking performance by 50-100%, potentially yielding thousands of additional life-years annually.
Retrieval-Augmented Personalization with Foundation Models for Wearable Stress Detection
This paper introduces a retrieval-augmented personalization method for wearable stress detection using frozen foundation models, achieving near-supervised fine-tuning performance without requiring labeled user data.
When Offline Selectors Cannot Beat the Best Single Model: A Diagnostic Study on edX Dropout Prediction
This paper proposes a three-stage diagnostic framework to identify why offline model selectors fail to beat the best single model, applying it to dropout prediction on edX clickstream data. The study finds that the bottleneck is local representational ambiguity rather than learner choice or distribution shift, recommending state redesign or new data collection over further algorithm tuning.
Supervised Fine-tuning with Synthetic Rationale Data Hurts Real-World Disease Prediction
This paper demonstrates that supervised fine-tuning with synthetic rationale data consistently harms prediction performance for Alzheimer's disease detection compared to label-only fine-tuning, across many configurations and model families. The degradation persists despite high-quality rationales and is attributed to a conflict between narrative plausibility and discriminative optimization.