WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
Summary
WearableQA is a benchmark for evaluating large language models' reasoning over real-world wearable health data, using multiple-choice questions derived from longitudinal measurements.
View Cached Full Text
Cached at: 09/10/26, 06:09 AM
Paper page - WearableQA: A Benchmark for Health Reasoning over Real-World Wearable Data
Source: https://huggingface.co/papers/2609.05405
Abstract
WearableQA is a benchmark of multiple-choice questions derived from real longitudinal wearable data that evaluates large language model reasoning across data and health dimensions.
Recent advances in wearable sensing enable continuous monitoring of physiological and behavioral signals, yet existing benchmarks rarely evaluate whether AI systems can reason over a real user’s longitudinal wearable record. We introduce WearableQA, a benchmark comprising 4,084 10-option multiple-choice questions constructed from the wearable time series, blood biomarkers, and demographics of 200 real users, each with up to 500 days of daily measurements. WearableQA preserves authentic wearable distributions that include device noise and inter-individual variability. To evaluate distinct reasoning capabilities, we introduce 16 question types organized along two complementary axes: data versus health reasoning, which distinguishes computation over longitudinal measurements from physiological interpretation; and single- versuscross-signal reasoning, which separates reasoning about individual signals from the integration of multiple signals. To construct reliable questions at scale, we adopt adual-grounding frameworkthat combines literature-grounded physiological findings with statistically validated population-grounded physiological patterns. This enables the capture of meaningful relationships observed in real-world wearable data. Evaluation of 14 proprietary and open-sourceLLMsdemonstrates that WearableQA effectively differentiates model capabilities, with performance ranging from 19.6% to 72.9% against a 10% chance baseline. Moreover, WearableQA remains far from solved: most models achieve accuracies below 60%. Overall, WearableQA provides a realistic and diagnostic benchmark for evaluating LLM reasoning over real-world wearable data.
View arXiv pageView PDFGitHub6Add to collection
Get this paper in your agent:
hf papers read 2609\.05405
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.05405 in a model README.md to link it from this page.
Datasets citing this paper1
#### facebook/WearableQA Viewer• Updated6 days ago • 20.4k • 10
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.05405 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Meta's WearableQA Health Reasoning Benchmark (GitHub Repo)
WearableQA is a benchmark dataset of 4,084 multiple-choice questions for health reasoning over real-world wearable data, designed to evaluate AI models on longitudinal health data analysis.
Rewarding Reasoning, Not Answers: Fixing and Bounding Test-Time Reinforcement Learning on Medical QA
The paper diagnoses failures in test-time reinforcement learning for medical QA due to answer-space structure and introduces PROSE, which rewards reasoning quality to improve model performance without labeled data.
Introducing HealthBench
OpenAI introduces HealthBench, a new benchmark for evaluating AI systems in healthcare contexts, created with 262 physicians across 60 countries. The benchmark includes 5,000 realistic health conversations with physician-written rubrics to assess model performance on meaningful, trustworthy, and improvable metrics.
Revisiting a Pain in the Neck: A Semantic Reasoning Benchmark for Language Models
Researchers present SemanticQA, a benchmark for evaluating language models on semantic phrase processing tasks including idioms, noun compounds, and verbal constructions, revealing significant performance variation across model architectures and scales on semantic reasoning tasks.
OpenMHC: Accelerating the Science of Wearable Foundation Models
OpenMHC introduces the largest open-access wearable health dataset with over 60 million hours of data and open-source implementations of wearable foundation models, including a unified benchmark for prediction, imputation, and forecasting.