Tag
This paper finds systematic differences in outputs of large language models for health advice based on access modes like APIs and chatbot interfaces, undermining evaluation validity. It calls for model providers to enable faithful replication of consumer experiences for rigorous auditing.
Google highlights AI advancements to accelerate science and improve lives, including supporting 300 languages, mapping genetic changes, and enhancing weather prediction models.
WearableQA is a benchmark dataset of 4,084 multiple-choice questions for health reasoning over real-world wearable data, designed to evaluate AI models on longitudinal health data analysis.
WearableQA is a benchmark for evaluating large language models' reasoning over real-world wearable health data, using multiple-choice questions derived from longitudinal measurements.
This research paper explores the use of 24-hour wrist movement data to study and predict human health conditions and diseases, involving interdisciplinary collaborations across multiple universities.
The paper introduces a knowledge-guided agentic framework that identifies missing patient context in health queries and asks targeted follow-up questions to improve the accuracy and consistency of downstream language model responses.
Google Research introduces PhotoScan, a deep learning framework that estimates body composition metrics from standard smartphone photos to predict insulin resistance and cardiometabolic risk with near-DXA accuracy.
Google Research and Google DeepMind's AMIE medical AI system demonstrates expert-level real-time audio-visual clinical consultation capabilities in a first-of-its-kind randomized study, with clinical evaluators rating it favorably across core competencies.
Introduces a new paradigm called Combodied Agents that unify digital and embodied AI agents to model, predict, and support individual human-state trajectories over time, focusing on sustained human benefit rather than task completion.
This arXiv paper evaluates federated training of tokenized generative event models (GEMs) on ICU EHR data from three health systems, showing that federated learning preserves most centralized performance and improves cross-site transportability compared to conventional supervised models.
HealthClaw is an open-source agent architecture for longitudinal personal health management that uses self-evolving memory to improve support over repeated encounters, achieving higher accuracy and privacy compared to baselines across biomedical tasks.
RubricsTree proposes a scalable, expert-aligned evaluation framework for personal health agents using over 100 atomic Boolean rubrics, achieving up to 66% relative gains on HealthBench across Gemini, GPT, and Qwen model families.
This study evaluates the use of large language models (Gemini 3.0 Flash) with personal health records to answer patient health queries, finding significant improvements in helpfulness, safety, and personalization when PHR context is provided.
This article critiques Mark Kaplan's approach to fine-tuning medical LLMs via his platform healtthruth.ai, highlighting pitfalls in overriding foundational training for healthcare AI.
Perplexity has launched Perplexity Health, a feature that allows users to ask health questions across their medical records, lab results, and wearable device data.
OpenAI introduces ChatGPT Health, a dedicated experience with enhanced privacy and security features that allows users to securely connect medical records and wellness apps to receive more personalized health guidance. The feature addresses the common use case of health queries on ChatGPT (230+ million weekly users) while maintaining strict data isolation and declining to use health conversations for model training.