Tag
This study evaluates nine ECG foundation models for Brugada syndrome detection, finding that pre-training provides optimization stability but not transferable clinical knowledge, challenging assumptions about the benefits of large-scale pre-training for rare diseases.
CLExEval introduces a human-in-the-loop framework for evaluating LLM clinical reasoning under progressive information masking, revealing failure patterns such as verbosity bias, hidden knowledge paradox, and reasoning-to-output mismatch in models like GPT-4o-mini and HuatuoGPT-o1.
Researchers from Boston Children's Hospital, Harvard, and OpenAI used the OpenAI o3 Deep Research reasoning model to reanalyze 376 unsolved rare disease cases, leading to diagnoses in 18 additional cases (4.8% yield) after expert review and clinical confirmation. The study, published in NEJM AI, demonstrates how AI-assisted workflows can help experts revisit difficult cases as scientific knowledge evolves.
Google DeepMind's Co-Scientist is a multi-agent AI system that acts as a virtual team of scientists to search literature, generate hypotheses, and design experiments, compressing months of research into days and already yielding new scientific discoveries.