Tag
The paper proposes a method for monitoring web agents without access to internal model signals, using observable trajectories and key-step supervision to predict failures early. It demonstrates competitive performance with internal-signal baselines across benchmarks.
The paper proposes structured evidence routing, a router–predictor–reviewer workflow for incident risk prediction from multimodal longitudinal electronic health records, achieving competitive performance with interpretable patient-specific evidence trails.
This paper develops a multi-dimensional framework to evaluate discrimination, calibration, interpretability, and algorithmic fairness for machine learning-based type 2 diabetes risk prediction models, revealing significant performance degradation under real-world distribution shifts and biases by age and obesity.
UC Berkeley researchers trained an AI model on hundreds of thousands of EKGs to detect a previously unrecognized signal that predicts sudden cardiac death risk more accurately than current methods, potentially saving thousands of lives annually.
This paper presents LiverRisk, a machine learning framework for NAFLD risk prediction that combines gradient-boosted decision trees with conformal prediction to provide calibrated, distribution-free coverage guarantees on individual risk estimates, achieving high AUROC on internal and external cohorts.
This study evaluates five machine learning classifiers for chronic kidney disease risk prediction, finding that near-perfect internal performance fails under distribution shift. It emphasizes the need for calibration stability and conformal coverage transfer before clinical deployment.
This paper introduces yvsoucom-iterkit, a deterministic, log-driven AutoML framework for reproducible pipeline optimization in healthcare risk prediction, evaluated on diabetes and stroke datasets with over 18,000 pipeline configurations, achieving strong performance and revealing structured search spaces with component redundancy.