Classifying coronary heart disease risk from NHANES survey data (2011-2018), with a full leakage audit and calibration check [P]
Summary
A project analyzing NHANES survey data to classify coronary heart disease risk, emphasizing data leakage audit and calibration checks while comparing machine learning models.
Similar Articles
Machine learning prediction of obstructive coronary artery disease using opportunistic coronary calcium and epicardial fat assessments from CT calcium scoring scans
This paper presents a machine learning framework using CatBoost and SHAP to predict obstructive coronary artery disease from CT calcium scoring scans, achieving high accuracy by combining calcium-omics and epicardial fat features.
CardioMeta: Calibrated Multi-Task Prediction of Diabetes, Hypertension, and Cardiovascular Disease Across Population and EHR Data
CardioMeta is a calibrated multi-task framework for jointly predicting diabetes, hypertension, and cardiovascular disease across NHANES and MIMIC-IV data, emphasizing leakage control, calibration, and transparent reliability.
Accelerometry-Derived Digital Biomarkers for Cardiometabolic Risk: A Population-Representative Tabular Benchmark with Uncertainty Quantification
This paper introduces the NHANES Accelerometry Cardiometabolic Benchmark, a population-representative tabular dataset for predicting cardiometabolic risk from accelerometry data, and evaluates ridge regression, XGBoost, and TabPFN v2 with uncertainty quantification using conformal prediction.
LLMs for Cardiovascular Risk Prediction from Structured Clinical Data
This paper presents a hybrid framework that combines structured clinical data with LLM-generated narratives for coronary artery disease prediction, achieving high fidelity in variable extraction and comparing ML models with LLM-based zero-shot and few-shot classification.
Transforming Heart Disease Prediction with Advanced Machine Learning Techniques
This research paper compares various machine learning classifiers for heart disease prediction, finding that Support Vector Machine and Simple Cart achieve the best performance on UCI and Kaggle datasets respectively, highlighting ML's potential to aid in early clinical diagnosis.