Vollständiger Abstract
Worum geht es in dieser Arbeit?
Background/Objectives: Machine learning is increasingly applied to identify hypertension from broad predictor panels, but whether it outperforms a few routine measurements is unclear. We compared machine learning pipelines against three readily obtained variables in a national Chilean survey. Methods: We analysed 5516 participants aged 15 years and older from the Chilean National Health Survey 2016–2017 with two valid blood-pressure readings. Hypertension was defined as mean systolic pressure ≥ 140 mmHg or mean diastolic pressure ≥ 90 mmHg or current antihypertensive treatment, the mean taken over the second and third readings. A total of 43 candidate predictors plus 11 laboratory missingness indicators formed 54 model inputs, excluding all blood pressure and hypertension-related variables. Data were partitioned once by census segment. Seventeen algorithms were compared under segment-grouped cross-validation with preprocessing fitted within folds; twelve were tuned identically, nested specifications independently, thresholds locked on development data and intervals derived from segment bootstrap. Results: Prevalence was 38.3% unweighted and 29.2% design-weighted. Across seven metrics, the twelve tuned pipelines were closely comparable, and none ranked first on all. On held-out data, the full XGBoost model reached an area under the receiver operating characteristic curve (AUC) of 0.890 (95% confidence interval [CI] 0.872 to 0.906) and the full penalised logistic model 0.882 (0.864 to 0.901). A model containing age, sex and waist-to-height ratio alone reached 0.882 (0.864 to 0.900) under XGBoost and 0.881 (0.861 to 0.898) under penalised logistic regression, differing from the corresponding full model of the nested comparison, which was tuned independently and reached 0.881 under penalised logistic regression and 0.891 under XGBoost, by −0.0004 (−0.0087 to +0.0081) under penalised logistic regression and −0.0084 (−0.0148 to −0.0024) under XGBoost, computed on unrounded values. Removing age cost 0.027 AUC. Findings held under design weighting, in the laboratory subsample and among untreated participants. Conclusions: Three routine measurements provided most of the discrimination achieved by fifty-four inputs; the remainder added no detectable increment under a penalised logistic specification and approximately 0.008 AUC under XGBoost. Equivalence was not formally tested, validation was internal, and transportability is undemonstrated.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Rodrigo Yáñez-Sepúlveda, Boryi A. Becerra-Patiño, Felipe Montalva-Valenzuela, Rodrigo Olivares, Alejandra Uribe-Díaz, Eduardo Guzmán-Muñoz, Yeny Concha-Cisternas, Daniel Rojas-Valverde, José Francisco Tornero-Aguilera, Vicente Javier Clemente-Suárez, José Francisco López-Gil
- Quelle
- Diagnostics
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2075-4418
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Rodrigo Yáñez-Sepúlveda, Boryi A. Becerra-Patiño, Felipe Montalva-Valenzuela, Rodrigo Olivares, Alejandra Uribe-Díaz, Eduardo Guzmán-Muñoz, Yeny Concha-Cisternas, Daniel Rojas-Valverde, José Francisco Tornero-Aguilera, Vicente Javier Clemente-Suárez, José Francisco López-Gil (2026). Age, Sex, and Waist-to-Height Ratio Approach the Discrimination of Fifty-Four Model Inputs for Prevalent Hypertension: An Explainable Machine Learning Analysis of the Chilean National Health Survey. Diagnostics. https://doi.org/10.3390/diagnostics16172843
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1