EUVIMEDEuropean Health Evidence
Uhr 10/10Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

Lung cancer risk prediction using interpretable ensemble models on lifestyle and clinical data

Shahid Mohammad Ganie, Pijush Kanti Dutta Pramanik, Zhongming Zhao

PLOS One · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Lung cancer remains one of the leading causes of cancer-related mortality worldwide, where early detection and reliable risk assessment are critical for improving outcomes. This study develops and evaluates a range of ensemble learning models for lung cancer prediction using lifestyle and clinical indicators, with an emphasis on both predictive performance and interpretability. Five base learners—logistic regression, k-nearest neighbors, naïve Bayes, support vector machine, and linear discriminant analysis—were used to construct multiple boosting and bagging models. Building on these, voting and stacking ensembles were designed by selectively combining high-performing models. All approaches were evaluated on the original dataset as well as on balanced and upsampled variants derived through synthetic augmentation. Model performance was assessed using accuracy, precision, recall, F1-score, Matthews correlation coefficient (MCC), and AUC-ROC. The results show that ensemble approaches consistently outperform individual models, with voting and stacking demonstrating superior performance over boosting and bagging methods. The stacking model achieved the strongest overall performance across all evaluated models. On the original dataset, which provides a more realistic representation of practical deployment conditions, it attained an accuracy of 93.53%. Performance further improved on the balanced and upsampled datasets on the upsampled dataset. To enhance transparency, SHAP and LIME were employed to provide global and local interpretability, respectively, enabling identification of key clinical factors and patient-specific risk drivers. The analysis highlights both alignment with known clinical patterns and dataset-driven variations, supporting informed interpretation of model outputs. The results suggest that stacking-based ensembles can improve risk prediction from lifestyle and clinical indicators while maintaining model transparency through SHAP and LIME explanations. These findings highlight the potential of interpretable ensemble learning as a decision-support tool for lung cancer risk assessment.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Shahid Mohammad Ganie, Pijush Kanti Dutta Pramanik, Zhongming Zhao
Quelle
PLOS One
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
1932-6203
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Shahid Mohammad Ganie, Pijush Kanti Dutta Pramanik, Zhongming Zhao (2026). Lung cancer risk prediction using interpretable ensemble models on lifestyle and clinical data. PLOS One. https://doi.org/10.1371/journal.pone.0357291
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1