Vollständiger Abstract
Worum geht es in dieser Arbeit?
Men with de novo metastatic prostate cancer (mPCa) face a poor but heterogeneous prognosis, yet interpretable, transparently reported diagnostic-time tools remain limited. Using the Surveillance, Epidemiology, and End Results (SEER) 17-registry database (2010–2020), we developed and temporally validated machine-learning models to predict fixed 3-year cancer-specific mortality (CSM) and overall mortality (OM) in newly diagnosed mPCa. From 42,691 exported cases, we defined a cohort of 41,057 (survival ≥1 month) and split it temporally into training (2010–2016, n = 21,584) and validation (2017–2020, n = 19,473) sets. Seven algorithms were benchmarked on an identical diagnostic-time feature space [age, year, race, marital status, prostate-specific antigen (PSA), Gleason, and metastatic sites], excluding treatment and American Joint Committee on Cancer (AJCC) T/N; CatBoost was selected as the primary model. In temporal validation, it achieved an area under the curve (AUC) of 0.701 [95% confidence interval (CI) 0.693–0.708] for net CSM, a cause-specific quantity excluding other-cause deaths, not an absolute competing-risks cumulative incidence, and 0.704 (0.697–0.712) for OM, with good calibration (slopes 1.12 and 1.08) and positive decision-curve net benefit for flagging high-risk patients across roughly 30%–70% thresholds; discrimination was comparable across leading algorithms and to logistic regression. Temporal internal–external cross-validation confirmed stable discrimination (pooled AUC 0.701/0.703; 95% prediction intervals from 0.672 to 0.725); because diagnosis year is a model feature, this establishes stability within 2010–2020, and the model should not be applied to later diagnosis years without local recalibration. Full validation-set SHapley Additive exPlanations (SHAP) analysis identified age, Gleason score, PSA, diagnosis year, and metastatic sites as dominant predictors, and training-derived risk tertiles separated validation event rates monotonically (net CSM 28%/49%/67%; OM 34%/57%/74%). Transportability was limited: in an independent non-SEER single-center cohort ( n = 150), discrimination was attenuated (cancer-specific AUC 0.59, 95% CI 0.49–0.69; overall 0.59, 0.50–0.68) and calibration drifted markedly [slope 0.47; observed-to-expected ratio (O:E) 1.23 and 1.19], with only the OM risk ordering preserved as a monotone gradient. This interpretable model therefore provides calibrated diagnostic-time risk stratification within SEER 2010–2020, intended not for treatment selection but for prognostic counseling, closer follow-up, and trial enrichment; local intercept-and-slope recalibration and larger multicenter validation are required before use outside that setting or window.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Xin Wang, Guanglei Yao, Heqian Liu, Wei Ding
- Quelle
- Frontiers in Oncology
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2234-943X
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Xin Wang, Guanglei Yao, Heqian Liu, Wei Ding (2026). Interpretable machine-learning risk stratification at the time of diagnosis for 3-year mortality in de novo metastatic prostate cancer: development and temporal validation in SEER. Frontiers in Oncology. https://doi.org/10.3389/fonc.2026.1904589
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1