Vollständiger Abstract
Worum geht es in dieser Arbeit?
Cancer incidence exhibits substantial spatial disparities associated with environmental, behavioral, built-environment, healthcare access, and socioeconomic conditions, yet the extent to which these county-level associations vary geographically and across spatial scales remains insufficiently understood. This study evaluates an explainable spatial epidemiology workflow that integrates Random Forest (RF), SHapley Additive exPlanations (SHAP), permutation importance, Ordinary Least Squares (OLS), Geographically Weighted Regression (GWR), and Multiscale Geographically Weighted Regression (MGWR) to examine county-level incidence for all-site cancer as a composite benchmark, colorectal cancer, female breast cancer, and melanoma of the skin across Texas, USA. All outcomes were screened from a common leakage-safe pool of 44 predictors. RF-SHAP/permutation screening was conducted only in spatially separated discovery counties, and the retained predictors were subsequently evaluated using OLS, GWR, and MGWR in independent confirmation counties. The screening procedure retained 19–23 predictors across outcomes, reducing model dimensionality by approximately 48–57%. Although screening substantially reduced AICc, in-sample also decreased, indicating improved parsimony and complexity-adjusted fit rather than improved explanatory performance; same-cardinality random, correlation-based, and LASSO benchmarks further showed that the advantage of RF-based screening was outcome- and model-dependent. In independent confirmation analyses, GWR was preferred by AICc for all-site cancer (AICc = 442.79), whereas OLS was preferred for colorectal cancer (356.48), female breast cancer (367.50), and melanoma (234.56); MGWR was not AICc-preferred for any outcome. However, where estimation was feasible, MGWR provided complementary multiscale information by distinguishing fitted associations characterized by near-global versus more localized spatial bandwidths. Five-fold nested spatial-block cross-validation showed stronger geographic predictive performance for RF, with pooled values of 0.422, 0.235, 0.442, and 0.341 for all-site, colorectal, breast, and melanoma outcomes, respectively, whereas GWR produced negative held-out for all four outcomes. These findings support explainable-ML screening primarily as a transparent dimensionality reduction strategy and demonstrate complementary roles for spatial modeling: AICc evaluates whether additional spatial complexity is justified, MGWR characterizes predictor-specific spatial scales where feasible, and spatial-block validation evaluates geographic predictive generalization. The resulting associations and spatial scales are interpreted as ecological and descriptive rather than causal.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Yuhang Xie, Zhe Zhang, Chanam Lee, Marcia G. Ory, Ipek Nese Sener, Bahar Dadashova, Gisou Salkhi Khasraghi, Jinsil Hwaryoung Seo, Galen Newman, Chunwu Zhu, Wenjin Wang, Xuemei Zhu
- Quelle
- Applied System Innovation
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2571-5577
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Yuhang Xie, Zhe Zhang, Chanam Lee, Marcia G. Ory, Ipek Nese Sener, Bahar Dadashova, Gisou Salkhi Khasraghi, Jinsil Hwaryoung Seo, Galen Newman, Chunwu Zhu, Wenjin Wang, Xuemei Zhu (2026). Spatial Heterogeneity in Cancer Incidence: Assessing Behavioral and Environmental Associations Using Machine Learning and Multiscale Geographically Weighted Regression. Applied System Innovation. https://doi.org/10.3390/asi9090181
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1