Vollständiger Abstract
Worum geht es in dieser Arbeit?
Accurate differentiation of lung cancer subtypes from computed tomography (CT) images remains challenging because malignant classes may exhibit overlapping radiographic characteristics, while conventional convolutional neural networks frequently rely only on final-layer semantic representations and may underuse informative intermediate features. This study proposes MLHNet-Lung, an attention-guided multi-level CNN–Transformer framework for image-level classification of adenocarcinoma, large cell carcinoma, normal, and squamous cell carcinoma CT slices. The study-specific novelty lies in the coordinated integration of intermediate block-4 and final EfficientNet-B0 feature maps, their independent refinement through Convolutional Block Attention Modules, projection into a shared 256-dimensional space, adaptive generalized mean pooling, compact three-token Transformer interaction, and three-source feature fusion. The experiments used 7,800 CT images, comprising 3,900 unique original images and 3,900 training-only augmented copies, divided into 5,200 training, 1,300 validation, and 1,300 independent test images. Validation and test sets contained only unique, non-augmented originals, and duplicate screening was applied to reduce image-level leakage. On the independent test set, MLHNet-Lung achieved an accuracy of 0.9415, a macro F1-score of 0.9414, and a macro ROC-AUC of 0.9821. Under the same experimental protocol, the proposed framework improved macro F1-score by 0.0104 over the strongest competing baseline and by 0.0175 over the standard EfficientNet-B0 classifier. The improvement over the strongest baseline was statistically significant according to McNemar’s test ( p = 0.028 ). Grad-CAM analysis further indicated that the model primarily emphasized pulmonary regions associated with its predictions. These findings demonstrate the value of jointly modeling complementary intermediate and high-level CT representations; however, patient-level multicenter validation and clinically verified labels remain necessary before diagnostic application.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Fareeha Hanif, Amira Elsir Tayfour Ahmed, Ali Raza
- Quelle
- Frontiers in Medicine
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2296-858X
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Fareeha Hanif, Amira Elsir Tayfour Ahmed, Ali Raza (2026). MLHNet-Lung: an attention-guided multi-level CNN–transformer fusion framework with CBAM and GeM pooling for explainable multiclass lung CT image classification. Frontiers in Medicine. https://doi.org/10.3389/fmed.2026.1924216
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1