Vollständiger Abstract
Worum geht es in dieser Arbeit?
Purpose This study introduces a novel multi-modal transformer-based framework that integrates clinical data, radiomic features, and deep learning representations for automated classification of laryngeal lesions from contact endoscopy images. Methods We retrospectively enrolled 300 patients with laryngeal lesions from three independent medical centers (Centers A, B, and C), acquiring 7,847 contact endoscopy images with narrow band imaging. Clinical variables (18 features), radiomic features extracted using PyRadiomics (156 features), and deep learning representations from three state-of-the-art architectures, ConvNeXt V2, Swin Transformer V2, and EfficientNet V2 (3,328 combined features), were systematically integrated through a six-layer transformer encoder with multi-head cross-attention mechanisms. We evaluated three feature selection strategies (LASSO, mutual information, ReliefF) and three advanced classification architectures (Vision Transformer, Attention-MLP, Graph Neural Network). The framework was developed using data from Centers A and B ( n = 230 patients) with rigorous five-fold cross-validation, and validated on an independent external test set from Center C ( n = 70 patients). Results The optimal configuration (LASSO feature selection with Vision Transformer classifier) achieved accuracy of 96.1 ± 1.0% (AUC-ROC: 0.992 ± 0.007) on training, 93.8 ± 1.4% (AUC-ROC: 0.981 ± 0.012) on internal validation, and 91.4 ± 2.1% (AUC-ROC: 0.967 ± 0.019) on external testing, with sensitivity of 90.0 ± 2.8% and specificity of 92.5 ± 2.5% on the external cohort. The multi-modal fusion approach significantly outperformed all single-modality Methods clinical features alone (71.3% validation accuracy), radiomic features alone (84.2%), and the best individual deep learning model (87.6%), with improvements of 22.5, 9.6, and 6.2 percentage points respectively (all p < 0.001. Attention mechanism analysis revealed that the model dynamically weighted modality contributions, allocating 47.0% attention to deep features, 34.7% to radiomic features, and 18.3% attention to clinical features for correctly classified malignant cases. Conclusion This study presents the first comprehensive framework for laryngeal lesion classification that systematically integrates clinical data, radiomics, and deep learning through transformer-based fusion.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Yunjing Duan, Wei Li
- Quelle
- Frontiers in Surgery
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 2296-875X
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Yunjing Duan, Wei Li (2026). A transformer-based multi-modal fusion framework for laryngeal lesion classification using contact endoscopy. Frontiers in Surgery. https://doi.org/10.3389/fsurg.2026.1840300
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1