EUVIMEDEuropean Health Evidence
Uhr 7/7Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

A transformer-based multi-modal fusion framework for laryngeal lesion classification using contact endoscopy

Yunjing Duan, Wei Li

Frontiers in Surgery · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Purpose This study introduces a novel multi-modal transformer-based framework that integrates clinical data, radiomic features, and deep learning representations for automated classification of laryngeal lesions from contact endoscopy images. Methods We retrospectively enrolled 300 patients with laryngeal lesions from three independent medical centers (Centers A, B, and C), acquiring 7,847 contact endoscopy images with narrow band imaging. Clinical variables (18 features), radiomic features extracted using PyRadiomics (156 features), and deep learning representations from three state-of-the-art architectures, ConvNeXt V2, Swin Transformer V2, and EfficientNet V2 (3,328 combined features), were systematically integrated through a six-layer transformer encoder with multi-head cross-attention mechanisms. We evaluated three feature selection strategies (LASSO, mutual information, ReliefF) and three advanced classification architectures (Vision Transformer, Attention-MLP, Graph Neural Network). The framework was developed using data from Centers A and B ( n = 230 patients) with rigorous five-fold cross-validation, and validated on an independent external test set from Center C ( n = 70 patients). Results The optimal configuration (LASSO feature selection with Vision Transformer classifier) achieved accuracy of 96.1 ± 1.0% (AUC-ROC: 0.992 ± 0.007) on training, 93.8 ± 1.4% (AUC-ROC: 0.981 ± 0.012) on internal validation, and 91.4 ± 2.1% (AUC-ROC: 0.967 ± 0.019) on external testing, with sensitivity of 90.0 ± 2.8% and specificity of 92.5 ± 2.5% on the external cohort. The multi-modal fusion approach significantly outperformed all single-modality Methods clinical features alone (71.3% validation accuracy), radiomic features alone (84.2%), and the best individual deep learning model (87.6%), with improvements of 22.5, 9.6, and 6.2 percentage points respectively (all p < 0.001. Attention mechanism analysis revealed that the model dynamically weighted modality contributions, allocating 47.0% attention to deep features, 34.7% to radiomic features, and 18.3% attention to clinical features for correctly classified malignant cases. Conclusion This study presents the first comprehensive framework for laryngeal lesion classification that systematically integrates clinical data, radiomics, and deep learning through transformer-based fusion.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Yunjing Duan, Wei Li
Quelle
Frontiers in Surgery
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
2296-875X
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Yunjing Duan, Wei Li (2026). A transformer-based multi-modal fusion framework for laryngeal lesion classification using contact endoscopy. Frontiers in Surgery. https://doi.org/10.3389/fsurg.2026.1840300
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1