Vollständiger Abstract
Worum geht es in dieser Arbeit?
Objectives: Positron Emission Tomography-Computed Tomography (PET-CT) is integral to breast cancer staging. However, manual interpretation is time-consuming and subject to inter-observer variability. While artificial intelligence has shown promise in lesion detection and standardised uptake value (SUVmax) quantification, existing approaches are fragmented and do not generate complete TNM-aligned structured reports. Consequently, there remains a need for an integrated framework capable of translating PET-CT data into structured clinical reports. Material and Methods: We developed a Multimodal Report Generator (MMRG) as an exploratory pilot study using data from 85 breast cancer patients (439 clusters; 59/15/11 train/validation/test). The framework combines tri-planar Swin3D feature extraction, multimodal feature fusion, multi-task lesion analysis, and (Generative Pre-trained Transformer) GPT-2-based report generation with Facebook AI Similarity Search (FAISS) retrieval-augmented generation. The entire pipeline was trained on a single consumer-grade graphics processing unit (GPU), supporting a resource-efficient design. Results: Model Training demonstrated stable convergence without evidence of train-validation divergence. At the selected validation checkpoint (15 patients, 50 clusters), lesion detection achieved a receiver operating characteristic area under the curve (ROC-AUC) of 0.861 and an F1 score of 0.757, while descriptor classification achieved a ROC-AUC of 0.787 and an F1 score of 0.757. SUVmax and lesion-size estimation showed Pearson correlations of 0.44 and 0.64, respectively. On the held-out internal test set ( n = 11 patients, 61 clusters), lesion detection achieved an ROC-AUC of 0.857, an F1 score of 0.652, and a Cohen’s κ of 0.466, while descriptor classification achieved an ROC-AUC of 0.759, an F1 score of 0.546, and a κ of 0.381. Quantitative estimation yielded SUVmax mean absolute error (MAE) = 6.44 (Pearson r = 0.49) and lesion-size MAE = 0.73 cm (Pearson r = 0.34). Tumour-Node-Metastasis (TNM) staging accuracy was 18.2%, with M-stage accuracy of 63.6%. The organ-domain hallucination rate was 35.7% on breast-cancer-relevant test clusters. The framework successfully generated structured PET-CT reports with automated lesion description and TNM staging. Conclusion: This exploratory pilot study demonstrates the feasibility of an end-to-end multimodal AI pipeline for PET-CT interpretation in breast cancer, combining lesion detection, quantitative measurement, structured report generation, and TNM staging within a single framework. While lesion- and descriptor-classification showed encouraging performance, TNM staging results remain preliminary in this small cohort. The framework is presented as a resource-efficient prototype for future larger-cohort studies rather than a clinically validated deployment-ready system. Future work will focus on larger multi-centre datasets, improved quantitative estimation, and direct or joint TNM-stage prediction architectures.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Dibya Prakash, Venkatesh Rangarajan, Manoj Kumar
- Quelle
- Indian Journal of Nuclear Medicine
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 0974-0244, 0972-3919
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Dibya Prakash, Venkatesh Rangarajan, Manoj Kumar (2026). A Resource-Efficient End-to-End Multimodal AI Framework for Automated PET-CT Report Generation and TNM Staging in Breast Cancer. Indian Journal of Nuclear Medicine. https://doi.org/10.25259/ijnm_97_2026