Vollständiger Abstract
Worum geht es in dieser Arbeit?
Objectives: Positron emission tomography-computed tomography (PET-CT) reporting is cognitively demanding, requiring integration of quantitative metabolic data and anatomical findings across multiple body regions. Existing segmentation and analysis models provide anatomical and physiological parameters, but not cohesive physician-style reports; however, they can supply a structured clinical context for a dedicated PET-CT language model. Existing large language model (LLM) approaches for radiology report generation predominantly target chest radiography and lack domain-specific grounding for nuclear medicine, and unconstrained free-text generation from generic prompts introduces clinically unacceptable hallucination risk. Material and Methods: We assembled 836 section-level PET-CT reports from 170 patients (85 breast cancer, 85 controls) across four primary anatomical regions (Brain excluded due to stereotyped language) plus a whole-body report for each patient. The biomedical language model BioMedLM (2.7B) was fine-tuned using Low-Rank Adaptation (LoRA) with a faithfulness-augmented objective penalising hallucination and rewarding negation preservation. Structured clinical context, including maximum standardised uptake value (SUVmax), laterality, and lesion localisation, served as grounded input. Evaluation included five-fold patient-level cross-validation, held-out test performance, temporal and cross-group validation, clinical safety metrics, and blinded independent review by two nuclear medicine physicians across 28 test cases. Results: Physician review showed factual correctness 4.59/5, completeness 4.77/5, and clinical usefulness 4.12/5, with no unsafe hallucinations identified (one-sided 95% CI: 0-10.7%). Clinical acceptability (Grade A) was 53.6% (Physician 1) and 57.1% (Physician 2; combined 55.4%), with moderate inter-rater agreement (κ = 0.61). BioMedLM-LoRA achieved Recall-Oriented Understudy for Gisting Evaluation (ROUGE-L) 0.545 and Bidirectional Encoder Representations from Transformers Score (BERTScore-F1) 0.911 on the held-out test set, outperforming template and retrieval-augmented baselines (all p < 10 -10 ); SUVmax faithfulness was 1.000 and negative-finding preservation 0.790. Term frequency-inverse document frequency (TFIDF) retrieval achieved higher ROUGE-L (0.814), reflecting lexical-overlap advantage in this closed, single-site corpus. Automated laterality assessment remained proxy-level (accuracy 0.271), underscoring rule-based evaluation limitations. Not all flagged laterality (33/120) and hallucination-proxy (14/120) cases underwent physician adjudication; these findings represent supportive evidence from a limited subset, not a generalisable safety claim. Conclusion: Parameter-efficient fine-tuning of biomedical language models can generate coherent physician-style PET-CT report sections from structured clinical findings, with high physician-assessed factual correctness and no unsafe hallucinations in this limited reviewed subset, supporting the feasibility of structured-finding-to-narrative report generation for clinical AI pipelines. However, the structured findings used as model inputs were derived from reference reports rather than independently generated by an upstream image-analysis system, representing a limitation. Larger external, multi-institutional studies using independently extracted image-derived context are required to establish generalisability and assess suitability for clinical deployment under physician supervision.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Dibya Prakash, Pramukh Kulkarni, Venkatesh Rangarajan, Manoj Kumar
- Quelle
- Indian Journal of Nuclear Medicine
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 0974-0244, 0972-3919
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Dibya Prakash, Pramukh Kulkarni, Venkatesh Rangarajan, Manoj Kumar (2026). Context-Grounded PET-CT Report Generation Using LoRA-Fine-Tuned BioMedLM with Physician Validation and Safety Evaluation. Indian Journal of Nuclear Medicine. https://doi.org/10.25259/ijnm_104_2026