EUVIMEDEuropean Health Evidence
Uhr 7/7Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

Multimodal medical diagnosis: a mini review of LLM–vision fusion models in low-resource healthcare settings

Kahakashan Ashraf, Md. Hamid Hosen, Nuzhat Tabassum Farah, Md Kishor Morol, Tze Hui Liew, Dip Nandi, Mashiour Rahman, Abdullah Al Jubair

Frontiers in Digital Health · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Recent advances in large language models (LLMs) and vision transformers have enabled multimodal systems that integrate clinical text with medical imaging for diagnostic decision-making. While these systems show promising results on benchmark datasets in well-resourced research settings, their applicability in low-resource healthcare environments where diagnostic disparities are most severe remains limited and poorly understood. This mini review synthesizes key developments in LLM–vision fusion architectures from 2018 to 2026, with a focus on radiology-oriented visual question answering (VQA) and report generation systems viewed from a deployment perspective. Rather than comprehensively cataloguing multimodal medical AI, we synthesize the evolution of LLM–vision fusion architectures and discuss complementary deployment-enabling strategies, including parameter-efficient adaptation, post-training quantization, federated learning, and multilingual support, where they directly improve the feasibility of radiology AI in resource-constrained healthcare settings. Rather than focusing solely on performance benchmarks, we examine these approaches through a deployment-oriented lens, highlighting trade-offs between representational capacity, computational efficiency, interpretability, and memory footprint. We argue that current progress remains substantially shaped by model scaling and benchmark optimization, which often do not address the memory, connectivity, and annotation constraints of low-resource healthcare systems. While cross-modal transformer architectures provide strong representational alignment, their computational demands and reliance on large curated datasets limit real-world deployment. In contrast, emerging directions including parameter-efficient fine-tuning, post-training quantization, federated learning, and modular agent-based systems offer more tractable pathways toward clinical integration under hardware and data constraints. To bridge the gap between benchmark performance and clinical utility, we identify concrete challenges in data scarcity, multilingual coverage, and calibration, and propose a shift toward lightweight, interpretable, and hardware-aware multimodal AI. This perspective highlights the need to move beyond scaling-centric design toward models that can run on 4–8 GB VRAM, operate offline, and generalize across languages and imaging equipment.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Kahakashan Ashraf, Md. Hamid Hosen, Nuzhat Tabassum Farah, Md Kishor Morol, Tze Hui Liew, Dip Nandi, Mashiour Rahman, Abdullah Al Jubair
Quelle
Frontiers in Digital Health
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
2673-253X
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Kahakashan Ashraf, Md. Hamid Hosen, Nuzhat Tabassum Farah, Md Kishor Morol, Tze Hui Liew, Dip Nandi, Mashiour Rahman, Abdullah Al Jubair (2026). Multimodal medical diagnosis: a mini review of LLM–vision fusion models in low-resource healthcare settings. Frontiers in Digital Health. https://doi.org/10.3389/fdgth.2026.1862236
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1