EUVIMEDEuropean Health Evidence
Uhr 10/10Sources Journal Tree
Easy Demo

Lokaler Crossref-Datenbestand · journal-article

Large language models in the vertical integration of spine surgery workflow: a scoping review

Edwin H. Y. Lui, Ralph J. Mobbs

European Spine Journal · 2026

Vollständiger Abstract

Worum geht es in dieser Arbeit?

Abstract Purpose To evaluate current evidence regarding the clinical reliability and reasoning capabilities of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) within the spine surgery workflow. This scoping review utilizes a novel Vertical Integration Maturity Scale (VIMS) to map model maturity across five stages of the vertical workflow, identifying persistent research gaps and technical prerequisites for clinical implementation. Methods A systematic search was conducted across PubMed, Embase, and Scopus for peer-reviewed studies published between January 2023 and January 2026. Utilizing the Population, Concept, and Context (PCC) framework, studies were selected based on their application of generative AI to the clinical evaluation and surgical management of spinal pathologies. Data were analyzed via independent dual-coding using a two-dimensional framework mapping five VIMS levels across five functional workflow stages. Synthesis included categorization by model architecture, input modality, and performance benchmarks, with inter-rater reliability calculated to validate the novel framework. Results Of 351 identified records, 40 studies met the inclusion criteria. Research density peaked in Stage III (Evaluative Decision Logic) at VIMS Level 1 ( n = 15), indicating a methodological focus on evidence-based guideline retrieval over patient-specific synthesis. While ChatGPT-4 showed high concordance (61.1% to 88.2%) with clinical guidelines, diagnostic precision fell to 11.1% in complex spinal deformity. A critical evidence gap persists in Stage IV (Tactile Procedure Planning) with a total absence of studies evaluating autonomous workflow agency (VIMS 4–5). This absence likely reflects a limitation of retrospective study designs reliant on proprietary, non-version-controlled models. While qualitative synthesis suggests hybrid architectures and Retrieval-Augmented Generation (RAG) may bridge existing semantic gaps, the inferential strength of these trends is currently limited by heterogenous benchmarking metrics. Conclusion Current LLMs demonstrate proficiency in foundational medical knowledge, but a significant evidence gap remains regarding their inferential validity in the non-deterministic surgical scenarios. Advancing clinical integration requires modular multimodal ensembles that process radiographic data alongside domain-specific reasoning scaffolds. Crucially, future research must transition from retrospective algorithmic benchmarking toward standardized, prospective clinical trials to evaluate algorithmic decision-making against real-world surgical outcomes.

Bibliografischer Nachweis

Publikationsdaten

Autor:innen
Edwin H. Y. Lui, Ralph J. Mobbs
Quelle
European Spine Journal
Publikation
2026-01-01
Band / Ausgabe
Nicht angegeben
Seiten
Nicht angegeben
ISSN / ISBN
0940-6719, 1432-0932
Zitationen
0 laut Crossref
Referenzen
0 hinterlegt

Zitieren

Zitierfähiger Nachweis

Edwin H. Y. Lui, Ralph J. Mobbs (2026). Large language models in the vertical integration of spine surgery workflow: a scoping review. European Spine Journal. https://doi.org/10.1007/s00586-026-10348-x
RIS BibTeX CSL-JSON

Kontext

Themen, Förderung und Nutzung

Lizenzhinweise: Lizenz 1