Vollständiger Abstract
Worum geht es in dieser Arbeit?
While large language models (LLMs) offer promising support for injury coding, their use raises concerns related to reliability and data privacy in healthcare. This study examines whether LLMs can reliably support injury coding and the efficacy of locally deployed offline LLMs. A benchmark dataset of 100 injury narratives, sampled from the National Electronic Injury Surveillance System, was used to examine performance. Three model families were evaluated: (1) cloud-based LLMs (GPT 5 and LLAMA 4), (2) a locally deployed offline LLM-Mistral 7B, and (3) deep learning baselines (LSTM/RNN) trained on labeled NEISS data. GPT 5 achieved the highest recall (0.82) and precision (0.81), with LLAMA 4 showing comparable results across most injury categories. The Mistral 7B model achieved moderate performance (0.62 recall; 0.71 precision), outperforming traditional deep learning baselines. These results indicate that cloud-based LLMs can effectively support single diagnosis injury coding, while locally deployed LLMs offer a privacy preserving alternative.
Bibliografischer Nachweis
Publikationsdaten
- Autor:innen
- Rishab Ranjan Chakravarty, Mustak Ahmad, Gaurav Nanda
- Quelle
- Proceedings of the Human Factors and Ergonomics Society Annual Meeting
- Publikation
- 2026-01-01
- Band / Ausgabe
- Nicht angegeben
- Seiten
- Nicht angegeben
- ISSN / ISBN
- 1071-1813, 2169-5067
- Zitationen
- 0 laut Crossref
- Referenzen
- 0 hinterlegt
Zitieren
Zitierfähiger Nachweis
Rishab Ranjan Chakravarty, Mustak Ahmad, Gaurav Nanda (2026). Automated Injury Diagnosis Coding Using LLMs: Performance and Privacy Trade-offs. Proceedings of the Human Factors and Ergonomics Society Annual Meeting. https://doi.org/10.1177/10711813261485840
Kontext
Themen, Förderung und Nutzung
Lizenzhinweise: Lizenz 1