Timely recognition of sepsis remains difficult when early physiological abnormalities are subtle or incomplete. This study examined whether general-purpose large language models could discriminate sepsis risk from an initial ICU vital-sign snapshot as effectively as established clinical scoring approaches. We performed a retrospective benchmark using the de-identified PhysioNet Sepsis Prediction Dataset. The eligible cohort included 39,234 adults, of whom 2,733 were sepsis-positive. Two zero-shot language models and two modified clinical scores were evaluated on the same patients. Discrimination was measured by the area under the receiver operating characteristic curve (AUROC), with bootstrap confidence intervals and DeLong tests for paired comparisons. AUROC was 0.597 for GPT-5.5, 0.592 for modified NEWS2, 0.591 for Claude Sonnet 5, and 0.570 for modified SOFA. Neither language model differed significantly from modified NEWS2, whereas both produced higher AUROC values than modified SOFA. A pre-onset-only sensitivity analysis, restricted to patients whose snapshot preceded their first positive sepsis label (achieved median lead time 39 hours), showed all four estimators converging (AUROC 0.579–0.591) with no significant pairwise differences, indicating the primary-analysis gap over SOFA was partly attributable to post-onset records. A stricter analysis limited to a minimum 6-hour lead time reversed the ranking: the SOFA-derived score obtained the highest AUROC (0.597), with the LLMs and NEWS2-derived score converging lower (0.573–0.578), again with no significant pairwise differences. Their alerting behaviour was not interchangeable: GPT-5.5 produced a more balanced sensitivity-specificity profile, while Claude Sonnet 5 identified more positive cases at the cost of additional false alerts. These findings show that zero-shot language models can provide discrimination comparable to the strongest modified clinical score in this dataset. The modest absolute AUROC values, use of modified scores, and retrospective single-dataset design mean that the models should be considered complementary research tools, not stand-alone clinical detectors.
Introduction
Sepsis is a life-threatening condition caused by an abnormal immune response to infection, making early detection essential for improving survival. Traditional bedside scoring systems such as NEWS2 and SOFA help assess patient risk but require specific clinical measurements and have limitations when data are incomplete. Recent advances in large language models (LLMs) suggest they may assist in clinical risk assessment, but their real-world effectiveness requires rigorous evaluation.
This study compared GPT-5.5, Claude Sonnet 5, and modified versions of NEWS2 and SOFA using the same cohort of 39,234 adult ICU patients from the public PhysioNet Sepsis Prediction Dataset. Each model received a single early patient assessment containing vital signs and demographic information and generated a sepsis risk score. Performance was evaluated using AUROC, sensitivity, specificity, predictive values, and subgroup analyses.
Results showed that GPT-5.5 achieved the highest AUROC (0.597), followed closely by the NEWS2-derived score (0.592) and Claude Sonnet 5 (0.591), while the SOFA-derived score (0.570) performed the worst. Both LLMs significantly outperformed the modified SOFA score but showed no statistically significant improvement over the NEWS2-derived score or over each other. Although AUROC values were similar, the models differed in their operating characteristics: Claude Sonnet 5 achieved higher sensitivity but lower specificity, whereas GPT-5.5 provided a more balanced trade-off between detecting sepsis and reducing false alarms.
Overall, the study demonstrates that general-purpose LLMs can rank sepsis risk at a level comparable to simple bedside scoring systems using only limited clinical information, but their predictive performance remains modest. The findings suggest that LLMs may serve as supportive clinical decision aids rather than standalone diagnostic tools, and further research using richer clinical data, prospective validation, and specialized model training is needed before deployment in routine healthcare settings.
Conclusion
In 39,234 ICU records, two zero-shot LLMs assigned ordinal scores that discriminated record-level sepsis status about as well as a NEWS2-derived score and better than a SOFA-derived score. A pre-onset-only sensitivity analysis on 38,658 records (median achieved lead time 39 hours) showed all four estimators converging, with no significant pairwise differences, indicating that the primary-analysis gap over the SOFA-derived score was driven in part by post-onset records rather than a stable advantage under stricter early-prediction conditions. A further, stricter sensitivity analysis restricted to a minimum 6-hour lead time reversed the ranking, with the SOFA-derived score obtaining the highest point estimate; this instability across cohort definitions is itself an informative finding, indicating that the relative ranking of these estimators is not robust to the prediction horizon considered and should not be treated as a fixed property of any single model. Across both analyses, the two LLMs’ overall AUROCs were similar to each other, but their sensitivity-specificity trade-offs differed enough to imply very different alert burdens.
These results support further evaluation of LLM-derived signals within critical-care research, not stand-alone bedside use. Future studies should define a fixed, prespecified pre-onset prediction horizon applied uniformly across all patients, use complete and feature-matched comparators, lock model versions, publish scoring and missing-data rules, assess repeatability and calibration, and validate prespecified thresholds prospectively before considering clinical integration.
References
[1] K. E. Rudd et al., “Global, regional, and national sepsis incidence and mortality, 1990–2017: Analysis for the Global Burden of Disease Study,” Lancet, vol. 395, no. 10219, pp. 200–211, 2020, doi: 10.1016/S0140-6736(19)32989-7.
[2] A. Kumar et al., “Duration of hypotension before initiation of effective antimicrobial therapy is the critical determinant of survival in human septic shock,” Crit. Care Med., vol. 34, no. 6, pp. 1589–1596, 2006, doi: 10.1097/01.CCM.0000217961.75225.E9.
[3] Royal College of Physicians, National Early Warning Score (NEWS) 2: Standardising the Assessment of Acute-Illness Severity in the NHS: Updated Report of a Working Party. London, U.K.: RCP, 2017. [Online]. Available: https://www.rcp.ac.uk/resources/national-early-warning-score-news-2/
[4] M. Singer et al., “The Third International Consensus Definitions for Sepsis and Septic Shock (Sepsis-3),” JAMA, vol. 315, no. 8, pp. 801–810, 2016, doi: 10.1001/jama.2016.0287.
[5] K. Singhal et al., “Large language models encode clinical knowledge,” Nature, vol. 620, no. 7972, pp. 172–180, 2023, doi: 10.1038/s41586-023-06291-2.
[6] Z. Kanjee, B. Crowe, and A. Rodman, “Accuracy of a generative artificial intelligence model in a complex diagnostic challenge,” JAMA, vol. 330, no. 1, pp. 78–80, 2023, doi: 10.1001/jama.2023.8288.
[7] J. Achiam et al., “GPT-4 Technical Report,” arXiv:2303.08774, 2023, doi: 10.48550/arXiv.2303.08774.
[8] G. B. Haim et al., “Evaluating large language model-assisted emergency triage: A comparison of acuity assessments by GPT-4 and medical experts,” J. Clin. Nurs., published online Nov. 28, 2024, doi: 10.1111/jocn.17490.
[9] A. Zaboli et al., “Chat-GPT in triage: Still far from surpassing human expertise: An observational study,” Am. J. Emerg. Med., vol. 92, pp. 165–171, 2025, doi: 10.1016/j.ajem.2025.03.028.
[10] O. A. Usman, A. A. Usman, and M. A. Ward, “Comparison of SIRS, qSOFA, and NEWS for the early identification of sepsis in the emergency department,” Am. J. Emerg. Med., vol. 37, no. 8, pp. 1490–1497, 2019, doi: 10.1016/j.ajem.2018.10.058.
[11] R. Goulden et al., “qSOFA, SIRS and NEWS for predicting inhospital mortality and ICU admission in emergency admissions treated as sepsis,” Emerg. Med. J., vol. 35, no. 6, pp. 345–349, 2018, doi: 10.1136/emermed-2017-207120.
[12] J. L. Vincent et al., “The SOFA (Sepsis-related Organ Failure Assessment) score to describe organ dysfunction/failure,” Intensive Care Med., vol. 22, no. 7, pp. 707–710, 1996, doi: 10.1007/BF01709751.
[13] M. A. Ansari Khoushabar and P. Ghafariasl, “Advanced meta-ensemble machine learning models for early and accurate sepsis prediction to improve patient outcomes,” arXiv:2407.08107, 2024, doi: 10.48550/arXiv.2407.08107.
[14] Q. Mao et al., “Multicentre validation of a sepsis prediction algorithm using only vital sign data in the emergency department, general ward and ICU,” BMJ Open, vol. 8, no. 1, e017833, 2018, doi: 10.1136/bmjopen-2017-017833.
[15] G. S. Collins et al., “TRIPOD+AI statement: Updated guidance for reporting clinical prediction models that use regression or machine learning methods,” BMJ, vol. 385, e078378, 2024, doi: 10.1136/bmj-2023-078378.
[16] M. A. Reyna et al., “Early prediction of sepsis from clinical data: The PhysioNet/Computing in Cardiology Challenge 2019,” Crit. Care Med., vol. 48, no. 2, pp. 210–217, 2020, doi: 10.1097/CCM.0000000000004145.
[17] E. R. DeLong, D. M. DeLong, and D. L. Clarke-Pearson, “Comparing the areas under two or more correlated receiver operating characteristic curves: A nonparametric approach,” Biometrics, vol. 44, no. 3, pp. 837–845, 1988, doi: 10.2307/2531595.