Peptides DB
Research-centric peptide and protocol reference hub
Study 18 of 25HCG (Human Chorionic Gonadotropin) literatureDEN open · Observational2023

Evaluation of Large Language Models in the Clinical Management of Patients With Upper Gastrointestinal Bleeding: Insights From Real-World Patient Data.

Large language models showed moderate performance in risk stratification for upper gastrointestinal bleeding but did not surpass established clinical scores like the Glasgow-Blatchford Score.

Read at DEN openAdd to compare

Where it sits

this study against the rest of the hcg (human chorionic gonadotropin) corpus
5
Preclinical
18
Observational · this one
0
Open-label
1
Randomised
1
Reviews

Summary and findings

This study evaluated large language models (LLMs) for pre-endoscopy risk stratification and prediction of endoscopic findings in upper gastrointestinal bleeding (UGIB) among 384 patients. The performance of LLMs was compared with clinical risk scores and conventional machine learning models. The findings indicate that LLMs showed moderate performance but were inferior to established clinical scores.

How much of this paper we could read: full text read (0.80). We had a clear abstract, so the summary below closely tracks the paper. What this means →
GBS showed the best discriminative performance among clinical scores (AUROC 0.73).n=3842023

Abstract

The authors’ words, as DEN open supplied them

<h4>Objective</h4>To evaluate large language models (LLMs) for pre-endoscopy (PE) risk stratification and prediction of endoscopic findings in upper gastrointestinal bleeding (UGIB), and compare their performance with clinical risk scores, conventional machine learning (ML) models, and hybrid approaches.<h4>Methods</h4>This multicenter retrospective study included 384 patients with UGIB who underwent endoscopy. Five LLMs (GPT-5, Gemini-2.5-Flash, Llama 4, Grok, and DeepSeek R1) were tested using structured zero-shot prompts based on PE clinical and laboratory data. Their performance in identifying high-risk patients was compared with the Glasgow-Blatchford Score (GBS), AIMS65, PE Rockall score, and ML models. Hybrid LLM-score models were also assessed. Two gastroenterologists evaluated the quality of LLM-generated justifications.<h4>Results</h4>GBS showed the best discriminative performance among clinical scores (area under the receiver operating characteristic curve [AUROC] 0.73). Among LLMs, GPT-5 achieved the highest accuracy (0.66), while Grok showed the best-balanced performance (0.59; F1 0.47). Gemini-2.5-Flash had the highest sensitivity (0.89) but low specificity (0.21). Endoscopic prediction performance was modest, with Gemini-2.5-Flash achieving the highest exact-match accuracy (0.34) and micro-F1 (0.38). Hybrid models improved performance over standalone LLMs but did not outperform GBS alone (best: GBS+GPT-5, AUROC 0.670. LLMs showed higher numerical performance than conventional ML models, but a statistical comparison was not possible due to unavailable instance-level data. Grok received the highest human evaluation score for explanation quality.<h4>Conclusions</h4>LLMs showed moderate performance in UGIB risk stratification and endoscopic prediction but were inferior to clinical scores, especially GBS. Hybrid models modestly improved over standalone LLMs but not GBS, supporting their use as adjunct tools rather than clinical decision-support systems.<h4>Trial registration</h4>N/A.

Background

This paper addresses the clinical question of how effectively large language models can assist in risk stratification and prediction of endoscopic findings in patients with upper gastrointestinal bleeding. Previous studies have established clinical risk scores for this purpose, but the potential of LLMs remains underexplored. Understanding the performance of LLMs in this context could inform their utility as adjunct tools in clinical settings.

Methods

This multicenter retrospective study included 384 patients with UGIB who underwent endoscopy. Five LLMs were tested using structured zero-shot prompts based on pre-endoscopy clinical and laboratory data. The performance of LLMs was compared with clinical risk scores, including the Glasgow-Blatchford Score (GBS), AIMS65, and PE Rockall score, as well as conventional machine learning models.

Results

The Glasgow-Blatchford Score (GBS) showed the best discriminative performance among clinical scores with an area under the receiver operating characteristic curve (AUROC) of 0.73. Among the LLMs, GPT-5 achieved an accuracy of 0.66, while Grok had a balanced performance with a score of 0.59 (F1 0.47). Gemini-2.5-Flash had the highest sensitivity at 0.89 but a low specificity of 0.21. The highest exact-match accuracy for endoscopic prediction was 0.34, and hybrid models did not outperform GBS alone.

Interpretation

The findings suggest that while LLMs can provide some insights into UGIB risk stratification, their performance is moderate and inferior to established clinical scores like GBS. The statistical comparisons were limited due to the lack of instance-level data, which raises questions about the robustness of the findings. The modest improvements seen with hybrid models indicate potential for LLMs as supportive tools rather than replacements for traditional clinical decision-making processes.

Key findings

  • GBS showed the best discriminative performance among clinical scores (AUROC 0.73).
  • GPT-5 achieved the highest accuracy among LLMs (0.66).
  • Grok showed the best-balanced performance (0.59; F1 0.47).
  • Gemini-2.5-Flash had the highest sensitivity (0.89) but low specificity (0.21).
  • The highest exact-match accuracy for endoscopic prediction was achieved by Gemini-2.5-Flash (0.34).
  • Hybrid models did not outperform GBS alone (best: GBS+GPT-5, AUROC 0.67).

Limitations

  • Statistical comparison limited due to unavailable instance-level data.
  • LLMs did not outperform established clinical scores.
  • Retrospective study design may introduce bias.
  • Multicenter data may have variability affecting results.

Elsewhere in the HCG (Human Chorionic Gonadotropin) corpus

DComparison of inhalational methoxyflurane, intranasal fentanyl, and intravenous morphine for treatment of prehospital acute pain in Norway (PreMeFen): a randomised, non-inferiority, three-arm, phase 3 trial.Lancet (London, England) · 2026DPreoperative mFOLFIRINOX versus PAXG for stage I-III resectable and borderline resectable pancreatic ductal adenocarcinoma (PACT-21 CASSANDRA): results of the first randomisation analysis of a randomised, open-label, 2 × 2 factorial phase 3 trial.Lancet (London, England) · 2026DRussia's health workers complicit in abducting Ukraine's children.Lancet (London, England) · 2026AEndoscopic Thrombin Injection for Gastric Variceal Bleeding: A Systematic Review and Meta-Analysis of Observational and Trial Data.DEN open · 2023 · n=417 · Initial hemostasis rate of 93% (95% CI: 0.89-0.95) across 13 studies, n=417.HumanBEfficiency and Safety of Endoscopic Injection Sclerotherapy With Ligation for Esophageal Varices: A Retrospective Study.DEN open · 2023 · n=148 · Procedure duration: 13.1 ± 7.2 min for EISL vs 20.7 ± 8.0 min for EIS, p < 0.0001.HumanBCase Report of a Rapidly Progressive Indolent T-Cell Lymphoma of the Gastrointestinal Tract.DEN open · 2023 · n=1 · Not reported in abstract.Human