External validation reveals poor calibration despite preserved discrimination for a vasopressin treatment-signal prediction model across ICU databases
The prediction model for vasopressin treatment signals shows some ability to discriminate but lacks accurate calibration for absolute risk, limiting its clinical utility.
Where it sits
this study against the rest of the vasopressin corpusSummary and findings
This study developed and validated a prediction model for a vasopressin treatment signal using ICU databases. The model was assessed for its ability to discriminate treatment signals but showed poor calibration in absolute risk. No treatment benefits were claimed or demonstrated.
Abstract
<title>Abstract</title> <p>Background Prediction models for treatment-dependent targets may retain discrimination after database transfer while misrepresenting absolute risk. We developed and externally validated a model for a near-term documented vasopressin treatment signal across MIMIC-IV and eICU-CRD. The outcome is a treatment-documentation signal, not evidence of treatment benefit. Methods Using 15 shared variables, we developed a LASSO logistic model in MIMIC-IV v3.1 with a patient-level split and grouped cross-validation, then tested the frozen model in held-out MIMIC-IV and unrecalibrated eICU-CRD v2.0. We assessed discrimination and calibration with clustered uncertainty. Additional analyses included a post hoc probability-oriented penalty selection, local updating with bootstrap uncertainty propagated through the update step, and a 500-repetition recalibration learning curve evaluated in disjoint patients or hospitals. Results Training, internal-test, and eICU cohorts contained 3,692, 1,582, and 1,801 stays (509, 219, and 166 signals). AUROC was 0.768 (95% CI 0.734–0.804) internally and 0.746 (0.713–0.782) externally. Internally, mean predicted risk was aligned (calibration-in-the-large 0.007), but a slope of 1.501 indicated under-dispersed predictions; externally, the calibration errors were opposite in direction. Mean overprediction (calibration-in-the-large − 0.712) occurred with a slope of 0.446. External AUPRC was 0.200, approximately 2.2 times the 9.2% prevalence baseline. Offset-only updating retained a slope of 0.437. At the largest same-system target of 125 events, median out-of-sample slope was 0.877 (93.6% of repetitions within 0.8–1.2); hospital-held-out slope was 0.789 (38.4% within 0.8–1.2), and only 2 of 67 hospitals had at least 10 events. Conclusions The model retained cross-database discrimination but not calibrated absolute risk. The 125-event result was a conditional observation at the eICU data boundary, not a deployment threshold; additional pooled events did not correct unseen-hospital transport, and no single-hospital threshold was estimable. The model does not estimate treatment benefit or support direct treatment recommendations.</p>
Background
This paper addresses the calibration and discrimination of a prediction model for vasopressin treatment signals in ICU settings. Previous models have shown varying degrees of success in transferring predictive capabilities across different databases. Understanding the calibration of such models is crucial for their application in clinical settings.
Methods
The study employed a LASSO logistic model developed using 15 shared variables from the MIMIC-IV database. The model was validated using internal and external cohorts, with n=3,692 for training, n=1,582 for internal testing, and n=1,801 for external validation. The primary outcome was the treatment-documentation signal, and various statistical methods were used to assess model performance.
Results
The primary endpoint showed an AUROC of 0.768 (95% CI 0.734–0.804) internally and 0.746 (0.713–0.782) externally. Calibration-in-the-large was 0.007 internally, indicating alignment, but the slope of 1.501 suggested under-dispersed predictions. Externally, the model exhibited a mean overprediction with a calibration-in-the-large of -0.712 and a slope of 0.446.
Interpretation
The model retained discrimination across databases but failed to provide calibrated absolute risk estimates. While the AUROC values indicate some predictive capability, the clinical significance is limited due to poor calibration. The findings suggest that while the model may identify treatment signals, it cannot be relied upon for direct treatment recommendations, especially given the small number of events in some hospitals.
Key findings
- AUROC was 0.768 (95% CI 0.734–0.804) internally and 0.746 (0.713–0.782) externally.
- Mean overprediction (calibration-in-the-large − 0.712) occurred with a slope of 0.446.
- External AUPRC was 0.200, approximately 2.2 times the 9.2% prevalence baseline.
- At the largest same-system target of 125 events, median out-of-sample slope was 0.877 (93.6% of repetitions within 0.8–1.2).
- Hospital-held-out slope was 0.789 (38.4% within 0.8–1.2).
- Only 2 of 67 hospitals had at least 10 events.
Limitations
- The model does not estimate treatment benefit or support direct treatment recommendations.
- Only 2 of 67 hospitals had at least 10 events.
- Short follow-up period limits assessment of long-term applicability.
- Post hoc analyses may introduce bias.
- Single-site and small n limitations affect generalizability.