Brief

Closing the Loop: A Brief Report on Clinician Trust After Transparent Model Explanations

Dr. Sofia Rinaldi (Riverstone Children's Hospital), Dr. Fatima Al-Rashid (Sable Digital Health)

Volume 5, Number 3 · June 2026 · pp. 196–208 · doi:10.59821/jhds.2026.0307

Received January 9, 2026 · Accepted April 22, 2026 · Published June 15, 2026

Abstract

We added plain-language explanations to a deterioration-risk alert and measured whether clinicians acted on it differently. Trust and appropriate use both rose modestly, but only when the explanation named the specific factors driving each individual alert.

clinical risk modelsmodel monitoringhealth equitymachine learning

Introduction

Health systems increasingly depend on predictive models embedded in electronic health records. Yet most published evaluations describe performance at a single moment in time, leaving practitioners with little evidence about how these tools behave months or years after go-live.

Methods

We conducted a retrospective cohort study across participating sites between 2023 and 2025. Performance was assessed monthly using discrimination (AUROC), calibration slope, and subgroup-specific false positive rates. The study was approved by each site's institutional review board with a waiver of consent.

Results

Of the models studied, a substantial share showed statistically significant degradation, most commonly in calibration rather than discrimination. Degradation was concentrated in periods following documentation template changes and shifts in patient mix.

Discussion

Our findings suggest that routine, low-cost monitoring can detect meaningful performance changes well before they surface through clinician complaints. We recommend health systems assign a named owner to every deployed model.

Figure 1. Median AUROC by quarter after deployment, with recovery following recalibration in Q7.

References

  1. Finlayson SG, Subbaswamy A, Singh K, et al. The clinician and dataset shift in artificial intelligence. N Engl J Med. 2021;385(3):283-286.
  2. Davis SE, Lasko TA, Chen G, et al. Calibration drift in regression and machine learning models for acute kidney injury. J Am Med Inform Assoc. 2017;24(6):1052-1061.
  3. Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366(6464):447-453.