Review

The State of Clinical NLP: A Review of Methods, Benchmarks, and Open Problems

Dr. Benjamin Tran (Cascadia University), Dr. Mei-Ling Wong (Pinewood Regional Medical Group)

Volume 5, Number 3 · June 2026 · pp. 154–166 · doi:10.59821/jhds.2026.0304

Received January 9, 2026 · Accepted April 22, 2026 · Published June 15, 2026

Abstract

We survey a decade of clinical natural language processing, from rule-based extraction to transformer models, and trace which problems have genuinely been solved and which have merely been renamed. We argue the field's biggest remaining barrier is not modeling but the scarcity of shared, well-annotated clinical text.

clinical risk modelsmodel monitoringhealth equitymachine learning

Introduction

Health systems increasingly depend on predictive models embedded in electronic health records. Yet most published evaluations describe performance at a single moment in time, leaving practitioners with little evidence about how these tools behave months or years after go-live.

Methods

We conducted a retrospective cohort study across participating sites between 2023 and 2025. Performance was assessed monthly using discrimination (AUROC), calibration slope, and subgroup-specific false positive rates. The study was approved by each site's institutional review board with a waiver of consent.

Results

Of the models studied, a substantial share showed statistically significant degradation, most commonly in calibration rather than discrimination. Degradation was concentrated in periods following documentation template changes and shifts in patient mix.

Discussion

Our findings suggest that routine, low-cost monitoring can detect meaningful performance changes well before they surface through clinician complaints. We recommend health systems assign a named owner to every deployed model.

Figure 1. Median AUROC by quarter after deployment, with recovery following recalibration in Q7.

References

  1. Finlayson SG, Subbaswamy A, Singh K, et al. The clinician and dataset shift in artificial intelligence. N Engl J Med. 2021;385(3):283-286.
  2. Davis SE, Lasko TA, Chen G, et al. Calibration drift in regression and machine learning models for acute kidney injury. J Am Med Inform Assoc. 2017;24(6):1052-1061.
  3. Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366(6464):447-453.