AI Drift: The Patient Safety Risk Your Security Tools Can't See

Shantanu Nigam
CEO
Oct 2, 2026
9 minutes
CEO
Physician with tablet showing AI

AI models don’t always perform the way they did on go-live day. As patient populations, clinical practice and data change, AI drift can quietly shift a model to become less reliable even when its code remains unchanged. Drawing on real-world cases, Vitea Founder and CEO Shantanu Nigam explains why continuous monitoring is a core governance responsibility for health system leaders.

An AI failure that puts patients at risk may not set off an alarm. It may not crash the system or show up in a vendor report. It can, on the surface, look like a tool that’s working as intended.

Consider a readmission model that slowly starts flagging the wrong patients, diverting care managers from those who need them most. Or a deterioration score that grows less reliable for patients over 75 while its overall numbers look fine. Neither model has been modified, and aggregate dashboards still look reassuring.

That is silent drift, and I believe it’s one of the most underestimated risks in healthcare AI today. Worse yet, most governance models aren’t built to catch it.

AI governance needs to extend beyond the moment of approval. Procurement and validation establish a baseline, but go-live cannot be the finish line. From that day forward, the world a model learned from can change.

Medicine already understands this. No hospital would credential a surgeon once and never review her performance again. Ongoing professional practice evaluation provides a framework for reviewing clinicians’ performance over time. AI that influences care deserves the same discipline: initial approval followed by ongoing evaluation.

What is drift in AI?

AI drift refers to changes over time in the data a model encounters or the relationships it relies on, which can affect its performance after deployment. These changes can be gradual or abrupt and do not always reduce performance. When degradation does occur, it can be difficult to spot: summary metrics may hold steady while risk estimates become less reliable or performance gaps widen across patient groups.

Research on AI “aging” by Harvard Medical School, MIT, and partner institutions found that model quality degrades as time passes since training, even without major shifts in the underlying data. By one widely cited count, degradation appeared in 91% of the cases tested, using data from healthcare operations, finance, transportation, and weather.

Part of the problem is what gets measured. Predictive risk models, for example, are often judged using discrimination, or how well they separate higher-risk patients from lower-risk ones. That measure can hold steady while calibration, whether a predicted 30% risk corresponds to an observed rate of about 30%, slips. When it does, the thresholds, staffing plans and quality benchmarks built on those numbers can quietly become unreliable.

I think of AI drift in three forms:

  • Performance drift: a model’s accuracy or calibration erodes as the world around it changes.
  • Fairness drift: performance erodes unevenly, so some patient groups are failed before others.
  • Use drift: a model’s role in real decisions expands beyond what anyone approved.

The cases and research below illustrate each of them.

1. When COVID-19 changed the patients and the alert burden

In April 2020, the University of Michigan Hospital turned off its Epic sepsis alert model after nursing reports of overalerting. The model predated the pandemic, and when COVID-19 changed the hospital's patient mix, it began firing a flood of false alarms, raising concerns how the alerts fit the demands of clinical care.

Michigan wasn’t alone. A study of 24 hospitals across four health systems found that daily sepsis alerts were 43% higher in the three weeks after each hospital’s first COVID-19 case than in the three weeks before, even as hospital census fell 35%.

Why it matters: This was drift at its most visible, and it still took a pandemic to force action. The study’s authors concluded that health systems must be able to rapidly assess and disable AI alerts. That capability has to exist before the crisis.

2. Performance drift: When validated models decay on their own

Vanderbilt researchers analyzing national VA data tracked seven types of models predicting acute kidney injury over several years. Their ability to distinguish higher-risk from lower-risk patients remained relatively stable. But as actual rates of kidney injury and death fell, the models increasingly overstated risk while their headline discrimination measures remained stable.

A later VA analysis warns that this kind of miscalibration can distort risk-adjusted quality measures. A drifting model can make a hospital’s outcomes look better than they are relative to an accurate benchmark.

The VA studies tested models against historical data. A July 2026 study in PLOS Digital Health retroactively examined four AI systems that had been used in clinical workflows at one large healthcare organization, including EHR-based risk prediction models and a diagnostic support tool. The models were not intentionally modified during the study period, yet all four showed degradation. Calibration deterioration often preceded declines in discrimination, and operational signals such as missing inputs and delayed data provided early warning.

Why it matters: If you wait for readmissions, deaths or confirmed diagnoses to reveal a failing model, you'll learn too late. The first signals are already in your data. Is anyone watching them?

3. Fairness drift: When the average looks fine and some patients don’t

A model that performs equitably at launch can develop disparities without changes to the model itself. A 2025 study in the Journal of the American Medical Informatics Association tracked VA surgical models across nearly 1.74 million cases, measuring performance by race and sex every quarter from 2014 to 2023. Fairness was defined as the performance gap between patient groups, and that gap shifted over time in both the original and retrained models.  

In practice, a model that works equally well for men and women at launch can, years later, be noticeably less reliable for one of them while its overall performance looks fine. Retraining closed the gap in some cases and widened it in others.

Why it matters: A bias review before go-live cannot establish that performance will remain equitable. This research shows that disparities can change over time and that model updating alone won’t reliably resolve them. Organizations that monitor only overall performance can miss widening gaps between patient groups. For health systems with equity commitments, that’s an accountability gap.

4. When no one can prove what an algorithm is actually doing

In November 2023, Medicare Advantage members and their families filed a class action, Estate of Lokken v. UnitedHealth Group, alleging their post-acute care coverage was cut off because of reliance on nH Predict, an AI tool from UnitedHealth subsidiary naviHealth. It isn’t a textbook case of statistical drift. It’s a warning about use drift: an algorithm whose influence on decisions allegedly grew beyond what was promised.

Before UnitedHealth acquired naviHealth in 2020, nH Predict was promoted as a tool to help clinicians build personalized post-acute care plans. For 2023, according to STAT, naviHealth set a target to keep patients’ rehab stays within 1% of the days the algorithm projected.

According to the complaint, 91-year-old Gene Lokken fractured his leg and ankle and later received approximately 19 days of covered post-acute rehabilitation before coverage was terminated. The complaint describes subsequent out-of-pocket costs of $12,000 to $14,000 per month from July 2022 until his death in July 2023. The lawsuit alleges nH Predict has a 90% error rate, based on the share of denials reversed on appeal.

The shift showed up in the aggregate numbers, too. A 2024 Senate investigation found that UnitedHealth's prior authorization denial rate for post-acute care more than doubled, from 10.9% in 2020 to 22.7% in 2022, after naviHealth began managing those decisions in 2019.

A March 2026 discovery order required UnitedHealthcare to produce records on the tool’s development, implementation, and oversight, including whether it was designed to supplant physician decision-making. UnitedHealth disputes the allegations, and Optum has said physicians, not AI, make coverage decisions.

Why it matters: Whatever the court decides, health system leaders should be prepared to answer a similar question. Can you document what your algorithm produced, how its outputs influenced decisions and whether that use matched its approved purpose? Validation results alone do not answer those questions.

Why is AI drift a leadership problem?

Because the consequences of AI drift, from clinical harm to inequity to legal exposure, affect the people and organizations using AI. Expectations for evidence are expanding beyond initial validation to ongoing performance. Leadership must fund and support the oversight needed to provide that evidence.

Many health systems aren’t there yet. According to a September 2025 federal data brief, 79% of hospitals using predictive AI reported some post-deployment monitoring in 2024, but only 58% did so for all or most of their models. That means roughly four in ten couldn’t say they were monitoring most of the AI they’d deployed.

Scrutiny is increasing. The FDA has asked for input on managing performance drift in AI-enabled medical devices. The nH Predict litigation has brought discovery into algorithmic oversight. In July 2026, senators pressed UnitedHealth, Humana, and CVS again over AI-driven coverage decisions. The question is no longer whether health systems will be asked to prove their AI still works. It’s when, and by whom.

6 questions every health system leader should be able to answer

Continuous AI assurance means giving each deployed model an appropriate baseline, ongoing measurement, clear triggers for action and a named owner. Leaders need clear answers to these six questions, with monitoring tailored to each model’s purpose and clinical risk.

  1. What is this model’s baseline on our patients? Document relevant performance measures, including calibration for risk prediction models and subgroup results, at go-live. Drift assessment requires a reference point.
  1. Are we tracking calibration where the model estimates risk? Ranking metrics can hold steady while risk estimates go wrong, as both the VA and PLOS research show.
  1. Which early signals are we watching? Alert volume, override rates, missing inputs and data delays can prompt investigation while outcome data are still pending. Use these signals alongside outcome-based evaluation.
  1. Is performance holding across relevant patient groups? Assess results by groups such as race, sex, age, payer and site where appropriate, accounting for sample size and uncertainty so aggregate results do not obscure meaningful differences.
  1. Who can turn it off, and what triggers that? Define in advance which thresholds prompt review, recalibration, or deactivation, and who has the authority to act, and how care will continue safely if the tool is paused.
  1. Do we know how outputs are actually used, and when vendors change the model? Track how recommendations influence decisions. Review vendor model updates and AI features added to tools you already approved.

Final thoughts

Healthcare has spent years asking whether AI works. The next question is whether it still works, today, for the patients it serves. AI drift means that answer can change without anyone noticing.  

Vitea’s AI governance platform for healthcare supports ongoing oversight through continuous AI testing and monitoring.

Drift is often silent. Your governance shouldn’t be. Contact our team for an in-depth look at Vitea, the leading platform for AI governance in healthcare.

Suggested for You

Inspired by what you’ve recently viewed.

Bring AI under control
without slowing innovation.
We're here to help you innovate and transform
Discover every AI in use, including shadow AI
Enforce 100+ out-of-the-box policies in real time
Stop risky AI activity before sensitive data is exposed
Continuously monitor AI performance and prove governance on demand
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.