AI Is in 90% of Health Systems. 56% Are Flying Blind After Go-Live, New Report Finds

Vitea Newsroom
Editorial team
Aug 18, 2026
5 minutes
Editorial team
Physician with tablet showing AI

A new UPMC–KLAS study puts hard numbers to a risk technology leaders have sensed for months: nearly every health system is running third-party AI, but far fewer have built the healthcare AI governance, validation, and oversight infrastructure to prove it is safe.

For most of the past two years, the healthcare AI conversation has been about promise. A new report from the Center for Connected Medicine at UPMC and KLAS Research — drawn from interviews with more than two dozen health system leaders — moves it toward reckoning. The headline numbers:

  • More than 90% of health systems now deploy third-party AI. AI adoption is effectively universal; the experimentation phase is over.
  • 92% say they test AI tools before deployment — but only 44% have a dedicated environment to do it in. The rest lean on vendor-reported results, limited pilots, or informal review.
  • 63% describe their AI strategy as “developing” or “ad hoc,” rather than established or advanced.
  • There is no consensus on how to measure whether AI is actually working once an AI tool live.
  • Clinical documentation (52%) is the most common use case, followed by revenue cycle, coding, and billing (36%), which means the heaviest concentration of AI sits closest to the patient record.

Together, these data points describe a single condition: AI deployment has scaled, and AI governance has not.

What the numbers actually mean

For CIOs and CISOs, the value of this data is in the gaps it exposes — not the adoption rate, which surprises no one, but the distance between “we deployed it” and “we can prove it is safe.”

“Implementation is only the first step,” Rob Bart, MD, chief medical information officer at UPMC, told researchers. The industry, he noted, is now turning to the governance structures, testing capabilities, and organizational strategies that turn deployment into measurable value.

“Testing” and “validation” are not the same thing

The most revealing pair in the study is the gap between 92% and 44%. Almost everyone tests; fewer than half have a dedicated platform to do it well. In practice, that means most “testing” is a vendor demo, a short pilot, or a review of performance metrics the vendor supplied... not independent, reproducible validation against the health system’s own patients, workflows, and edge cases.

A model that performs well on a vendor’s benchmark population can behave very differently on yours. Without a controlled environment to test that, a CISO cannot verify the claims in the contract, and a CIO cannot demonstrate to the board or a regulator that the tool is safe in local conditions.  

“Ad hoc” governance does not survive production

When 63% of leaders call their strategy developing or ad hoc, they are describing healthcare AI governance that is largely point-in-time: an approval at go-live, a policy in a static PDF, a signed attestation. That model assumes AI behaves the same on day 300 as it did on day one.

It does not. Models drift as local data shifts. Vendors push silent updates that change behavior without notice. Users find off-label uses no one anticipated. And a single hallucinated detail can be signed into the record, then repeat and compound across future encounters. The risk in healthcare AI lives after go-live, which is precisely the moment ad hoc governance stops watching. We’ve written before about why balancing innovation with governance is a post-deployment discipline, not a launch checklist.

The biggest AI footprint is also the highest-stakes one

That clinical documentation tops the list at 52% is not a neutral fact. Ambient scribes and note-generation tools sit at the intersection of patient safety, consent, and legal exposure, as the recent ambient listening lawsuits against Sharp and Sutter have made clear. When the largest category of deployed AI is also the one writing into the medical record, informal validation is a patient-safety gap.

If you cannot measure it, you cannot defend it

The absence of consensus on success metrics is the quiet finding with the loudest consequences. Without agreed measures, health systems cannot prove ROI to finance, cannot detect performance degradation before it reaches patients, and cannot answer a regulator, a plaintiff’s attorney, or a board asking a simple question: How do you know this is working?

How to govern AI in healthcare – What leaders need to know

None of this argues for slowing adoption, but rather for closing the distance between adoption and assurance. The following six moves matter most:

  • Inventory everything — including what you didn’t approve. You cannot govern what you cannot see. Build a continuously updated map of every AI system in use, including vendor tools with AI quietly embedded and the shadow AI running on personal accounts.
  • Validate on your own data, not the vendor’s. Establish (or partner for) an environment to test tools against your population and workflows before rollout. Vendor benchmarks are a starting point, not evidence.
  • Replace point-in-time approval with continuous monitoring. Treat go-live as the beginning of oversight, not the end. Watch for drift, silent model updates, and out-of-policy behavior in production.
  • Define success and failure before you deploy. Agree on the metrics that mean “working,” “degrading,” and “must be pulled.” Ambiguity at deployment becomes a crisis at incident time.
  • Make policy enforceable, not aspirational. Healthcare AI governance that lives in a PDF depends on everyone reading and remembering it. Governance that lives in the network can actually intervene when a tool steps out of bounds.

The barriers the study names — limited resources, time, and specialized talent — are real, and they are exactly why 56% of systems have not built a validation platform on their own. Most cannot staff a governance function from scratch. The practical path for many is to buy the capability rather than build it, an approach we’ve detailed for teams trying to move from pilot to production without stalling.

Closing the gap on healthcare AI governance

The UPMC–KLAS data describes three gaps: health systems cannot fully see the AI running across their environment, cannot validate it against their own reality, and cannot watch it after it goes live. Vitea was built to close exactly those three.

The industry has proven it can adopt AI. The next phase — the one this research makes unavoidable — is proving it can govern it. If that is the gap is something your organization is hoping to fill this year, we’re happy to talk.

Suggested for You

Inspired by what you’ve recently viewed.

Bring AI under control
without slowing innovation.
We're here to help you innovate and transform
Discover every AI in use, including shadow AI
Enforce 100+ out-of-the-box policies in real time
Stop risky AI activity before sensitive data is exposed
Continuously monitor AI performance and prove governance on demand
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.