

This week in healthcare AI: a federal program began using AI to decide Medicare coverage in six states, two studies documented what AI systems leave out of clinical decisions and records, the AMA and a JAMA analysis staked out opposing positions on the physician's role, and the FDA opened public comment on how it may regulate generative-AI medical devices.
Read on for the five healthcare AI news stories your team should have on its radar this week.
Under a federal pilot called the Wasteful and Inappropriate Service Reduction (WISeR) Model, AI now screens Medicare prior-authorization requests for 15 procedures — including epidural steroid injections and knee arthroscopy — in Washington, Arizona, New Jersey, Ohio, Oklahoma and Texas.
According to BBC Science Focus, physicians and patients report delays that have stretched treatment timelines from roughly a day to weeks, along with repeat denials. CMS says decisions should be returned within 72 hours and that any denial must be reviewed by a human clinician before it is final.
Physicians interviewed described the models as difficult to interrogate, and noted that vendor compensation is tied in part to the savings generated from avoided procedures, a structure CMS says is offset by adjustments based on customer satisfaction and overturned-denial rates. A U.S. senator's snapshot report cited average waits of 15 to 20 days in one state.
Why it matters: The pilot places AI at the center of consequential coverage decisions without a clear mechanism for clinicians to see or challenge the model's reasoning. It is an early, large-scale test of what happens when AI acts on patients faster than the transparency, explainability, and audit mechanisms around it can keep up.
Source: BBC Science Focus
Researchers at Stanford, Harvard Medical School and other institutions released a benchmark called NOHARM that evaluated 24 AI tools — 20 general-purpose large language models and four clinical tools — across 1,100 case-based tasks spanning 10 specialties, as reported by Telehealth.org.
Specialized clinical systems produced the lowest rates of potentially severe errors (ranging from about 2.9% to 5.4%), while general-purpose models ranged from roughly 8.9% to 24.6%.
The key finding: more than 80% of severe errors were omissions, clinically important recommendations the model failed to make, rather than overtly wrong ones. A companion randomized study of 101 physicians found AI assistance improved performance, but AI-assisted physicians still missed useful recommendations the AI had surfaced. The work is currently a preprint and has not yet completed peer review.
Why it matters: Omissions are harder for a reviewer to catch than obvious mistakes, because spotting a missing test or follow-up requires already knowing it should be there. The study indicates that strong medical-knowledge scores do not guarantee safe recommendations, and that human review alone is an incomplete safeguard, pointing toward pre-deployment validation and continuous monitoring.
Source: Telehealth.org (reporting on the NOHARM study)
A TechTarget feature examined how ambient AI scribes, now used in an estimated 30% of physician practices, handle social determinants of health.
Experts including Columbia University's Maxim Topaz noted that social needs (i.e. housing instability, food insecurity, transportation) are often raised quietly, indirectly, or in a second language, and are frequently dropped from the generated note. Because most systems discard the audio and transcript shortly after the note is created, there is typically no audit trail to verify what was captured.
Research also shows transcription accuracy is lower for patients with accents or limited English proficiency. The article notes a counter-view: Kristine Lee of The Permanente Medical Group said scribes can be configured to capture and flag social needs, and cited data showing scribes freed up significant clinician time. Recommended safeguards include structured screening steps, retained audit trails, and mandatory clinician review of notes.
Why it matters: The information most likely to be lost tends to concern the patients already at highest risk, which can widen documentation and equity gaps rather than close them. Without a retained record of what was said versus charted, organizations cannot measure or correct the gap.
Source: TechTarget (xtelligent Health IT)
The American Medical Association, in partnership with the Digital Medicine Society, released a framework arguing that physicians must remain central to patient care and that the profession's core responsibilities — clinical judgment, human connection, stewardship of technology — remain stable even as tasks change, according to Axios.
The framework was published the same week as a JAMA analysis, co-authored by former White House health policy adviser Zeke Emanuel, arguing that autonomous AI will likely be ready for many real-world cognitive medical tasks by around 2030 and that keeping humans in the loop could, in some workflows, worsen outcomes. AMA CEO John Whyte pushed back, distinguishing performing a diagnostic task from the broader practice of medicine. The piece also flagged unresolved questions around liability and reimbursement for AI-delivered care.
Why it matters: With prominent voices openly divided on where the boundary between human and machine judgment belongs, there is no industry consensus for health systems to defer to. Each organization is left to define — and operationalize — where AI may act and where a clinician is required.
Source: Axios
The FDA's Digital Health Center of Excellence, within the Center for Devices and Radiological Health, issued a discussion paper on August 18 outlining a potential approach to regulating generative-AI-enabled medical devices, per Healio and the FDA.
The framework proposes a two-axis risk assessment and a competency-based premarket evaluation — modeled loosely on how clinicians are trained and credentialed — pairing non-clinical benchmarking with clinical confirmation, followed by risk-proportionate postmarket monitoring. It also addresses foundation models and agentic AI systems specifically. The agency is accepting public comment through October 19 under docket FDA-2026-N-7874. The FDA emphasized that the paper is a framework for discussion, not draft guidance or a proposed rule.
Why it matters: The paper signals the direction of federal thinking — validate capability before deployment, monitor performance after — but most freestanding clinical AI in use today falls outside the device pathway, and no rule is yet in place. In the near term, oversight remains an internal responsibility for health systems.
Taken together, the week's stories point to a widening distance between how quickly AI is being given consequential roles in care and how quickly oversight is catching up. The recurring theme focuses on the quieter gaps of AI in healthcare — decisions that can't be explained, recommendations that are never made, records that are never kept, and boundaries that haven't been defined.
For health systems, the practical takeaways are consistent across all five: know where AI is operating, validate tools before they reach patients, monitor their behavior in production, and set clear boundaries for where AI can and cannot act.
Those are the disciplines Vitea helps health systems put in place. If your team is working through where oversight is thin, we're happy to talk.
Follow Vitea on LinkedIn for more of the latest news and views on AI governance in healthcare, including our weekly roundup of the stories healthcare leaders need to know.