In March 2026, Hologic recalled 1,200 monitors used with its AI-enabled cervical cancer screening system after some sites changed manufacturer-installed display settings. The AI itself had not stopped working. The problem was that the system was no longer operating within the configuration in which its clinical performance had been validated.
That distinction matters. A model can remain technically functional while the clinical environment around it changes enough to affect whether its outputs should still be trusted.
For years, the healthcare AI conversation has focused on whether a model works: whether it is accurate enough, integrates into workflow, improves an outcome, or convinces clinicians to use it. Once that model moves from a closely watched pilot into routine care, a different question becomes more important.
How does an organization know that it is still working the way everyone thinks it is?
This is not only a model-monitoring problem. It is also a problem of distributed observability. The vendor sees one set of signals, the hospital sees another, clinicians see another, and regulators may see only the subset of events that cross a reporting threshold. The same fragmentation that makes AI difficult for health systems to buy can also make it difficult to recognize when a deployed system is beginning to behave differently.
The model can stay the same while the system changes
Healthcare has already seen how much AI performance can depend on the environment in which a model is used.
One of the best-known examples is Epic’s Sepsis Model. Epic had reported substantially stronger performance than researchers later observed when they independently evaluated the model at the University of Michigan. In more than 38,000 hospitalizations, researchers found an AUROC of 0.63. At the threshold they studied, the model alerted on 18% of hospitalizations while failing to identify 67% of patients who eventually developed sepsis.
The Michigan results showed something important: performance established in one context cannot simply be assumed to transfer unchanged into another.
FDA is studying this problem directly in AI-enabled medical devices. The agency has noted that patient populations, acquisition systems, clinical protocols, and data distributions can all change over time and across sites, affecting performance even when the underlying model has not changed.
Often, the cause is ordinary operational change. An EHR update alters an input. A patient population shifts. A workflow changes. A local configuration affects how an output is displayed.
The model can be exactly the same on Monday as it was on Friday while the system in which it operates has become meaningfully different.
The Hologic recall makes that point concrete. The affected monitors were part of an FDA-cleared digital cervical cytology system using the Genius Cervical AI algorithm. The field correction was triggered because users had modified display and calibration settings, placing some systems outside their validated configuration.
Calling this simply an “AI failure” would be misleading. The model, software, hardware configuration, and human use all contributed to whether the system was operating as intended.
Postmarket monitoring therefore has to look beyond the model itself.
Pilots are often the period when AI is watched most closely
Pilots tend to involve a small number of users, an engaged clinical champion, close vendor participation, and a team actively trying to determine whether the product works.
Production is different.
The product expands into additional sites. New clinicians use it. Workflows change. Software is upgraded. Local workarounds emerge. The people who originally championed the project move on.
A pilot establishes whether a technology can create value under a particular set of conditions. Production requires an organization to recognize when those conditions have changed enough that the original evidence may no longer tell the whole story.
That is difficult because responsibility is distributed.
A vendor may see model telemetry but have little visibility into clinician behavior. IT may know that an integration is functioning without knowing whether predictions remain clinically useful. Quality teams may investigate incidents without seeing model-performance signals. Clinicians may notice that recommendations feel less reliable but assume each case is isolated.
An organization can possess nearly every signal required to recognize an emerging problem while still lacking a mechanism for connecting them.
Adverse-event reporting solves only part of the problem
For regulated medical devices, healthcare already has a mechanism for turning failures into collective learning: postmarket surveillance and adverse-event reporting.
Historically, medical-device reports involving deaths, serious injuries, and malfunctions flowed into FDA systems including MAUDE. In 2026, FDA began moving reporting into its Adverse Event Monitoring System, or AEMS, intended to improve standardization and surveillance across regulated products.
AEMS can improve how reports are collected and analyzed, but AI raises another problem: whether the reports describe the failure in enough detail to be useful.
Suppose several hospitals experience clinically important problems involving similar AI systems.
At one institution, there is a software defect.
At another, the patient population has shifted enough to affect performance.
At a third, an EHR change alters the inputs feeding the model.
At a fourth, clinicians have gradually begun relying on the recommendation differently.
All could eventually surface as some version of an incorrect recommendation.
But “incorrect result” describes what was observed. It does not explain why it happened.
If very different mechanisms collapse into the same broad incident category, regulators, manufacturers, and health systems can accumulate reports without gaining enough information to recognize the pattern behind them.
Much of healthcare AI will never appear in an FDA device database
FDA adverse-event requirements apply to regulated medical devices. They do not cover the full range of AI being introduced into healthcare.
Health systems are also deploying predictive models, ambient documentation, administrative automation, generative AI, and other applications that may fall outside the medical-device framework depending on their intended use.
For FDA-regulated AI, the question is partly how postmarket infrastructure can evolve to describe AI-related incidents more meaningfully.
Outside that framework, the health system itself may be the primary surveillance system.
Internal quality processes, clinical governance, vendor oversight, model monitoring, and patient-safety reporting become the mechanisms through which emerging problems are detected.
This also means not every change in AI behavior is an adverse event. Drift is not automatically an adverse event. A clinician override is not necessarily evidence that anything is wrong. A shift in data distribution may be entirely benign.
These are better understood as signals that may warrant investigation.
And the warning signs and the adverse event are not the same thing.
The warning signs and the adverse event are not the same thing
Healthcare needs two connected layers of information.
The first is a signal layer: evidence that the circumstances around an AI system may be changing. That might include declining performance, shifts in patient populations or inputs, increasing clinician overrides, or changes in workflows and configurations.
None of these necessarily indicates patient harm. They tell us where to look.
The second is an incident layer, used when something clinically meaningful has occurred. That record should preserve enough context to reconstruct the system around the event.
At minimum, investigators should be able to answer:
Human interaction matters too: whether the system was misunderstood, used outside the intended workflow, or trusted in a way that changed clinical decision-making.
This is not a proposed FDA reporting form. It is the kind of context that makes an AI-related incident useful for learning.
A patient-safety report becomes more informative if investigators can determine that clinicians had been overriding recommendations more frequently for several weeks and that the event followed a workflow or configuration change. An apparently benign signal becomes more meaningful if several institutions later associate the same pattern with clinically significant failures.
The purpose is to make early warning signs useful before they become clinically consequential.
Monitoring the model is not enough
Software offers one advantage: it can generate continuous information about its operating environment, including inputs, outputs, version history, configurations, and user interactions.
But monitoring the model is not the same as monitoring the clinical system.
A manufacturer may detect a change in data without knowing that clinicians changed their workflow. A hospital may see a rise in incident reports without knowing that another customer is seeing the same pattern. A clinician may distrust a recommendation without knowing that a new model version went live days earlier.
The important information exists across organizational boundaries.
Postmarket safety therefore depends on a feedback loop between the people who build the AI, the institutions that operate it, the clinicians who use it, and, when the product is regulated, the agencies responsible for surveillance.
Without that loop, more monitoring simply creates more disconnected data.
Postmarket observability should be part of the product
Enterprise readiness should include post-deployment observability.
Today, a product is often considered ready for enterprise deployment when the evidence is persuasive, the integration works, the security review is complete, and regulatory obligations have been addressed.
Those are necessary conditions. They are not enough.
A mature AI product should also make it possible for the customer to understand how it is behaving after deployment and to reconstruct what happened when something changes.
Before scaling, the vendor and healthcare organization should know who owns ongoing performance, which signals each side monitors, what constitutes a meaningful change, how clinician concerns are aggregated, and how an incident can be traced back to the relevant model version, configuration, and operating environment.
These questions are usually grouped under governance, but they are really questions of operating ownership.
If ownership is unclear, the organization may not know who is responsible for acting on an emerging signal.
Postmarket observability should be part of the product.
A better database won’t help if we describe the wrong thing
FDA is already modernizing the infrastructure through which adverse-event information flows. AEMS is intended to make reports more standardized and easier to analyze across regulated products.
Better infrastructure could be paired with richer AI-specific context: what changed, which version was involved, how the issue was detected, and whether the same pattern is appearing elsewhere.
Several hospitals may discover that the same workflow change precedes degraded performance. Manufacturers may see a pattern across customers that no single institution can see locally. Regulators may identify a class of failures that cuts across otherwise unrelated products.
The point is to make the experience of one deployment informative to the next.
Healthcare may have to solve this before everyone else
The underlying problem will not remain confined to medicine.
A lending model can stay operational while economic conditions change its behavior. An AI agent can successfully execute individual actions while errors accumulate across a larger process. Autonomous systems can encounter environments unlike those represented during validation. Generative systems can continue producing plausible outputs even when those outputs are wrong.
Across these domains, technical uptime will become an increasingly poor proxy for trustworthy performance.
Healthcare is better positioned than many industries to address this because it already has decades of experience with adverse-event reporting, quality systems, clinical governance, and postmarket surveillance. Those systems were not designed for probabilistic, data-dependent software, but they provide a foundation.
Those systems now need to capture not only that something went wrong, but what changed in the model, environment, or clinical use before it happened.
So far, healthcare AI governance has concentrated heavily on the deployment decision: Is the model accurate enough? Is the data appropriate? Is it secure? Should we approve it?
Mature governance will also have to account for what happens months or years after that decision.
The hardest healthcare AI failure to detect may not be the system that crashes. It may be the one that keeps running and keeps producing plausible outputs while the evidence that something has changed remains scattered across clinicians, vendors, quality teams, and monitoring systems that were never designed to see the same problem.
Author : Arvita Tripati , MBA | Founder & Principal, Vahana Labs AI



