logo making sense

Latest posts

Explore our categories

How Should Mid-Market Healthcare Firms Choose an AI Implementation Partner?

AI is already in daily clinical use. The open question for mid-market providers is who builds it, and the criteria that answer it are not the ones used to evaluate a software vendor.

Sep 17, 2026

AI adoption is already well underway in U.S. healthcare. According to the American Medical Association's 2026 physician survey, 81% of U.S. physicians now use these tools at work, up from 38% in 2023. What has not changed is how mid-market healthcare organizations evaluate an AI implementation partner. The same scorecard that works for hiring someone to develop a patient portal still gets used here: portfolio, stack, references, price, delivery timeline. Every vendor clears it, which is why it separates no one.

physician-ai-use-2023-vs-2026.png

 

That evaluation needs to account for factors that go beyond software delivery: model performance, clinical adoption, PHI, regulation, and ongoing monitoring all affect whether a pilot can scale.

Those risks often become visible only after the initial pilot looks successful.

Take a three-site clinic group evaluating an ambient documentation system. The physician talks with the patient, the system listens, drafts the visit summary, and files it in the right section of the chart without anyone typing. The demo goes well, and the clinical committee approves a pilot at one location.

Nine months later, it is still running in that one location because several things got in the way: access to the EHR integration environment took longer than expected, physicians at the other two sites were hesitant to trust AI-generated summaries, no one had defined an acceptable error rate before the pilot began, and the vendor had estimated development time without accounting for what it would take to scale the system.

Several of those issues can be identified before the project begins, if the evaluation includes the right questions. 

Whether to build an internal team or bring in outside help is a separate decision, covered in this piece. What follows assumes it has been made. 

Why AI projects cannot be measured like software projects

A polished demo and a development estimate can show that a partner is ready to build a pilot. For an AI system, the evaluation also needs to cover how the model will be judged once it is in use: what level of error is acceptable, how performance will be measured, and who will review the output.

That difference comes from how traditional software and AI systems are evaluated:

  • Traditional software is evaluated against a specification. A feature either works as expected or it does not. If it does not, it can be fixed.
  • An AI model is evaluated against an acceptable level of error. It gets most cases right and still make mistakes in others because it generates outputs based on the patterns it learned and what it determines is the most likely response. That is why the evaluation needs to define what level of error is acceptable for that specific use case.
spec-versus-acceptable-error-rate.png

 

Better data narrows the gap, and so does limiting the system to cases where being wrong costs little, or keeping a person in the loop before any text reaches the chart. What none of that does is bring the miss rate to zero, and if the project requires zero errors, it will never move into production.

How working well gets defined

Someone has to decide, before any code is written, how anyone will know the system is doing its job. Five questions worth asking during the evaluation: 

  1. What cases will you use to test the model’s performance? This should be defined before launch. Waiting to see how the model performs in production is not a reliable evaluation method.
  2. Who will review those cases and define the expected results? This requires time from clinical staff, and that time needs to be included in the project plan and estimate.
  3. What is the acceptance threshold and who signs off on it?
  4. How does measurement work once the system is live, how often, and who reads the report?
  5. What happens in cases where the model is unsure? Letting the system say it does not know and route the case to a person is a design decision, not a limitation.

How PHI is handled when using an external AI model

When an AI system processes clinical information, PHI may be sent to an external model provider. That makes it important to understand what happens to the data once it leaves the healthcare organization’s systems.

HIPAA compliance alone does not answer that question. During the evaluation, organizations should understand:

  • Who is covered by the BAA. Does it include the AI model provider, or only the implementation partner?
  • How long the provider keeps the data. The contract should clarify whether data is retained and how long it remains in logs.
  • Whether the data can be used for model training. This should be explicitly addressed in the provider’s contract.
  • What data is de-identified and what activity is logged. The organization should know what information is removed before it reaches the model and what records are kept about its use.

The same AMA survey found that 86% of physicians consider data privacy assurances important for AI adoption. It also found that clear liability frameworks are the regulatory action physicians most often say would increase their trust.

When an AI tool becomes a regulated system 

An assistant that drafts a visit summary is not regulated. A system that suggests a course of treatment is, because it falls under clinical decision support rules, and that changes what the partner has to document. 

The ONC's HTI-1 rule requires certified health technology to be transparent about the AI models it uses in clinical decision-making. For predictive models, that means documenting 31 specific attributes, including what data trained the model, how it was tested, which patient populations it was evaluated on, and whether it was validated externally. The full requirements are available at healthit.gov.

These disclosures support the FAVES framework, which ONC uses to evaluate models involved in clinical decision-making. FAVES focuses on five areas: whether the model is fair across different patient groups, appropriate for the population where it will be used, which is rarely the training population , valid when tested on new data, effective in real-world use, and safe when errors occur.

when-clinical-ai-becomes-regulated.png

 

A CEO does not need to know all 31 attributes in detail. During partner evaluation, check whether the team understands these requirements, can explain which ones apply to the proposed use case, and knows how to address them before development begins. That is especially important when the AI system may influence clinical decisions.

What data an AI system needs from your EHR 

Traditional software integrations usually rely on structured EHR data, such as HL7 v2 interfaces, FHIR resources, and defined data fields.

AI systems often need information that was never structured in the first place, , which at most healthcare organizations is the larger share of the clinical record. In healthcare organizations, that can include free-text clinical notes, scanned referrals, faxes, and lab results stored as attachments. Accessing and preparing that information can become one of the biggest challenges in an AI implementation.

That is why evaluating a partner requires more than confirming that they can connect to the EHR. Ask what they did the last time a client had no access to a test environment, how they handled inconsistent document formats, and how long it actually took to get permissions granted. Those issues can have a direct impact on the project timeline.

Who is responsible for the AI modelafter launch?

An AI implementation does not end when the system goes live. Performance can change without anyone touching the code, as clinical workflows, patient populations, and the underlying model evolve. That means someone needs to monitor how the system is performing and decide when adjustments are needed.

ai-system-cost-after-launch-hold-period.png

Ongoing support also includes the people using the system. Clinical teams need to understand when they can rely on the output, when they should review it more closely, and how to report problems. This is why AI adoption and enablement should be part of the implementation plan, rather than limited to training at launch.

When evaluating healthcare technology partners, ask what happens after launch. Who monitors performance? How often is it reviewed? Who is responsible if the model needs to be updated or retrained? And which of those responsibilities are included in the contract?

How to measure ROI on a healthcare AI pilot

To measure the ROI of an AI implementation, healthcare organizations need a baseline before the pilot begins. Without it, there is no reliable way to show how much the process improved or build a business case for further investment.

The right metric depends on the use case. It could be documentation hours per clinician per week, first-pass denial rate, days in accounts receivable, or average prior authorization time. Defining that baseline during Discovery makes it possible to compare results once the system is in use.

Revenue cycle workflows are often a practical place to start because they already have measurable operational and financial metrics, with less regulatory complexity than many clinical use cases. You can read more about how healthcare organizations are using technology to drive impact and about our approach to intelligent workflow automation

What PE-backed healthcare companies should require from an AI partner 

For PE-backed healthcare companies, two additional factors matter when evaluating an AI implementation: ongoing cost and documentation.

First, the cost of the system continues well beyond the initial implementation. Monitoring, model updates, and retraining create recurring expenses that run through the entire hold period and belong in the investment case from the start. 

Second, the system needs to be well documented. When a company is sold, a buyer will want to understand how the model was tested, what data was used, who is responsible for it, and what ongoing costs are required to operate it. Gaps in that documentation can become an issue during a technology due diligence review.

How to evaluate an AI implementation partner 

Healthcare organizations that get AI into production tend to pick someone who understands the process before proposing a tool, and who can clearly explain where AI can add value, where human review is still needed, and what risks need to be addressed.

The evaluation looks more like a joint discovery process than an RFP. Working through the problem together helps define the requirements, risks, and priorities before development begins.

Collaborative Workflow Planning Meeting.png

 

The same approach shaped our work with Opya, an autism treatment clinic in the United States, well before AI was part of the conversation. The organization operated almost entirely offline. Patient information sat across separate systems, and coordination among therapists, clinicians, and families depended on manual work. 

The first step was mapping how each of those groups actually worked, before deciding what to build. That led to a suite of three HIPAA-compliant applications that centralized treatment history and improved coordination between families and the clinical team. Understanding how people actually worked, before choosing the technology is what made the rest possible.

That Discovery process becomes even more important in an AI implementation, where decisions about data, acceptable error rates, human oversight, adoption, and ongoing monitoring need to be made early.

We have spent more than twenty years working with mid-market companies in the United States, including healthcare organizations operating in highly regulated environments.

For organizations evaluating healthcare technology partners for an AI project, our healthcare page covers how we approach the sector. You can also book a conversation with our team.


Sep 17, 2026

Say Hello!

Get the latest news and updates
logo footer making sense

|

Technology Fueling Growth