How to Evaluate AI Capability During Technical Due Diligence
A working demo says little about provider dependency, data exposure, or what the system costs to run at scale. Diligence needs evidence for each.
Sep 29, 2026
A target describes its product as AI-powered, the demo runs cleanly, and the deal team is working against a compressed timeline before the investment committee meets. Everything on screen works.
What the buyer still does not know is how much of it the company actually built, how much belongs to a provider whose pricing can move, and what it would cost to keep the whole thing running a year after close. Technical due diligence answers those questions with evidence. The AI side of a business is often where that evidence is hardest to assemble, which is worth knowing before the clock starts.
Why AI findings are harder to verify
Most of what a technical review scores leaves a trail someone can hand over: a document, a repository, a running process. The AI criterion often does not.
Every finding in a review carries two separate judgments: the score, and how solid the evidence behind it is. The evidence side has three levels:
- Observed. Someone saw the code, the documentation, or the process running.
- Self-reported. The company described how it works, and nothing contradicted the description.
- Unknown. There was not enough information to conclude either way.
Infrastructure and release practices are often easier to verify directly, because they leave artifacts behind whether anyone intended them to or not: configuration files, deployment history, commit logs.

AI adoption behaves differently. The documents that would settle it are a formal evaluation suite, drift and bias metrics, and a record of what changed in the prompts and when. Those tend to get built after a system has been in production long enough to cause someone a problem. A product whose AI features are still relatively new often has not reached that point, so the strongest available answer is the team's own account of how the system behaves. That account is useful, but it carries a different level of confidence than evidence someone can review independently.
An M&A technology assessment that marks an AI finding as self-reported has done its job: it tells the investment committee how much to trust it.
What the review examines inside the AI criterion
A technical review examines four things inside the AI criterion: the quality of the data, how that data moves, what is already serving customers, and how the company handles bias and harmful output.
Data quality
Whether the datasets feeding the system are consistent, where they come from, and who maintains them. A model that performs well on curated internal data can stumble once real volume and unusual cases arrive.
Data flows
How information moves through the product, which parts leave the company, and under what agreement. Tracing these flows can surface privacy exposure, particularly when customer or proprietary data is sent to external model providers.

The AI and ML products already in production
What is serving customers today, as opposed to what is still in testing or in a slide. The gap between the two is what a review can size, and the release history shows it.
Ethical AI and bias mitigation
How the company handles outputs that affect people: what gets measured, what triggers a review, and who signs off when a model's behavior changes. The concrete ask is a bias evaluation with results, and a documented threshold for acting on them. What draws attention is a product making decisions about a customer's price, eligibility or risk level with no record of anyone having checked the outcome by segment.
How much the product depends on an outside provider
Products described as AI-powered cover a wide range. At one end, a thin layer on top of someone else's model. At the other sits a system the company built and runs itself, with its own data and its own testing, and plenty of solid businesses operate somewhere between. What separates them for a buyer is margin, switching cost, and what happens when the provider changes terms.
Provider disclosure has limits of its own. Stanford HAI's 2026 AI Index reports that the Foundation Model Transparency Index average fell from 58 to 40 out of 100 between 2024 and 2025. A buyer assessing a product built on those models is working with what the provider chose to publish.
Pricing the dependency means looking at inference cost at current volume and under realistic growth scenarios, not just today's number. On Doppler's ECO IA, a customer-facing assistant running on Anthropic's Claude, prompt caching was designed in from the start to keep inference costs predictable as volume grew: a 97% cache hit rate at USD 0.066 per interaction. Those are the kinds of figures diligence should be able to surface when inference cost is material to product economics.
It also means reading what the contract says about pricing changes, and checking whether the architecture allows a provider swap without a rewrite. AI cost governance covers the spend side in more depth.
What evidence moves a finding from self-reported to observed
The request list is short, and the answers are informative whether or not the documents exist:
| What the review is establishing | Evidence to request | What the answer tells you |
| AI performance is actually measured | An evaluation suite, with results from more than one run | How current the picture is: today's behavior, or the behavior at launch |
| Someone is watching the system after launch | Monitoring for drift, bias and hallucination rate, plus the threshold that triggers action | Whether a degradation gets caught internally or by a customer |
| Changes are recorded | Prompt versioning, with a record of what changed and why | How quickly a bad change can be undone, and who decides |
| The system has been tested against misuse | Results from adversarial and prompt injection testing | Whether the team found weaknesses, or only looked for them |
| Sensitive data stays inside the company | PII detection and redaction before data leaves the company | The size of the privacy exposure, and whether it was designed for or discovered late |
| AI-generated code gets extra review | Review controls that apply specifically to AI-generated code | Whether review practice kept pace with how much code the team now generates |
| More than one person can run it | The name of the person who owns the system six months after go-live | How much of the capability walks out if one person leaves |
Not all of it carries the same weight. A bias evaluation matters far more for a model that sets a customer's price or eligibility than for an internal search assistant, and a review that treats both the same is not reading the product.
Incident tracking has become more relevant as deployment has widened. The same AI Index reports 362 incidents logged in the AI Incident Database during 2025, against an annual figure that stayed under 100 until 2022. A company keeping an internal record of its own near-misses has something to show you. AI governance sets out what that structure looks like when it is working.
Turning the finding into the first year's plan
Diligence identifies the gap and the confidence behind the finding. The next step is translating that gap into a remediation plan with an owner, a cost, and a timeline. Those priorities then enter the value creation plan according to their effect on risk, scalability, margin, or product growth.

Building an evaluation suite and closing a data privacy exposure both carry a cost, a timeline, and someone who has to own the work. A missing drift dashboard is an engineering task. A system that only one person understands is a different problem, and it usually determines whether the acquired product keeps improving after close or simply keeps running. AI adoption and enablement is the work of closing that kind of gap, and it tends to start from the same findings the review produced.
The sequencing matters for the hold period. Technology value creation from diligence to exit covers how those findings carry forward, and the tech stack signals buyers notice covers the same review from the sell side.
Scoping the AI side of a review
In a PE technical assessment, AI earns its own line of inquiry. It gets its own interview, its own audience inside the target (usually the CTO alongside whoever leads data science), and its own set of documents to request.
Making Sense runs independent technical reviews for investors and acquirers. The AI portion comes back as a scored criterion with a confidence level on each finding, the evidence behind it, and what closing each gap would cost in the first year. Talk to our technical due diligence team about a specific target, or see how the private equity practice supports the full investment cycle.
Frequently asked questions
What is the difference between AI adoption and AI readiness?
Adoption describes what a company is already running: models in production, tools the team uses, products customers touch. Readiness describes whether the foundations underneath can support more of it, mainly data quality, data flows, and governance. A company can score well on one and poorly on the other, and the combination tells a buyer whether growth in this area needs new investment or just more of the same.
Can a buyer rely on the target's own AI benchmark results?
They are a starting point. What settles it is whether the evaluation can be reproduced, whether it ran against data the model had not already seen, and when it last ran. A benchmark from the quarter the feature launched says little about how the system performs on today's traffic. Ask for the harness and the date alongside the number.
What can realistically be assessed during a compressed diligence window?
Enough to support a decision, provided the scope says so upfront. A short window is enough to map what is in production, trace the data flows, price the provider dependency, and establish which findings rest on documents and which rest on interviews. It is not enough to independently validate model performance claims. A review that says so is more useful than one that pretends otherwise.
What are the most common AI-related findings in technical due diligence?
The recurring ones are absence of formal evaluation and drift monitoring, privacy handling for data sent to external providers, dependency of the product core on a single model vendor, and quality controls for AI-generated code that have not kept pace with how much of it the team now produces.
Sep 29, 2026