AI can read faster than a deal team. It cannot decide what the deal team should believe. That difference is the starting point for responsible AI in due diligence.
By early 2025, generative AI was already useful for searching large document sets, extracting clauses, summarizing recurring themes and drafting first-pass issue lists. The opportunity is real. So are the failure modes: plausible but unsupported answers, missed exceptions, loss of confidentiality, and reviewers who trust a polished summary more than the underlying evidence.
The first wave of value is practical, not autonomous
The most reliable use cases reduce repetitive work while keeping judgment with domain experts:
- classifying and routing incoming documents;
- extracting defined fields from contracts, policies and reports;
- identifying clauses that differ from an agreed standard;
- finding references to a customer, entity, jurisdiction or obligation across files;
- summarizing a topic with links back to the supporting passages;
- drafting diligence questions from missing or inconsistent evidence.
These tasks can compress the time between receiving information and asking the next good question. They do not remove the need to understand materiality, credibility, legal interpretation or commercial consequence.
Search, extraction and judgment are three different jobs
AI discussions often group them together, which hides the risk.
- Search asks where relevant information may be located.
- Extraction converts information into a defined field, issue or comparison.
- Judgment determines whether the evidence is complete, reliable and material to the transaction.
A system can perform well at search and still create a poor diligence conclusion. It may find the main contract but miss an amendment, extract a termination clause but ignore the revenue concentration behind it, or summarize management’s position without testing it against operational data.
The workflow therefore needs separate quality gates for each job rather than one generic “AI reviewed” status.
The evidence rule
No material AI-generated diligence finding should reach a decision forum without a visible source, a named human owner and an explicit assessment of what remains uncertain.
Five risks that deserve design controls
- Hallucination: the model produces a confident statement that is not supported by the documents.
- Omission: the answer looks complete but excludes an exception, attachment, image, table or later amendment.
- Context collapse: facts from different entities, periods or jurisdictions are combined incorrectly.
- Confidentiality leakage: deal data is exposed to an inappropriate model, user, log or downstream service.
- Automation bias: reviewers reduce their skepticism because the output is fluent, structured and fast.
NIST’s 2024 Generative AI Profile emphasizes that generative systems introduce risks across the full lifecycle and should be governed, measured and managed rather than treated as ordinary software output. In M&A, this is especially important because the information is incomplete by design and decisions are time-sensitive.
A better operating flow for AI-supported diligence
- Define the question. Specify scope, entity, period, document set and the decision the analysis will inform.
- Retrieve the evidence. Limit the system to approved sources and retain passage-level links.
- Extract into structure. Use agreed fields, issue categories and confidence indicators.
- Cross-check. Compare related documents, amendments, data tables and management responses.
- Escalate exceptions. Route low-confidence, conflicting or high-materiality findings to the right expert.
- Decide and record. A human owner assesses materiality and records the conclusion, rationale and downstream action.
The output is no longer a free-standing summary. It becomes one controlled step in a diligence process.
Human review must be designed, not assumed
“Human in the loop” is weak if the reviewer receives a 20-page AI report and a deadline to approve it. Effective review requires a usable interface and clear decision rights.
- Show the source passage beside the extracted finding.
- Highlight contradictions and missing evidence, not only conclusions.
- Require stronger review for high-value, irreversible or legally sensitive decisions.
- Make corrections visible so recurring errors can become tests.
- Sample apparently low-risk outputs to detect silent quality deterioration.
The reviewer should not repeat all the machine’s work. The system should focus human attention where judgment is most valuable.
Regulation raises the importance of literacy and accountability
The EU AI Act’s provisions on prohibited practices, definitions and AI literacy began to apply on 2 February 2025. For M&A teams, the immediate practical message is broader than legal classification: users need enough understanding to recognize limitations, challenge outputs and operate the system within its intended purpose.
Organizations should document which models and services process deal information, where data is stored, whether prompts or files are retained, who can access outputs and how incidents are handled. These questions belong in the M&A operating model, not only in an enterprise AI policy.
How smartmerger.com can turn AI output into controlled deal work
AI is most useful when its output enters the same environment that manages the diligence process. In smartmerger.com, findings can be linked to source documents, structured through Smart Fields, assigned to responsible reviewers, routed through approvals and connected to risks, decisions or integration actions.
This avoids a common failure: generating insights in one AI tool and then copying them into another tracker without their evidence or confidence context. Purpose-built permissions also help keep analysis inside the appropriate deal team, clean team or workstream boundary.
Start with a pilot that can fail safely
A good first pilot is narrow, repeatable and easy to verify—for example, extracting change-of-control clauses from a defined contract population or triaging a diligence request list.
Measure more than elapsed time:
- precision and recall against a human-reviewed sample;
- percentage of findings with valid evidence links;
- correction and escalation rates;
- review time by risk category;
- security or access exceptions;
- whether the output changed a decision or only saved administration.
Do not scale until the organization understands both the performance and the failure pattern.
Evaluate the operating environment, not only the model
Before using an AI capability on live deal information, test the full environment around it. A strong model can still be unsafe or ineffective if data retention is unclear, access is too broad, source links are weak or administrators can see confidential content.
- Confirm where prompts, files, embeddings and logs are stored.
- Understand whether deal information is used for model training.
- Test permission inheritance with representative restricted users.
- Verify how deleted or superseded documents are handled.
- Check whether outputs can be exported without their evidence links.
- Define an incident process for incorrect, leaked or inappropriate output.
This review should involve M&A, legal, information security and the business owner. Model accuracy is only one part of production readiness.
Questions AI should not answer alone
Some diligence questions combine incomplete evidence with legal, commercial or ethical judgment. They should not be delegated to an AI system as final decisions:
- whether a red flag is acceptable at the proposed price;
- whether management is credible enough to support the forecast;
- how a legal clause should be interpreted in the relevant jurisdiction;
- whether a cultural issue can be mitigated through integration design;
- whether the team has seen enough evidence to stop asking questions.
AI can organize the record, compare scenarios and make uncertainties visible. Accountability for these judgments remains with experienced people who understand the deal thesis and consequences.
Better diligence comes from better questions
The strongest case for AI is not that it can produce a faster report. It is that it can give experts more time to challenge assumptions, investigate inconsistencies and connect findings across functions.
Used well, AI shortens the path from information to the next high-value question. Used carelessly, it creates an illusion of completeness. The difference lies in evidence, process design and accountable human judgment.