Smartmerger Blog

Governed Agentic AI Requires Evidence, Controls and Measurable Outcomes

Written by Michael Klawon | 25.May 2026

Generative AI governance focused on what a model could say. Agentic AI governance must also control what a system can do. An agent may retrieve restricted information, call tools, create records, send messages and change workflow state over several steps.

That expands the value opportunity, but it also changes the control question. M&A leaders need evidence that the system completed the right business task, under the right authority, without introducing hidden work or uncontrolled consequences.

Output review is no longer enough

A reviewer can read a summary and judge whether it looks sensible. That does not reveal whether the agent:

  • used an unauthorized source;
  • ignored a superseding agreement;
  • created duplicate actions during a retry;
  • shared restricted information with the wrong workstream;
  • changed a record before approval;
  • failed to stop when evidence was incomplete.

Governance must therefore cover the full action chain: purpose, context, tools, decisions, state changes and downstream effects.

An eight-layer control stack

  1. Purpose: approved use case, process owner and prohibited uses.
  2. Identity: unique agent identity and accountable human sponsor.
  3. Access: least-privilege retrieval, tool and disclosure permissions.
  4. Evidence: authoritative sources, provenance and version status.
  5. Action policy: allowed changes, thresholds, approvals and stop conditions.
  6. Evaluation: realistic tests for task success, consistency and security.
  7. Observation: logs, monitoring, sampling and exception review.
  8. Response: rollback, incident handling, notification and learning.

Controls should be proportionate to materiality and reversibility. A drafting assistant and an agent that changes a critical milestone do not need the same governance burden.

The governance objective

Constrain authority tightly enough that failure is detectable, reversible where possible and owned by a named person—while preserving enough freedom for the agent to create real operational value.

Evidence must travel with the action

Every material agent-created record should answer four questions:

  • What source supports this result?
  • Which transformation or reasoning produced it?
  • Who or what approved the change?
  • Which downstream decision or object now depends on it?

This is particularly important in M&A because source status changes quickly. A draft agreement may be superseded, management data may be corrected and a clean-team estimate may be replaced after closing. Evidence links and version history allow downstream users to recognize when an earlier conclusion must be revisited.

Evaluate agents as non-deterministic processes

Agents can take different paths on repeated runs. A useful evaluation program therefore tests more than average answer quality.

Anthropic’s January 2026 guidance on evaluating AI agents emphasizes multi-turn behavior, task-level outcomes and the difficulty created by flexible tool use. For M&A, an evaluation suite should measure:

  • task success: was the business outcome correct and complete?
  • evidence validity: do material claims link to authoritative sources?
  • policy compliance: were only permitted data and actions used?
  • consistency: how often does the workflow succeed across repeated runs?
  • recovery: what happens when a source, tool or human response is unavailable?
  • state integrity: are retries and partial failures handled without duplicates?

Track both “can it succeed at least once?” and “does it succeed reliably every time?” A high-stakes workflow needs the latter.

Controls should attach to actions, not labels

Calling a system “human-in-the-loop” says little. Specify controls for each action:

  • read a restricted document;
  • write a structured finding;
  • change issue severity;
  • assign or notify a person;
  • publish information across an access boundary;
  • alter a baseline, milestone or financial record;
  • send an external message.

Each action may have different approval and logging requirements. A single agent can be autonomous for low-risk classification while approval-bound for disclosure and high-impact state changes.

Incident response belongs in the operating model

Agent incidents will not always look like cybersecurity incidents. They may include:

  • materially wrong recommendations reaching a decision forum;
  • unauthorized disclosure through a generated summary;
  • duplicate or missing actions after retry;
  • model or tool changes causing silent quality deterioration;
  • users bypassing controls because the workflow is too slow;
  • an agent continuing after a human correction should have stopped it.

Define severity, containment, rollback, evidence preservation and notification before launch. The organization should be able to disable one tool or action without shutting down every AI-supported workflow.

European transparency is becoming more concrete

On 8 May 2026, the European Commission opened consultation on draft guidance for the AI Act’s Article 50 transparency obligations. The obligations were expected to apply from 2 August 2026, after this article’s publication date.

The immediate lesson for M&A deployers is broader than labeling content. Organizations should know when users interact with AI, which outputs are machine-generated or machine-modified, how that fact is communicated and which records must preserve the distinction. Transparency should be designed into the workflow rather than added to the final document.

Productivity claims need a critical reading

BCG’s 19 May 2026 analysis of AI in the M&A technology workstream estimated about 15% lower total execution effort, with higher impact in selected tasks. It also made an important qualification: transaction timelines are often fixed by regulation, contracts and change-management realities.

That distinction matters. An agent may reduce manual reconciliation without shortening sign-to-close. The freed capacity can be:

  • banked as lower cost;
  • reinvested in better Day 1 readiness and risk resolution;
  • used to expand analytical scope;
  • absorbed by new review and control work.

Governance should measure net value after the cost of evaluation, supervision, exception handling and incidents—not gross automation hours.

smartmerger.com can make governance part of execution

smartmerger.com’s roles, permissions, approvals, structured Smart Fields, purpose-built apps and audit trails provide a controlled environment for human M&A work and the process foundation for future or connected agent-executed workflows. In an agent-enabled deployment, outputs could remain linked to source evidence and move through the same accountable workflows used by deal teams.

This architecture supports differentiated control: an agent could prepare a finding while a workstream owner approves materiality; monitor a Day 1 plan while a steering decision changes the critical path; or draft a synergy variance while finance approves recognition. Governance becomes visible process rather than a separate policy document.

Use a governance scorecard that executives can challenge

  • task success and repeat-run consistency;
  • material correction and override rates;
  • evidence coverage and source freshness;
  • permission denials, exceptions and attempted boundary crossings;
  • human review time and hidden rework;
  • duplicate, missing or late state changes;
  • incident severity and time to contain;
  • deal outcome affected: cycle time, readiness, cost, risk or value realization.

Review the scorecard by use case and risk tier. Aggregate averages can hide one workflow that performs well and another that should not be in production.

Apply three lines of defense without creating three parallel workflows

Business ownership, control oversight and independent assurance should be distinct, but they should work from the same evidence.

  • First line: the M&A process owner defines the use case, monitors outcomes and owns exceptions.
  • Second line: legal, security, privacy and AI-governance functions define policy, challenge controls and review material incidents.
  • Third line: internal audit or independent assurance tests whether controls operate as described.

Avoid asking each line to maintain its own spreadsheet and evidence pack. The agent’s identity, evaluations, approvals, versions and incidents should be available through one governed record, with access appropriate to each role. This reduces compliance theatre and makes assurance part of the operating process.

Governance is what allows useful authority

The choice is not between autonomous agents and no agents. It is between undefined delegation and governed delegation.

Organizations that can prove purpose, permission, evidence, evaluation and response will be able to grant useful authority more confidently. Those that cannot will either accept unmanaged risk or keep their most promising AI confined to demonstrations.