All posts

Scope & Governance

How Should AI Detect Scope Changes in Due Diligence?

Chris Stefaner11 min read
How Should AI Detect Scope Changes in Due Diligence?

How should AI detect scope changes in due diligence? It should create a reviewable candidate, never an approved change. A defensible system can point to an advisor update, compare it with the agreed workstream scope, connect the signal to cost evidence and explain why the item needs attention. It must not alter a fee cap, committed spend or the instruction to proceed. Those are commercial decisions for a named deal lead.

That boundary follows the National Institute of Standards and Technology's Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (2024). NIST calls for defined human-AI roles, documented provenance, testing in conditions similar to deployment and monitoring of human overrides. Applied to diligence, the machine may say, “this looks different from the baseline, and here is the evidence”. A person decides whether the difference is an in-scope clarification, a genuine advisor change request or noise.

The commercial thesis is unchanged by AI. Diligence budgets usually break because scope drifts without being priced, not because an advisor's rate suddenly becomes unreasonable. Automation can shorten the time between an update and a warning. Only a control architecture can stop that warning from becoming unexamined authority.

Key Takeaway

AI scope change detection should identify and evidence a possible scope delta, then route it to human review. The safe architecture keeps the engagement letter as a versioned baseline, separates detection from approval, records provenance and learns from false positives without allowing a model to commit spend.

How should AI detect scope changes in due diligence?#

AI scope change detection should detect tension between the approved commercial baseline and new evidence. It is looking for a possible delta, not declaring that scope has changed.

The baseline is more than a fee total. It includes the engagement-letter deliverables, assumptions, exclusions, jurisdictions, workstream owner, fee structure, cap and approved changes. A legal update mentioning a second jurisdiction may be a new mandate, or it may sit inside an agreed multi-country review. A QoE update proposing another customer-cohort cut may be additional analysis, or it may be an expressly listed deliverable. Without the baseline, language classification is guesswork.

Four signal families are useful because no single one is reliable enough on its own:

Signal familyWhat the system comparesWhy it may be wrong
Scope languageNew advisor wording against agreed deliverables, assumptions and exclusions“Additional” may describe detail inside an existing deliverable
Deliverable changeA proposed output, jurisdiction, entity or analysis against the baseline listThe engagement letter may define the deliverable broadly
Cost movementCommitted spend, forecast or run-rate against the approved workstream capBilling cadence or a retainer may create a temporary step-up
Sequence conflictWork reported as started against the change-request and approval logApproval may exist in another channel and not yet be recorded

A good detector combines those signals into a candidate record. It should show the exact source passage, the baseline clause in tension, the cost movement, the relevant advisor and workstream, and the reason for the flag. Confidence without that evidence is not useful to a deal lead.

The distinction from existing governance is important. Our guide to managing diligence advisor change requests begins once a recognisable ask has landed and explains how to price and decide it. Detection sits one step earlier. Its job is to find the update or cost pattern that might otherwise never enter the change-request process at all.

The competitive gap is equally narrow. AI discussion in M&A software tends to centre on document summaries, risk extraction, tasks and deal workflow. Those are legitimate product categories. They do not answer the commercial control question: did this advisor update imply work beyond the scope that was priced, and who is authorised to decide?

Why put controls before automation?#

Controls should come first because adoption is moving faster than the evidence needed to trust an automated commercial classification. A plausible label is not enough when the downstream action could expand an advisor mandate.

Deloitte's 2025 GenAI in M&A Study, based on 1,000 senior US corporate and private-equity leaders, found that 86% had integrated GenAI into M&A workflows, while 83% expected a moderate or significant impact on M&A decision-making. The same study found 67% identified data security as a leading concern. Adoption, consequence and control pressure are rising together.

Adoption and control pressure around GenAI in M&A

Source: Deloitte, 2025 GenAI in M&A Study; 1,000 senior US corporate and private-equity leaders

The study is evidence of organisational intent, not proof that an AI scope classifier is accurate. Its respondents were US-based, and “integrated” covers many use cases that do not commit advisor spend. A UK mid-market team should read the figures as a reason to design the approval boundary now, not as a benchmark for buying software.

Erik Dilger, Managing Director at Deloitte Financial Advisory Services LLP, and Will Engelbrecht, Principal at Deloitte Consulting LLP, frame GenAI in the study as support that enhances professional work rather than replacing the professional. Scope control makes that principle concrete. The model can compress the evidence; the deal lead retains the mandate.

Assurance is becoming a field in its own right. The UK Department for Science, Innovation and Technology's Assuring a Responsible Future for AI (2024) estimated that 524 firms supplied AI assurance goods and services in the UK. That market count says nothing about M&A detector quality, but it reinforces a practical point: deployment needs measurement, documentation and review, not just a model and a confidence score.

What control architecture works before automation?#

A workable control architecture has five separated stages: baseline, evidence intake, detection, human review and authorised write-back. Separation matters because the detector should not be able to change the commercial record it is testing.

Version the commercial baseline

Baseline
Record the scope, deliverables, assumptions, exclusions, fee structure, cap and approver for each workstream. Preserve prior versions and approved changes rather than overwriting the original agreement.

Detection quality is capped by baseline quality. A vague engagement letter cannot produce a precise scope delta.

Snapshot every source before interpretation

Evidence
Keep the original advisor update, cost record, invoice line or meeting note, with its source, timestamp, author and deal access controls. Normalisation may create a working copy, but the reviewer must be able to retrieve the original.

A summary is an aid to review, not a substitute for the source.

Layer deterministic and model signals

Detection
Use rules for hard conflicts such as a forecast moving without an approved change, then use a model to propose semantic matches and classify ambiguous language. Keep each signal visible so the reviewer can see why the candidate exists.

Rules provide reproducibility; a model provides coverage where advisor language varies.

Build an evidence pack, not an alert

Review queue
Package the source excerpt, baseline clause, proposed delta, estimated cost effect, confidence and detector version into one review object. Deduplicate repeated updates before they reach the queue.

The fastest review is the one that does not require the deal lead to search three systems for context.

Keep approval on a separate write path

Approval
Allow an authorised person to mark the candidate in scope, approved change, descope, decline, duplicate or insufficient evidence. Only that human decision may update scope, forecast, cap or committed spend.

Never let a classifier's output instruct an advisor to proceed or modify the budget record automatically.

The architecture supports existing diligence controls rather than replacing them. The agreed baseline should follow the same workstream discipline used to prevent diligence advisor scope creep. Cost signals should preserve the distinction between advisor fee estimates and actuals, including fixed fee vs. T&M treatment. A model cannot reconcile a number correctly if the underlying system treats a fee estimate, a commitment and an invoice as the same state.

Provenance needs to survive the whole route. NIST's 2024 profile recommends tracking source and modification metadata, evaluating system performance under real deployment conditions and documenting operator overrides. For AI diligence cost control, that translates into a record of what the detector saw, which baseline version it used, what it proposed, what the reviewer decided and why the cost record did or did not move.

How should false positives and missed changes be handled?#

False positives should be treated as labelled review outcomes, while missed changes require active sampling and retrospective testing. The objective is not to eliminate all alerts; it is to manage the cost of review without hiding commercially material drift.

A flag can be wrong for ordinary reasons. An advisor may repeat a previously approved instruction. A higher run-rate may reflect invoice timing. A new phrase may describe an existing deliverable. The reviewer should be able to reject the candidate as in scope, duplicate or unsupported, with a short rationale. Those distinctions matter because “false positive” alone does not tell the team whether the baseline, data connection, rule or model failed.

False negatives are harder because they never enter the queue. A team needs periodic review of a sample of unflagged advisor updates and later invoices, especially after a new workstream, advisor or engagement-letter format is introduced. Compare the sample with the approved change log and final cost position. A detector that generates a quiet queue may be precise, or it may simply be blind.

John Sotiropoulos, Senior Security Architect at Kainos and co-lead of the OWASP Top 10 for LLM Applications, wrote the UK government's Implementation Guide for the AI Cyber Security Code of Practice (2025). The guide recommends testing whether operators can distinguish true cases from false positives, measuring the accuracy of human oversight and designing interfaces that do not condition reviewers to click approve. The examples are cross-sector rather than M&A-specific, which is a limitation. The operating lesson is directly relevant: review quality is part of system performance.

The associated Code of Practice for the Cyber Security of AI (2025) reports that 80% of respondents supported DSIT's proposed approach. Consultation support does not validate a diligence detector, but it gives the controls an authoritative policy context: organisations asked for implementation detail, and the resulting guide makes human oversight measurable rather than ceremonial.

The review dashboard should therefore track precision, recall, duplicate rate, time to decision, overrides and missed changes by workstream. A legal detector and a QoE detector may need different thresholds because their language, cost cadence and consequence differ. Corrections should inform testing, but an individual approval should not automatically become training truth. Deal leads can be inconsistent too, particularly under signing pressure.

The most dangerous false positive is not an unnecessary alert. It is an alert that looks so polished that the reviewer assumes the commercial conclusion has already been reached.

Where does live cost-and-scope control fit today?#

The live Advilink product is the cost-and-scope control layer, not automated AI scope change detection. Deal teams can structure advisor budgets by workstream, track committed and actual spend against agreed scope, keep changes and approval status visible, and prepare an IC-ready cost view.

Automated extraction from advisor updates and engagement letters, AI drafting, classification and AI scope-change detection are in development. They must not be treated as live or proven capabilities. The intended design boundary is the one described here: detection produces an evidenced candidate; a named human reviews it; only an authorised decision changes scope or money.

That distinction keeps this topic narrower than an AI due diligence software category. It does not analyse deal-room documents, produce a QoE opinion or automate legal diligence. It sits beside those tools and controls the commercial consequences when new findings create more advisor work. Once a candidate is approved, its rationale belongs in the same record as the IC cost report and scope-change log.

The useful test for any future detector will be whether a deal lead can reconstruct a disputed flag months later, from original update to final approval, without asking the model to remember what it meant.

Frequently Asked Questions#

What is AI scope change detection in due diligence?#

AI scope change detection compares advisor updates and cost evidence with an approved workstream baseline to flag a possible scope delta. It should provide source evidence and a reason for review, not declare that the mandate changed or approve the associated fee.

Can AI approve an advisor change request?#

No. Approval changes the commercial mandate and may commit spend, so it belongs to a named deal lead or other authorised person. AI can assemble the baseline clause, proposed work and cost effect into a review pack.

How do you reduce false positives in AI scope detection?#

Use a versioned baseline, combine language and cost signals, deduplicate repeated updates and record why reviewers reject flags. Test thresholds by workstream and sample unflagged items as well, because reducing alerts without measuring misses can create a falsely reassuring queue.

Is AI scope change detection live today?#

No. Advilink's live product is advisor cost-and-scope control, including workstream budgets, committed and actual spend tracking, visible scope changes and IC-ready reporting. Automated extraction, drafting, classification and AI scope-change detection are in development.

Sources#

  1. Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile — Chloe Autio, Reva Schwartz, Jesse Dunietz, Shomik Jain, Martin Stanley, Elham Tabassi, Patrick Hall and Kamie Roberts, National Institute of Standards and Technology, 2024. Human-AI roles, provenance, evaluation, monitoring and feedback controls.
  2. 2025 GenAI in M&A Study — Deloitte, 2025. Survey of 1,000 senior US corporate and private-equity leaders; adoption, decision impact, security concerns.
  3. Assuring a Responsible Future for AI — Department for Science, Innovation and Technology, 2024. UK assurance-market research and the estimate of firms supplying AI assurance goods and services.
  4. Implementation Guide for the AI Cyber Security Code of Practice — John Sotiropoulos for the Department for Science, Innovation and Technology, 2025. Human-oversight testing, false positives, review interfaces and accountability.
  5. Code of Practice for the Cyber Security of AI. Department for Science, Innovation and Technology, 2025. Consultation support and baseline controls for AI systems.

Catch scope creep before it becomes an overrun

Advilink flags work that drifts beyond the agreed scope, so you can approve or push back before the next invoice — not after.

See scope tracking