4. Deployment and Governance

Adoption and Business Value

Test whether recommendations are used and whether the policy causes worthwhile outcomes

Adoption and Business Value

Deployment is not impact

A model can be available, accurate and unused. It can also be used frequently and destroy value. Measure the full adoption funnel:

StageExample measureFailure revealed
eligibleorders for which a score should existpoor coverage
deliveredscore arrived before checkout confirmationpipeline or latency failure
viewedplanner saw the recommendationinterface failure
acceptedplanner chose the proposed actiontrust or feasibility problem
executedcapacity or promise actually changeddownstream process failure
effectiveintended outcome improvedweak or harmful intervention

Do not call acceptance “value.” It is an intermediate behaviour.

Value is incremental

A compact value model is

net value=incremental outcomes×value per outcomeintervention costsystem costharm.\text{net value} =\text{incremental outcomes}\times\text{value per outcome} -\text{intervention cost}-\text{system cost}-\text{harm}.

All terms require a time period and counterfactual. Revenue attributed to customers who would have purchased anyway is not incremental value.

Worked pilot

HarborMart pilots a staffing recommendation in comparable stores:

StoresBefore late rateAfter late rateChange
pilot8.0%6.2%−1.8 points
comparison7.8%7.5%−0.3 points

A simple difference-in-differences estimate is

(1.8)(0.3)=1.5 percentage points.(-1.8)-(-0.3)=-1.5\text{ percentage points}.

At 12,000 eligible monthly orders, this corresponds to about 180 fewer late orders. If a prevented late order is worth £7 and monthly operating cost is £900, estimated net value is

180(£7)£900=£360.180(£7)-£900=£360.

The arithmetic is easy; the identification is not. Were trends parallel? Did stores share staff? Did order mix or weather change? Report an interval for the effect and sensitivity to the £7 valuation.

Choose an evaluation design

SituationStrong practical optionMain threat
units can be randomisedcluster or individual randomised trialspillovers and non-compliance
rollout must be gradualrandomised or justified stepped rolloutcalendar trends
threshold determines actionregression discontinuity near cutoffmanipulation and limited scope
only observational history existsmatched or adjusted comparisonunmeasured confounding
no credible counterfactualdescriptive monitoring onlycausal claim is not justified

Pre-specify the unit, assignment, estimand, primary outcome, guardrails, analysis and stop rules. Analyse by assigned policy where possible; adoption analysis can then explain why the effect was diluted.

Diagnose low adoption with evidence

Suppose 10,000 recommendations were eligible:

Funnel stepCountConditional rate
delivered on time9,60096.0% of eligible
viewed7,20075.0% of delivered
accepted5,40075.0% of viewed
executed3,24060.0% of accepted

The largest absolute loss after delivery is execution, not acceptance. Interviewing users about “trust in AI” may miss a capacity-system integration failure. Join event logs to short, purposeful qualitative inquiry.

Roll out with stopping rules

Before exposure, state:

  • success: a minimum decision-relevant effect over a fixed horizon;
  • guardrails: service, safety, fairness and customer measures that must not deteriorate materially;
  • pause: data or process conditions that make inference unreliable;
  • rollback: trigger, owner and tested fallback;
  • review: who sees adverse events and unresolved overrides.

Avoid repeatedly checking significance and stopping when a favourable result appears. Sequential methods are possible, but the rule must be designed before observing the result.

Current organisational example

Walmart's fiscal-2026 10-K describes AI-powered tools, supply-chain automation and changes intended to combine capabilities across Walmart and Sam's Club. It illustrates that analytics adoption involves operating processes, physical capacity and organisational design—not only model selection.

Treat an annual report as management disclosure, not an independent causal evaluation. Ask which baseline, costs, failed pilots and distributional effects are not visible.

Quick check

The pilot lowers late deliveries but also rejects more orders. Which metric decides whether to scale?

Answer
Neither rate alone. Compare incremental contribution after delivery costs, rejection opportunity cost, service harm and capacity constraints. Report both outcomes and subgroup effects. The objective and guardrails should have been agreed before the pilot.

Next: Capstone

Copyright © 2026