Adoption and Business Value
Adoption and Business Value
Deployment is not impact
A model can be available, accurate and unused. It can also be used frequently and destroy value. Measure the full adoption funnel:
| Stage | Example measure | Failure revealed |
|---|---|---|
| eligible | orders for which a score should exist | poor coverage |
| delivered | score arrived before checkout confirmation | pipeline or latency failure |
| viewed | planner saw the recommendation | interface failure |
| accepted | planner chose the proposed action | trust or feasibility problem |
| executed | capacity or promise actually changed | downstream process failure |
| effective | intended outcome improved | weak or harmful intervention |
Do not call acceptance “value.” It is an intermediate behaviour.
Value is incremental
A compact value model is
All terms require a time period and counterfactual. Revenue attributed to customers who would have purchased anyway is not incremental value.
Worked pilot
HarborMart pilots a staffing recommendation in comparable stores:
| Stores | Before late rate | After late rate | Change |
|---|---|---|---|
| pilot | 8.0% | 6.2% | −1.8 points |
| comparison | 7.8% | 7.5% | −0.3 points |
A simple difference-in-differences estimate is
At 12,000 eligible monthly orders, this corresponds to about 180 fewer late orders. If a prevented late order is worth £7 and monthly operating cost is £900, estimated net value is
The arithmetic is easy; the identification is not. Were trends parallel? Did stores share staff? Did order mix or weather change? Report an interval for the effect and sensitivity to the £7 valuation.
Choose an evaluation design
| Situation | Strong practical option | Main threat |
|---|---|---|
| units can be randomised | cluster or individual randomised trial | spillovers and non-compliance |
| rollout must be gradual | randomised or justified stepped rollout | calendar trends |
| threshold determines action | regression discontinuity near cutoff | manipulation and limited scope |
| only observational history exists | matched or adjusted comparison | unmeasured confounding |
| no credible counterfactual | descriptive monitoring only | causal claim is not justified |
Pre-specify the unit, assignment, estimand, primary outcome, guardrails, analysis and stop rules. Analyse by assigned policy where possible; adoption analysis can then explain why the effect was diluted.
Diagnose low adoption with evidence
Suppose 10,000 recommendations were eligible:
| Funnel step | Count | Conditional rate |
|---|---|---|
| delivered on time | 9,600 | 96.0% of eligible |
| viewed | 7,200 | 75.0% of delivered |
| accepted | 5,400 | 75.0% of viewed |
| executed | 3,240 | 60.0% of accepted |
The largest absolute loss after delivery is execution, not acceptance. Interviewing users about “trust in AI” may miss a capacity-system integration failure. Join event logs to short, purposeful qualitative inquiry.
Roll out with stopping rules
Before exposure, state:
- success: a minimum decision-relevant effect over a fixed horizon;
- guardrails: service, safety, fairness and customer measures that must not deteriorate materially;
- pause: data or process conditions that make inference unreliable;
- rollback: trigger, owner and tested fallback;
- review: who sees adverse events and unresolved overrides.
Avoid repeatedly checking significance and stopping when a favourable result appears. Sequential methods are possible, but the rule must be designed before observing the result.
Current organisational example
Walmart's fiscal-2026 10-K describes AI-powered tools, supply-chain automation and changes intended to combine capabilities across Walmart and Sam's Club. It illustrates that analytics adoption involves operating processes, physical capacity and organisational design—not only model selection.
Treat an annual report as management disclosure, not an independent causal evaluation. Ask which baseline, costs, failed pilots and distributional effects are not visible.
Quick check
The pilot lowers late deliveries but also rejects more orders. Which metric decides whether to scale?