3. Prescriptive Analytics

Causal Targeting and Policy Learning

Target actions by incremental effect and value rather than untreated risk alone

Causal Targeting and Policy Learning

Risk is not uplift

For customer features X=xX=x, conditional treatment effect is

τ(x)=E[Y(1)Y(0)X=x].\tau(x)=E[Y(1)-Y(0)\mid X=x].

A high-risk customer may remain inactive whether contacted or not. A moderate-risk customer may respond strongly. If the action is a retention offer, rank expected incremental value, not churn risk alone.

Four response types

TypeOutcome without offerOutcome with offerTargeting implication
persuadableinactiveactivepotential value
sure thingactiveactiveoffer may be wasted
lost causeinactiveinactiveoffer may be wasted
adverse responderactiveinactivetreatment can harm

Individual types are not directly observed because only one potential outcome appears. Models estimate group-conditional effects under assumptions.

Net-value rule

If a retained customer is worth ViV_i, offer cost is cic_i and estimated probability improvement is τi\tau_i, act when

τiVici>0,\tau_iV_i-c_i>0,

subject to eligibility, capacity and fairness constraints.

Worked pair

CustomerChurn riskEstimated risk reductionValueOffer costExpected net value
A80%2 points£100£50.02(100)5=£30.02(100)-5=-£3
B35%15 points£80£50.15(80)5=£70.15(80)-5=£7

Risk targeting chooses A; uplift value chooses B.

Evidence requirements

Prefer randomised treatment data with:

  • treatment and outcome definitions fixed;
  • sufficient overlap across relevant profiles;
  • no post-treatment features;
  • honest train/evaluation splits;
  • treatment-cost and outcome-value data;
  • policy evaluation at the intended capacity.

With observational data, justify exchangeability, positivity and consistency; use propensity or outcome models as tools, not substitutes for the assumptions.

Evaluate a policy, not only a CATE model

Useful checks include:

  • incremental outcome/value versus random or current targeting;
  • uplift or policy-value curves on held-out data;
  • sensitivity to costs, capacity and nuisance models;
  • overlap and subgroup support;
  • off-policy assumptions if evaluating a new policy from old-policy data;
  • effects after deployment feedback.

The 2025 contextual optimisation survey places direct policy learning, predict-then-optimise and conditional stochastic optimisation in a common decision setting. A 2024 fair multiobjective predict-then-optimise study illustrates that downstream objectives can include fairness and robustness, not only efficiency.

Quick check

A treatment-effect model is trained using customers selected by the old high-risk policy. What problem arises?

Answer
Treatment and outcome evidence may have poor overlap outside the old policy’s selected region. The model cannot reliably learn effects for rarely treated profiles without design or strong extrapolation assumptions. Preserve exploration or run a suitable experiment.

Next: Deployment and Governance

Copyright © 2026