Causal Targeting and Policy Learning
Causal Targeting and Policy Learning
Risk is not uplift
For customer features , conditional treatment effect is
A high-risk customer may remain inactive whether contacted or not. A moderate-risk customer may respond strongly. If the action is a retention offer, rank expected incremental value, not churn risk alone.
Four response types
| Type | Outcome without offer | Outcome with offer | Targeting implication |
|---|---|---|---|
| persuadable | inactive | active | potential value |
| sure thing | active | active | offer may be wasted |
| lost cause | inactive | inactive | offer may be wasted |
| adverse responder | active | inactive | treatment can harm |
Individual types are not directly observed because only one potential outcome appears. Models estimate group-conditional effects under assumptions.
Net-value rule
If a retained customer is worth , offer cost is and estimated probability improvement is , act when
subject to eligibility, capacity and fairness constraints.
Worked pair
| Customer | Churn risk | Estimated risk reduction | Value | Offer cost | Expected net value |
|---|---|---|---|---|---|
| A | 80% | 2 points | £100 | £5 | |
| B | 35% | 15 points | £80 | £5 |
Risk targeting chooses A; uplift value chooses B.
Evidence requirements
Prefer randomised treatment data with:
- treatment and outcome definitions fixed;
- sufficient overlap across relevant profiles;
- no post-treatment features;
- honest train/evaluation splits;
- treatment-cost and outcome-value data;
- policy evaluation at the intended capacity.
With observational data, justify exchangeability, positivity and consistency; use propensity or outcome models as tools, not substitutes for the assumptions.
Evaluate a policy, not only a CATE model
Useful checks include:
- incremental outcome/value versus random or current targeting;
- uplift or policy-value curves on held-out data;
- sensitivity to costs, capacity and nuisance models;
- overlap and subgroup support;
- off-policy assumptions if evaluating a new policy from old-policy data;
- effects after deployment feedback.
The 2025 contextual optimisation survey places direct policy learning, predict-then-optimise and conditional stochastic optimisation in a common decision setting. A 2024 fair multiobjective predict-then-optimise study illustrates that downstream objectives can include fairness and robustness, not only efficiency.
Quick check
A treatment-effect model is trained using customers selected by the old high-risk policy. What problem arises?