Decision Metrics and Thresholds
Decision Metrics and Thresholds
A score does not act
A threshold creates false positives and false negatives. Under simple constant costs, choose positive when
so
This formula assumes calibrated , mutually exclusive actions and correctly specified costs.
Intervention value needs an effect
Suppose retaining a would-be churner is worth £50 and a message costs £8.
- If the message certainly prevents churn, act when , or .
- If it prevents only 20% of churn, act when , or .
Risk is not treatment responsiveness. Targeting the highest-risk customers can waste capacity if they cannot be influenced.
Threshold audit
Compare thresholds by decision cost
The sample is deliberately small. In practice, choose the threshold on validation data and report uncertainty on a later test period.
Capacity changes the policy
If HarborMart can contact only 1,000 customers, the policy may be “select the top 1,000 eligible expected incremental values,” not “score above 0.5.” Capacity, contact fatigue, fairness and channel constraints enter the ranking.
Compare policies at equal capacity:
| Policy | Customers contacted | Expected incremental contribution |
|---|---|---|
| random eligible | 1,000 | baseline |
| highest churn risk | 1,000 | depends on effect overlap |
| highest uplift | 1,000 | depends on credible causal model |
| constrained uplift | 1,000 | effect minus cost with policy constraints |
Decision curves, not metric shopping
Plot or table value across plausible thresholds, costs and capacities. A model that wins at one threshold may lose elsewhere.
AUROC averages ranking across thresholds the business may never use. F1 assigns a particular symmetric trade-off without pounds, capacity or intervention effects. Use them as diagnostics, not objectives by default.
Quick check
A model’s probabilities are perfectly calibrated overall but systematically too low for one customer group. Can the same economic threshold be applied safely?