3. Policy Evaluation Designs

Matching, Weighting and Overlap

Construct observable comparison groups while keeping exchangeability, balance and positivity distinct

Matching, Weighting and Overlap

Adjustment cannot repair an unmeasured assignment process

Scholarship recipients have higher prior scores and stronger adviser support than non-recipients. Matching can align observed scores and support. It cannot align unrecorded motivation merely because the propensity-score model predicts receipt well.

The identifying claim is conditional exchangeability:

(Y(1),Y(0))DX,(Y(1),Y(0))\perp D\mid X,

plus positivity: every covariate profile of interest has a non-zero chance of each treatment. These are substantive claims about assignment, not consequences of balance.

Choose the target before the weights

Let e(X)=P(D=1X)e(X)=P(D=1\mid X).

TargetTreated weightControl weightQuestion
ATE1/e(X)1/e(X)1/[1e(X)]1/[1-e(X)]effect for the combined target population
ATT1e(X)/[1e(X)]e(X)/[1-e(X)]effect for treated units
overlap1e(X)1-e(X)e(X)e(X)effect where treatment choice was most uncertain

Changing weights changes the population. It is not merely a technical stabilisation.

A four-student example

StudentReceipt DDCompletion YYPropensity e(X)e(X)ATE weight
A110.801.25
B100.551.82
C010.451.82
D000.051.05

Student B and C receive more weight because their observed treatment was less predictable. If a treated student had e(X)=0.02e(X)=0.02, the ATE weight would be 50: one observation could dominate the result. That is a lack-of-overlap warning, not a request for a more elaborate classifier.

Design stage: outcomes stay hidden

  1. define pre-treatment covariates from an assignment story;
  2. estimate or construct a distance/propensity score;
  3. inspect common support and extreme weights;
  4. match, subclassify or weight;
  5. assess balance and effective sample size;
  6. revise the design without looking for a favourable outcome effect;
  7. estimate effects and uncertainty for the retained target population.

For weights wiw_i, a useful information diagnostic is

neff=(iwi)2iwi2.n_{eff}=\frac{(\sum_i w_i)^2}{\sum_i w_i^2}.

Ten thousand records with neff=420n_{eff}=420 do not provide ten thousand equally informative comparisons.

Balance is not a propensity-score pp-value

Use standardised mean differences, distribution plots and substantively important interactions. A large sample can make a negligible imbalance statistically significant; a small sample can hide important imbalance.

DiagnosticUseful question
standardised differenceare covariate means close on a scale-free metric?
variance/quantile comparisondo distributions align beyond the mean?
propensity overlapare both treatments represented in the target region?
maximum weightcan one unit control the estimate?
effective sample sizehow much information remains after weighting?

Doubly robust does not mean assumption-free

An augmented inverse-probability estimator combines an outcome model and a treatment model. Under regularity conditions it can remain consistent when one nuisance model is correct. It does not survive unmeasured confounding, failed positivity, post-treatment controls or two badly misspecified models.

Stuart (2010) reviews matching as a design strategy. Li, Morgan and Zaslavsky (2018) develop overlap weights that target the region with greatest covariate overlap.

A defensible conclusion

Weak: “After propensity-score matching, treatment is as good as random.”

Stronger: “Among applicants in the shared-support region, weighted pre-treatment covariates are closely balanced. The estimate is causal if the recorded covariates suffice to control joint causes of receipt and completion; adviser motivation remains a plausible unmeasured confounder.”

Quick check

Trimming removes all high-need applicants because every high-need applicant received the scholarship. What changed?

Answer
The data contain no untreated comparison for high-need applicants. Trimming may improve internal credibility, but the estimand now excludes that group. State the new target population; do not generalise the trimmed estimate back without extra assumptions or evidence.

Next: Synthetic Control

Copyright © 2026