Matching, Weighting and Overlap
Matching, Weighting and Overlap
Adjustment cannot repair an unmeasured assignment process
Scholarship recipients have higher prior scores and stronger adviser support than non-recipients. Matching can align observed scores and support. It cannot align unrecorded motivation merely because the propensity-score model predicts receipt well.
The identifying claim is conditional exchangeability:
plus positivity: every covariate profile of interest has a non-zero chance of each treatment. These are substantive claims about assignment, not consequences of balance.
Choose the target before the weights
Let .
| Target | Treated weight | Control weight | Question |
|---|---|---|---|
| ATE | effect for the combined target population | ||
| ATT | 1 | effect for treated units | |
| overlap | effect where treatment choice was most uncertain |
Changing weights changes the population. It is not merely a technical stabilisation.
A four-student example
| Student | Receipt | Completion | Propensity | ATE weight |
|---|---|---|---|---|
| A | 1 | 1 | 0.80 | 1.25 |
| B | 1 | 0 | 0.55 | 1.82 |
| C | 0 | 1 | 0.45 | 1.82 |
| D | 0 | 0 | 0.05 | 1.05 |
Student B and C receive more weight because their observed treatment was less predictable. If a treated student had , the ATE weight would be 50: one observation could dominate the result. That is a lack-of-overlap warning, not a request for a more elaborate classifier.
Design stage: outcomes stay hidden
- define pre-treatment covariates from an assignment story;
- estimate or construct a distance/propensity score;
- inspect common support and extreme weights;
- match, subclassify or weight;
- assess balance and effective sample size;
- revise the design without looking for a favourable outcome effect;
- estimate effects and uncertainty for the retained target population.
For weights , a useful information diagnostic is
Ten thousand records with do not provide ten thousand equally informative comparisons.
Balance is not a propensity-score -value
Use standardised mean differences, distribution plots and substantively important interactions. A large sample can make a negligible imbalance statistically significant; a small sample can hide important imbalance.
| Diagnostic | Useful question |
|---|---|
| standardised difference | are covariate means close on a scale-free metric? |
| variance/quantile comparison | do distributions align beyond the mean? |
| propensity overlap | are both treatments represented in the target region? |
| maximum weight | can one unit control the estimate? |
| effective sample size | how much information remains after weighting? |
Doubly robust does not mean assumption-free
An augmented inverse-probability estimator combines an outcome model and a treatment model. Under regularity conditions it can remain consistent when one nuisance model is correct. It does not survive unmeasured confounding, failed positivity, post-treatment controls or two badly misspecified models.
Stuart (2010) reviews matching as a design strategy. Li, Morgan and Zaslavsky (2018) develop overlap weights that target the region with greatest covariate overlap.
A defensible conclusion
Weak: “After propensity-score matching, treatment is as good as random.”
Stronger: “Among applicants in the shared-support region, weighted pre-treatment covariates are closely balanced. The estimate is causal if the recorded covariates suffice to control joint causes of receipt and completion; adviser motivation remains a plausible unmeasured confounder.”
Quick check
Trimming removes all high-need applicants because every high-need applicant received the scholarship. What changed?