Synthetic Control
Synthetic Control
The donor pool writes the counterfactual
Clearborough is the only city introducing a clean-air zone in 2024. No single city matches its earlier pollution path. Synthetic control chooses non-negative donor weights that reproduce Clearborough before treatment.
| Donor city | Weight |
|---|---|
| Alder | 0.50 |
| Birch | 0.30 |
| Cedar | 0.20 |
| all others | 0.00 |
If 2025 pollution is 30 in Clearborough and in the synthetic control, the estimated effect is μg/m³.
The calculation is easy. The credibility lies in why those donors represent untreated Clearborough.
Five design decisions
- treated unit and intervention date: no anticipation or hidden earlier exposure;
- donor pool: no spillovers, competing policies or structurally impossible comparators;
- predictors: determined before treatment and justified substantively;
- pre-period: long enough to learn outcome dynamics, not chosen for a preferred result;
- estimand: a time path for this treated unit, not automatically an average effect.
Pre-treatment fit is necessary, not sufficient
Suppose the pre-treatment root mean squared prediction error is 0.8. Good fit shows that the weighted donors reproduce observed history. It does not prove they would follow Clearborough after 2024 absent treatment.
Inspect:
- the full treated and synthetic paths;
- gaps in every pre-period, not only a fit statistic;
- donor weights and leave-one-donor-out estimates;
- whether important predictors are balanced;
- events affecting donors after treatment;
- spillovers into neighbouring cities.
Poor pre-fit usually means the donor pool cannot construct the desired counterfactual. A large post-gap built on poor pre-fit is not persuasive.
Placebos create a reference distribution
Reassign the intervention to each donor city, rebuild its synthetic control and compare post-treatment gaps. One useful statistic is
If Clearborough has and most well-fitted placebos have , its break is unusual relative to the donor pool. Exclude or visibly flag placebos with extremely poor pre-fit; otherwise the comparison is mechanically easy to win.
This exercise provides design-based perspective, not a conventional large-sample -value by default.
Synthetic DID changes the target of balancing
Synthetic difference-in-differences combines unit weights with time weights, then estimates a DID-style contrast. It can reduce sensitivity to imperfect pre-fit under its model, but does not excuse contaminated donors or anticipation. Arkhangelsky et al. (2021) develop the method. The current synthdid documentation also makes implementation restrictions visible; consult live documentation rather than assuming every adoption pattern is supported.
The classic lesson
Abadie, Diamond and Hainmueller (2010) use synthetic control to study California’s tobacco-control programme. The durable contribution is not “weighted averages are causal.” It is that comparative case studies can make the construction and quality of the counterfactual unusually transparent.
Method-choice checkpoint
| Setting | Better starting point |
|---|---|
| many treated and untreated units, credible common trends | DID |
| one treated unit, long pre-period, rich donor pool | synthetic control |
| threshold determines exposure | RDD |
| assignment explained by observed individual covariates | matching/weighting |
Quick check
One donor receives weight 0.82. Removing it changes the estimated effect from −6.3 to −1.0. What should the report say?