Module 5 — Modern Causal Analysis

External Validity and Transport

Move causal effects across populations, treatment versions, outcomes and contexts with explicit support and stability assumptions

External Validity and Transport

Internal validity has an address

The historical Pathways lottery identifies an ITT for applicants in participating districts, in score band 68–72, under that year’s offer and university-capacity conditions. “The programme works” silently removes every qualifier.

External validity has at least four dimensions:

DimensionTransport question
population XXdo effect modifiers differ in the target applicants?
treatment TTis the expanded scholarship the same treatment version?
outcome YYare definitions, follow-up and measurement comparable?
context CCdo institutions, prices, peers or implementation differ?

This population–treatment–outcome–context framework is developed empirically by Egami and Hartman (2023).

Reweight only supported effect modifiers

Suppose the pilot contains 80% high-preparedness and 20% lower-preparedness applicants. Effects are 4 and 12 points respectively:

ATEpilot=0.8(0.04)+0.2(0.12)=0.056.ATE_{pilot}=0.8(0.04)+0.2(0.12)=0.056.

The expansion target is 30% high and 70% lower preparedness:

ATEtarget=0.3(0.04)+0.7(0.12)=0.096.ATE_{target}=0.3(0.04)+0.7(0.12)=0.096.

The transported effect is 9.6 rather than 5.6 points if preparedness captures the relevant effect modification and group-specific effects remain stable.

More generally, study observations can be weighted by target-versus-study participation odds. This requires overlap: target profiles absent from the study cannot be recovered by large weights.

A transport estimate needs three audits

  1. selection diagram or mechanism: why study participation differs;
  2. effect-modifier set: why conditioning makes effects stable across study and target;
  3. support: whether every important target stratum appears in the study.

Report weighted and unweighted covariate distributions, weight concentration, effective sample size and sensitivity to omitted effect modifiers. Degtiar and Rose (2023) review generalizability and transportability assumptions and methods.

Scale can change the treatment itself

A 500-person pilot may provide intensive advising. A 20,000-person rollout can produce:

  • adviser congestion;
  • university-capacity constraints;
  • tuition or housing price responses;
  • peer spillovers;
  • applicant behavioural changes;
  • lower implementation fidelity.

These are not sampling differences. They violate stable treatment versions or no-interference assumptions. Model saturation, collect rollout evidence or define cluster/market-level effects rather than reweighting individuals and declaring transport complete.

Separate three recommendations

EvidenceDefensible recommendation
internally valid pilot onlyexpand evidence collection
supported population transporttarget similar applicants under same delivery
scale and equilibrium evidenceconsider system-wide rollout

The final policy memo should distinguish “estimated effect,” “transported effect under assumptions” and “decision after costs and constraints.”

Quick check

The target population has excellent covariate overlap, but the scholarship value is halved. Can population reweighting transport the original effect?

Answer
Not by itself. The treatment version changed. You need a dose-response or mechanism argument, evidence at the new value, or a deliberately bounded conclusion.

Next module: Research Workflow

Copyright © 2026