0. Causal Questions and Designs

Causal Diagrams and Controls

Distinguish confounders, mediators, colliders and proxies before adjustment

Causal Diagrams and Controls

A control set is a causal claim

Suppose family resources XX affect both scholarship receipt DD and degree completion YY:

Rendering diagram…

The backdoor path DXYD\leftarrow X\rightarrow Y creates confounding. Conditioning on a well-measured pre-treatment XX may block it. The diagram does not prove that all common causes were measured.

Three variables that look like controls

Confounder

DXYD\leftarrow X\rightarrow Y

A pre-treatment common cause can belong in an adjustment set.

Mediator

DMYD\rightarrow M\rightarrow Y

If mentoring hours MM are caused by the scholarship, controlling for MM removes part of the total effect and may introduce further bias. It answers a different direct-effect question under stronger assumptions.

Collider

DSUYD\rightarrow S\leftarrow U\rightarrow Y

If survey response SS is affected by the scholarship and unobserved motivation UU, analysing responders only opens a non-causal association between DD and UU.

HarborMart-style “more controls” logic fails here

For Pathways, consider:

VariableTimingDefault roleAudit question
prior exam scorebefore offerpossible confounder or precision covariatedid it affect assignment and outcome?
application essay ratingbefore offer but assessor-dependentpossible confounder/proxywas it measured before assignment and consistently?
advising attendanceafter offermediatoris the target total or direct effect?
first-year GPAafter offermediator and selection variabledoes conditioning discard treatment pathways?
completion-record availabilityafter treatment/outcome processselection/collider riskdid treatment affect observation?

“Pre-treatment” is necessary for a standard confounder, but not sufficient. A pre-treatment variable can be an instrument, proxy or collider of earlier causes.

Adjustment cannot create overlap

Under conditional exchangeability and positivity,

{Y(1),Y(0)}DX,0<P(D=1X)<1.\{Y(1),Y(0)\}\perp D\mid X, \qquad 0<P(D=1\mid X)<1.

If every high-scoring applicant receives the scholarship and no low-scoring applicant does, outcome regression outside overlap depends on functional-form extrapolation. A rich control set can make the absence of comparison units harder to see.

Design before data-driven control selection

Use this order:

  1. define the total or direct effect;
  2. draw treatment, outcome, timing and plausible common causes;
  3. exclude descendants of treatment from a total-effect adjustment set;
  4. identify the smallest defensible sets and measurement limitations;
  5. inspect overlap and sensitivity to unmeasured causes;
  6. use prediction tools only within this causal design.

LASSO can select variables that predict YY while omitting weak outcome predictors that strongly affect DD, or include post-treatment variables. Algorithmic selection is not a substitute for a timing and causal audit.

DAGs clarify assumptions; they do not certify them

Two researchers can draw different graphs because institutional knowledge differs. Make contested arrows explicit, derive alternative adjustment sets and test whether the conclusion depends on those choices.

A missing arrow is an assumption. It should be defended in prose, not hidden by a clean diagram.

Quick check

An offer increases advising attendance, and advising raises completion. Should advising be controlled when estimating the total effect of the offer?

Answer
Normally no. Advising is a treatment-induced pathway. Conditioning on it removes part of the total effect and may open bias if advising and completion share unmeasured causes. A mediation estimand is possible, but it is a different question with stronger assumptions.

Next: Regression and Inference

Copyright © 2026