Microeconometrics — Designing Credible Counterfactuals
Microeconometrics — Designing Credible Counterfactuals
Microeconometrics is not a catalogue of estimators. It is a disciplined evidence chain:
The estimator comes after the counterfactual argument. A precisely estimated coefficient can still answer the wrong question.
Two running cases
Northbridge Pathways Scholarship
Northbridge offers a scholarship to some school-leavers. Eligibility, offer, take-up, university entry, completion and later earnings are different variables.
| Research object | Example |
|---|---|
| treatment | scholarship offer, receipt or university enrolment |
| outcome | entry within one year, completion within five years, or earnings |
| assignment | lottery, score cutoff, phased district rollout or self-selection |
| possible estimand | ITT of an offer, LATE of enrolment, ATT for recipients |
Clearborough Clean Air Zone
Municipalities adopt a clean-air policy at different dates. Pollution, traffic, retail activity and health visits are observed over time. This case develops panel, difference-in-differences and synthetic-control reasoning.
Both cases are synthetic. Real studies are named and linked explicitly.
Learning outcomes
By the end, you should be able to:
- define treatment, outcome, population, horizon and estimand before selecting a model;
- represent selection and timing with potential outcomes and causal diagrams;
- distinguish a linear projection from a causal effect and choose defensible controls;
- explain IV, fixed effects, DID, RDD, matching and synthetic control through their counterfactual source;
- diagnose weak instruments, non-parallel trends, manipulation, poor overlap and failed pre-treatment fit;
- match standard errors and uncertainty to sampling and treatment assignment;
- interpret binary, multinomial, count, censored and selected outcomes without confusing coefficients with effects;
- use machine learning for nuisance estimation and heterogeneity without claiming that prediction creates identification;
- report sensitivity, external validity, data provenance and a reproducible evidence chain.
The main route suits advanced undergraduates and taught postgraduates. Graduate extensions add formal identification, weak-identification-robust inference, staggered-treatment estimands, partial identification and policy learning.
Preparation
You need probability, expectations, confidence intervals, matrix-free OLS intuition and basic Python or R reading. Prior causal inference is not required.
Ten-minute readiness check
- Why can the same person’s outcome under treatment and no treatment not both be observed?
- Does adding more controls always reduce omitted-variable bias?
- Is a strong first-stage relationship enough to validate an instrument?
- Does an insignificant pre-trend prove parallel trends?
- If an RDD estimate is credible at a cutoff, does it identify an effect for everyone?
Answers: the missing potential outcome is counterfactual; no—mediators and colliders can introduce bias; no—the exclusion and assignment arguments remain; no—pre-tests may have low power and the post-treatment counterfactual is still unobserved; no—the canonical estimand is local to the cutoff.
Course route
| Module | Central question | Main deliverable |
|---|---|---|
| 0. Causal questions and designs | What effect, for whom, under which assignment process? | identification memo |
| 1. Regression and inference | What does OLS estimate, and when can it support a causal claim? | regression and uncertainty audit |
| 2. Instruments and panel data | Which variation isolates treatment, and what remains unidentified? | IV or within-unit design brief |
| 3. Policy evaluation designs | How do time, thresholds and comparison units construct ? | DID/RDD/weighting design |
| 4. Choice and limited outcomes | How does the outcome’s support change modelling and interpretation? | marginal-effect report |
| 5. Modern causal analysis | How should flexible prediction, sensitivity and policy learning enter? | cross-fitted or sensitivity analysis |
| 6. Research workflow | Can another researcher audit the complete evidence chain? | replication-ready research record |
| Capstone | Which design best evaluates the Pathways expansion? | policy memo plus technical appendix |
| Appendix | Which formula, diagnostic or source is needed? | revision and evidence map |
The identification memo
Complete this before running a regression:
| Field | Pathways example |
|---|---|
| decision | expand, redesign or stop scholarship offers |
| unit | one eligible school-leaver |
| treatment | offer received by 31 August 2024 |
| outcome | degree completion within five academic years |
| estimand | ITT among applicants in the 2024 eligibility window |
| assignment | lottery within score band 68–72 |
| counterfactual | outcomes of comparable applicants assigned no offer |
| interference | peer and university-capacity spillovers |
| missingness | completion not mature; migration may hide outcomes |
| inference | assignment at applicant level, schools may induce dependence |
| scope | applicants in participating districts and score band |
Changing “offer” to “enrolment” changes the estimand and usually the identification argument.
Four claims that must stay separate
| Claim | Example | Evidence needed |
|---|---|---|
| descriptive | recipients complete at a higher rate | transparent sample and denominator |
| predictive | prior scores predict completion | time-valid out-of-sample performance |
| causal | an offer raises completion by 6 points | credible treatment assignment and counterfactual |
| policy | expand offers to another score band | causal effects, costs, capacity, distribution and transportability |
A causal estimate need not answer a policy question if the affected population, treatment version or capacity changes.
A 13-week teaching route
| Week | Preparation | Workshop output |
|---|---|---|
| 1 | potential outcomes and estimands | one-page identification memo |
| 2 | experiments and causal diagrams | assignment and control audit |
| 3 | OLS, FWL and selection | projection versus causal interpretation |
| 4 | uncertainty and clustering | inference plan tied to design |
| 5 | IV and LATE | first-stage, exclusion and weak-IV audit |
| 6 | panel and fixed effects | within-unit comparison and threat map |
| 7 | 2×2 and staggered DID | group-time estimand and event-study plan |
| 8 | RDD | cutoff graph, bandwidth and falsification plan |
| 9 | matching and weighting | overlap and balance report |
| 10 | synthetic control | donor-pool and placebo audit |
| 11 | discrete, count and selected outcomes | marginal-effect interpretation |
| 12 | DML, heterogeneity and sensitivity | cross-fit or robustness-value exercise |
| 13 | capstone defence | recommendation, limitations and replication record |
Suggested workload is 90 minutes of preparation, two hours of workshop activity and 90 minutes of follow-up per week. Every executable cell has a static worked result; students may complete equivalent calculations in a spreadsheet or on paper. Tables carry the key values, and colour is never the only evidence channel.
Assessment alignment
This is a teaching template, not an institutional grading rule:
| Task | Suggested share | Evidence assessed |
|---|---|---|
| identification memo | 20% | estimand, assignment, assumptions and scope |
| method replication | 25% | data construction, estimator, inference and diagnostics |
| design comparison | 25% | counterfactual quality, sensitivity and communication |
| integrated capstone | 30% | complete evidence chain and policy boundary |
Postgraduate work should distinguish identification from estimation formally, justify the asymptotic or randomisation reference, and report a sensitivity or partial-identification analysis.
Current evidence window
The reading map was checked on 1 August 2026. Three updates shape this edition:
- Modern DID work separates group-time effects, aggregation and sensitivity rather than interpreting a single staggered two-way-fixed-effects coefficient automatically; see Callaway and Sant’Anna (2021), Borusyak, Jaravel and Spiess (2024) and Rambachan and Roth (2023).
- IV practice now includes quantitative sensitivity to exclusion and ignorability violations, not only first-stage strength; see Cinelli and Hazlett (2025).
- The AEA Data and Code Availability Policy, revised in February 2026, requires provenance and computational detail sufficient for replication, including when source data are restricted.
The reading and software map distinguishes foundational results, current methods, applied illustrations and live software documentation.
Study discipline
For every result, write six lines:
- estimand: which counterfactual contrast;
- assignment: where treatment variation came from;
- assumption: what makes the comparison causal;
- estimator: how the contrast was calculated;
- uncertainty: what varies in repeated assignment or sampling;
- scope: for whom, where, when and under which treatment version.
Start with causal questions and estimands.