Module 4 — Choice and Limited Outcomes
Module 4 — Choice and Limited Outcomes
An outcome’s support is part of the research design. A binary variable cannot fall below zero, a count has an exposure period, and observed wages exist only after an employment decision. The model should respect these facts without turning a causal question into a distributional exercise.
Prepare
Before the workshop, classify each outcome:
| Outcome | Support | First modelling question |
|---|---|---|
| completed degree | 0 or 1 | probability, risk difference or odds? |
| institution chosen | unordered alternatives | whose attributes vary: person or alternative? |
| emergency visits | non-negative integers | what population and exposure time generated the count? |
| observed wage | continuous but selected | is zero meaningful, censored or unobserved? |
Write one sentence defining the outcome, observation window and denominator. Many apparent “model failures” are outcome-definition failures.
Route
- Binary Choice: probabilities, nonlinear effects and calibration.
- Multinomial Choice: unordered, ordered and alternative-specific decisions.
- Count Outcomes: rates, exposure, overdispersion and excess zeros.
- Censoring and Selection: observed limits, participation and missing counterfactual outcomes.
Workshop: one policy, four outcomes
For the Pathways programme, compare:
- enrolment within one year: binary;
- university selected: multinomial;
- modules completed in year one: count;
- earnings five years later: continuous, but missing for people outside linked tax records.
For each, specify the estimand first. A nonlinear likelihood can improve fit; it cannot make scholarship receipt exogenous.
Follow-up deliverable
Submit a two-page marginal-effect report containing:
- outcome support and observation process;
- target contrast on an interpretable scale;
- model and identifying assumptions;
- one worked prediction or marginal effect;
- one diagnostic and one conclusion boundary.
Postgraduate extension: derive the likelihood contribution and distinguish structural distributional assumptions from assumptions needed for causal identification.