Module 6 — From Data to Auditable Evidence

Paper and Evidence Audit

Review an empirical paper claim by claim, from estimand and assignment through diagnostics, uncertainty and scope

Paper and Evidence Audit

Audit claims, not page count

Create one row for every abstract, conclusion and policy claim:

ClaimTypeEstimandEvidenceKey assumptionStress testScope
offer raises completioncausallottery ITTmean contrastcorrect random assignmentrandomisation inferencescore band 68–72
receipt raises completioncausalcomplier LATEoffer IVexclusion, monotonicityweak-IV/exclusion sensitivityoffer compliers
expansion is cost-effectivepolicynet social valueeffects + coststransport and valuationcapacity/cost scenariosnamed target districts

If a policy sentence cannot be linked to a row, it is unsupported or underspecified.

Seven-pass review

  1. question: are treatment, outcome, population and horizon fixed?
  2. design: where does counterfactual variation come from?
  3. data: do units, timing, joins and exclusions match the design?
  4. estimation: does the estimator target the declared parameter?
  5. inference: what sampling or assignment variation justifies uncertainty?
  6. diagnostics: do tests probe the design’s actual vulnerabilities?
  7. communication: do title, abstract, tables and conclusion respect scope?

Do these in order. A polished coefficient table cannot rescue an undefined treatment.

Read tables as executable arguments

For every primary table or figure, verify:

  • denominator and sample restrictions;
  • treatment and outcome units;
  • reference categories and omitted coefficients;
  • standard-error or randomisation procedure;
  • number of clusters and treatment-assignment level;
  • whether controls are pre-treatment;
  • support behind subgroup/event-time estimates;
  • whether notes permit independent interpretation;
  • consistency with the text and generated source values.

A ten-point coefficient can mean 10 percentage points, a 10% multiplicative change or 10 log points. Units belong in the title or note.

Organise robustness by threat

ThreatTargeted analysis
confoundingnegative controls, alternative adjustment, calibrated sensitivity
weak IVfirst-stage diagnostics and weak-IV-robust intervals
non-parallel trendscohort-time estimates, placebos and trend sensitivity
RDD manipulationinstitutional/density audit and predetermined covariates
poor overlapsupport, weight concentration and target-population change
attritionarm-specific response, bounds and missingness sensitivity
spilloversexposure mapping or cluster/market-level estimand

Twenty specifications that all ignore the same threat are not twenty independent confirmations.

Record the garden of forking paths

Maintain a decision log for outcome definitions, samples, controls, bandwidths, horizons and subgroup choices. Distinguish:

  • pre-specified primary analyses;
  • planned secondary analyses;
  • exploratory analyses discovered after seeing outcomes;
  • corrections made after validation failures.

Exploration is valuable when labelled. Hidden exploration turns inferential uncertainty into false certainty.

Final red-team questions

  1. What observation would most weaken the identification argument?
  2. Which estimate has the least support but the strongest wording?
  3. Does one coding decision drive the result?
  4. What treatment version and population are silently assumed in the conclusion?
  5. Can an independent researcher regenerate every primary value?

The AEA’s May 2026 report on data and code policy is a current example of treating reproducibility requirements as an evolving research institution. Check dates and venue-specific rules rather than copying an old checklist.

Quick check

All robustness estimates have the same sign, but every confidence interval includes substantively important harm and benefit. Is “robustly positive” defensible?

Answer
No. Sign stability of point estimates does not remove uncertainty. Report the interval, power or identified set and state that the evidence does not distinguish the decision-relevant alternatives.

Next: Capstone

Copyright © 2026