Paper and Evidence Audit
Paper and Evidence Audit
Audit claims, not page count
Create one row for every abstract, conclusion and policy claim:
| Claim | Type | Estimand | Evidence | Key assumption | Stress test | Scope |
|---|---|---|---|---|---|---|
| offer raises completion | causal | lottery ITT | mean contrast | correct random assignment | randomisation inference | score band 68–72 |
| receipt raises completion | causal | complier LATE | offer IV | exclusion, monotonicity | weak-IV/exclusion sensitivity | offer compliers |
| expansion is cost-effective | policy | net social value | effects + costs | transport and valuation | capacity/cost scenarios | named target districts |
If a policy sentence cannot be linked to a row, it is unsupported or underspecified.
Seven-pass review
- question: are treatment, outcome, population and horizon fixed?
- design: where does counterfactual variation come from?
- data: do units, timing, joins and exclusions match the design?
- estimation: does the estimator target the declared parameter?
- inference: what sampling or assignment variation justifies uncertainty?
- diagnostics: do tests probe the design’s actual vulnerabilities?
- communication: do title, abstract, tables and conclusion respect scope?
Do these in order. A polished coefficient table cannot rescue an undefined treatment.
Read tables as executable arguments
For every primary table or figure, verify:
- denominator and sample restrictions;
- treatment and outcome units;
- reference categories and omitted coefficients;
- standard-error or randomisation procedure;
- number of clusters and treatment-assignment level;
- whether controls are pre-treatment;
- support behind subgroup/event-time estimates;
- whether notes permit independent interpretation;
- consistency with the text and generated source values.
A ten-point coefficient can mean 10 percentage points, a 10% multiplicative change or 10 log points. Units belong in the title or note.
Organise robustness by threat
| Threat | Targeted analysis |
|---|---|
| confounding | negative controls, alternative adjustment, calibrated sensitivity |
| weak IV | first-stage diagnostics and weak-IV-robust intervals |
| non-parallel trends | cohort-time estimates, placebos and trend sensitivity |
| RDD manipulation | institutional/density audit and predetermined covariates |
| poor overlap | support, weight concentration and target-population change |
| attrition | arm-specific response, bounds and missingness sensitivity |
| spillovers | exposure mapping or cluster/market-level estimand |
Twenty specifications that all ignore the same threat are not twenty independent confirmations.
Record the garden of forking paths
Maintain a decision log for outcome definitions, samples, controls, bandwidths, horizons and subgroup choices. Distinguish:
- pre-specified primary analyses;
- planned secondary analyses;
- exploratory analyses discovered after seeing outcomes;
- corrections made after validation failures.
Exploration is valuable when labelled. Hidden exploration turns inferential uncertainty into false certainty.
Final red-team questions
- What observation would most weaken the identification argument?
- Which estimate has the least support but the strongest wording?
- Does one coding decision drive the result?
- What treatment version and population are silently assumed in the conclusion?
- Can an independent researcher regenerate every primary value?
The AEA’s May 2026 report on data and code policy is a current example of treating reproducibility requirements as an evolving research institution. Check dates and venue-specific rules rather than copying an old checklist.
Quick check
All robustness estimates have the same sign, but every confidence interval includes substantively important harm and benefit. Is “robustly positive” defensible?