Reading and Evidence Map
Reading and Evidence Map
This selective teaching map was checked on 1 August 2026. It is not a systematic review. Sources were chosen for conceptual importance, methodological clarity, recency or direct relevance to the HarborMart decisions. Follow the links and check later amendments before applying them.
Read by question
What kind of claim am I making?
- Shmueli (2010), To Explain or to Predict?, distinguishes explanatory and predictive modelling aims.
- Bertsimas and Kallus (2020), From Predictive to Prescriptive Analytics, formalises decisions that use contextual data.
- Kleinberg et al. (2018), Human Decisions and Machine Predictions, shows why evaluating predictions requires attention to the decisions and selectively observed outcomes they affect.
Use: chapters 1, 8, 13 and 20. None of these licenses a causal claim from predictive accuracy.
How should evidence be displayed and evaluated?
- Cleveland and McGill (1984), graphical perception, provides experimental foundations for comparing visual encodings.
- Hullman et al. (2019), a survey of uncertainty visualisation evaluation, maps what uncertainty displays are intended to achieve.
- Saito and Rehmsmeier (2015), precision–recall versus ROC, explains why class imbalance changes evaluation interpretation.
- Guo et al. (2017), calibration of modern neural networks, motivates reliability checks after training.
Use: chapters 4, 7, 11 and 13. A display or metric is evidence only for its defined task, population and period.
How do experiments support decisions?
- Kohavi et al. (2009), controlled experiments on the web, gives practical experimentation principles and failure modes.
- Kohavi, Tang and Xu (2020), Trustworthy Online Controlled Experiments, develops design, metrics and organisational practice.
Use: chapters 8, 19, 20 and 23. Randomisation solves assignment confounding, not bad measurement, interference or non-compliance automatically.
How should models be interpreted and operated?
- Rudin (2019), interpretable rather than post-hoc explained models, argues for transparent models in high-stakes settings where feasible.
- Sculley et al. (2015), Hidden Technical Debt in Machine Learning Systems, catalogues system-level maintenance risks.
- Breck et al. (2017), The ML Test Score, offers a production-readiness rubric.
- Gebru et al. (2021), Datasheets for Datasets, structures documentation of dataset motivation, composition, collection and use.
- Mitchell et al. (2019), Model Cards for Model Reporting, structures intended use, evaluation and limitations.
Use: chapters 2, 14, 21 and 22. Documentation supports accountability only when it is accurate, reviewed and tied to controls.
Current research extensions
These sources connect established analytics to current decision and system questions:
- Longo et al. (2024), Explainable Artificial Intelligence (XAI) 2.0, surveys the move from isolated explanations toward broader, human-centred evaluation.
- Smyth et al. (2024), AI and prescriptive analytics for supply-chain resilience, reviews links among descriptive, predictive and prescriptive capabilities while identifying fragmented evidence.
- Sadana et al. (2025), a survey of contextual optimisation methods, connects prediction with downstream decisions under uncertainty.
- Dinh, Kotary and Fioretto (2024), fair multi-objective predict-then-optimise, shows that downstream performance may combine efficiency, robustness and fairness objectives.
- Cerreia-Vioglio et al. (2026), Making Decisions Under Model Misspecification, develops decision criteria that explicitly recognise models as approximations.
Teaching boundary: these works are not interchangeable. A review maps a field; a theoretical paper establishes results under assumptions; an empirical study estimates within its design. Translate each into the claim it can support.
Official governance sources
| Source | What it contributes | Boundary |
|---|---|---|
| NIST AI RMF | voluntary govern–map–measure–manage structure | NIST says version 1.0 is being revised |
| NIST Generative AI Profile | GenAI-specific risk actions, released July 2024 | profile, not a universal legal rule |
| ISO/IEC 42001:2023 | requirements for an AI management system and continual improvement | standard text and certification scope require careful interpretation |
| OECD AI Principles, updated 2024 | human-centred values and policy recommendations | principles, not an implementation test suite |
| EU Regulation 2024/1689 and 2026/1744 amendment | binding EU framework and current amendments | identify jurisdiction, role, system category and date; obtain legal advice |
The Commission's AI Omnibus update was last updated 31 July 2026. It is included to demonstrate why compliance slides need dates.
Organisational cases
- Airbnb Chronon illustrates feature lineage, temporal backfills, online/offline consistency and drift monitoring.
- Uber's 2024 investor update describes matching, routing, dispatch, pricing and incentives as linked marketplace decisions.
- Walmart's fiscal-2026 10-K describes AI tools, automation and supply-chain investment within a large operating system.
These are primary accounts of organisational practice. They can illustrate architecture and managerial claims; they are not independent estimates of causal impact.
Emerging evidence
Sun et al. (June 2026), Decision-Centered Learning in Closed-Loop Optimization and Control, is an emerging working paper. It is included to show a live research direction—learning evaluated through downstream decisions and feedback—not as settled or peer-reviewed evidence.
When citing emerging work, label its status, version and access date. Prefer established results for core teaching and use preprints to formulate questions or replication exercises.
A five-line reading note
For any source, record:
- question and claim;
- data or formal setting;
- identification or assumptions;
- main result and uncertainty;
- what it changes—and does not change—in a HarborMart decision.