Site Reliability Engineering- Beyer, Jones, Petoff, Murphy (Google, 2016), free online.Introduced SLOs, error budgets and toil - the vocabulary this map is built on.
The Site Reliability Workbook- Beyer, Murphy, Rensin, Kawahara, Thorne (2018), free online.Worked examples for putting the first book into practice.
Seeking SRE- Blank-Edelman, ed. (O'Reilly, 2018).Essay collection on SRE practice outside Google.
Implementing Service Level Objectives- Hidalgo (O'Reilly, 2020).Book-length treatment of SLIs, SLOs and error budgets: definition, measurement,
negotiation.
Observability and operations
Observability Engineering- Majors, Fong-Jones, Miranda, Parker (O'Reilly, 2nd ed.).The case for high-cardinality events over dashboards of predefined metrics in
distributed systems.
Distributed Systems Observability- Sridharan (O'Reilly, 2018), short report.Logs, metrics and traces and their trade-offs, in about 30 pages.
Practical Monitoring- Julian (O'Reilly, 2017).Alert design and monitoring anti-patterns; matches the map's observability
line.
How Complex Systems Fail- Cook (1998), free online.Eighteen short theses on how failures develop in complex systems.
PagerDuty Incident Response- PagerDuty (open documentation).PagerDuty's internal incident command process, published for adoption as is.
Engineering practice and delivery
Accelerate- Forsgren, Humble, Kim (IT Revolution, 2018).The research behind the delivery practices on the map's development line.
Release It!- Nygard (Pragmatic Bookshelf, 2nd ed., 2018).Stability patterns - timeouts, bulkheads, circuit breakers - that many map nodes
assume.
The DevOps Handbook- Kim, Humble, Debois, Willis (IT Revolution, 2nd ed., 2021).Flow, feedback and continual learning behind the map's cultural practices.
Chaos Engineering- Rosenthal, Jones (O'Reilly, 2020).Failure injection with hypotheses, steady-state metrics and blast-radius
control.