Skip to content
Incident Bridge · Standby
INCIDENT BRIDGE: STANDBY PAGER DUTY · ALL CLEAR · 02:14:09 UTC

O'REILLY · 2023 Now in second printing

The Incident Response Playbook

The on-call manual now required reading at 230+ engineering organizations.

Co-authored by two ex-Google SREs and a former AWS principal engineer, this 312-page hardcover distills 14,600+ resolved production incidents into the runbook we wish someone had handed us in 2014. No theory. No vendor slides. Just the sequence of decisions, severity calls, and postmortem templates that actually shorten an outage.

  • ISBN978-1-4920-XXXX-X
  • FORMATPrint · ePub · PDF
  • LENGTH312 pages · 7 chapters

ORIGIN STORY · 2019 → 2023

Why we wrote it: 14,600 production incidents, one methodology.

SteelRats started in a WeWork on East 6th Street, in Austin, in 2019. Two of us had shipped Search and Ads reliability at Google. The third had run the principal engineer's seat at AWS for RDS. Within our first eighteen months running a client services shop, we realized the same incident was happening three, four, sometimes five times a year at different companies — same root cause, same severity confusion, same postmortem that pointed fingers and changed nothing.

By the end of 2022, we had resolved 14,600 production incidents across 184 client environments. We had a methodology. We had runbooks. We had a hiring rubric for on-call rotation. So we sat down and wrote the book we wished existed in 2014: a field manual, not a manifesto.

Published by O'Reilly in October 2023, the hardcover sold through its first printing in eleven weeks. The second printing shipped in Q1 2024. It is now required reading at 230+ engineering organizations, used in staff-engineer promo packets at three FAANG-adjacent firms, and assigned as the week-one read for every new platform engineer we hire.

/adoption-telemetry

Adopted faster than we expected.

Numbers as of November 2024, pulled from the O'Reilly catalog feed, GitHub API, and our internal training-program rosters.

/orgs_required_reading 230+ ↑ +47 in last 90d engineering organizations
/oss_stars_combined 41,000 ↑ +1,204 / 30d across 3 CNCF Sandbox projects
/first_printing 11w → sold out in second printing 2024-Q1
/ycombinator_grads 11 → W23 cohort using as staff-eng promo packet

TABLE OF CONTENTS · 7 CHAPTERS

Seven chapters, one on-call rotation you'll actually survive.

No marketing in this list — these are the actual chapter titles from the 312-page hardcover, in order. Skim them. If three of them solve a problem on your team's roadmap this quarter, the book is for you.

  1. 01

    Defining Severity Without a Committee Meeting

    A four-tier severity matrix that fits on one page and resolves 80% of "is this a Sev2 or Sev3?" Slack threads. Includes the exact incident-classifier script we ship to every client in week one.

  2. 02

    Building the Bridge: Roles, Channels, and Decision Logs

    Who joins, who commands, who takes notes, who calls the rollback. Template Slack channels, the on-call commander's checklist, and the live decision log we keep in plain text.

  3. 03

    Incident Triage: The First Eight Minutes

    This is the chapter we give away. A minute-by-minute script for what to do between page-ack and your first comms update — including the four questions you should ask before opening any dashboard.

    ↑ Sample chapter available below
  4. 04

    Communication Cadence: Status Pages, Stakeholders, and the 15-Minute Update

    The comms template, the status-page commit grammar, and the rule for when the CEO gets texted. Plus a worked example from a 47-minute database failover in 2022.

  5. 05

    On-Call Rotation Design: Compensation, Handoffs, and Burnout Math

    The rotation math nobody puts in the SRE handbook: page volume, follow-the-sun handoffs, comp parity across regions, and the empirical threshold at which senior engineers start quitting.

  6. 06

    Blameless Postmortems That Actually Change Code

    Our postmortem template, the five-question timeline reconstruction, and the action-item tracker that has a 78% close-rate at 90 days (vs. the industry norm of 31%).

  7. 07

    Cutting MTTR: The Four Levers That Move the Number

    Observability, rollback latency, runbook quality, and pager hygiene — the four forces that, when fixed in order, cut mean time to resolution by 71% inside 90 days across our client environments.

FIELD TESTIMONIAL · Q3 2024

"We bought three copies for every platform engineer in the org. The severity matrix in chapter one replaced a 14-page internal Google Doc that everyone had stopped reading in 2021. Chapter three — the eight-minute triage script — is taped to the side of every on-call monitor in the office. If you run production, you need this book. If you don't run production, you need to read it before your next interview."

Priya Raman VP of Engineering · mid-stage fintech, Series C Required-reading cohort, Q2 2024

FREE SAMPLE · CHAPTER 3

Read Chapter 3 free, then book your platform audit.

Drop your work email. We'll send you the 38-page Chapter 3 PDF (Incident Triage: The First Eight Minutes) instantly — no drip campaign, no sales sequence. You'll get one follow-up from a SteelRats engineer asking if the chapter held up to your environment, and one link to book a 30-minute platform audit if you want one. Nothing else.

  • 38-page PDF, no DRM, print-friendly
  • Includes the incident-classifier script (chapter 1 bonus)
  • One human follow-up, no drip

We don't sell or share your email. SOC 2 Type I audited deliverability. Unsubscribe lives in every email we send, and we honor it the same hour.

CAPTURE · /chapter-3-gate

Send me Chapter 3 (PDF).

READY TO TALK TO AN ENGINEER?

Book Your Free Platform Audit →

30 min · with a SteelRats solutions engineer · no slide deck