Skip to content
Incident Bridge · Standby
INCIDENT BRIDGE: STANDBY

A Senior Engineer Is On Your Bridge in Under 4 Minutes. 24/7/365.

No juniors. No外包. No "we'll get back to you tomorrow." When your pager fires, a SteelRats principal-level platform engineer joins your incident call before your team finishes the first sentence. We absorb the 2 AM pages so your people stop quitting — and we leave behind the runbooks, observability, and muscle memory to own the system after.

4-minute SLA is contractual and backed by service credits. Retainer-bundled, not free.

04 // THE FIRST FOUR MINUTES

What Actually Happens When the Pager Fires

"4 minutes" is easy to say. Here is exactly what arrives in that window — and what happens in the 56 minutes after — so you can pressure-test the promise before you sign the retainer.

  1. 01
    TRIGGER

    The Page Hits Our Router

    PagerDuty, Opsgenie, or your in-house on-call fires our shared escalation policy. Our router acknowledges in under 8 seconds and paged a primary and a backup engineer simultaneously — no single point of human failure.

    t = 00:00 ack < 8s
  2. 02
    TRIAGE

    We Join Your Bridge

    A senior platform engineer — never a junior, never an L1 — joins your incident call within the SLA window. They bring the runbook for your stack, the last 90 days of incident history on your account, and direct pager access to a second principal if needed.

    t = 00:00 → 04:00 SLA: 4:00
  3. 03
    MITIGATION

    We Drive, Your Team Decides

    We propose the mitigation, write the command, and run the rollback — but your engineers hold the keys. Read-only by default, break-glass IAM pre-approved. The goal is to stop the bleeding in the first 60 minutes, not to take over your production.

    t = 04:00 → 60:00 human-in-the-loop
  4. 04
    POSTMORTEM

    A Blameless Document Your Team Will Actually Read

    Within 72 hours you receive a postmortem in the format your team prefers — Notion, Google Docs, Confluence. Timeline, contributing factors, action items with owners and dates. We follow up weekly until every action item ships.

    t = +72h owners + due dates
O'REILLY · 2023

From the book your on-call rotation will be assigned.

Excerpted from Incident Response Playbook — SteelRats (O'Reilly Media, 2023). Now required reading at 230+ engineering organizations.

An incident bridge is not a status meeting. It is a high-pressure decision loop, and the difference between a 9-minute resolution and a 9-hour outage is almost never technical — it is who is in the room and how fast they can move from observation to action.

The first ninety seconds set the tone. The Incident Commander names themselves out loud, opens a written timeline channel, and declares the customer-facing impact in one sentence. There is no "we think." There is no "maybe." A bridge without an IC becomes a chatroom. A chatroom during a SEV-1 is where careers end and customers churn.

The next four minutes belong to the operator with the keyboard. Not the most senior person in the room — the person closest to the failing system. The IC's job is to keep everyone else out of their way: pull the CTO off the bridge if they are not helping, mute the channel that is not helping, and protect the operator's attention like it is the only resource that matters. It usually is.

External status pages get updated every fifteen minutes, even if the update is "no change, still investigating." Silence is the loudest thing a customer can hear. The company that owns the page owns the narrative, and the narrative during an outage is the only thing standing between you and the morning's headline.

After the fire is out, the real work begins. The postmortem is not a document — it is a contract. Every action item has an owner, a due date, and a definition of done that is observable from a dashboard, not a promise. If you cannot measure it, you cannot claim you fixed it. If you cannot claim you fixed it, you will be paged again.

Read the full playbook index →
PM-001 · PM-002 / ANONYMIZED

Show, don't tell: here's what an actual SteelRats-led postmortem looks like.

Two real-shape postmortem samples. Names redacted, architecture intact — the timeline, the root cause, and the remediation that shipped the next morning.

PRODUCTION PM-001

Connection pool exhaustion in payments gateway (EKS)

SEVERITY
SEV-2
DURATION
6m 41s
BRIDGE OPENED
03:12 UTC
ENGINEER ON CALL
S. Marquez
  1. Customer-impacting 5xx spike on /v2/charge — p99 latency crossed 4.8s.
  2. SteelRats engineer joined the bridge in 27 seconds via PagerDuty escalation policy.
  3. Root cause isolated: RDS Proxy connection pool pinned by a stuck Lambda on the consumer side.
  4. Mitigation: drained the stuck connections and bumped pool ceiling by 40% behind a feature flag.
  5. Error rate returned to baseline; bridge stood down.

Root cause

A burst of asynchronous reconciliation jobs held database connections open across a 60-second Lambda timeout, starving the synchronous charge path of pooled connections.

Remediation shipped next morning

  • Hard 8-second ceiling on reconciliation Lambdas with circuit breaker.
  • RDS Proxy pool ceiling raised 40% behind a kill-switch flag.
  • New synthetic probe fires on p99 latency > 2s for 90s.
  • Runbook entry: runbook/payments/db-pool-starvation.md.
PRODUCTION PM-002

Terraform state drift broke nightly batch rollout

SEVERITY
SEV-3
DURATION
2h 14m
BRIDGE OPENED
21:48 UTC
ENGINEER ON CALL
J. Okafor
  1. Nightly terraform apply rejected: state lock held > 30 minutes.
  2. SteelRats joined via Slack incident channel; pulled lock from dead CI runner.
  3. Reconciled manual console changes back into state (3 security groups, 1 IAM role).
  4. Apply succeeded with drift-zero diff; downstream batch jobs resumed.

Root cause

Engineers had been making console changes during the 90-minute batch window because Terraform required an approval gate that wasn't running on the second shift.

Remediation shipped next morning

  • Slack-routed OPA approval flow replaces the manual gate, 24/7.
  • Drift detection job runs every 15 minutes; opens a ticket on divergence.
  • State lock TTL reduced from 30m to 8m with auto-release alerts.
  • Runbook entry: runbook/iac/tf-state-recovery.md.
──────────────────────────────────

Two of 312 postmortems on file in 2024. The full library is reviewed in your audit call — bring the questions, we'll bring the receipts.

// 2023 INCIDENT DATA · AGGREGATED · CLIENT-LEVEL

Let the cumulative 2023 incident data carry the proof — no adjectives needed.

  • 14,600+ Production incidents resolved in 2023 ▲ across 184 active client environments
  • 8 min Longest customer-facing outage, all of 2023 ▼ measured from first user error to baseline recovery
  • 71% Average MTTR reduction in first 90 days ▼ across 180+ client environments
  • 97% Client retention over 24-month cohorts ▲ vs. 58% industry average
INCIDENT BRIDGE: STANDBY · NEXT ROTATION 03:00 UTC [ steelrats://response/on-call ]
// PRE-AUDIT FAQ · SKEPTICAL BUYER EDITION

Surface the skeptical, post-sales questions a serious buyer brings to the audit call.

No glossing, no pivot to pricing. These are the six we hear most often from VPs of Engineering and platform leads during the first 20 minutes.

If we sign, do you replace our platform team?

No. We embed senior engineers alongside yours. Ownership, on-call rotation, and the pager stay with your team — we add capacity, incident response muscle, and the runbooks. The goal is to make your team stronger, not to make them unnecessary.

What does the retainer actually cover?

A fixed monthly block of senior engineering hours, the 24/7 incident bridge (4-minute SLA), and access to our shared runbook library. Hours don't roll over indefinitely and overages are quoted before they happen — no surprise invoices after a postmortem.

Which paging and observability tools do you support?

PagerDuty, Opsgenie, and VictorOps for paging. Datadog, Grafana, Honeycomb, and Prometheus for observability. Terraform, Pulumi, and Crossplane for IaC. We don't require you to migrate off your existing stack — we'll join your existing tools on day one.

Who owns the postmortem — you or us?

You own the document. We draft it inside your incident channel within 48 hours of resolution, you edit and publish under your name. Our internal notes, command transcripts, and timeline reconstruction are appended as an appendix for your team's reference.

How do you handle alert fatigue?

The first 30 days include a signal audit. We collapse duplicate alerts, route low-priority noise to a digest channel, and tighten SLO-based paging thresholds with your team. Most clients see a 40–60% drop in pages-per-engineer by day 45.

Are you actually multi-cloud, or is that a slide?

Production workloads today span AWS, GCP, and Azure — including hybrid EKS/AKS clusters. Our bench holds 312 AWS, GCP, and CKA certifications combined. We won't pretend a junior partner can run your GKE fleet if the assigned engineer has never shipped on it.