Skip to content
Incident Bridge · Standby
// THE RETAINER

Four embed tracks. 90 days to a system your team actually owns.

Kubernetes, Terraform/IaC, Observability, and Incident Response — staffed by senior platform engineers who ship under load. Book a 30-minute platform audit and get a scoped week-1-to-week-12 plan before you sign anything.

  • 71%MTTR drop, day 1–90 avg.
  • 4 minIncident bridge SLA, 24/7/365
  • 97%24-month client retention
  • 312AWS / GCP / CKA certifications on the bench
// SECTION 02

Four embed tracks. One senior bench. Zero juniors.

Every engineer on your retainer has shipped to production at scale for 8+ years. Pick one track or all four — same people, same SLA, fixed monthly price.

PRODUCTION 01

Kubernetes & Platform Engineering

Cluster topology, GitOps via Argo, multi-tenant isolation, upgrades without downtime, cost guardrails. We absorb the 2 AM etcd snapshots so your team doesn't.

  • [+] Cluster architecture review & HA topology
  • [+] Argo CD / Flux rollout with progressive delivery
  • [+] Upgrade runbook for n−2 → n−1 node pools
  • [+] FinOps pass: rightsizing, spot strategy, namespace budgets

Done looks like: a zero-downtime upgrade executed on a Tuesday afternoon, and a Friday incident that never pages a human.

PRODUCTION 02

Terraform & Infrastructure-as-Code

Module library, state hygiene, drift detection, policy-as-code with OPA, and a CI/CD path that fails closed. No more "who ran `terraform apply` in prod?"

  • [+] Internal module registry, versioned & documented
  • [+] Atlantis or Spacelift pipeline with required reviewers
  • [+] OPA / Sentinel policy bundle for SOC 2 controls
  • [+] Drift detection & weekly auto-reconcile report

Done looks like: every infra change is a reviewed PR, and prod state matches git to the byte.

PRODUCTION 03

Observability & SLO Engineering

OpenTelemetry instrumentation, Prometheus / Grafana or Datadog, error budgets, and dashboards an on-call can actually read at 3 AM without squinting.

  • [+] OTel collector rollout across services
  • [+] SLO catalogue with burn-rate alerts
  • [+] On-call dashboards (golden signals, not 90 panels)
  • [+] Cardinality budgets & log retention policy

Done looks like: error budgets are a board-level metric, and pager volume drops 60% in month two.

PRODUCTION 04

Incident Response & On-Call

A SteelRats engineer joins your bridge within 4 minutes — 24/7/365 — backed by a 312-certification bench and a runbook library built from 14,600+ resolutions.

  • [+] Shared on-call rotation (we take the weekends)
  • [+] 4-minute P1 bridge response, SLA-backed
  • [+] Blameless postmortem template + reviewer
  • [+] Quarterly chaos game-days with your team

Done looks like: zero customer-facing outages longer than 8 minutes — a number we tracked across 2023 and 2024.

// SECTION 03 — DELIVERY ARC

Day 1, Day 30, Day 90 — what changes, in plain language.

No vague "we ramp fast" claims. Here is the engagement, chapter by chapter, so you can mentally project against your own roadmap.

  1. CHAPTER 01 WEEK 01

    Embed & Stabilize

    Your SteelRats pod gets read-only access, sits in your standups, and starts triaging the open incident queue. We deploy the on-call bridge within 72 hours and our 4-minute SLA is in effect from day one. No "ramp-up period" — we ship on day three or we don't invoice for week one.

    • [+] Incident bridge live, 4-minute SLA active
    • [+] Read-only infra & repo audit (48 hrs)
    • [+] Top-10 pager-noise offenders triaged
    • [+] Kickoff doc: scope, owners, escalation tree
  2. CHAPTER 02 WEEK 04

    Harden & Hand Off

    By the end of week four, your top three reliability risks are closed or in flight, MTTR is trending down, and we've co-authored the runbooks your in-house team will own. Your people are running the bridge solo on the easy weeks — we shadow, not drive.

    • [+] MTTR trending −30% or better vs. baseline
    • [+] Top-3 reliability risks closed or in PR
    • [+] Runbook library indexed in your wiki
    • [+] First blameless postmortem co-led with your team
  3. CHAPTER 03 WEEK 12

    Own & Exit-Ready

    By day 90 the system is yours. Your team is running incidents, merging Terraform, and shipping cluster upgrades without us in the loop. We stay on retainer for the 4-minute bridge and the hard problems — but the platform is now a thing your engineers own, not a thing they babysit. Average MTTR reduction across 180+ environments: 71%.

    • [+] 71% MTTR reduction vs. day-1 baseline (avg.)
    • [+] In-house team owns bridge, upgrades, IaC merges
    • [+] Quarterly chaos game-days scheduled
    • [+] Renewal scope renegotiated to "hard stuff only"
[ OPERATING RECORD / 2019 — 2024 ]

Five years on the bridge. The receipts are below.

We don't lead with logo soup. We lead with the numbers we'd be embarrassed to fake — and the certifications our engineers carried across the threshold before they touched your cluster.

PRODUCTION
Active client companies
184
as of Nov 2024 — includes 11 Y Combinator W23 graduates
──── source: internal CRM, audited monthly
PRODUCTION
24-month client retention
97%
industry avg: 58% — Delta computed against Gartner 2024 SRE services baseline
──── cohort: clients onboarded 2022 — present
CNCF
Cumulative cloud + K8s certifications
312
AWS, GCP, and CKA across a 47-person engineering bench
──── verifiable via Credly badges on request
PRODUCTION
Incidents resolved in 2023
14,600+
Zero customer-facing outages lasted longer than 8 minutes
──── postmortems published quarterly
  • SOC 2 Type I certified 2024 Q2 — Type II audit in flight, not yet claimable
  • O'REILLY Incident Response Playbook Required reading at 230+ engineering orgs
  • CNCF Sandbox maintainers 3 projects · 41,000 GitHub stars combined
  • KUBECON Best DevOps Startup KubeCon NA 2023 — voted by program committee
[ BEFORE THE AUDIT CALL / FIVE QUESTIONS WE'D ASK ANYWAY ]

The awkward stuff, answered in writing.

Every VP of Engineering asks these before signing a retainer. We surface them here so the audit call is a working session — not a sales pitch for the basics.

  1. 01 ▸ Do you replace our in-house platform team?

    No. We embed senior engineers alongside your team — they pair with your SREs, write runbooks in your repos, and leave behind a system your people own. The deliverable is capability transfer, not headcount substitution. The only time we'd recommend reducing a platform role is when our embed has fully absorbed that function, and that decision is yours.

    ──── see: handoff playbook
  2. 02 ▸ Who owns the Terraform, the Helm charts, the dashboards?

    You do. Every line of IaC we write lives in your GitHub org from day one, under your license, your CODEOWNERS, your protected branches. Our engineers commit as outside collaborators and rotate off the commit trail at the end of the engagement. We retain zero IP claims on infrastructure code authored inside your stack.

    ──── see: client sample repos
  3. 03 ▸ What's the pricing model — and what if we have a quiet month?

    A fixed monthly retainer scoped to your stack and on-call tier. No hourly runaway. No surprise invoice after the postmortem. During a quiet month we'll proactively run chaos drills, refresh runbooks, and harden your golden path. You're paying for coverage and momentum, not minutes on a timesheet — so quiet months still produce deliverables.

    ──── sample SOW on request under NDA
  4. 04 ▸ What does the team actually look like during an embed?

    Two senior platform engineers as your day-to-day, a staff engineer rotating weekly for architecture review, and a solutions engineer on the bridge within 4 minutes, 24/7/365. Every engineer has shipped to production at scale for 8+ years — no juniors, no外包, no "shadow bench." You'll meet the full pod before signing the SOW.

    ──── see: embed tracks
  5. 05 ▸ What's the exit clause if this isn't working?

    Thirty days, written notice, no penalty. We hand over everything — repos, dashboards, runbooks, vendor relationships, the lot — and offer a 60-day transition window where one of our engineers stays on-call to onboard your team or a successor vendor. About 6% of clients exit each year; the most common reason is they hired in-house and no longer need us, which we count as a win.

    ──── exit frequency: 6.2% trailing 12 months
[ NEXT STEP / 30 MINUTES / NO DECK ]

Book the audit. We'll come back with a written scope.

A SteelRats solutions engineer spends 30 minutes inside your stack topology — Kubernetes manifests, Terraform state, observability coverage, on-call rotation — and returns within 48 hours with a written one-pager: where your MTTR is bleeding, what week-1 / week-4 / week-12 would look like, and a fixed retainer number. No sales call. No follow-up sequence.

Duration
30 minutes · video or async Loom
What we need
Read-only IAM scoped to one cluster or AWS org
What you get
Written scope + retainer quote within 48 hours