Cloud-native principles for any SaaS team that runs production systems.
8 domains · 28 rules.
DETECT · DECLARE · MITIGATE · REVIEW
In complex distributed systems, incidents are inevitable. The engineering quality of a team is not measured by whether incidents happen — it is measured by how fast and how well they respond when they do.
A database running at 100% CPU is not an incident if users are unaffected. A 2% failure rate on checkout affecting 500 paying customers is a SEV1. Customer impact is the only axis that matters. Attach concrete criteria and response SLAs to each level — vague thresholds cause delayed escalation.
All hands. Acknowledge within 5 min, IC assigned within 15 min. Examples: checkout down, authentication broken, data breach in progress, full service outage.
Acknowledge within 15 min. SLO burning fast. Workaround may exist. Examples: elevated error rate on a core flow, intermittent failures, major feature unavailable for a subset of tenants.
Acknowledge within 1 hour. Business hours response acceptable. Examples: slow page, non-critical feature degraded, single tenant impacted with workaround.
Normal sprint process. No on-call page. Examples: a dashboard shows wrong data, a background job is slower than usual, an internal tool is unavailable.
Every incident needs clear, named roles before the incident starts. When roles are invented in the moment, coordination collapses, communication stops, and the most senior engineer ends up doing everything.
Silence during an incident causes customers to file tickets, executives to escalate, and support queues to overflow — all of which add load at the worst moment. Communication is not a post-fix courtesy. It is an active mitigation tool.
The priority during an active incident is returning customers to a working state — not finding root cause. Understanding comes in the post-mortem. A rollback you don't fully understand is better than a fix you're still designing while users are impacted.
Root cause analysis is a post-incident activity. During the incident, the only question is: what action restores service fastest? Understanding what caused the failure comes after customers are unblocked, never during.
If the last deployment correlates with the incident, roll it back immediately — before investigating. Feature flags, blue/green deploys, and canary releases exist specifically to make rollbacks fast and cheap. Maintain a deployment timeline inside every incident timeline.
Feature flags that disable expensive or failing functionality in seconds are a tier-1 incident tool. When a payment provider is failing, you need to disable it without a deploy. Kill switches must be tested outside of incidents — they are useless if discovered broken during one.
If the on-call hasn't made progress toward mitigation within 30 minutes, they escalate — not continue debugging in isolation. Time-boxing forces escalation paths to be real, tested, and known. A single engineer blocking a SEV1 is an organizational failure, not a personal one.
An alert without a runbook sends the on-call engineer into an incident with no map. Runbooks don't need to be long — they need to exist, be accurate, and be linked directly from the alert so they are found in under 10 seconds at 3am.
A post-mortem that produces no action items is documentation theater. A post-mortem that names an individual as the root cause is counterproductive. The goal is systemic understanding that prevents recurrence — and learnings that the whole organization can use.
People are never the root cause. Systems, processes, tooling, and conditions are. When engineers fear appearing in a post-mortem as the cause of an incident, they write sanitized timelines. Blameless culture is the only way to get the honest account needed to prevent recurrence.
The timeline runs from the first observable anomaly to full resolution. It is built from logs, deploy records, alert timestamps, and chat history — not from memory. Gaps in the timeline are gaps in observability. Every gap is a follow-up action item.
A post-mortem with no owners and no dates is a list of good intentions. Every contributing factor maps to a concrete remediation: a runbook update, a monitoring gap closed, a kill switch added, a process change. No owner and deadline means it does not happen.
The learnings from one team's incident are often directly relevant to three other teams running similar systems. Post-mortems should be indexed, searchable, and accessible company-wide. Tribal knowledge locked inside one team prevents the organization from learning.
A rotation that requires heroism to sustain will eventually fail. Alert fatigue, burnout, and turnover are direct, measurable costs of unsustainable on-call. On-call toil is a product defect — it belongs in the sprint backlog, not the "we'll get to it" list.
Declare early. Assign roles immediately.
Mitigate before you understand.
Communicate before customers ask.
Post-mortems find system failures, not people failures.
On-call toil is a bug. Fix it in the sprint.