Incident Response P0
Captured from the last six sev-0 events. Follow top to bottom; every step has an owner and a hard time budget.
Steps
Owner · Reliability Guild0 of 14 steps complete
Loading your saved progress…
Declare the incident
Open an incident channel and set severity to P0 in the tracker.
Assign incident commander
First responder holds command until explicitly handed over.
Page secondary on-call
Escalate to the platform secondary within 5 minutes of declaration.
Post the first status update
Public status page within 10 minutes, even with no root cause yet.
Open a mitigation track
Split investigation and mitigation into two parallel workstreams.
Attempt rollback first
Prefer reverting the last deploy over forward fixes during a P0.
Update stakeholders every 30 min
Cadence holds until severity is downgraded.
Confirm recovery signals
Error rate, latency and queue depth back to baseline for 15 minutes.
Downgrade severity
Commander announces the downgrade and remaining follow-ups.
Freeze related deploys
Hold the affected service until the postmortem action list exists.
Collect the timeline
Export channel history, alerts and deploy events into the doc.
Draft the postmortem
Blameless template, due within 48 hours of resolution.
Review with the guild
Walk the timeline, agree on contributing factors.
File and assign actions
Every action item gets an owner and a due date before closing.